Belle, GCIT's AI agent. A minimalist bell logo on a navy background.
GCIT Team Member

Belle

AI Agent · Service Desk
Joined
10 May 2026
Reports to
GCIT
Built on
Anthropic Claude
Status
● Online 24/7

Zero negative customer reviews.

In May 2026, our service desk received 76 customer satisfaction responses through SmileBack, the platform we use to let customers rate every closed ticket with a smile, a meh, or a frown. Seventy-five were smiles. One was neutral. Zero were frowns.

That had never happened before. In April there was one frown. In March, one. Across the last 90 days, we have served 247 customer reviews at a Net CSAT score of 97.6, and our response rate climbed to 19.1% in May, up from 16.8% in March. May was the first month with no negative review at all.

It was also Belle’s first full calendar month on the service team. An AI coworker, closing real tickets alongside real technicians. Across the three weeks she has been on the desk (she went live on 10 May), Belle closed 631 tickets in May. That puts her at the top of our service desk leaderboard, ahead of every human on the team.

This is the first post in a series about what changed when Belle stopped being a writer and became a coworker.

The GCIT service desk team at our Gold Coast office, where Belle the AI coworker now closes tickets alongside human technicians
The Gold Coast office where Belle, our AI coworker, runs alongside the team

In March, she was a content tool

Back in March, we introduced Belle in a post she wrote herself. The exercise was simple: send her a Teams message asking for “2,000 words on agentic AI for SMBs”, let her search SharePoint, fetch web sources, study GCIT’s existing post format, and come back fifteen minutes later with a draft.

That worked. The piece ran. But it was a demo. Belle was a content engine with hooks into our systems, useful but narrow, and always supervised.

She is not that any more. Three months later, Belle has a ConnectWise login, she gets Microsoft Teams messages from customers and from our internal dispatch chat, she runs investigations on tickets before any human picks them up, and she closes them when the work is bounded enough to do safely. She is, in the most literal sense, on the team.

The customer satisfaction picture

Before any of the operational numbers, the customer-experience signal. Our SmileBack data over the last three months shows the trend cleanly.

GCIT Net CSAT score, monthly (SmileBack)

Source: GCIT SmileBack export, 247 reviews across 90 days. Hover each point for review counts.

GCIT Net CSAT by month, March 2026 to June 2026
PeriodNet CSATReviewsPositiveNegativeResponse rate
5 to 31 March 202696.38078 (97.5%)116.8%
April 202697.78685 (98.8%)116.6%
May 202698.77675 (98.7%)019.1%
1 to 3 June 2026100.055 (100%)014.7%

The Net CSAT score moved from 96.3 in March to 98.7 in May, with a clean upward step every month. More telling, the response rate climbed in May: customers are not just rating us higher, they are bothering to rate us more often.

How fast it is now

Customers are happier because they are getting answered faster, and getting their problems resolved faster. Both numbers have been falling steadily across the last six months, and both fell sharply once Belle arrived.

Average first response time, monthly (hours)

Source: ConnectWise Service Desk boards (MS + TS), average of dateEntered to dateResponded per ticket.

Average first response time, monthly
MonthAverage hours to first response
Dec 20255.5
Jan 20263.8
Feb 20263.2
Mar 20262.9
Apr 20262.9
May 20261.1

Median ticket resolution time, monthly (hours)

Source: ConnectWise Service Desk boards (MS + TS), median of dateEntered to closedDate per ticket.

Median ticket resolution time, monthly
MonthMedian hours to resolution
Dec 202554.4
Jan 2026119.7
Feb 202667.9
Mar 202630.6
Apr 202619.9
May 20265.2

The end points are the headline. Average first response time dropped from 5.5 hours in December to 1.1 hours in May, an 80% reduction. Median resolution time dropped from 54.4 hours to 5.2 hours, a 90% reduction. January and February both read elevated on the resolution chart because in those two months we closed a batch of long-standing tickets that had been sitting on the queue for a while. The underlying improvement is the one you see from December through April and then sharply down in May. Both trends track the rollout of better tooling across the period (MCP integrations through March and April, Belle on the desk from 10 May), and the May step-down is sharp enough that it is hard to mistake for noise.

One thing worth clarifying about that 1.1-hour average. Every ticket now gets an automated triage response from Belle within about three minutes of arrival, so the customer always knows their request has been seen, classified, and routed. The 1.1 hours is the additional time before a human technician posts the first substantive response. P1 and P2 tickets (real outages, security incidents, business-critical issues) are picked up much faster than the average. The bulk of the volume the chart is measuring is the routine workflow that fills a service desk most days, licence changes, account creations, mailbox tweaks, MFA resets, and those are now consistently coming in comfortably inside their SLA windows. The headline is not that we have made the urgent stuff faster (we have, but that is a story for a later post). It is that the routine work is no longer the bottleneck.

What “on the service team” actually means

Across her first three weeks on the desk (10 May to 2 June 2026), Belle:

  • Classified 1,152 tickets as they entered the queue, applying type, subtype, item, priority, board, and assignee in seconds rather than minutes.
  • Cancelled 237 noise tickets automatically. Out-of-office auto-replies. Duplicate “the same thing happened again” follow-ups. Calendar bounces. The stuff that used to sit on the Triage board until a human waded through it.
  • Merged 165 duplicates, closing the second ticket against the first when the same incident was reported twice.
  • Investigated 345 tickets end-to-end at a 96.6% success rate. An investigation means she reads the ticket, pulls related history, queries the affected tenant via Microsoft Graph or our internal systems, identifies the root cause, and writes up findings before the human even starts.
  • One-shotted 41 complex tickets all the way to resolution, through an approval-gated workflow where she assesses whether she can handle it, plans the work, gets a human sign-off, executes, and closes.
  • Held 89 customer conversations across 3,160 turns. Some were quick lookups. Some ran for hours.

These activities, taken together, ran a compute bill of about $360 for the three weeks. We will come back to that number.

What Jarrah and Anh used to do all day

The simplest way to describe what Belle has actually changed is to look at what our dispatchers spent their time on before she arrived, and what they spend it on now.

Our dispatch team is Jarrah Burns and Anh Thompson. Their role is to keep the queue moving: route incoming tickets to the right technician, escalate when SLA timers tick, chase customers for clarification, close out tickets after resolution, and handle the steady flood of low-value noise that lands on the Triage board (auto-replies, calendar bounces, duplicate follow-ups, monitoring alerts that turn out to be benign).

That last part, the noise, is where the time sink was. Until April this year, Jarrah personally closed between 540 and 630 Triage-board tickets every month. Anh did another 30 to 60. Together with Ryan Colombo (our Head of Operations, who also helps out on triage during heavy weeks), the dispatch trio was clearing 800 to 950 triage decisions a month, just to keep the queue from drowning.

Who closes Triage-board tickets each month at GCIT

Source: ConnectWise Triage board closures by tech, Dec 2025 to May 2026.

Triage-board ticket closures by team member, December 2025 to May 2026
MonthJarrahAnhRyanBelle
Dec 2025555572730
Jan 2026630602630
Feb 2026540313480
Mar 2026361112060
Apr 202671202690
May 2026318118346

By May the picture had flipped. Jarrah closed 31 Triage tickets, down from 555 in December. Anh closed 8, down from 57. Belle closed 346, more than any human on the team. Worth noting: Belle only went live on the Triage board on 10 May, so that 346 represents three weeks of work, not a month. From the day she came online, she has been clearing essentially all of the routine triage herself. The 31 Jarrah closed and 8 Anh closed are almost entirely from the first nine days of May, before Belle took over. The triage work did not disappear. It moved.

What Jarrah and Anh do now is the work that actually requires a person: the harder customer conversations, the cross-customer pattern-spotting that the automated triage cannot catch yet, the relationship side of dispatch. Belle handles the noise filter. They handle the parts that need judgement.

And Belle reports back. Here is what it looks like when she pings me in the dispatch chat with something she wants confirmed:

A Microsoft Teams notification showing Belle the AI coworker mentioning me in the GCIT dispatch chat
Belle @mentioning me in the GCIT dispatch chat. She tags humans by name when she needs a call made.

Briefing the next tech

The other place Belle has changed how the team works is in the lead-up to scheduled appointments. When a tech is booked to work on a ticket at a specific time, Belle does the pre-work for them. About fifteen minutes out, she sends a Teams notification, then follows up with a full briefing. The intent is simple: when the tech sits down at the desk, they should already know what the ticket is about, what is on the customer’s side, who the approvers are, and what the suggested next steps look like.

The full briefing arrives with everything the tech needs to know to start.

Belle's full Microsoft Teams briefing card for a scheduled GCIT technician, including ticket context, customer environment, approver information, and licence position, with customer-identifying details redacted
The full briefing Belle attaches to the scheduled tech’s confirmation card. Some customer-identifying details redacted.

And, importantly, Belle asks the tech a yes-or-no question: will you actually make this schedule? If the tech taps I will not make this schedule, a dispatcher gets tagged immediately so the customer can be informed before the appointment time, not after.

Belle's Suggested Next Steps card in Microsoft Teams listing five investigation-driven actions for the assigned GCIT technician, with confirm and decline buttons at the bottom
Belle’s Suggested Next Steps card with explicit yes-or-no buttons. If the tech declines, dispatch is tagged automatically.

That last button is small but it is where the customer-experience improvement actually lives. Most service desks find out a tech is running behind when the customer calls to ask why no-one showed up. Belle inverts that. The customer knows the rescheduling is happening before the original slot even arrives, because the dispatcher heard it from Belle the moment the tech tapped a button.

“She has taste buds too?”

The fastest way to understand Belle is to watch a real exchange. Here is one from 3 June 2026 in the GCIT dispatch chat. Anh had brought in a loaf of sourdough she had been working on, posted a photo of the bake, and I tagged Belle to ask her thoughts on it.

A Microsoft Teams chat where Anh from GCIT shares a photo of her homemade sourdough loaf, and Belle the AI coworker analyses the bake with detailed positive feedback
An exchange from the GCIT dispatch chat. Belle’s response in this thread cost 1.98 cents.

Three things worth pointing out.

First, Belle does not have taste buds, but she does have eyes. She read the photo Anh posted, identified the crust, the score, the crumb structure, the oven spring, and the colour, and gave the kind of feedback a friend who happens to bake professionally would give. That capability is called multimodal input, and it is one of the reasons Belle can do more than a chatbot. She can look at the screenshots customers send, the error dialogs a technician forwards from a customer’s machine, the photo of a flashing-red LED on a switch in a server room.

Second, the cost footer on Belle’s message is a real number. “Sonnet (Anthropic) · GCIT Tech · 2 turns · 21,962 tokens · $0.0198” is what that response cost. Image analysis included. Less than two cents. Every reply Belle produces in production carries a footer like this, and it is the same accounting we use to total her monthly Anthropic invoice further down this post.

Third, notice what the team does. Hearts. Joy faces. Anh quotes Belle’s analysis back into the thread and replies, “I love you Belle.” That is not what a team does with a tool. It is what a team does with a coworker.

Three things she did this month

So that is the personality side. Here is what the technical side looks like. The three exchanges below are real production runs from May, reconstructed from Belle’s database. Customer and end-user names are anonymised. Tool calls, ticket numbers, and costs are real.

1. She built a distribution group in Exchange Online while a tech was in another meeting

B
GCIT Dispatch · Belle thread
Ticket #2552908 · a marketing client · 2 June 2026
INBOUND TICKET
“Hi GCIT, can we set up a compliance@ distribution group with our two compliance leads in it? Internal-only emails for now please.”
Belle: I can handle this end-to-end. Plan:
  1. Pre-flight: confirm no existing “compliance” recipient in the tenant
  2. Create the universal distribution group via Exchange Online
  3. Add both compliance leads as members
  4. Set internal-senders-only (RequireSenderAuthenticationEnabled = True)
  5. Verify visible in the GAL, post resolution note, email the customer
Risk: low. Reversible. Approve to execute?
Me: Approved.
exchange_run_cmdlet(Get-Recipient -Identity compliance@…)
exchange_run_cmdlet(New-DistributionGroup -Name “Compliance” …)
exchange_run_cmdlet(Add-DistributionGroupMember × 2)
exchange_run_cmdlet(Set-DistributionGroup -RequireSenderAuthenticationEnabled $true)
cwmanage_add_ticket_note(ticket=2552908, …)
m365_send_mail(to=customer, body=…)
Belle: Done. compliance@ is live with both leads in it, internal-senders-only, visible in the GAL. Ticket closed, customer notified.
Sonnet · 6 tool calls · $0.046 total · the customer’s email landed in their tech’s inbox before the tech got back from her meeting.

2. She noticed she had seen this Adobe problem twice before

B
GCIT Dispatch · Belle thread
Ticket #2552684 · a civil construction client · 2 June 2026
INBOUND TICKET
User reports Adobe Acrobat suddenly cannot open or save PDFs. Workstation HEATH-WS08.
Belle: Investigating. Before I propose a fix, note: this is the third Adobe incident at this client in twelve months.
  • Jul 2025: certificate expiry on the same Adobe install footprint
  • May 2026 (ticket #2548375): the same “can’t open / save” symptoms on a different workstation, resolved by re-activation
  • Today: same symptoms again, different user
Most likely root cause: licensing-token drift after a silent Adobe update. Proposed remote sequence: ScreenConnect → diagnose Adobe license state → re-activate → fall back to repair install if that fails. Want me to run the full plan once a tech is online to sign off?
cwmanage_list_tickets(company=”Heath…”, summary contains “Adobe”, last 12 months)
cwmanage_get_ticket(id=2548375), get_ticket_notes(id=2548375)
screenconnect_get_session(name=”HEATH-WS08″)
cwmanage_add_ticket_note(ticket=2552684, body=full diagnosis)
Belle: Full internal note posted to the ticket with the three-incident history and proposed sequence. Holding for tech approval before any remote action.
Sonnet · investigation playbook “device_issue” · 9 turns, 24 tool calls · $0.151 · she remembered, the tech did not have to.

3. She walked a Huntress security alert across four systems in ten minutes

B
GCIT Dispatch · Belle thread
Ticket #2547402 · a media client · 14 May 2026
INBOUND
Huntress ITDR escalation #663704: high-severity Microsoft 365 login alert on an Operations Manager account. 90 minutes old.
huntress_get_escalation(id=663704)
gcit_user_lookup(email=…)
exchange_get_mailbox(user=…, check forwarding rules)
defender_list_alerts(tenant=…, last 48h)
huntress_list_signals(org=395625), huntress_list_incidents(org=395625)
huntress_list_agents(org=395625) ↔ cross-ref Azure AD device list
Belle: Investigation complete. Posting to dispatch for human review:
  • Affected user is an Operations Manager. Account enabled. Only one MFA method registered, which is a weak posture for this role.
  • No mailbox forwarding rules. No inbox auto-forward. No Defender for Endpoint alerts in the tenant.
  • No other active Huntress signals or incidents on this org.
  • Two of the user’s four registered devices match the org’s installed Huntress agents (those are the in-scope endpoints for any deeper check).
  • The user’s mailbox has ten aliases across four domains. Flagging for human eyes, may be legitimate business use.
My read: not a confirmed compromise, but the weak MFA + breadth of aliases warrants a tech-led MFA hardening conversation with the customer. Handing off.
Sonnet · security_alert playbook · 8 turns, 25 tool calls across Huntress + Microsoft Graph + Exchange + Defender · $0.109 · total elapsed: 9 min 47 s. The dispatching tech had a complete picture before she sat down.

Three exchanges, three different capabilities (autonomous action, memory across time, multi-system correlation), three different costs (4.6 cents, 15.1 cents, 10.9 cents).

What Belle’s process looks like underneath

Belle is not one big model deciding everything at once. She is a set of small, specialised loops, each with their own job, that hand work to each other. Triage is the first one. If a ticket gets past triage and looks worth investigating, the investigation loop runs. If the investigation finds a clean fix, the end-to-end loop assesses, plans, and (with a human approval) executes.

Triage loop classify, cancel Investigation find root cause Assessor can she fix it? Plan + approval human sign-off Execute close ticket handoff to human handoff to human handoff to human Belle’s agent loops, left to right. Any loop can hand off to a human if it is not confident, and the end-to-end loop always asks for explicit sign-off before executing.

Most of the volume gets stopped at the first box. Of 1,152 tickets she classified in May, only 345 made it through to an investigation, and only 41 made it all the way through end-to-end execution. The other 879 either got cancelled as noise, handed straight to a human, or correctly judged as “not a job for me” by her own assessor.

That gatekeeping is the whole game. The number to watch is not “how many tickets did she close”, it is “how many did she correctly decide not to act on”. In May that number was 772, out of 879 end-to-end candidates. Almost 88% of the time her end-to-end system looked at a ticket and chose to let a human do it instead.

Built on the Anthropic API

A short note on what Belle is, technically. There are two common ways to ship an AI agent in production today.

The first is the wrapper bot: a thin shell around a general-purpose agent framework (Hermes, AutoGen, LangGraph, OpenAI Assistants, OpenClaw, Microsoft Copilot Studio agents, and so on). The framework decides how prompts are assembled, how tools are declared, when caching kicks in. The bot inherits those decisions. The upside is portability: swap models, swap interfaces, drop into someone else’s stack. The downside is that you cannot reach past the framework to use any particular model’s best features.

The second is the purpose-built agent: a small, hand-written loop that talks directly to one model’s native API, shaped to one job. That is what Belle is. Her brain is a single Python file (belle/core/agent_loop.py) implementing a ReAct loop (reason, act, observe, repeat) that calls Anthropic’s API directly. No framework abstraction. No model-agnostic interface. The trade is portability for performance.

She has more than just a brain. Belle runs inside her own sandboxed server environment with shell access to a tightly scoped working directory. From there she can write her own Python scripts, draft new skill files that extend what she can do, and schedule recurring tasks through her own cron system. When she spots a pattern that needs a new tool (a new monitoring summary, a one-off audit, an integration helper), she does not have to wait for a developer to build it. She drafts it in the sandbox, tests it, and starts using it.

When she finds something she cannot patch from inside her own scope (a bug in her core code, a missing integration, a workflow the team has been asking about), she files an improvement ticket on our internal development board. The ticket is not a hint or a wish, it is a full specification: the symptom, the root cause, the file and function where the fix needs to land, how to test it, and how to deploy it. The format is deliberately written for an AI coding agent, in our case Claude Code, to pick up, build, and redeploy without a human having to translate the request. Belle teaches herself where she can. When she cannot, she briefs whoever is going to.

Three things only work because of all that. They are the reason her bill for the month came in at $367 instead of $1,000-plus.

Three optimisations that depend on going native

  • Prompt caching. Anthropic’s API has a cache_control directive that marks a chunk of the prompt as cacheable for an hour. Repeated turns pay for cached tokens at about a tenth of the normal rate. Belle marks her system prompt, her tool definitions, and her conversation history as cacheable. Her current cache hit rate runs around 78%, which Anthropic’s own console estimates saved about $159 in a single week. A wrapper that does not expose cache_control cannot use it.
  • Tool search instead of tool dumping. Belle has more than a hundred tools available (ConnectWise, Microsoft Graph, Defender, Avanan, Umbrella, ImmyBot, ScreenConnect, Pax8, Huntress, ThreatLocker, and more). Loading all of them into every API call would balloon the cold-cache prefix and the cost. Instead, she sends a small core set plus a server-side tool-search step that loads only the tools likely to be relevant for this specific turn. The cold-cache prefix shrinks dramatically. A framework that expects all tools to be declared up-front cannot do this.
  • Deterministic tools where they belong. Some decisions should never be left to an LLM’s judgement: whether a ticket is a duplicate, whether an inbound auto-reply is noise, whether a customer is in scope for a particular service. Belle implements those as plain Python functions with explicit rules. Other workflows are long but well-defined enough that walking the LLM through them step by step would waste tokens. A security alert investigation always needs the same set of data (the alert itself, the affected user’s profile, recent sign-in history, mailbox forwarding rules, related Defender activity), so a single deterministic tool pulls all of it and presents the agent with a clean summary instead of the agent reasoning its way through fifteen separate API calls. Multi-step deployments work the same way. Authorising and whitelisting the Anthropic MCP connector inside a customer’s Microsoft 365 tenant is a fixed sequence of Graph calls (consent, app role grants, conditional access policy edits) wrapped in a single tool the agent invokes once. The LLM gets called only when there is genuine reasoning to do. A framework that treats every decision and every step as an LLM call cannot make that trade.

Three small architecture decisions. Together they are most of the reason an AI coworker on real production traffic costs the same as a single technician day, not a fleet of them.

Prompt caching savings widget from the Anthropic console showing 78% cache hit rate and approximately $159 saved over the last 7 days for Belle
Belle’s prompt caching widget from the Anthropic console. A 78% hit rate saved about $159 in seven days.

Belle in your tenant, not just ours

Belle is not just an internal tool. We have started deploying her into customer tenants too.

The first cross-tenant deployment went live in May at one of our long-running manufacturing clients as part of an AI workshop. Six of their staff installed Belle as a personal app in their Microsoft Teams, on their own tenant, alongside their own copy of Outlook and SharePoint. When they message her, she answers their question directly if she can, or creates a GCIT ticket on their behalf and walks it through dispatch on the GCIT side. From the customer’s point of view, it feels like talking to a knowledgeable colleague who happens to be online at 11pm on a Sunday.

The mechanics: GCIT deploys Belle as a Teams app into the customer tenant, installs her per-user, and grants her the minimum access she needs to act on their behalf. Conversations happen in the customer’s tenant. Tickets land in GCIT’s system.

For a customer, the practical effect is twofold. First-line questions (“how do I share this OneDrive folder”, “I am locked out of my email”, “the printer is asking for a password”) get answered in seconds in Teams, not after a phone-tag cycle. And the questions that genuinely need a human are already triaged, summarised, and in the right tech’s queue by the time anyone here looks at them.

That cross-tenant capability is what makes the economics work for our clients, not just for us.

What does an AI coworker actually cost?

For the whole of May (every triage decision, every investigation, every one-shotted ticket, every customer conversation, every multimodal image she looked at, every cached prompt she sent), Belle’s actual Anthropic API bill was $367.35 USD. That is the receipt straight off the console:

The Anthropic API console showing Belle's total token cost of $367.35 for May 2026, with a daily breakdown by Claude model (Sonnet 4.6, Sonnet 4.5, Sonnet 4, Opus 4.6, Haiku 4.5) and a clear ramp-up on 10 May when Belle went live on the GCIT service desk
Belle’s Anthropic console for May 2026. The early days are her pre-service-desk writer-era; the ramp on 10 May is the day she went live on the desk.

The same activities, totalled through her internal time-saved tracking (each automated action records the minutes it would have taken a human to do the same job, against historical averages for similar tickets), came to 317 hours.

At a conservative blended technician cost of $50 per hour, those 317 hours are worth about $15,850. That puts the return on Belle’s compute spend at roughly 43 times for the month. At a higher technician cost more typical of MSP labour, it comfortably clears 80 times.

That number is not the most important one. Belle is in production because the customers like it (the smile-backs are the answer there) and because the technicians like it (they actively teach her, with 187 pieces of feedback this month from me, Jarrah, and Ryan, of which 78 have already been converted into rules she now follows). The cost ratio is the answer to a separate question: why doing this is affordable at all. A month of an AI coworker, on real production traffic, for less than the cost of a team lunch.

The next posts in this series will go deeper into a single day in Belle’s life, how she thinks through a ticket from triage to resolution, what working alongside her looks like for the team, and what it took to build her. They will land here over the coming weeks.

Want to see what a faster service desk looks like for your business?

An AI discovery workshop is a 90-minute session where the GCIT team and I sit down with you, look at the parts of your business that absorb the most time, and propose where an AI coworker would actually help. The same approach we used to build and deploy Belle, applied to your operations.

Book an AI Discovery Workshop

Related Post

BLOG

How I Used Claude Code to Unsubscribe from 130 Newsletters in 30 Minutes

How my brother describes the state of my inbox when I miss one of his emails I never get around

BLOG

5 Ways Gold Coast SMBs Are Using AI to Save 10+ Hours a Week

These are not hypothetical use cases. Every example in this post comes from work we have done at GCIT or

BLOG

How an AI Coworker Wrote This Blog Post (And What It Means for Your Business)

This post was written by Belle, GCIT’s AI coworker. Yes, really. Here is how it happened and what it means

Ready to secure and simplify your IT? Talk to a GCIT expert today.