Agents were reading a question they had answered four hundred times, then editing an order number into a macro. Response times slipped every peak.
Customer service4 min read
Sector, size, and function are stated. The client is not.
A written never-automate list
At a glance
Sector
eCommerce
Function
Customer service
Channels
Email, web chat, marketplace messaging
Data cleaned first
Order management, carrier tracking, returns system, and a policy library that contradicted the website
Never automated
Payment disputes, damage claims, legal or health matters
Engagement
Build, then a tuning period
The short version
First response time fell sharply, and held through peak
The seasonal slippage that used to arrive with volume did not, which was the point of doing this at all.
Automated resolution covers the intents it was built for
Order status, returns, sizing, invoices, delivery exceptions. Nothing outside that list is auto-answered.
Peak was absorbed with fewer temporary agents
The permanent team spent the season on difficult cases instead of editing order numbers into macros.
Why the volume mix made this worth doing
Volume was seasonal and the peaks were brutal. The question mix was neither. Order status, returns, sizing, invoices, and delivery exceptions accounted for most of it, month after month.
Response times slipped during peaks. Exactly when they matter.
The repetitive share of the volume was large, well understood, and answerable from systems the company already had. That combination is what made automation worth doing here.
Two weeks watching before anything was designed
Support work does not survive being described in a meeting.
Sessions with the support leads, then two weeks alongside agents watching how tickets were actually handled.
Three months of tickets were sampled and classified, over and over, until the intent list stopped changing.
Tone rules were agreed. Then escalation rules, written conservatively and shown to the whole support team before anything was built, because a team that fears the tool will route around it. Payment disputes, damage claims, anything with a legal or health dimension, and anything from a customer already complaining goes to a person immediately and is never auto-answered.
The data job connected order management, carrier tracking, and the returns system. That is what lets a reply state a fact instead of promising to check.
The macro library was cleaned out and outdated policy text retired. Policy was consolidated into one source. Replies and the website had been contradicting each other for months, and no one had noticed, because no one read both.
What was built, and what it never touches
Three tiers, and the third one matters most.
A triage and drafting agent working in three tiers.
Above the confidence threshold, low-risk intents
Classified, order and shipment status pulled, reply drafted in brand voice with the correct policy language, sent.
Above the threshold, everything else
Queued for one-click agent approval.
Below the threshold
Routed to a person with no draft attached, because a bad draft is worse than none.
The red banner is not decoration. Naming what will never be automated is what made the rest of it acceptable to the support team.Operating rule, set before go-live
A fixed share of automated sends is sampled and reviewed every week, by a person, against the same rubric agents are held to. Automation without sampling is an unmonitored liability.
bearingbridge.ai project record, phase Run
Where it stands after a few months
First response time fell. Automated resolution settled at a share of total volume concentrated in the intents it was built for.
Peak season was absorbed with fewer temporary agents than the previous year, and satisfaction scores moved with it.
Agents spend their time on the difficult cases, which is what they were hired for and what satisfaction actually responds to.
Months later the sampled accuracy had held. Two intents were pulled back out of automation after the first peak, and stayed out.
Three things that made it work
One policy source, first
Replies could not be correct while the site and the macro library disagreed with each other, and no amount of model quality fixes that.
The escalation list was deliberately conservative
Naming what will never be automated is what made the rest of it acceptable to the support team, and their cooperation was the adoption story.
Sampling runs weekly
Quality gets measured on a schedule rather than discovered through a complaint, and it is the difference between a system you trust and one you hope about.
Figures on this page are client-verified and published with permission. Where a figure is absent, the outcome is described in operational terms instead.
A conversation with the senior team about your markets, your data, and where AI would actually pay back for you. No slides, no obligation, and if the honest answer is that AI is not your next move, you will hear that too.