The pilot was working. At the end of month three the report was clean: the AI assistant took most of the incoming contacts and ended 70 of every 100 without transferring to a human. Cost per contained contact, twenty cents. A human contact costs several dollars. The slide put those two numbers side by side and the room agreed to scale.
Other numbers moved over the same three months. The AI bill went up, which everyone had budgeted for. Total contacts into the center went up too, which nobody had. And the average contact that still reached a person cost more than it used to.
Nothing in the report was false. The AI really did end 70 of every 100 contacts it took without a transfer, and twenty cents really was the model and speech-processing cost divided by the contacts it contained. The number was right for the narrow set of costs it counted. The business case had two problems. It left real costs out of the sum, and it divided what was left by the wrong outcome.
This is a composite, assembled from ordinary patterns rather than from one company. Every number attached to it, the containment rate and the dollar figures alike, is this article’s worked example, not an industry benchmark.
The pilot’s number and the customer’s experience
I argued in From Deflection to Resolution that contact centers have to stop counting the contacts they avoided and start counting the issues they finished. Treat that as settled. This piece is about what usually does not change next: the business case still divides its costs by the old unit.
Here is what containment looked like from the customer’s side. The rest of the article prices this one issue.
A customer asks where their order is. The AI reads the tracking record and gives a plausible answer: it is on its way, expected Thursday. The contact ends without a human. It is marked contained.
Thursday passes and nothing arrives. Or the answer was never specific enough to settle the question in the first place.
The customer comes back. The AI reads the same record and says the same thing in slightly different words. Contained again.
The customer contacts a third time and asks for a person.
The human opens the order system, reads the two previous conversations, sees that the shipment was split across two warehouses, corrects the expectation, and resolves it.
The pilot counted two contained contacts. The customer experienced three contacts. The business paid for all three and resolved one issue.
Two units, and only one of them is what you buy
Containment is the share of contacts that end without a human agent. It is a real operational measure. But it counts contacts, and a business does not buy contacts. It buys resolved customer issues. One issue can span three contacts, two channels and a week.
Gartner surveyed 5,728 customers in December 2023 and found that only 14% of customer service and support issues are fully resolved in self-service. Even among issues customers themselves described as very simple, the figure was 36% of those issues. That survey asked customers whether their issue was fully resolved, not whether a contact closed. It is measuring the unit that matters. Read it for the unit, not as a target: it is a self-service resolution finding, not a benchmark for AI-agent containment.
Two other measures come close and still miss. First contact resolution, or FCR, asks whether the customer’s issue was resolved on the first interaction, with no follow-up needed. That is a real outcome measure, and a useful one. But it answers only half the question. FCR tells you how many issues succeeded on the first attempt. It does not tell you what the issues that took several attempts ended up costing. Deflection is contact-level too.
Contact-level accounting hides the gap in a specific way. The first contact closes and is booked as a win. The second arrives as fresh volume, and the third as fresh volume again. Nothing in the ledger ties them back to the same customer question. The agent looks efficient per contact while the issue behind those contacts stays open and keeps spending.
| Unit | What it counts | What it hides |
|---|---|---|
| Deflected contact | A contact that never reaches an agent, human or AI, because the customer found the answer in self-service. | Whether the customer's issue was actually resolved. Gartner found only 14% of customer service issues are fully resolved in self-service. |
| Contained contact | A contact the AI ends without escalating to a human. | Repeat contacts on the same issue. Each one arrives as fresh volume, never tied back to the contact before it. |
| Resolved contact | A contact that resolves the issue within that single contact, on whichever channel handled it. Also called first contact resolution, or FCR. | The total cost of issues that take more than one contact, unless those contacts are joined and priced as one case. |
| Resolved customer issue | The customer's stated reason for contacting is addressed, and no further contact on that same reason arrives in the days that follow. | Little, as long as repeat contacts are tracked, given enough time to appear, and matched back to the original issue. Track too briefly or too loosely and repeat contacts slip back into looking resolved. |
A failed attempt is not a free miss
A failed AI attempt is usually described as a free miss. The AI tried, it did not work, the customer reached a human, and nothing was lost but a few cents of inference.
That is not always wrong. A failed attempt often returns real value. It can collect the order number, verify the customer, and capture the reason for the contact, which shortens the human conversation that follows. It also sits alongside the contacts the AI genuinely finished, which is the whole point of deploying it.
But it is not automatically free either. A failed attempt can cost more than sending that customer to a human at the start. Three things push it there: the repeat contacts it caused, the work it duplicated, and the minutes a human spent undoing an expectation the AI had set. Not always, though. It depends on how many repeat contacts it generated and how much usable context it handed over.
The point is narrower than “AI is expensive.” The cost of a failure is a number you have to calculate, not a rounding error you get to assume away. So stop pricing attempts and start pricing outcomes.
The formula, and the window it needs
Here is the whole of it.
Cost per resolved issue = every AI and human cost attached to a group of customer issues, divided by the number of those issues resolved inside a defined window
The window is the part people skip, and skipping it is exactly what makes containment misleading. Without a resolution window, containment tells you only that the session ended without a human. It does not tell you whether the customer’s issue stayed resolved.
Seven rules make the formula usable. Agree on them before you measure anything. Each is a place where a flattering answer can appear by accident.
- How contacts join into one issue. Customer identity plus contact reason, across every channel. Three contacts from one customer about one order are one issue, whether they arrive by voice, chat or email.
- How long the window is. Seven days is a workable default. Choose yours to be long enough to catch the repeat contact and short enough to close the books. If a promised outcome takes ten days to fail, a seven-day window will report it as resolved.
- What counts as resolved. The customer’s stated reason for contacting is addressed, and no further contact on that reason follows inside the window. Not “the session ended.”
- How a channel switch is treated. A customer who gives up on chat and calls has not created a second issue. Same issue, one more contact, more cost.
- What a repeat contact does. It reopens the customer issue. The earlier interaction stays counted as contained, and that is fair, because the session really did end without a human. It just stops counting as a resolved issue, and every cost stays attached to the same case.
- Which issues are old enough to score. Only matured cohorts get scored. An issue cannot be called resolved until its whole window has passed. An issue opened on the last day of the month cannot be scored under a seven-day rule in that month’s report. Count it anyway and the reported result flatters you.
- How shared cost is allocated. Platform fees, retrieval, observability and the engineering time to keep the agent working have to land somewhere. Divide them by resolved issues in the period, not by contacts.
The simple version of the sum
Now price the order question using only the lines a pilot usually counts. Every dollar figure the worked example produces from here to the end of the mix-shift section is illustrative, anchored to published ranges, and not a benchmark.
One note on the arithmetic, so you can check it with a calculator. Every worked-example figure below is computed from the rounded numbers printed on this page, not from longer ones held behind them. Those numbers are 11.6 minutes for a human contact, 52 cents a staffed minute, and $0.0220 a telephony minute. Telephony is billed per call, rounded up to the next whole minute, which is how Twilio bills it. A contact the AI attempts runs 2 minutes of audio and costs 14 cents. The wage floor cited below is the one exception. It comes from the source arithmetic, so it will not quite reproduce from the rounded hourly range printed here.
The AI side of one contact runs about one cent of model inference and about thirteen cents of speech to text and text to speech. Call it fourteen cents for each contact the AI attempts. Containment is reported at 70 of every 100 contacts. Divide fourteen cents by that rate and the reported cost per contained contact is twenty cents.
That number moves a lot with the stack you choose, so here is the one it was priced on.
| Input | What this example assumes | Cost |
|---|---|---|
| Model and pricing tier | A small, fast model, the tier a latency-sensitive voice agent normally runs on. List price $1 per million input tokens and $5 per million output tokens (Claude Haiku 4.5). | See the two rows below. |
| Input tokens per attempted contact | 8,000 billed input tokens across the whole contact, at the standard input rate. The 8,000 is a stand-in token count for a two-minute contact, not a measurement. A real deployment splits its input across cache writes, cache reads and new text at three different rates, as the caching row below explains, and would land near this figure rather than on it. | $0.0080 |
| Output tokens per attempted contact | 300 output tokens, which is the agent's actual spoken replies. | $0.0015 |
| Prompt caching | Assumed on, at the five-minute setting. Caching makes repeated text cheap, not free: writing the system prompt and tools into the cache costs 1.25 times the standard input rate, each later read costs a tenth of it, and anything new in a turn costs the full rate. One limit to check on your own agent: Claude Haiku 4.5 caches nothing under 4,096 tokens, so a lean system prompt plus tools may never be cached at all. | Included above. |
| Contact duration | About 2 minutes of audio, of which the agent speaks about 1,200 characters. | See the two rows below. |
| Speech to text | Deepgram Nova-3, streaming, monolingual, pay as you go, $0.0048 per minute, on 2 minutes. | $0.0096 |
| Text to speech | ElevenLabs Multilingual v2 or v3, $0.10 per 1,000 characters, on 1,200 characters. This is the standard-fidelity tier, not the cheapest one. | $0.1200 |
| Pricing date | All three vendors' published list prices as of July 2026. Model and speech prices move often, so re-check them before you reuse this. | |
| Total per attempted contact | $0.0095 of inference plus $0.1296 of speech. | $0.1391, call it $0.14 |
Two lines deserve a closer look: text to speech, which is the biggest cost here, and inference, which is the smallest. The model everyone argues about only rises above ten cents a contact if you run a frontier tier and switch caching off, both at once.
Text to speech is the biggest line by far. It is 12 cents of the 14. Move it to the low-latency tier at $0.05 per 1,000 characters and the same 1,200 characters cost 6 cents. The whole contact then costs about 8 cents.
Inference is the smallest line, at about one cent. A small model, a short contact and prompt caching all keep it there. The exact figure depends on how one contact’s tokens split between cache writes, cache reads and new text. Work that split out on plausible assumptions and you land between about 0.8 and 1.2 cents. Either way it is about a cent. Turn caching off and run the same contact on a frontier-tier model at $5 per million input tokens and $25 per million output. The full context is now resent on every turn. Assume that pushes billed input to about 20,500 tokens for the contact, instead of 8,000. That token count is an assumption, the same as the two in the table above. It works out at about 10 cents of input and under a cent of output, so inference alone reaches about 11 cents. The whole contact then costs about 24 cents, about 1.7 times the figure above. It takes both changes to get there. Make only one of them and inference stays under six cents.
So published prices alone put one attempted contact anywhere between about 8 cents and about 24 cents. That is why the assumptions belong on the page and not in a footnote.
Set that against a human. The US Bureau of Labor Statistics puts the wage for a customer service representative at $21.53 an hour at the median and $22.40 at the mean. Those come from its May 2025 survey period, published on 15 May 2026. That is the latest occupational wage data BLS has published. Publication trails the survey by about a year. The next release is due around May 2027, and it will cover the May 2026 survey. So a 2025 survey month is the current figure in a 2026 article. Its March 2026 employer cost data, still the latest release, shows private-industry benefits adding about 43% on top of wages. Apply that loading to both wage figures and you get roughly $31 to $32 an agent-hour. At the average handle time SQM Group benchmarked for 2024, 697 seconds or 11.6 minutes, one human contact costs about $5.97 to $6.21. That is wages and benefits only. Treat it as a floor: it excludes facilities, technology and management overhead.
Twenty cents against six dollars. That is the slide. It is arithmetically correct and it prices the wrong thing.
What the simple version leaves out
The twenty cents covers model inference and speech processing, priced per contained contact. It does not cover the issue. Call this an AI-first customer issue, meaning one where the AI took the first contact and a person finished it. This section walks through six of the missing lines, and the two largest are human. The table below adds two more.
The second AI contact is missing. It was also marked contained, so it was booked as a separate win, not a cost on the same issue. The retrieval and platform cost is missing. That is the order lookups, the search over policy documents, the software that strings the steps together, and the vendor’s platform fee. The cost of watching the agent is missing too: recording what it did, reviewing a sample of runs, and running a test suite that catches it getting worse. Those run forever, so they are an operating line, not a project cost. The channel cost is missing: telephony minutes on all three contacts, including the two that resolved nothing.
Then the two large ones. Human recovery is missing, which is the full cost of the third contact, the one that actually resolved the issue. And the extra handling caused by the human arriving without context is missing. That is reading two prior conversations, working out what the customer was already told, and correcting an expectation the AI created.
That last line is the one teams argue about, so be specific. It is not the human’s normal handle time. It is the extra time spent because the AI already gave the customer an answer that turned out to be wrong.
| Cost line | Who owns it | Do pilots usually count it? |
|---|---|---|
| Model inference and speech processing | AI or ML engineering, or the line item on the platform vendor's bill. | Yes. It is usually the only number on the slide. |
| Repeat AI interactions on the same issue | Whoever owns the AI experience end to end. | Rarely. Each repeat is booked as a separate, new contact, not tied back to the issue it belongs to. |
| Retrieval and platform infrastructure | Platform engineering, plus the vendor's platform fee. | Sometimes, but usually as one flat fee, not divided by resolved issues. |
| Observability and evaluation | The team running the agent day to day. | Rarely at pilot stage. It reads as a one-time project cost, then becomes a permanent operating line once the agent is live. |
| Channel and telephony minutes | Whoever holds the telephony or messaging contract. | Sometimes, for the AI's own minutes. Rarely, for the extra minutes a failed attempt adds to the contacts that follow it. |
| Human recovery handling | Contact-center operations. | Rarely as an AI cost. It is booked as ordinary headcount, not attributed back to the attempt that failed first. |
| Extra handling caused by the human arriving without context | Contact-center operations, the same team as human recovery. | Almost never. It looks like normal handle time, not a cost the AI's earlier answer caused. |
| Integration maintenance | Platform or backend engineering, keeping the agent's connections to order, billing and other systems working. | Rarely. Folded into general engineering overhead instead of priced against the agent. |
| Knowledge and content upkeep | The content or knowledge team that owns the source documents the agent retrieves from. | Rarely. Treated as a one-time setup cost, not an ongoing line that keeps answers current. |
The last column reflects common implementation patterns observed in practice, not a formal industry benchmark.
The same order, priced properly
Same order question, same composite, same worked-example figures. Now with seven of the nine lines in it, which makes the total a lower bound rather than the whole bill. Integration maintenance and knowledge upkeep are named in the table above and left unpriced here, because there is no defensible generic figure to put against either one. The human side carries the same caveat: it is wages and benefits only, without facilities, technology or management overhead. Put your own numbers on all three and the total goes up, never down.
Here are the same lines in text, so they can be read, checked and added up without the chart.
| Cost line | How it is priced | This issue |
|---|---|---|
| First AI attempt | Inference and speech, from the assumptions table above. | $0.14 |
| Second AI attempt | The same stack again, on the repeat contact the pilot booked as a separate win. | $0.14 |
| Retrieval and platform | Order lookups, document search, orchestration and the vendor platform fee, divided by resolved issues. | $0.21 |
| Observability and evaluation | Run recording, sampled review and the regression test suite, divided by resolved issues. | $0.07 |
| Telephony, all three contacts | Inbound minutes on both AI attempts and on the human call, including the two contacts that resolved nothing. Twilio bills each call rounded up to the next whole minute, so contacts of 2, 2 and 15.6 minutes bill as 2, 2 and 16. That is 20 minutes at $0.0220 a minute. | $0.44 |
| Human handling of the contact that resolved it | 11.6 minutes at about 52 cents a staffed minute, wages and benefits only. | $6.03 |
| Extra handling because the human had no context | 4 further minutes at the same rate, spent reading the prior conversations and correcting the expectation the AI set. | $2.08 |
| Total cost per resolved customer issue, on the seven lines priced here | The seven lines above, added together. Integration maintenance and knowledge upkeep are named in the earlier table and are not priced here, so this is a lower bound. | At least $9.11 |
Two of those lines are assumed allocations rather than published prices: the 21 cents of retrieval and platform, and the 7 cents of observability. The same two become the $28 spread across 100 issues later in this article. Every other priced line traces to a price printed on this page, so those two are the ones to replace with your own first. Two more lines, integration maintenance and knowledge upkeep, are named in the earlier table and carry no value here at all.
That $9.11 is not a different scenario. It is the same three contacts the pilot already ran, priced a different way. The two contacts the pilot booked as wins are charged to the issue here, and so are the channel minutes and the human minutes it never charged to anything. On its own books the pilot prices these same events at $0.40, two contained contacts at twenty cents each. The full accounting reaches at least $9.11 for the one issue those three contacts belong to.
One note on the human line. It covers wages and benefits only; telephony is a separate line. It starts at the pre-AI average of 11.6 minutes and adds the extra handling caused by the missing context, so this issue carries 15.6 human minutes, $8.11 in all. That sits above the post-AI human average in the next section. So $9.11 is a floor, not a bet on the old average.
Compare it to the alternative honestly. Routing that customer to a human on the first contact would have cost about $6.29: the same wages-and-benefits floor, plus telephony. Both numbers exclude human-side facilities, technology and management overhead. Loading those onto both sides raises both, and it does not close the gap between $9.11 and $6.29, because the AI-first path carries more human minutes: 15.6 in this example against 11.6.
Two cautions before anyone screenshots the $9.11. It is one issue, and one of the expensive ones. The issues the AI genuinely finished on the first contact cost well under a dollar each. And it is a worked example, not a measurement of your operation.
The transferable part is not the number. It is this. Twenty cents per contained contact and $9.11 per resolved issue are two measurements of the same pilot, and they are not like-for-like. Three things differ between them, not one. The first counts model inference and speech only, and the second adds telephony, platform, observability and human time. The first divides by the contacts the AI contained, and the second divides by one resolved customer issue. And the first is a pilot-wide unit price, while the second is the whole bill for a single issue that ran to three contacts and was one of the expensive ones. All three together are why the two printed figures sit more than forty times apart. Hold the events fixed and that third difference drops out: the pilot’s own ledger books these same three contacts as two contained contacts, $0.40 in all, against at least $9.11 here. The gap is smaller measured that way, and it is still large. Nothing about the events changed. The accounting did.
Your human costs go up, and that is not a failure
When an AI absorbs contacts, it rarely takes a random sample. Which contacts it takes depends on your routing policy, the channels customers use, the intents the agent covers and how many customers choose it. Those can produce very different mixes, so measure yours.
The common case is that the AI absorbs more of the simpler contacts. Where that happens, what is left in the human queue is harder. Average handle time rises, and the cost of a human-handled contact can rise with it, even while the total cost per resolved issue falls.
Here is the effect on the same worked example, scaled to 100 customer issues.
Both sides of this comparison count the same three categories: human wages and benefits, telephony on every contact, and AI cost including its platform. Before the AI, the third category is zero. Keeping the categories matched is the whole point, because a before figure that leaves out telephony and an after figure that includes it will show a saving that is partly just a change of scope.
Before the AI, those 100 issues generated 125 contacts, all handled by humans: 75 short ones at about 8 minutes and 50 long ones at about 17 minutes. That is 1,450 staffed minutes, and the average works out at 11.6 minutes, which is the handle time SQM Group benchmarked. Agent wages and benefits run about 52 cents a staffed minute, which sits between the median-based and mean-based rates above. So a human-handled contact costs $6.03, inside the floor range, and the 100 issues cost $754 in wages and benefits. Telephony on those 1,450 minutes adds $31.90. Total $785.90, so the blended cost per resolved issue is $7.86.
After the AI, the same 100 issues generate 128 AI contacts and 52 human contacts. The AI finishes 55 of the issues on its own. Note the unit there: 55 of 100 issues. The pilot’s 70 was 70 of every 100 contacts, a different denominator.
The 52 contacts that still reach a human skew hard. Only 12 are simple and 40 are complex, where before there were 75 and 50. So average handle time rises to about 14.9 minutes, and a human-handled contact now costs about $7.76. Across all 52, that is 776 staffed minutes and wages and benefits of $403.52, down from $754.
Telephony comes to $22.70: $17.07 on the 776 human minutes, and $5.63 on the 256 minutes the AI spent across its 128 contacts, which run much shorter than a human call. AI cost comes to $45.92: $17.92 of inference and speech across those 128 contacts, plus $28 of retrieval, platform and observability allocated across the 100 issues.
Add the three categories. Total $472.14. Blended cost per resolved issue, $4.72.
Read those two lines together. The cost per human-handled contact went up, $6.03 to about $7.76. The cost per resolved issue went down, $7.86 to $4.72. Both are true at once, and only the second is what the business pays.
A rising human average is not evidence the AI failed. It is what success looks like on the human side of the ledger. The failure is a business case that assumed the old human average would hold. If your model priced the remaining human contacts at $6.03 and they now cost $7.76, the savings you signed for were never available.
Latency shows up in the bill
Latency, the pause before the agent answers, is usually filed as an experience problem. It is also a cost line, and the chain is short.
A slow agent means longer calls. Longer calls mean more telephony minutes and more lines open at once. Slow turns also test the customer’s patience, and some customers hang up. I have no benchmark for how often that happens inside an AI conversation, so treat it as a direction, not a rate. An abandoned contact is not a saved contact. The issue is still open, so the customer contacts again.
The minutes themselves are cheap. Twilio’s published US rate for inbound toll-free voice is $0.0220 per minute as of July 2026, billed per call and rounded up to the next whole minute. A few extra seconds per turn costs a fraction of a cent. The money is in what the delay causes, and that cost is already counted once above, in the human minutes and the repeat contacts. Why voice is hard to make fast is a separate argument, made in why voice AI agents are harder than chatbots.
Four numbers to demand from any business case
Asking separately for AI cost, containment, average handle time and human cost is not enough. Each of them can move the right way on its own while the cost per resolved issue gets worse. AI cost per contact falls. Containment rises. Handle time improves on the contact types the team keeps reporting. The human cost line falls with headcount. And the cost per resolved issue still climbs, because repeat contacts multiplied faster than unit costs fell.
These four numbers close that gap. Each one names an owner and the artifact that proves it, because a number with no owner and no artifact is an assertion.
| Number | Owner | Proving artifact |
|---|---|---|
| Verified resolution rate after the defined window | Customer operations or CX analytics. | Issue-level cohort report linking repeat contacts across channels. |
| Repeat-contact rate after reported AI containment | Contact-center analytics. | Contact-reason and customer-identity reconciliation report. I found no credible cross-industry benchmark for repeat contact after AI containment, so measure it in your own operation, by intent and by channel. |
| Fully loaded cost of a failed AI attempt followed by human resolution | Finance, with platform engineering. | Cost model covering AI, infrastructure, channel and human recovery. |
| Blended cost per resolved issue against the pre-AI baseline | Finance or the transformation office. | Signed baseline-versus-production benefits model. |
These belong before launch, not in the first quarterly review. Whether a cost ceiling exists at all, and who owns it, is one of the hard blocks in the AI agent production readiness checklist. This article is about how to calculate a ceiling that is worth enforcing.
When counting deflection is honest
There is a case where deflection is a defensible unit, worth naming precisely.
Deflection is honest when five things hold. The request is purely informational. The answer is complete the moment it is delivered. No follow-up action is expected from anyone. The risk of a repeat contact is negligible. And an incorrect self-service answer pushes no meaningful work downstream. Store opening hours. A policy statement. A balance the customer only wanted to read.
For that kind of contact, a closed contact really is a finished issue. One contact, one issue, so the two units say the same thing.
One qualifier keeps this honest. You do not establish that a contact type fits by deciding it looks simple. You establish it by taking a matured cohort of that contact type, with enough volume to trust the result, and finding its repeat-contact rate below a threshold you agreed in advance. Until that measurement exists, the fit is assumed, not shown. And the assumption fails in a familiar place: the contact that looks informational but sets an expectation, like a delivery date. That is where this article started.
The same arithmetic outside the contact center
None of this is really about contact centers.
A coding agent’s unit is a merged, working change, not a suggestion it produced. A claims agent’s unit is a settled claim, not a processed document. In both, the cheap unit to count is the attempt and the expensive real unit is the finished outcome. The retries and the rework land in a different column from the one that reports the win.
Back in customer service, Gartner expects the cost side to sharpen. It predicts that by 2030, cost per resolution for generative AI will exceed $3, higher than many B2C offshore human agents. That is a prediction, not a measurement of today. Note the unit Gartner chose: cost per resolution.
So, four things to do. Define the resolution window and the joining rule before the next pilot reports anything. Make repeat contacts reopen the issue in the data. Book every cost line against the issue, not the contact. And re-baseline the human side: the contacts that remain will cost more each than the ones you measured before.
Designing agents so those costs can be traced back to an issue at all is a system design problem. It is the subject of Designing Enterprise Agentic AI Systems.
The Agentic AI Readiness Review does this work on a live business case. It prices your agent per resolved issue rather than per contained contact, before the case is signed. Details at how I can help.
Frequently asked
Quick answers
- What is cost per resolved issue, and how is it different from cost per contact?
- Cost per resolved issue is all the AI and human cost attached to a group of customer issues, divided by the number of those issues resolved inside a defined window. Cost per contact divides much the same money by conversations instead. That sounds like a small difference and it is not, because one issue can span three contacts, two channels and a week. If a customer asks the same question three times and a person finally fixes it, cost per contact reports three cheap events. Cost per resolved issue reports one expensive outcome, which is the thing the business actually bought. In this article's worked example, a pilot reporting twenty cents per contained contact has one order question inside it that cost at least $9.11 to resolve. Those two figures are not like-for-like. The twenty cents counts model and speech cost only, spread across the contacts the AI contained. The $9.11 adds telephony, platform, observability and human time, then charges all of it to one issue that took three contacts.
- How long should the resolution window be, and how do I set one?
- Seven days is a workable default. Choose yours so it is long enough to catch the repeat contact and short enough to close the books. The risk of a short window is easy to see: if a promised outcome takes ten days to fail, a seven-day window will report it as resolved. You also need a joining rule that says which contacts belong to the same issue. Customer identity plus contact reason, across every channel, works well. Three contacts from one customer about one order are one issue, whether they arrive by voice, chat or email. Without a resolution window, containment tells you only that the session ended without a human. It does not tell you whether the customer's issue stayed resolved.
- Does a failed AI attempt always cost more than sending the customer straight to a human?
- No, not always. A failed attempt often returns real value. It can collect the order number, verify the customer and capture the reason for the contact, which shortens the human conversation that follows. It also sits alongside the contacts the AI genuinely finished, which is the point of deploying it. But it is not automatically free either. Three things can push it past the cost of routing that customer to a person at the start: the repeat contacts it caused, the work it duplicated, and the minutes a human spent undoing an expectation the AI had set. Which way it lands depends on how many repeat contacts it generated and how much usable context it handed over. The cost of a failure is a number you calculate, not a rounding error you assume away.
- Why does our cost per human-handled contact go up after we deploy an AI agent?
- Usually because the AI is not taking a random sample of contacts. Where it absorbs more of the simple ones, what is left in the human queue is harder, so both the average handle time and the cost of a human-handled contact rise. Your own mix depends on routing policy, the channels customers use, the intents the agent covers and how many customers choose it, so check it rather than assume it. That rise is not evidence the AI failed. It is what success looks like on the human side of the ledger. In this article's worked example, a human-handled contact goes up from $6.03 to about $7.76 while the blended cost per resolved customer issue goes down from $7.86 to $4.72. Both are true at once, and the second one is what the business pays. The real failure is a business case that assumed the old human average would hold. If your model priced the remaining human contacts at $6.03 and they now cost $7.76, the savings you signed for were never available.
- Which numbers should I ask for before signing an AI contact-center business case?
- Four, and each one needs a named owner and an artifact that proves it, because a number with no owner and no artifact is an assertion. First, the verified resolution rate after your defined window, owned by customer operations or CX analytics, proved by an issue-level cohort report that links repeat contacts across channels. Second, the repeat-contact rate after reported AI containment, owned by contact-center analytics. I found no credible cross-industry benchmark for repeat contact after AI containment, so measure it in your own operation, by intent and by channel. Third, the fully loaded cost of a failed AI attempt followed by human resolution, owned by finance with platform engineering. Fourth, the blended cost per resolved issue against the pre-AI baseline, owned by finance or the transformation office. Ask for all four before launch, not at the first quarterly review. Asking separately for AI cost, containment, average handle time and human cost is not enough, because each of those can move the right way while the cost per resolved issue gets worse.
- Is counting deflection ever a fair measure?
- Yes, in a narrow case. Deflection is honest when five things hold. The request is purely informational. The answer is complete the moment it is delivered. No follow-up action is expected from anyone. The risk of a repeat contact is negligible. And an incorrect self-service answer pushes no meaningful work downstream. Store opening hours, a policy statement, or a balance the customer only wanted to read. For that kind of contact, a closed contact really is a finished issue. One qualifier keeps this honest. You do not establish the fit by deciding a contact type looks simple. You establish it by taking a matured cohort of that contact type, with enough volume to trust the result, and finding its repeat-contact rate below a threshold you agreed in advance. The trap is the contact that looks informational but quietly sets an expectation, like a delivery date.