Every voice AI vendor leads with containment rate, and containment rate cannot distinguish between a caller who got what they needed and a caller who gave up. That single ambiguity is why so many voice deployments look good on the dashboard and do nothing for revenue. Here are the numbers that actually predict whether an inbound voice agent books work.
TL;DR
- Containment rate is the industry's favourite metric and its least honest one. Always pair it with an outcome rate.
- Latency is the strongest technical predictor of success. Under ~800ms end to end feels conversational; past ~1.5s abandonment spikes.
- Voice AI converts well on booking and capture, poorly on diagnosis, negotiation, and complaints.
- Your real baseline is usually voicemail, not a great receptionist. That changes the math enormously in AI's favour.
- Usage pricing runs roughly $0.05-$0.30/minute in 2026, plus a $50-$500/month platform fee. Custom builds are $15K-$50K.
- Measure five numbers for two weeks before you deploy anything, or you will never be able to prove the result.
Public benchmark data here is thin, and that matters
Unlike email or web conversion, there is no large neutral dataset for inbound voice AI performance. Almost every number in circulation comes from vendor marketing, measured on their own deployments, with the definitions they chose. This post gives ranges where they are defensible and, more usefully, gives you the measurement framework to generate your own numbers. Treat any precise-sounding industry figure on this topic with suspicion, including from vendors you like.
The metric that misleads everyone
Containment rate is the share of calls a voice agent handles without transferring to a human. It is the headline number on every voice AI pitch deck, and on its own it is close to meaningless.
Consider two deployments. The first contains 70 percent of calls and books appointments on 15 percent of them. The second contains 45 percent and books on 40 percent. The first has the better containment number and the second makes roughly two and a half times the revenue per hundred calls. Containment counted the hangups as successes.
The fix is to never look at containment without an outcome metric beside it:
Effective resolution = containment rate x outcome rate
Where outcome rate is whatever the call was supposed to produce: an appointment booked, a lead qualified and captured, an order status genuinely delivered. In the example above the first deployment scores 10.5 percent and the second scores 18 percent, which correctly ranks them.
Ask any vendor for outcome rate, not containment. The quality of the answer, and whether they have one at all, tells you a great deal about whether their existing customers are measuring anything real.

Latency: the number that decides everything else
Human conversation runs on turn gaps of roughly 200 milliseconds. That is the rhythm callers unconsciously expect, and voice agents are judged against it whether or not that is fair.
Practical thresholds, which are consistent across how speech interfaces behave rather than being one vendor's claim:
| End-to-end response latency | What callers do |
|---|---|
| Under 500ms | Reads as a normal conversation |
| 500ms - 800ms | Slightly slow but comfortable; most good deployments live here |
| 800ms - 1.5s | Callers begin interrupting and talking over the agent |
| Over 1.5s | Callers assume the call dropped; abandonment climbs sharply |
End to end means everything: speech recognition, model inference, response generation, and text to speech, plus network and telephony overhead. Vendors frequently quote model inference latency alone, which can be a fraction of the real number. Ask specifically for end-to-end p95 latency at production call volume, not average, and not in a demo. The p95 matters because the slow tail is where callers hang up, and averages hide it completely.
The second latency issue is barge-in handling: what happens when the caller starts talking while the agent is still speaking. Poor barge-in handling makes an agent feel robotic even at good latency, because real people interrupt constantly. Test this explicitly during evaluation by talking over the agent mid-sentence. A surprising number of production systems handle it badly.
Where voice AI converts, and where it does not
Call type predicts outcome far better than vendor choice does.
Converts well:
- Booking a standard appointment from a known service list
- Confirming, rescheduling, or cancelling an existing booking
- Capturing caller details for a callback when the office is closed
- Answering hours, location, pricing, and availability questions
- Qualifying an inbound lead against three or four clear criteria
Converts poorly:
- Diagnosing an unfamiliar problem before quoting
- Negotiating price or discussing a bespoke quote
- Handling a complaint or a cancellation with an unhappy caller
- Anything requiring reassurance rather than information
- Multi-part requests that change mid-call
The pattern is clear. Voice AI performs where the call has a known shape and a small decision tree. It degrades where the caller's goal has to be discovered through conversation.
Calls a voice agent should never take
Medical triage, legal or financial advice, emergency dispatch, and any call from a distressed person. The failure mode is not a lost booking, it is harm and liability. Design the agent to detect these fast and hand off immediately, and test that path harder than you test the happy path.
Your real baseline is probably voicemail
This is the reframe that changes most evaluations.
Businesses instinctively compare a voice agent against a good human receptionist, and it loses that comparison on nuance, warmth, and judgment. But that is rarely the actual alternative. For a large share of small service businesses, the real alternative for a meaningful slice of calls is nobody answering: after hours, during jobs, at lunch, during a rush, when the one person who answers the phone is on the other line.
Against voicemail, the comparison is completely different. An agent that books 35 percent of previously unanswered calls is not a mediocre receptionist. It is 35 percent of a revenue stream that was going to zero, and the callers who would have preferred a human were not getting one anyway.
So the honest evaluation question is not "is this as good as a person," it is "what share of my calls currently go unanswered, and what is a captured one worth?" Measure your answer rate for two weeks, split by hour and day. Most owners are surprised, and the surprise is where the business case lives. We work through that specific calculation for trades and home services in the missed-call revenue math.
The five numbers to track
Establish these for your current setup for two weeks before any deployment. Without the before, you cannot claim the after.
- Answer rate. Share of inbound calls answered by a human today, split by hour and weekday. This is your gap.
- End-to-end latency (p95). Post-deployment. The tail, not the average.
- Containment rate. Share handled without transfer. Necessary but not sufficient.
- Outcome rate. Bookings or qualified captures per contained call. This is the one that pays.
- Escalation satisfaction. How callers who did reach a human rate the experience. A voice agent that irritates everyone it fails to help is a net negative even with good outcome numbers.
Two derived numbers make the business case: effective resolution (containment x outcome) and revenue per hundred calls before and after. If a vendor cannot help you instrument these, that is informative.
What it costs
Usage-based voice pricing in 2026 generally lands between $0.05 and $0.30 per minute depending on the model, the telephony provider, and vendor margin. A typical four-minute service call therefore costs roughly $0.20 to $1.20 in usage. Platform or seat fees usually add $50 to $500 a month. A custom voice agent build, with real integrations into a scheduling or CRM system, runs $15,000 to $50,000 plus usage, and the integration work is nearly always the bulk of it rather than the voice layer itself.
Against those numbers, the break-even is usually low. If a captured call is worth $300 in booked work and you are missing 40 calls a month, capturing even a third of them is roughly $4,000 in monthly revenue against a few hundred dollars of cost. That is why voice does well in high-ticket service businesses and poorly in low-ticket ones, regardless of how good the technology is.
The bottom line
Voice AI agents convert when the call has a predictable shape, when latency stays under about 800 milliseconds end to end, and when the honest comparison is against calls nobody was answering. They disappoint when they are bought on containment rate, evaluated in a demo, and pointed at conversations that require diagnosis or reassurance.
Measure your answer rate first. If very few calls are going unanswered and most of your inbound requires real diagnosis, voice AI is not your next project. If a quarter of your calls hit voicemail and a booked one is worth hundreds of dollars, the arithmetic makes the decision for you.
Next step: For the vertical version of this math, see how home-service companies lose jobs to missed calls. For agent build costs generally, see what a custom AI agent costs.
Frequently asked questions
Do voice AI agents actually convert inbound calls?+
They convert well on simple, high-intent, transactional calls such as booking a standard appointment, confirming a time, or capturing details for a callback. They convert poorly on calls requiring diagnosis, negotiation, or reassurance. The honest framing is that a voice agent competes against your current alternative, which for most small businesses is voicemail. Converting 40 percent of calls that previously went unanswered is a large gain even though 40 percent would be a poor number against a good human receptionist.
What latency do voice AI agents need to feel natural?+
Natural human conversation has turn gaps around 200 milliseconds. Voice agents that respond end to end within roughly 800 milliseconds feel conversational to most callers. Past about 1.5 seconds, callers start talking over the agent or assume the line dropped, and both failure modes spike abandonment. Latency is the single most predictive technical metric for whether a voice deployment succeeds, and it is the one vendors are least willing to quote at your call volume rather than in a demo.
What is a good containment rate for a voice AI agent?+
Containment rate, the share of calls handled without a human, is the metric vendors lead with and the one you should trust least. A high containment rate can mean the agent resolved the call or that the caller gave up. Containment is only meaningful paired with an outcome: booked appointments, qualified leads captured, or issues resolved. A 70 percent containment rate with a 15 percent booking rate is worse than a 45 percent containment rate with a 40 percent booking rate.
How much does a voice AI agent cost?+
Usage-based voice AI typically prices somewhere between $0.05 and $0.30 per minute in 2026 depending on model, telephony, and vendor margin, which puts a four-minute call in the $0.20 to $1.20 range. Platform fees on top usually run $50 to $500 a month. A custom voice agent build runs $15,000 to $50,000 depending on integrations, plus usage. Compare against the fully loaded cost of the calls you currently miss, not against a receptionist's salary, because missed calls are the real baseline for most businesses.
What kinds of calls should a voice AI agent never handle?+
Anything where a wrong answer creates liability or harm: medical triage, legal advice, financial guidance, emergency dispatch, and anything involving a distressed caller. Also avoid complex diagnosis, price negotiation, and cancellation or complaint calls, where callers strongly prefer a human and being routed to a machine measurably worsens the outcome. Route these to a person immediately and design the agent to recognise them fast rather than attempting them.
How do I measure whether a voice AI agent is working?+
Track five numbers weekly: answer rate, average end-to-end response latency, containment rate, booked or qualified outcome rate, and escalation satisfaction. The pair that actually matters is containment against outcome rate, because containment alone rewards the agent for callers who hung up. Establish all five for your current setup for two weeks before deployment, otherwise you will have no baseline and every subsequent claim about improvement will be unfalsifiable.
Free PDF · No fluff
The 2026 AI Development Rate Sheet
Real build, agent, RAG, and consulting rates by tier — the numbers vendors quote behind NDAs, in one PDF.
Written by
Pankaj Kumar
Founder · Metageeks Technologies
Metageeks builds production-ready AI products for $1M–$15M companies — shipped in fixed-price sprints, not open-ended retainers. We write about what actually works in the field.
Connect on LinkedInThe AI Build Brief
Ship AI that actually works.
Practical playbooks on building, pricing, and shipping production AI — one email, every other week. No fluff.





