What Every Business Should Know Before Buying an AI Voice Agent

What Every Business Should Know Before Buying an AI Voice Agent

The latency figure you are shown when shopping for an AI voice agent and the latency figure somebody actually measured differ by a factor of two to four, and the language model you pick moves the quality score more than the platform you pick does. Both of those are verifiable from figures the vendors themselves publish, and neither appears on any page currently ranking for this search.

Six of the seven substantial articles ranking for it are published by a company selling a voice agent platform, and five of those rank their own product first. TextToolz sells no voice platform, no telephony, no speech synthesis and no model, and takes no referral fee from anything named here. That is the only reason this page can put the marketed numbers and the measured ones in the same table.

What follows is aimed at somebody about to spend money in this category. Every rate was read from the vendor’s own pricing page on 14 August 2026. Every benchmark figure belongs to Cekura, whose Voice Orchestration Benchmark is the only independent measurement anywhere on this search, and we read those figures on telnyx.com, which is the single page that cites them.

What an AI voice agent actually is

An AI voice agent is software that answers or places a phone call, understands what the caller says, decides what to do, and speaks back, without a person on the line. In practice it is four components stitched together: speech to text, a language model, text to speech, and telephony, orchestrated so the caller can interrupt mid-sentence and be understood.

That stack is worth holding in mind, because everything difficult about buying in this category comes from it. Each of the four layers costs money separately, which is why headline per-minute prices are so misleading. Each layer adds delay, which is why latency figures depend entirely on where you start and stop the stopwatch. And one of the four, the language model, is usually swappable, which turns out to matter more than anything else on the comparison tables.

A traditional phone menu maps keypresses to branches. A voice agent holds a conversation, handles being interrupted, and can call an API mid-sentence to check whether Tuesday at ten is free. The difference is not the voice. It is that the branches are no longer written in advance.

The latency number you are quoted is not the one that was measured

Latency is the thing every page in this category leads with, and it is the number a buyer is least equipped to check. Here are the figures in circulation for four platforms, beside the figures Cekura measured for the same four.

Marketed latency against measured median turn latency. Marketed figures as published on the pages named; measured figures from Cekura’s Voice Orchestration Benchmark, read on telnyx.com, 14 August 2026.
Platform Marketed figure Published by Cekura median turn latency Cekura P95
ElevenLabs ~700 to 900 ms arahi.ai 1.73 s 3.19 s
Retell AI ~600 to 800 ms, “sub-800 ms” arahi.ai, medium.com 1.96 s 3.79 s
Vapi ~500 to 700 ms arahi.ai 2.34 s 2.95 s
Synthflow ~800 to 1,000 ms, “sub-500 ms” arahi.ai, vellum.ai 3.16 s 5.08 s

The marketed numbers cluster between five hundred and one thousand milliseconds. The measured ones run from 1.73 to 3.16 seconds. Before anyone concludes that the vendors are lying, the important part: these are almost certainly measuring different things, and that is the actual finding.

An end-to-end model round trip measures how long the software takes to produce a response once it has decided to respond. A turn latency measures the gap the caller experiences, from the moment they stop speaking to the moment they hear a voice, and that gap includes deciding they have finished speaking at all. Silence detection is a meaningful part of the wait and it belongs to the caller’s experience, not to the model’s.

So both sets of numbers can be honest and still be useless side by side, which is the problem. Not one page on this search states which metric its number refers to. A buyer comparing a vendor advertising sub-500 milliseconds against a benchmark reporting 1.73 seconds is comparing two different units and has no way of knowing it.

The correct instruction is published on this search, once, buried in a FAQ answer on rasa.com: ask vendors for production benchmarks, not demo metrics. Credited to them, and worth more than most of the comparison tables it sits underneath. The practical version is shorter still: ask what the number measures, and ask for the P95 rather than the median, because the call that annoys your customer is not the average one.

Grouped bar chart of four products showing marketed figures between 500 and 900 milliseconds against measured median turn latencies between 1.73 and 3.16 seconds
Marketed figures against Cekura’s measured median turn latency for the same products.

The model you choose moves the score more than the platform does

Telnyx publishes something about itself that no other vendor in this category publishes: the same platform, measured twice by the same benchmark, running two different language models.

With GPT-4.1, Cekura scored Telnyx at 76.3 percent on repeatable task completion with a 2.46 second median turn latency. With Kimi K2.6, on the same platform running the same scheduling agent, it scored 88.1 percent with a 1.44 second median. That is a swing of nearly twelve points in reliability and a full second in latency, from changing one component.

Now compare that swing to the spread between platforms on Cekura’s main leaderboard, which runs from 96.6 percent at the top to 76.3 percent at the bottom. The gap produced by changing the model inside one platform is larger than the gap between most pairs of platforms.

Every comparison article on this search ranks platforms and either holds the model constant or does not mention it. If Telnyx’s figures are representative, that entire genre is arguing about the smaller variable. It also means a buyer who shortlists on a comparison table, picks a winner and deploys it with a default model may land well below the score that made them choose it.

The practical consequence for anyone running a pilot is straightforward and appears at the end of this page as an instruction: before you swap platforms, swap the model.

Three cards showing one platform scoring 76.3 percent at 2.46 seconds with one model and 88.1 percent at 1.44 seconds with another, and every ranking page comparing platforms instead
Two configurations of one platform, both published by that platform.

Only one page in this category cites an independent benchmark

Cekura's Voice Orchestration Benchmark ran the same scheduling agent across six products using fifty-nine evaluators. It is, as far as this research found, the only independent measurement of these platforms available, and exactly one of the pages ranking for this search cites it.

That page is telnyx.com, and the detail worth pausing on is that the benchmark does not flatter Telnyx. Its GPT-4.1 configuration scores 76.3 percent on repeatable task completion, tied last with ElevenLabs among the figures Telnyx itself chose to publish. A vendor citing an independent measurement that ranks its own default configuration at the bottom is doing the opposite of what the rest of this search does, and it deserves saying by name.

Cekura’s own conclusion, credited, is the one every list in this category contradicts: benchmark leadership depends on the metric. Retell leads repeatable task completion at 96.6 percent. ElevenLabs leads median turn latency at 1.73 seconds, the fastest of the set, while scoring 76.3 percent on completion. Vapi has the tightest latency range, with a 2.95 second P95 against a 2.34 second median, which for a live call is arguably the most useful property of the three.

Three metrics, three different winners, and every page on this search declares a single best. That is not a flaw in the pages so much as a structural feature of who writes them.

One caveat this page carries rather than hides: a single scheduling agent measured across six products is one task, not a general measure of quality. It is the best evidence available in this category, which is a statement about how little evidence there is.

Four figures showing one page citing a benchmark that ties it last, a worst case marketed to measured gap of 6.3 times, and no page stating which metric it quotes
What happens when a vendor cites a benchmark it does not win.

What a minute actually costs

Per-minute pricing looks like the one comparable number in this category and is not. Most headline rates are platform fees that exclude the model and the speech synthesis, which are the two components that actually consume money on a long call.

Headline rate against all-in cost, read from each vendor’s own pricing page on 14 August 2026.
Platform Headline rate What it excludes All-in as published
Vapi $0.05 per minute Model and speech, billed at cost, or $0 with your own API key Not published as one figure
Retell AI $0.07 to $0.31 per minute Nothing; the stack is itemised $0.115 as broken out: $0.045 model, $0.055 infrastructure, $0.015 speech
Telnyx $0.05 per minute Little; the model runs on its own hardware at about $0.004 About $0.056 in production
Bland AI $0.11 to $0.14 per minute Nothing; telephony is in-house $0.11 on the tier with a $299 monthly platform fee
ElevenLabs Monthly tiers from $0 to $990 Varies by tier From $0.12 per minute for conversational use, per arahi.ai

Retell is the only platform here that shows its working, and the breakdown is instructive: $0.045 a minute for the model, $0.055 for voice infrastructure and $0.015 for text to speech. The model and the speech together cost more than the platform layer that everyone advertises.

Telnyx makes the same point from the other direction, contrasting its production figure of roughly $0.056 a minute with what it calls the $0.12 to $0.42 of a stitched-together stack, credited to them as their characterisation. Running the model on its own hardware at around four tenths of a cent a minute is what produces that number, which is a genuine architectural difference rather than a discount.

Bland is the honest outlier at $0.11 a minute on its platform-fee tier, because it operates its own telephony and quotes all in. Compared against $0.05 per minute from Vapi it looks twice the price, and against Retell’s actual $0.115 it is level. That comparison is the entire argument for reading the unit before the number.

Where the published prices are simply wrong

This page cites vendor pricing pages rather than comparison articles for the same reason the previous section exists, and one case makes it unarguable.

arahi.ai lists Synthflow at “From $29/mo”. vellum.ai marks Synthflow’s pricing as very transparent, plans plus usage. Synthflow’s own pricing page, read on 14 August 2026, states that enterprise contracts start at $30,000 annually with final pricing scoped around call volume.

Both things can be true at once if a self-serve tier exists that the pricing page no longer leads with, and this page states it that way rather than calling anyone dishonest. What cannot be true is that a buyer reading either comparison page would arrive at the right expectation before their first sales call.

Voiceflow shows a smaller version of the same problem. arahi.ai gives it a free tier with paid plans from $60 a month in one passage and $49 in another. voiceflow.com/pricing yielded no extractable price when read on 14 August 2026. Millis AI is quoted on reasonable infrastructure-grade pricing by the same source; millis.ai/pricing returned a 404 on the same date.

The rule this page follows, and the reason it is worth the effort: a price is quoted only from the vendor’s own page, with the date it was read attached, and where that page publishes nothing, this page says so instead of borrowing a number from a competitor’s summary.

Compliance is a line item, not a checkmark

Four pages on this search carry a compliance list. Every one of them presents SOC 2, HIPAA and GDPR as attributes a platform either has or does not. None of them carries a price, and for a regulated buyer the price is the point.

Vapi publishes the numbers on its own pricing page, read 14 August 2026: HIPAA at $2,000 a month and Zero Data Retention at $1,000 a month, as separate line items on top of the per-minute rate and the concurrency charges.

Set that against the per-minute differences these comparison tables argue over. Three thousand dollars a month of compliance configuration is worth roughly fifty thousand minutes of calling at Vapi’s own headline rate. For any buyer in healthcare, finance or insurance, the compliant version of a platform is a materially different product at a materially different price from the one in the comparison table, and choosing on the table means choosing on the wrong number.

The question to take into a sales call follows from that: not whether the platform is HIPAA compliant, but what the compliant configuration costs and which features it disables.

The twelve platforms

Each entry below states the build model, the published rate where one exists, the measured figures where Cekura published any, who it suits and one real limitation. What no entry states is how any of them sound, because nobody publishes a measurement of that and we called none of them.

The twelve platforms at a glance. Rates from vendor pricing pages; benchmark figures from Cekura via telnyx.com. Read 14 August 2026.
Platform Build model Published rate Cekura pass^3 / median latency
Vapi API first $0.05/min platform fee 94.9% / 2.34 s
Retell AI Visual flows plus API $0.07 to $0.31/min 96.6% / 1.96 s
ElevenLabs Agents Workflows and SDKs Tiers $0 to $990/mo 76.3% / 1.73 s
Bland AI Agent builder and API $0.11 to $0.14/min Not on the leaderboard
Telnyx Portal and APIs $0.05/min, about $0.056 in production 76.3% / 2.46 s, or 88.1% / 1.44 s
Synthflow No-code builder Enterprise from $30,000/year 81.4% / 3.16 s
Voiceflow Design canvas No price extractable Not on the leaderboard
PolyAI Managed, enterprise Custom quote only Not on the leaderboard
LiveKit Open-source SDK Cloud or self-hosted Recorded by Cekura
Rasa Voice Self-hosted platform Not published Not on the leaderboard
Cognigy Enterprise CCaaS Not published Not on the leaderboard
Millis AI Component-level control Pricing page returned 404 Not on the leaderboard

Vapi

Vapi is the developer-first option and prices like infrastructure rather than like software. Its pricing page, read 14 August 2026, publishes a $0.05 per minute platform fee with the model and speech providers billed at cost, and at zero if you bring your own API key. Ten concurrent lines are included and additional lines cost $10 each per month.

The compliance line items sit on that same page and are worth reading before shortlisting: HIPAA at $2,000 a month, Zero Data Retention at $1,000 a month. For a healthcare or financial buyer those two lines change the economics more than any per-minute comparison in this article.

On Cekura’s leaderboard it scores 94.9 percent on repeatable task completion with a 2.34 second median turn latency and a 2.95 second P95, credited. That P95 is the tightest range of any platform measured, which on a live call matters more than a good median: it means the worst turns are not much worse than the typical ones, and a caller never sits through an unexplained four-second silence.

arahi.ai markets it at roughly five hundred to seven hundred milliseconds, credited, and the distance between that and Cekura’s 2.34 seconds is the clearest single illustration of the metric problem this page opened with.

Best for engineering teams that want to choose every component in the stack and tune it, and for anyone who already holds their own model API keys and would rather not pay twice. The limitation is stated plainly by its own positioning: it is API first, and a team without engineering capacity will not get value out of the flexibility they are paying for.

Retell AI

Retell is the only vendor in this comparison that publishes its full cost stack rather than a headline, and that transparency is the reason it is worth reading even if you buy elsewhere. Its pricing page, read 14 August 2026, publishes a range of $0.07 to $0.31 a minute, then breaks the typical case into $0.045 for the model, $0.055 for voice infrastructure and $0.015 for text to speech, totalling $0.115. It starts at zero with ten dollars of free credits.

Seen next to a $0.05 headline elsewhere, that $0.115 looks expensive. It is not; it is the same number the other vendors reach once the model and the speech are added, and Retell is simply the one showing it.

Cekura scores it highest of the six on repeatable task completion, at 96.6 percent, with a 1.96 second median turn latency, credited. Its P95 of 3.79 seconds is the widest among the leaders, which is the counterweight: the typical turn is fast, and the tail is long.

One fact about the source rather than the product: retellai.com publishes the article ranking first for this query, and that article ranks Retell first. The benchmark independently supports the ranking on one metric, which is a happier coincidence than most of this search offers.

Best for support and sales teams that need natural turn-taking and interruption handling, and for buyers who want to know what they are paying for. The limitation is that tail latency, and the general point that leading on completion does not mean leading on speed.

ElevenLabs Agents

ElevenLabs built the speech synthesis that several competitors in this comparison use, and its agents product is the logical extension of that. Voice quality is the pitch, and it is the one claim in this category with a structural argument behind it rather than only a marketing one.

Its pricing page, read 14 August 2026, publishes monthly tiers at $0, $6, $11, $22, $99, $299 and $990. arahi.ai quotes conversational use from $0.12 a minute, credited, which would make it the most expensive per-minute option among the platforms with published rates.

The benchmark result is the most interesting on the leaderboard because it splits cleanly. Cekura measures ElevenLabs at the fastest median turn latency of the entire set, 1.73 seconds, while scoring 76.3 percent on repeatable task completion, tied last among the six. Credited to Cekura.

That combination is exactly what Cekura’s own conclusion warns about. A page declaring ElevenLabs the fastest would be correct. A page declaring it the best would be reading one column. A buyer whose priority is a caller never waiting, and whose task is simple enough that completion is not at risk, is looking at the right product; a buyer with complex multi-step tasks is not.

Best for use cases where the caller experience is explicitly what is being bought: premium support, concierge lines, brand-forward experiences. The limitation is in the numbers above, and it is that fastest is not most reliable. The multilingual coverage is the other reason buyers land here, since the same voice work that produces the English quality extends across a wide language range, and a brand running one voice across several markets has few alternatives.

Bland AI

Bland is built for outbound at volume and is the only platform in this comparison publishing a genuinely all-in per-minute rate, because it operates its own telephony rather than assembling it. Its pricing page, read 14 August 2026, publishes $0.14 a minute on the entry tier with no platform fee, $0.12, and $0.11 on a tier carrying a $299 monthly platform fee, with transfer minutes at $0.04 to $0.05.

Reading those against a $0.05 headline from Vapi or Telnyx makes Bland look like the expensive option, and comparing like with like it is not: $0.11 all in sits alongside Retell’s itemised $0.115. The difference is presentation, and Bland’s is the more honest of the two conventions.

Its claim of handling up to one million concurrent calls is its own, named as such and not restated here as verified fact. It does not appear on Cekura’s leaderboard, so no independent measurement of it exists in this research.

The vertical integration argument is real regardless. Owning the telephony removes a vendor, a contract and a failure mode from an outbound campaign, and for anybody running lead qualification or appointment reminders at scale that consolidation is worth more than a cent a minute.

Best for high-volume outbound: qualification, reminders, surveys, collections. The limitation is stated by its own positioning, that it is optimised for outbound, and a buyer whose priority is answering inbound calls well is better served by the platforms above. The $299 platform fee is also worth doing the arithmetic on before signing, because it only pays for itself somewhere above ten thousand minutes a month.

Telnyx

Telnyx publishes the page worth reading in this category even if you never buy from it, for the reason set out in the third section of this article. Its pricing page, read 14 August 2026, publishes $0.05 a minute, states that a production agent lands near $0.056, and itemises the components: the model on Telnyx hardware at roughly $0.004 a minute, carrier cost from $0.0032. It publishes a worked example reaching about $12,968 a month at 230,000 minutes.

The architectural claim behind that number is that running the model on its own GPUs removes the largest per-minute cost from the stack, and the itemisation supports it. Telnyx contrasts its figure with what it calls the $0.12 to $0.42 of a stitched-together stack, credited as their characterisation of the alternative.

Its benchmark results are the most valuable data on this whole search. Cekura evaluated Telnyx separately from the main leaderboard: 76.3 percent and a 2.46 second median with GPT-4.1, against 88.1 percent and 1.44 seconds with Kimi K2.6. That is the model-swap finding, published by a vendor about its own product, including the configuration that ties last.

Best for high-volume phone agents where having telephony and compute inside one vendor removes a layer of integration and a layer of cost. The limitation is named in its own comparison, credited: fewer visual simulation and testing tools than the testing-led products, which for a team that wants to rehearse conversations before launch is a real gap. Given how much of this article rests on measuring things properly, that is a notable thing to be missing.

Synthflow

Synthflow is where the pricing problem in this category becomes impossible to ignore, and the honest treatment requires stating three facts and letting the reader weigh them.

arahi.ai lists Synthflow at “From $29/mo”. vellum.ai marks its pricing as very transparent, plans plus usage, and describes it as affordable at scale. Synthflow’s own pricing page, read 14 August 2026, states that enterprise contracts start at $30,000 annually with final pricing scoped around call volume.

These can be reconciled if a self-serve tier exists that the pricing page no longer foregrounds, and that is the likeliest explanation rather than anybody misleading anybody. It remains the case that a buyer shortlisting from either comparison page would arrive at their first sales call with an expectation off by three orders of magnitude.

The performance figures are contested in the same way. vellum.ai describes Synthflow as targeting sub-500 millisecond responses. Cekura measures a 3.16 second median turn latency and a 5.08 second P95, the slowest of the six measured, with 81.4 percent on repeatable task completion. Credited to both sources, which is all this page can honestly do.

Best for no-code teams and agencies that want a visual builder and are buying at a scale where an enterprise contract makes sense. The limitation is unusual and worth stating: the two numbers a buyer would normally use to shortlist a platform, its price and its latency, are both contested in the published record. The remedy is the same in both cases, which is to get the figure from Synthflow directly and in writing rather than from any page describing it.

Voiceflow

Voiceflow is design-first, built around a visual canvas for mapping conversation flows before anyone implements them, with governance and collaboration features aimed at teams large enough to need both. For an organisation with dedicated conversation designers, that canvas is the product and it has no real equivalent among the API-first platforms.

Pricing is the frustrating part of the entry. arahi.ai gives it a free starter tier with paid plans from $60 a month in one passage and $49 in another, credited. voiceflow.com/pricing yielded no extractable price when read on 14 August 2026, so this page quotes none rather than borrowing one.

It does not appear on Cekura’s leaderboard, so there is no independent measurement of it here. arahi.ai’s own assessment, credited, is that it is less focused on raw latency than the specialists, which is consistent with what the product is for: a design and governance layer rather than a real-time performance play.

The trade-off is legible and reasonable. A team that maps a complex branching flow properly before building it will ship something better than a team that does not, and the canvas is how that happens. A team building a single appointment-setting agent does not need any of it.

Best for enterprise teams with conversation designers, and for regulated industries that need approval workflows and version control over what the agent is allowed to say. The limitation is scale of need: the features that justify the price only start mattering above a certain organisational size.

PolyAI

PolyAI sits at the enterprise contact centre end of the market and is sold rather than signed up for. vellum.ai records custom enterprise quotes only, with a buying cycle measured in weeks, credited, and this page quotes no rate because none is published.

Its published strength is call containment, which vellum.ai reports above eighty percent, credited as PolyAI’s own claim rather than an independent measurement. Containment is the right metric for a contact centre: the share of calls resolved without reaching a person, which is what the investment is actually bought to move.

It does not appear on Cekura’s main leaderboard, which is a limitation of the benchmark’s scope rather than a comment on the product; the benchmark ran a scheduling agent, and PolyAI’s territory is multilingual high-volume service traffic plugged into existing contact centre infrastructure.

The reason it belongs in this list despite publishing nothing a buyer can check is that it defines one end of the category. Everything above it in this article can be bought with a credit card and evaluated in an afternoon. PolyAI cannot, and knowing which of those two buying motions you are in is more useful than any feature comparison.

Best for large inbound contact centres with existing CCaaS infrastructure, multilingual requirements and a procurement process that can absorb a weeks-long cycle. The limitation is that same cycle: nothing about it, including the price, can be learned without a sales conversation. For a buyer running a genuine evaluation, that asymmetry has a practical cost, because the platforms publishing rates can be tested in parallel while this one is still being scheduled.

LiveKit

LiveKit is open-source real-time infrastructure rather than a packaged agent builder. It provides an Agents SDK deployable either on its cloud or self-hosted, and what you get is the plumbing for real-time audio rather than a product that answers phones out of the box.

telnyx.com ranks it fourth and states the trade-off precisely, credited: more engineering ownership than a packaged agent builder. That is the entire evaluation in one line. Cekura records it in the benchmark set, and telnyx discusses it separately in the benchmark commentary.

The case for it is control and cost at scale. A team that operates its own infrastructure does not pay a platform fee per minute forever, and a team with unusual requirements is not waiting for a vendor’s roadmap. The case against is that everything a managed platform absorbs silently becomes yours, including the latency tuning that this article has spent several sections showing is hard to measure and easy to get wrong.

The honest framing is that choosing LiveKit is not choosing a cheaper platform. It is choosing to employ the people who would otherwise be the platform’s engineers, and that is a sound decision at some volumes and a poor one at others.

Best for engineering teams that want to own the real-time layer, have unusual requirements, and are prepared to operate what they build. The limitation is that operational burden, and it is not small. Worth noting too that the open-source licence removes the vendor from the pricing conversation without removing the underlying costs, since the model, the speech synthesis and the telephony all still bill by the minute regardless of who orchestrates them.

Rasa Voice

Rasa’s pitch is sovereignty: run the speech stack yourself rather than rent it, with cross-channel continuity between voice and chat handled through its own orchestrator. For an organisation with data residency requirements, or one unwilling to send customer audio to a third party, that is not a preference but a constraint, and it narrows the field to almost nothing else in this article.

It targets sub-second latency from end of speech to agent response, with barge-in handling built in. That is a claim rather than a measurement here; Rasa does not appear on Cekura’s leaderboard and this page has run no test.

rasa.com publishes the article ranking seventh for this query and ranks Rasa Voice first in it, which is worth knowing when reading it. It also publishes, in a FAQ answer beneath that ranking, the most useful sentence anywhere on this search: ask vendors for production benchmarks, not demo metrics. That is credited prominently in this article precisely because it undercuts the page it appears on, and pages that publish standards their own claims can be judged against are rare enough to name.

Its own comparison also makes the right argument about why latency matters, noting that the difference between five hundred milliseconds and two seconds is the difference between a conversation and an exchange of messages.

Best for enterprises with data residency or sovereignty requirements and the engineering capacity to run the stack. The limitation is the same one LiveKit carries: owning it means operating it.

Cognigy

Cognigy, now part of NICE, is enterprise contact centre automation at scale, with deep integration into NICE CXone including Agent Copilot and Agent Assist. It is the option for organisations already inside that ecosystem, and largely not an option for anyone else.

retellai.com’s comparison describes it as configurable, capable of complex long flows, and strong on omnichannel consistency and enterprise governance, with latency that varies by deployment, all credited to them. That last point is the honest one: a platform deployed into an existing contact centre estate performs according to how it was deployed, which is why a benchmark figure would be of limited use even if one existed.

No public per-minute rate is published, and it does not appear on Cekura’s leaderboard.

What Cognigy offers that the per-minute platforms in this article do not is governance: approval workflows, versioning, role-based control over what the agent may say, and audit trails. In a regulated contact centre those are not features but preconditions, and their absence is why several cheaper platforms are simply not eligible.

Best for organisations already running NICE CXone with complex omnichannel flows and a governance requirement. The limitation is that the flexibility making it powerful is the same flexibility making it slow to change, and the buying process matches the product. The acquisition by NICE cuts both ways for a prospective buyer: it brings the contact centre integration much closer, and it ties the product’s direction to a larger company’s roadmap. A team choosing Cognigy today is choosing a position inside that ecosystem rather than a standalone voice platform, and that is the decision to be comfortable with before the governance features are ever evaluated.

Millis AI

Millis is positioned on ultra-low latency and component-level control, aimed at engineering teams optimising each layer of the stack rather than buying an assembled one. arahi.ai reports sub-500 millisecond latency in the right configuration and describes the pricing as reasonable for infrastructure-grade voice, both credited.

Recorded honestly: millis.ai/pricing returned a 404 when fetched on 14 August 2026, so this page quotes no rate for it at all rather than repeating the characterisation from a comparison article.

It does not appear on Cekura’s leaderboard either, which leaves it as the least verifiable entry in this list on both of the axes this article cares about.

That matters more here than it would elsewhere, because of the first section of this page. A sub-500 millisecond claim from a vendor, on a search where every marketed latency figure sits two to four times below the measured one and no page states which metric it means, should be read as exactly what it is: an unverified claim in a category with a documented measurement problem. That is not an accusation against Millis specifically. It is the standard this article applies to every marketed figure, including the ones from vendors it has praised.

Best for engineering teams chasing latency who intend to measure it themselves. The limitation is the thinnest published record of the twelve, on both price and performance. For a team with the capacity to benchmark a shortlist properly that may not matter at all, since they will produce their own numbers regardless of what the vendor publishes. For anyone else it is disqualifying, because there is nothing here to shortlist on.

What the enterprise end of the market costs

Two anchors, both from vellum.ai and credited, are useful for sizing the top of this market against everything above. Sierra AI is recorded at roughly $150,000 a year, enterprise only. Leaping AI is recorded from $2,500 a month per digital agent on a bespoke basis, with a deployment window of two to four weeks.

Set those beside Vapi’s five cents a minute and the category spans three orders of magnitude for products that all, at the simplest description, answer a telephone. That spread is not irrational. It reflects who carries the integration work, who carries the governance obligation, and who carries the risk when the agent says something wrong to a customer.

The useful question for a buyer is therefore not which platform is best but which of those burdens they intend to carry themselves. A team willing to own integration, tuning and compliance configuration can operate at the per-minute end. A team that wants a vendor accountable for outcomes is shopping at the annual-contract end, and the price difference is largely the price of that accountability.

How to run a pilot that actually tells you something

Everything above points at three instructions, and they are worth more than any shortlist this article could offer.

First, define the task before you shortlist. Cekura’s benchmark measured one scheduling agent, and its results reordered depending on the metric applied. Your own ranking will do the same, so decide whether you are optimising for completion, for median latency or for the tail before you read anybody’s table.

Second, run the same task across your shortlisted platforms rather than comparing their published numbers. This is rasa.com’s point, credited: production benchmarks, not demo metrics. A demo is configured by the vendor to succeed and tells you almost nothing about your third-hardest call type.

Third, and this is the instruction this page contributes: before you swap platforms, swap the model. Telnyx’s own published figures show a twelve-point swing in repeatable task completion and a full second of latency from that single change on one platform, which is larger than most of the platform-to-platform gaps everyone is arguing about. A disappointing pilot may be a model problem wearing a platform’s name.

Measure repeatable task completion and P95 latency rather than the median. The average turn is not the one that loses you a customer.

What the tables cannot tell you

Naturalness, accent quality and how a voice actually sounds to your customers are absent from every published source in this research, and this page will not fill that gap with adjectives. No vendor publishes a measurement. Cekura measures latency and task completion, not warmth. We called none of these systems.

One page on this search does report first-person measurements: medium.com’s author reports measuring Retell between six hundred and seven hundred and fifty milliseconds and Bland at seven hundred to nine hundred. Those figures are credited, and a caveat is attached rather than the numbers being dropped: the same article promotes a product called Alexor at sub-300 milliseconds that appears nowhere else in this research, including in the benchmark. That does not make the other measurements wrong. It does mean they carry the same unverified status as every vendor’s own claim, which is exactly the standard this page has applied throughout.

Setup and deployment time is the other gap. vellum.ai publishes vendor-supplied ranges from hours to months, credited, and no vendor publishes a commitment. For a buyer, the cost of getting to production is frequently larger than the first year of per-minute charges, and it is the least documented number in the category.

How this comparison was built

Eight ranking pages were extracted whole on 14 August 2026, and seven vendor pricing pages were read the same day: Vapi, Retell, ElevenLabs, Bland, Synthflow, Voiceflow and Telnyx. millis.ai/pricing returned a 404 and voiceflow.com/pricing yielded no extractable price, and both facts appear in the entries rather than being papered over.

No system on this list was called. No latency was timed by us, no voice was rated, no deployment was run. Every benchmark figure belongs to Cekura and was read on telnyx.com, and both are named every time a figure appears, because we did not run the benchmark either.

Three of the twelve results for this query are discussion threads, two on Reddit and one on Quora, and none could be extracted: Reddit served its interstitial and Quora a bot challenge. A quarter of this search is discussion that this page has not read, and that is stated rather than quietly omitted.

Prices and benchmark results both change. Every figure here carries the date it was read. The durable findings are the structural ones: that a marketed latency and a measured one are different quantities, that the model is a larger variable than most comparisons admit, and that a headline per-minute rate in this category is a platform fee rather than a price.

If you build in this category and something here is out of date or wrong, the editorial contact page is the fastest route to a correction.

Frequently asked questions

These are the questions the pages ranking for this search answer in their own FAQ blocks, answered here from the same published sources used above.

What is an AI voice agent?

Software that answers or places phone calls, understands the caller, decides what to do and speaks back without a person on the line. It is four components working together: speech to text, a language model, text to speech and telephony, orchestrated so the caller can interrupt and still be understood.

How much do AI voice agents cost?

Published rates run from about $0.05 a minute as a platform fee to roughly $150,000 a year for enterprise contracts. The headline per-minute figure usually excludes the model and the speech synthesis; Retell publishes the full stack at $0.115 a minute, which is the more useful comparison point.

How should I think about latency when choosing a platform?

Ask which metric the number describes. Marketed figures in this category sit between 500 and 1,000 milliseconds while Cekura’s measured turn latencies run from 1.73 to 3.16 seconds for the same products, because they measure different things. Ask for the P95, not the median.

Can AI voice agents replace human agents?

For contained, repeatable tasks, largely yes, and containment above eighty percent is claimed by vendors at the enterprise end. The pricing model is the tell: platforms billing per minute assume the agent handles the call, and platforms selling human backup price per call because a person sometimes answers.

How accurate are AI voice agents?

On Cekura’s benchmark, repeatable task completion ranged from 96.6 percent for Retell down to 76.3 percent for ElevenLabs and Telnyx running GPT-4.1, credited. That measured one scheduling agent across six products, so it is the best available evidence rather than a general accuracy figure.

Do AI voice agents sound human?

This page cannot answer that. We called none of these systems, no vendor publishes a measurement of naturalness, and no independent test of it exists on this search. Every claim in circulation is either a vendor’s own or comes from an article promoting a product this research could not otherwise find.