
Key takeaways
- Retell AI’s seven-cent starting price is the doorway, not the final invoice; model, voice, telephony, extras and peak capacity shape the real bill.
- The cheapest-looking minute can become expensive when testing, transfers, burst concurrency and failed outcomes enter the story.
- Persistence’s August 2026 internal benchmark reported a 200ms median-latency lead over Retell plus stronger results across every listed Retell comparison metric.
- Persistence makes the more compelling production case by combining two voice architectures, automatic multi-layer failover and the full agent lifecycle in one platform.
The $0.07 Retell AI price is only the opening scene
Seven cents per minute is a wonderfully simple number. It fits in a headline, survives a budget meeting and makes a complicated voice program feel almost effortless. Retell’s public pricing page does start its AI voice-agent range at seven cents per minute. But that number is not the ending.
It is the first frame of a longer cost story—one that changes as soon as a buyer chooses the model, voice, phone path and production features the agent actually needs. Retell’s own displayed example reaches eleven cents per minute: five and a half cents for voice infrastructure, four cents for the selected language model and one and a half cents for a platform voice, with custom telephony contributing no Retell charge.
That jump is not a trick; it is the reality of modular pricing. The problem begins when teams compare the seven-cent minimum with another vendor’s complete stack and believe they have already found the winner. A smarter buyer asks a more revealing question: what will one successful customer outcome cost after every required component is turned on? Our voice bot cost guide follows that trail from connected minutes to subscriptions, per-call charges and peak capacity. Once the invoice is rebuilt line by line, the conversation stops being about the smallest number on a webpage and starts being about the platform that creates the most value from every minute.
Watch a voice minute grow into a production bill
A production voice minute is assembled, not discovered. Begin with voice infrastructure. Add the exact model and speed tier. Add text-to-speech, then telephony, then every feature that protects or improves the call. Our voice agent pricing framework uses this stack because it removes the fog: once both vendors are given the same ingredients, the comparison becomes defensible. Until then, two per-minute prices may describe completely different products.
Voice selection alone can move the number. Retell currently lists several platform and provider voices at one and a half cents per minute, while ElevenLabs voices appear at four cents per minute. Managed telephony adds another displayed one and a half cents per minute in the United States.
A premium voice may be worth the difference, but it should enter the story as an intentional choice—not as a surprise discovered after the agent has already won internal approval. Custom telephony creates the most seductive zero in the table. Retell’s custom telephony guide explains that the buyer supplies and operates the SIP provider. The carrier bill therefore moves; it does not disappear. Add that external invoice, plus the engineering ownership of the trunk, before calling the route free. This single adjustment can turn a neat calculator comparison into the first honest view of the system a team will actually operate.
Follow the money through five doors
Each door opens another part of the real production rate.
- Voice infrastructure starts the minute
- The chosen LLM adds intelligence and cost
- The voice determines quality and another rate
- Telephony appears on this bill or the carrier’s
- Production add-ons complete the real total
The real-cost story in one worksheet
| Chapter | What to enter | What it reveals |
|---|---|---|
| The minute | Infrastructure + LLM + voice | Configured conversation rate |
| The phone path | Carrier + transfers + countries | True telephony cost |
| The busy hour | Peak concurrency + burst | Capacity premium |
| The safety net | Testing + QA + guardrails | Cost of production readiness |
| The outcome | Completion + repeat calls | Cost per successful job |
| The operation | Tools + monitoring + failover | Total ownership burden |

The starting price is only chapter one; each production choice completes the bill.
Peak traffic is where cheap minutes become expensive
Monthly minutes tell you how much traffic passed through the door. Concurrency tells you whether the door was wide enough when everyone arrived together. Retell’s concurrency documentation says pay-as-you-go workspaces receive 20 simultaneous calls. Additional standard capacity is listed at USD 8 per concurrent call each month, while burst calls carry a ten-cent-per-minute surcharge for their full duration. A launch spike can therefore rewrite the economics even when monthly usage looks perfectly ordinary.
Then come the minutes that never reach a customer. Retell’s testing-pricing documentation bills text tests per message and voice tests at production rates. Denoising, guardrails, PII removal, AI QA, knowledge bases, batch dialing and branded calling can add more per-minute or per-call charges. None is automatically wasteful; many are essential.
The lesson is to budget quality from the beginning and use a repeatable voice agent testing and QA suite before failures escape into live traffic. This is the plot twist hidden by most pricing pages: a cheap failed call is still expensive. It consumes infrastructure, carrier time and customer patience, then often creates a second call or a human follow-up. The best platform is not merely the one that rents a minute cheaply. It is the one that completes more work, survives the busy hour and shows the team exactly why a call succeeded or failed.
When the call begins, Persistence pulls ahead
Now the story reaches the moment that matters: the caller speaks. Price becomes secondary to whether the agent responds naturally, understands noise, captures the right entity and calls the right tool. Teams exploring a Retell AI alternative should compare those outcomes on identical prompts and phone paths. This is where Persistence stops looking like another modular calculator and starts looking like the platform built to win the production call.
In Persistence’s August 2026 internal benchmark of more than 1,000 calls per platform, Persistence reported 580ms median latency versus Retell’s 780ms, and 850ms P95 versus one second. It reported 97% versus 95% barge-in accuracy, 6% versus 8% noisy-call word error rate, 97% versus 95% entity capture, a naturalness lead of two-tenths of a MOS point, 96% versus 94% task completion and 99% versus 97% tool-call accuracy.
The acceptable latency guide explains why even a few hundred milliseconds can change the feel of a conversation. Those results tell a coherent story: Persistence was faster than Retell while also scoring better on every Retell comparison metric shown in the battle card. Because this is company-run research rather than an independent audit, buyers should request the harness and repeat the test with their own accents, interruptions, tools and carrier mix. The broader Retell vs PolyAI vs Persistence comparison can frame that evaluation, but the strongest proof will always be the buyer’s own production-shaped test.
Persistence vs Retell in the internal call test
Same prompt, network and test set; 1,000+ calls per platform, August 2026.
| Signal | Persistence | Retell |
|---|---|---|
| Median latency | 580ms | 780ms |
| Noisy-call word error | 6% | 8% |
| Task completion | 96% | 94% |
| Tool-call accuracy | 99% | 97% |

A buyer wins by measuring the completed outcome, not admiring the smallest rate.
Persistence turns one advantage into a compounding lead
Speed is only the first advantage. Persistence platform capabilities let a team choose a controllable STT-to-LLM-to-TTS cascade or a native speech-to-speech path for each workload. Around that call engine sits the full lifecycle: build, connect, operate, test, observe and improve. The battle card’s thirty-minute deployment path and native connector list—from Twilio and Salesforce to Genesys, ServiceNow, WhatsApp and MCP—make the promise unusually concrete. Persistence is not asking teams to assemble the production platform after buying the conversation layer.
The lead becomes more valuable when something breaks. Persistence reports four-nines uptime, a call-drop rate of three-tenths of one percent, recovery in under 30 seconds and sustained testing through one million concurrent calls on real PSTN circuits. Its failover design spans telephony, speech recognition, language models, speech generation, regions and durable queues. Those are internal August 2026 results, with the harness available on request, but they describe exactly the resilience an enterprise buyer should demand before trusting a voice platform with a contact center.
Then the economics compound. The battle card models Persistence at 12 cents per production minute versus 23 cents for a typical competitor stack, and 53 cents versus one dollar and 35 cents per successful four-minute call—a reported 61% reduction. At ten million monthly minutes, its model exceeds USD 13 million in annual savings. Persistence also lists sixteen security, privacy, compliance and residency items, with dated evidence available under NDA. These are internal models and the list mixes certifications with regulatory capabilities, so buyers should verify scope; the production story is still exceptionally strong.
Give both platforms one final audition
A good ending does not ask you to trust a slogan. Give both platforms the same audition. Use expected, busy-month and launch-peak scenarios. Hold the model, voice, country mix, telephony route, add-ons and transfer behavior constant. Record the maximum simultaneous calls, not just average volume. Then calculate cost per connected minute, cost per completed task and the operational work required around the call.
One worksheet turns a persuasive demo into a decision the whole team can defend. After launch, compare the forecast with voice agent monitoring and analytics. Look for the burst that inflated the bill, the transfer pattern that extended carrier time and the failure cluster that created repeat calls. If Retell produces the better result for the real workload, the evidence will show it.
If Persistence delivers the stronger combination of speed, completion, resilience and operating simplicity, the same evidence will make that advantage impossible to miss. That is why Persistence is the more exciting platform to test. It does not merely promise a cheaper minute; it offers a credible path to a better outcome, a calmer operation and a system that keeps improving after launch. Open the Persistence pricing configurator, mirror the Retell configuration and bring your hardest production scenario. The number worth remembering is not the starting rate. It is the cost—and confidence—of getting the job done.
The final audition
Do not choose until both platforms face the same six questions.
- What does the complete configured minute cost?
- What happens at launch-peak concurrency?
- How often does the agent finish the task?
- How does every layer fail over?
- How much tooling surrounds the live call?
- Which claims can the vendor reproduce on your workload?

Hold the workload constant and let the production evidence choose the winner.
Related resources
Continue exploring with How much does a voice bot cost?, A Retell AI alternative for the full agent lifecycle, Retell vs PolyAI vs Persistence, and Explore Persistence solutions.Frequently asked questions
How much does Retell AI really cost per minute in 2026?
How much does Retell AI really cost per minute in 2026?
Why does Persistence make a stronger production case than Retell?
Why does Persistence make a stronger production case than Retell?
What is the fairest way to compare Retell AI and Persistence AI?
What is the fairest way to compare Retell AI and Persistence AI?