A production AI voice agent cost per minute in 2026 runs $0.07 to $0.21 for the all-in connected minute, depending on architecture and call volume. Cascaded STT-plus-LLM-plus-TTS stacks sit at the cheaper end of the band ($0.07 to $0.13), speech-to-speech models sit at the pricier end ($0.18 to $0.21) and trade variable rate for native audio expressivity. Developer platforms advertise headline rates of $0.05 to $0.31 per minute but the loaded cost lands 2x to 3x higher once telephony, LLM tokens, STT, and TTS stack on top. Done-for-you solutions charge $500 to $2,000 a month and trade variable rate for predictable billing plus zero engineering overhead.
How much does an AI voice agent cost per minute in 2026?
A production AI voice agent cost per minute in 2026 sits in a $0.07 to $0.21 band, with the headline platform rate almost never covering telephony, STT, LLM tokens, and TTS on top. Cascaded STT-plus-LLM-plus-TTS stacks are the cheaper end of the band at $0.07 to $0.13 per connected minute; speech-to-speech models buy native audio expressivity at $0.18 to $0.21 per minute. Developer platforms advertise $0.05 to $0.31 per minute but real loaded cost typically lands 2x to 3x higher once LLM, telephony, and transcription stack on. Subscription and done-for-you solutions run $500 to $2,000 a month and trade variable rate for predictable billing.
Published · Updated · Evidence-linked, not search-volume ranked.
Why this question is current
Exact query-volume data was unavailable, so RepoRadar uses these as current demand and intent signals rather than a claimed volume ranking.
- voice agent cost per minute 2026 · Inworld AI 2026-07-08 voice agent cost per minute worked model · global · checked 2026-08-20T21:55:00Z
Primary 2026-07-08 dated worked cost model showing six vendor stacks ranging from $0.0069 to $0.0912 per minute using public pricing. The strongest current same-year signal that per-minute cost is the central budget question. - voice agent cost pricing 2026 · Hacker News 2026-08-14 AI Agent News: The Calling Gap Starts to Close · global English-language developer community · checked 2026-08-20T21:55:00Z
Front-page Hacker News coverage of consumer AI voice agents placing phone calls, with explicit cost modeling and a $1.8 billion July 2026 AI agent funding total cited. Same-week active community signal that voice agent cost is the deciding question. - voice agent benchmark latency cost containment · DestiLabs 2026 AI Voice Agent Benchmark · global · checked 2026-08-20T21:55:00Z
Primary 2026 production benchmark across 10 plus deployments showing $0.07 to $0.21 per connected minute with median $0.105 per minute, plus latency and containment range. Confirms the band and the architecture split.
Who this helps
- founders scoping the budget for an AI voice product
- engineering leads building a cost model for an inbound or outbound call flow
- operations and procurement teams comparing vendor pricing pages against real production cost
- indie developers picking between a managed platform and a self-hosted cascaded stack
The four-band cost model
The 2026 market sits in four overlapping bands. Self-assembled cascaded stacks run $0.07 to $0.13 per connected minute for production traffic and can drop to $0.007 at the most aggressive committed-use tiers. Speech-to-speech models run $0.18 to $0.21 per minute and buy native audio expressivity with less per-stage control. Developer platforms advertise $0.05 to $0.31 per minute but the loaded cost lands 2x to 3x higher once LLM, STT, TTS, and telephony stack on. Done-for-you solutions charge $500 to $2,000 a month and trade variable rate for predictable billing plus zero engineering overhead.
The right band for a project depends on call volume, expected failure rate, and how much engineering time the team can absorb. The crossover between platform premium and self-assembly tends to land somewhere between 50,000 and 100,000 connected minutes per month, but the exact point depends on team cost, reliability requirements, and traffic shape.
- Cascaded self-assembled stack: $0.07 to $0.13 per minute, can drop to $0.007 at committed-use tiers.
- Speech-to-speech model: $0.18 to $0.21 per minute for native audio expressivity.
- Developer platform headline: $0.05 to $0.31 per minute, loaded cost 2x to 3x higher.
- Done-for-you solution: $500 to $2,000 a month for predictable billing.
What always stacks on top of the headline rate
Telephony is the floor cost on every call and is rarely included in headline per-minute pricing. Real numbers sit around $0.01 to $0.025 per connected minute depending on carrier, call type, and number geography. Inbound calls on a local French number are around $0.01 per minute; outbound calls to a mobile are around $0.04 per minute; a US inbound number carries an additional $1 to $2 monthly fee.
Speech-to-text typically bills per audio minute at $0.015 to $0.030 for streaming STT. LLM tokens are the biggest swing and run $0.020 to $0.080 per minute depending on model choice and conversation length. Text-to-speech is the dominant cascaded cost line and runs $0.020 to $0.060 per minute, with premium or cloned voices at the top of the range. Orchestration and platform fees add $0.005 to $0.020 per minute on platforms that bill them separately.
- Telephony: $0.01 to $0.025 per minute plus number rental.
- Speech-to-text: $0.015 to $0.030 per minute of streamed audio.
- LLM tokens: $0.020 to $0.080 per minute depending on model and history growth.
- Text-to-speech: $0.020 to $0.060 per minute of generated speech.
- Orchestration and platform: $0.005 to $0.020 per minute on platforms that bill separately.
The architecture split: cascaded vs speech-to-speech
Cascaded stacks (separate STT, LLM, and TTS) are cheaper at scale and give granular control. The same 2026 production benchmark shows cascaded deployments at $0.07 to $0.13 per minute with medians around $0.10. The cost advantage comes from being able to swap each component independently: tune STT endpointing for noisy environments, swap LLM per call type, cache TTS for fixed prompts.
Speech-to-speech models handle audio in and audio out through a single model, which skips the hand-offs between separate services and posts the lowest p50 latencies (540 to 580 milliseconds in the same benchmark). The trade is cost ($0.18 to $0.21 per minute) and less control over each stage. Native audio expressivity is the premium you are paying for.
How volume moves the per-minute rate
A typical 2026 voice agent deployment shows three volume bands. Around 1,000 minutes per month pays full retail rates at $0.16 to $0.21 per minute ($160 to $210 monthly). Around 10,000 minutes per month unlocks volume STT and TTS tiers plus model routing and lands at $0.11 to $0.15 per minute ($1,100 to $1,500 monthly). Around 100,000 minutes per month unlocks committed-use discounts plus caching plus smaller models for most turns and lands at $0.07 to $0.11 per minute ($7,000 to $11,000 monthly).
The crossover between buying a per-minute platform and running a self-assembled stack depends on team cost. At modest volume, paying a higher variable rate can still be cheaper overall once engineering time is included. At very large volume, saving a few cents across millions of minutes can justify owning more of the infrastructure.
- About 1,000 minutes per month: $0.16 to $0.21 per minute, full retail rates.
- About 10,000 minutes per month: $0.11 to $0.15 per minute, volume tiers active.
- About 100,000 minutes per month: $0.07 to $0.11 per minute, committed-use tiers active.
Cost per successful outcome beats per-minute cost
The right metric for budgeting is cost per successful outcome, not cost per connected minute. A system that needs 108 attempts to produce 100 successful outcomes burns the failed eight on transcription, model tokens, generated speech, telephony, and infrastructure before they end. If users retry, another full set of costs starts. A low headline price cannot make up for poor completion rates.
Cost per successful outcome is the formula that matters: (Voice stack plus failure and retry overhead plus human handling plus evaluation and operations) divided by successful outcomes. Engineering, observability, and evaluation infrastructure show up here even though they do not appear on any per-minute line. The cheapest headline rate can be the most expensive real cost when failure rates are high.
What this article does not publish
No vendor recommendation is published here. Vendor pricing pages and effective per-minute rates change every quarter as new speech-to-speech models ship and committed-use tiers move. The durable answer is to re-check the linked pricing pages at the moment of procurement.
No latency guarantee is published here. Production latency depends on model choice, geography, and the connection between the agent and the telephony provider. Run a small pilot on your real traffic shape before the cost model decides the architecture.
A useful next action
Estimate the monthly connected minutes the project will use in the first 90 days. Pick the volume band the estimate lands in and read off the per-minute range from the table above. Multiply by 2x to 3x to land on a realistic loaded cost for any developer-platform option, then compare against a self-assembled cascaded stack budget that includes engineering time.
Run a small pilot (1,000 to 10,000 minutes) on two candidate stacks before committing to a contract. Track per-minute cost, completion rate, and latency p50 and p95 on the same dashboard, then re-budget the production plan from the pilot data.
Sources checked
- Inworld AI: Voice Agent Cost Per Minute 2026 Worked Cost Model ↗ checked · global
Primary 2026-07-08 worked cost model across six stacks ranging $0.0069 to $0.0912 per minute using publicly listed vendor pricing. Establishes the cascaded vs speech-to-speech split and the per-component price range.
- DestiLabs: 2026 AI Voice Agent Benchmark (latency and cost per minute across 10 plus deployments) ↗ checked · global
Primary 2026 production benchmark across 10 plus live deployments showing $0.07 to $0.21 per connected minute with median $0.105 per minute, plus the cascaded vs speech-to-speech architecture tradeoff.
- The Next Web: Economics of a Voice Agent ↗ checked · global
Primary 2026 article walking through the five cost layers (STT, LLM, TTS, telephony, infrastructure) and the cost-per-successful-outcome model. Provides the framing for why the headline rate is rarely the real cost.
- Assindo News: AI Agent News August 2026 ↗ checked · global
Primary dated article covering the August 2026 voice-agent funding wave and the consumer calling gap closure, including ElevenLabs Agents, Twilio France telephony rates, and EU AI Act Article 50 disclosure obligations.
- Ringlyn: AI Voice Agent Pricing Per Minute in 2026 ↗ checked · global
Primary Q2 2026 cross-vendor pricing comparison across Vapi, ElevenLabs, Deepgram, Retell, GoHighLevel, xAI Grok, Ringg AI, voice.ai, Synthflow, Bolna, Molto, and Ringlyn AI with effective per-minute rates and what is included.
RepoRadar separates factual source claims from analysis. Recheck vendor docs before purchase, deployment, or policy decisions.