Text to Speech Price Comparison / August 2026

Every major TTS vendor quotes a different unit: characters, credits, audio tokens, minutes. Here is what 20 leading models actually cost once you put them all on the same axis - US dollars per minute of generated speech, at each vendor’s largest published plan.
Cheapest: $0.0070 per minute — Inworld TTS-2 Flash on the Growth plan, 42¢ per finished hour.
Median: $0.0290 per minute.
Most expensive: $0.1000 per minute — ElevenLabs v3 and MiniMax Speech 2.8 HD, both $6.00 an hour.
Spread: 14× from cheapest to dearest.
Why per-minute is the only comparison that works
Ask a dozen vendors what speech costs and you get a dozen incompatible answers. Amazon, Google, xAI and Inworld bill per million characters. Deepgram, Murf and ElevenLabs bill per thousand characters or per credit. OpenAI and Google’s Gemini-TTS models bill per million audio tokens. A handful quote minutes directly.
None of those numbers can be compared as printed. What a buyer actually budgets is finished audio, so every price below has been converted to one figure: USD per minute of generated speech.
How we normalized
1,000 characters ≈ 1 minute of speech. Roughly 150 words a minute at about 6 characters per word including spaces — the same ratio ElevenLabs uses for its own credit-to-minute estimates, and the one that reproduces Deepgram’s published per-minute figure from its per-character rate.
Audio tokens at 25 tokens per second. Google publishes this ratio for Gemini-TTS, giving 1,500 output tokens per minute. Text-input cost is under 1% of the total at typical lengths and is left out.
Biggest plan wins. Where a vendor publishes tiers, we used the largest published plan — which is also the cheapest per unit. Custom enterprise rates are excluded because they are not published. Where a plan carries a serious monthly minimum, the table says so.
Cost per minute of generated speech

Lower is cheaper. Figures are list prices at each vendor’s largest published plan, normalized at 1,000 characters or 1,500 audio tokens per minute of speech. Negotiated enterprise contracts are not shown.
The full table
Model | Plan used | List price | USD / min | USD / hour | Source |
|---|---|---|---|---|---|
TTS-2 Flash | Growth | $7 / 1M characters | $0.0070 | $0.42 | |
Neural TTS | Commitment tier 3 — $15,000/mo, 2B chars | $7.50 / 1M characters | $0.0075 | $0.45 | |
Falcon (conversational) | API, usage-based | $0.01 / 1K characters | $0.0100 | $0.60 | |
TTS-2 | Growth | $12.50 / 1M characters | $0.0125 | $0.75 | |
gpt-4o-mini-tts | Pay-as-you-go (single rate) | $12 / 1M audio tokens | $0.0150 | $0.90 | |
Gemini 2.5 Flash TTS | Pay-as-you-go (single rate) | $10 / 1M audio tokens | $0.0150 | $0.90 | |
Grok TTS | Pay-as-you-go (single rate) | $15 / 1M characters | $0.0150 | $0.90 | |
Polly Neural | Pay-as-you-go (single rate) | $16 / 1M characters | $0.0160 | $0.96 | |
Aura-2 | Growth (10% off pay-as-you-go) | $0.027 / 1K characters | $0.0270 | $1.62 | |
SILMA TTS v2 | Best published plan | $0.028 / minute | $0.0280 | $1.68 | |
Sonic | Scale — $239/mo, 8M credits | $239 / 8M credits | $0.0299 | $1.79 | |
Chirp 3: HD | Pay-as-you-go (single rate) | $30 / 1M characters | $0.0300 | $1.80 | |
Polly Generative | Pay-as-you-go (single rate) | $30 / 1M characters | $0.0300 | $1.80 | |
Studio-quality TTS | API, usage-based | $0.03 / 1K characters | $0.0300 | $1.80 | |
Gemini 3.1 Flash TTS | Pay-as-you-go, preview pricing | $20 / 1M audio tokens | $0.0300 | $1.80 | |
Octave | Business — $500/mo, 10M chars | $0.05 / 1K characters | $0.0500 | $3.00 | |
Flash v2.5 | Business — $990/mo | $0.05 / 1K characters | $0.0500 | $3.00 | |
Speech 2.8 Turbo | Pay-as-you-go (single rate) | $60 / 1M characters | $0.0600 | $3.60 | |
Eleven v3 | Business — $990/mo | $0.10 / 1K characters | $0.1000 | $6.00 | |
Speech 2.8 HD | Pay-as-you-go (single rate) | $100 / 1M characters | $0.1000 | $6.00 |
What the numbers actually say
There is now a sub-cent tier
Inworld’s TTS-2 Flash at $7 per million characters and Azure’s deepest commitment tier at $7.50 both land under a cent a minute. Neither is casually available: Inworld’s $7 rate needs the Growth subscription, and Azure’s $0.0075 is the third commitment tier — $15,000 a month for 2 billion characters, roughly 33,000 hours of audio. Miss the volume and you have paid for silence. Murf’s conversational Falcon model reaches $0.010 with no commitment at all.
The commodity band clusters at $0.015
OpenAI’s gpt-4o-mini-tts, Google’s Gemini 2.5 Flash TTS and xAI’s Grok TTS all land within a rounding error of a cent and a half, with AWS Polly’s neural voices just above at $0.016 and Inworld TTS-2 just below at $0.0125. Four of the largest model labs converging on the same number is not a coincidence — it is what commoditization looks like.
Three cents is the most crowded price on the board
Deepgram Aura-2 at $0.027, SILMA TTS v2 at $0.028, Cartesia Sonic at $0.0299, and Google Chirp 3: HD, Gemini 3.1 Flash TTS, AWS Polly Generative and Murf’s studio-quality API all at exactly $0.030 sit within three tenths of a cent of each other. At this tier price has stopped being the differentiator; latency, language coverage and voice quality decide it instead.
Google’s newest TTS model costs double its current one
Gemini 3.1 Flash TTS, still in preview, is priced at $20 per million audio tokens against $10 for Gemini 2.5 Flash TTS — $0.030 a minute versus $0.015. Its input text tokens also double, from $0.50 to $1.00 per million. Whether the newer model is worth twice the older one is a quality question, not a pricing one, but it is worth knowing that upgrading is not free and that preview pricing can move before general availability.
Expressive voices still carry a 2–7× premium
ElevenLabs Flash v2.5 and Hume’s Octave cost $0.05 a minute on their largest self-serve plans, MiniMax Speech 2.8 Turbo $0.06, and ElevenLabs v3 and MiniMax Speech 2.8 HD $0.10. That premium buys prosody, emotion control and voice cloning — worth it for audiobooks and character work, hard to justify for an IVR menu.
Read the unit, not the headline
A vendor quoting “$0.05 per 1,000 characters” and one quoting “5¢ per minute” are charging the same thing. A vendor quoting “$30 per million characters” sounds ten times more expensive than one quoting “$0.030 per 1,000” and is charging identically. Almost every pricing page in this comparison is technically honest and practically unreadable until you convert it.
Caveats worth stating
Speech rate moves everything. At 180 words per minute rather than 150, every character-billed model gets about 20% cheaper per minute and every token-billed one stays flat. Character-based pricing rewards fast voices.
Non-Latin scripts change the math. Arabic, Chinese and Japanese pack very different amounts of speech into a character than English does. Character-billed vendors can swing significantly on non-English text.
Subscription tiers are not the same as API rates. Murf’s $99 Studio plan works out near $0.21 a minute for its 8 included hours; its API rate for the same voices is $0.030. Where a vendor sells both, we used the API rate.
Preview pricing is provisional. Gemini 3.1 Flash TTS is in preview and its rate may change at general availability.
Enterprise pricing is negotiated. Inworld advertises enterprise rates “as low as $5” per million characters and every other vendor here will discount at volume. These are the published numbers only.
Sources
Azure AI Speech pricing — commitment-tier figures via TextToLab’s Azure breakdown
xAI model and pricing docs (Grok TTS)
Murf API pricing — rates via this Murf plan breakdown
MiniMax pricing — per-character rates via this MiniMax API cost breakdown
OpenAI API pricing — per-minute derivation via TextToLab’s OpenAI TTS breakdown
Prices collected 25 August 2026 from public pricing pages. TTS pricing changes often. Verify against the vendor before committing budget.