Research date: October 6, 2026. All prices are in US dollars per 1 million characters of text unless noted. Anything marked "not found" could not be confirmed from a public source. Prices and leaderboards change often, so re-check before you sign up.
The short version
Chance AI needs voices that sound natural on a phone, cost little, can be bought pay-as-you-go, and can legally be played to paying users, with a possible premium voice upgrade. With those needs in mind:
| Pick | Provider and model | Why |
|---|---|---|
| Best overall | xAI grok-tts (try the voice Castor) | $15 flat, pay-as-you-go. Its terms explicitly allow giving the output to your end users, and you own the output "in perpetuity". It scores #3 of 21 on Vapi's blind "humanness" test and has the best pronunciation accuracy measured on Coval. |
| Best budget | Speechify Simba 3.2 | 500,000 free characters every month with commercial use allowed. After that it's $10 for 1M, or $10/month for 1.9M. Blind-test results are near the top (#8 on Artificial Analysis, #2 on Vapi). English only. |
| Best raw quality | ElevenLabs Eleven v4 Turbo | #1 on the Artificial Analysis blind test. But putting it in a consumer app requires the $299/month Scale plan or higher; pay-as-you-go top-ups don't qualify. |
| Best premium upgrade | Inworld Realtime TTS-2 | Top-tier blind-test results (#7 on Artificial Analysis, #2 on the Hugging Face arena with its Max model, #4 on the cloned-voice board). Pay-as-you-go at $25, no monthly fee, very fast, and you can design voices from a text description. Read the licence caveats below. |
← swipe sideways to see all columns →
Final top 5 for Chance AI
- xAI grok-tts. It has the cleanest commercial terms of any provider checked, a simple $15 rate, 28 voices including Castor, 50 simultaneous streams, and Safari playback guidance in its docs. Its weak spots are a middling Artificial Analysis score (#26) and slower first audio (354 ms).
- Inworld TTS-2 / TTS-2 Flash. It ranks near the top on every blind test that includes it, and its Flash model is one of the fastest measured (61 ms). It costs $25 (or $15 for Flash) pay-as-you-go. Its licence wording ("internal business purposes", and deleting outputs when the contract ends) needs a written OK from Inworld.
- Speechify Simba 3.2. It's the cheapest of the top-quality voices and has a free tier that allows commercial use. Its terms require you to tell users the voice is AI and to pass that duty on to them. English only.
- Google Gemini 3.8 Flash TTS. #5 on Artificial Analysis, and Google's family leads Design Arena. You can steer the style with a plain-English prompt, and it costs about $16.50. It's still a preview, and Google has scheduled the price to roughly double on Jan 1, 2027.
- ElevenLabs v4 Turbo, as the luxury option. It's the best-sounding voice today, but the $299/month floor only makes sense once premium subscribers cover it.
Runner-up: Murf Falcon 2 ($10, with $10 of free credit every month, about 1M characters, pay-as-you-go).
Dropped from the earlier top 5: Cartesia. It's still excellent, but its free tier is non-commercial, it needs a subscription, and the $5 plan allows only 3 simultaneous requests.
Jargon, in plain English
- TTS: text-to-speech, turning the chatbot's text reply into spoken audio.
- Elo: a score from blind "which sounds better?" votes, like chess ratings. A gap of about 100 points means listeners prefer the higher one roughly 64% of the time.
- TTFA (time to first audio): how long until the first sound plays. Under about 300 ms feels instant in a chat app.
- WER (word error rate): how often the voice says a word wrong. Lower is better.
- Pay-as-you-go (PAYG): you pay only for what you use, with no monthly plan.
- Streaming / WebSocket: audio starts playing before the whole reply has been generated.
- SSML / audio tags: markup like
[laugh],<whisper>, or speed settings that control how a line is delivered. - Voice cloning / voice design: making a custom voice from a recording (cloning) or from a written description (design).
- OEM / end-user terms: rules for building a provider's service into your product for your customers.
How the voices compare in blind tests
No single leaderboard is the truth. Each one tests different voices, scripts and listeners, so look for models that do well on several.
| Model | Artificial Analysis (provider voices, Elo / rank of 94) | AA cloned-voice board (rank of 39) | Vapi Humanness (human = 100) | Hugging Face TTS Arena V2 (Elo) | Design Arena (Sept 2 snapshot) | Coval WER |
|---|---|---|---|---|---|---|
| Eleven v4 Turbo | 1334 / #1 | #2 (1173) | not tested | not listed | not listed | not found |
| Eleven v4 | 1321 / #2 | #3 (1160) | not tested | not listed | not listed | not found |
| Eleven v3 | 1174 / #18 | #9 (1072) | 97 (#1) | 1502 | 1290 (#4) | not found |
| Qwen-Audio 3.1 TTS Plus | 1292 / #3 | #1 (1181) | not tested | not listed | Qwen3 Flash 1040 | not found |
| Cartesia Sonic 3.6 | 1278 / #4 | #5 (1140) | not found | Sonic 2: 1514 | Sonic 3.5: 1149 | not found |
| Google Gemini 3.8 Flash TTS | 1275 / #5 | #14 (1050) | not tested | not listed | Gemini 3.1 Flash 1459 (#1) | Chirp 3 HD 1.6% (best) |
| Inworld Realtime TTS-2 | 1251 / #7 | #4 (1145) | TTS-1.5 Max: 79 | TTS MAX 1557 (#2) | 1.5 Max: 1149 | 4.5% (secondary source) |
| Speechify Simba 3.2 | 1242 / #8 | Simba 3.0: #15 | 96 (#2) | not listed | not listed | 4.3% (secondary source) |
| Luna TTS (VUI Labs) | 1227 / #10 | not listed | not tested | not listed | not listed | not found |
| Inworld TTS-2 Flash | 1214 / #12 | #10 (1071) | not tested | not listed | not listed | 5.4% (secondary source) |
| Murf Falcon 2 | 1155 / #21 | not listed | not found | not listed | Murf Gen2: 1103 | not found |
| Fish Audio S2.1 Pro | 1141 / #24 | #19 (1015) | 83 | OpenAudio S2: 1519 | not listed | not found |
| xAI grok-tts (listed as "SpaceXAI TTS") | 1136 / #26 | #13 (1054) | 93 (#3) | not listed | 1254 (#5) | 1.7% |
| MiniMax Speech 2.8 HD | 1173 / #19 | #16 | 93 | 1539 | 02 HD: 1130 | not found |
| Smallest Lightning v3.1 Pro | 1171 / #20 | not listed | not found | 1524 | 1149 | not found |
| Azure HD 2.5 | 1131 / #28 | not listed | not found | not listed | not listed | not found |
| OpenAI TTS-1 HD / gpt-4o-mini-tts | 1103 / #36 | not listed | not found | not listed | GPT-4o Mini TTS 1225 (#6) | not found |
| Polly Generative | 1065 / #53 | not listed | not found | not listed | Polly 989 | not found |
| Kokoro 82M (open source) | 1064 / #54 | not listed | not found | 1477 | 1040 | not found |
| Rime Coda | 1062 / #55 | not listed | not found | not listed | not listed | 11.6% |
| Google Chirp 3 HD | 1056 / #58 | not listed | not found | not listed | not listed | 1.6% |
| Hume Octave 2 | 1051 / #61 | not listed | not found | Octave: 1525 | 1162 | not found |
| Chatterbox (open source) | 1025 / #73 | #35 | not found | 1480 | 955 | not found |
| Deepgram Aura-2 | not listed | not listed | not found | not listed | 1015 | not found |
| Neuphonic | 940 / #82 | not listed | not found | NeuTTS Max 1439 | not listed | not found |
| Polly Neural | 894 / #90 | not listed | not found | not listed | not listed | not found |
← swipe sideways to see all columns →
What the boards mean:
- Artificial Analysis (pulled Oct 6, 2026; 94 models) has listeners compare each provider's own voices. Its separate "controlled voice" board uses the same 8 cloned voices for every model (4 US, 4 UK), so it measures the engine rather than the voice casting. Its Assistant, Customer-Service and US/UK-accent filters run in the browser, and I couldn't retrieve those filtered numbers (not found).
- Vapi Humanness Index (Oct 6; 21 models, 13,650 votes) measures how close each model gets to a real human recording of the same voice. It's the most relevant test for a chat app, because it rewards conversational naturalness over polished narration. An earlier Vapi post had grok-tts in first place at 95.
- Hugging Face TTS Arena V2 is a community blind test. Top entries: Vocu V3.0 1563, CastleFlow 1559, Inworld TTS MAX 1557, Papla P1 1548, Inworld TTS 1541. It doesn't include xAI, Google or OpenAI.
- Design Arena: its live board needs an API key, so these figures come from a public snapshot dated Sept 2, 2026.
- Coval measures speed and word error rate (pulled Oct 5, 2026, about 6:07 AM PT).
- MOS studies (lab-style 1–5 listening scores): I found no recent independent MOS study covering these 2026 models. Vendors publish their own MOS figures, but they aren't comparable across companies (not found).
Forum and review sentiment (Reddit blocks automated reading, so this comes from search summaries and is only a rough signal):
- Developers see Cartesia as the speed pick for voice agents and ElevenLabs as the expressive and cloning pick.
- Fish Audio is often named as a good-value ElevenLabs alternative.
- Inworld gets praise for expressive, context-aware delivery.
- Several third-party reviews still describe grok-tts as "$4.20 per million with 5 voices". That's out of date: xAI's own pages now show $15 and 28 voices.
Speed (time to first audio)
| Model | Coval median TTFA | Vendor's own claim |
|---|---|---|
| VUI Labs (maker of Luna TTS) | 43 ms | not found |
| Inworld TTS-2 Flash | 61 ms | P90 about 20 ms (vendor) |
| Speechify Simba 3.2 | 107 ms | not found |
| Inworld TTS-2 | 149 ms | P90 about 100 ms |
| ElevenLabs Flash v2.5 / v4 Turbo | 187 / 194 ms | about 75 ms / about 100 ms |
| Deepgram Aura-2 | 256 ms | not found |
| Murf Falcon 2 | 274 ms | about 100 ms |
| Rime Mist v3 / Coda | 279 / 319 ms | 37 / 96 ms (P50) |
| Cartesia Sonic 3.5 / 3.6 | 286 / 304 ms | under 90 ms |
| Fish Audio S2.1 Pro | 302 ms | not found |
| Smallest Lightning v3.1 Pro | 342 ms | under 100 ms |
| xAI grok-tts | 354 ms | not found. Vapi measured 460 ms (285 ms when streaming). |
| Google Chirp 3 HD | 536 ms | not found |
| Alibaba Qwen3 TTS Flash Realtime | 738 ms | not found |
| OpenAI gpt-4o-mini-tts | 2,461 ms | not found |
← swipe sideways to see all columns →
Vendor numbers are usually measured in ideal conditions. Independent numbers include real network time. For a chat app that shows the text first, anything under about 400 ms feels fine.
Voice selection
| Provider | Built-in voices | Languages | Voice design from a prompt | Cloning | Delivery controls | Standout voices to try |
|---|---|---|---|---|---|---|
| xAI grok-tts | 28 (Ara, Eve, Leo, Rex, Sal, Altair, Atlas, Aurora, Carina, Castor, Celeste, Cosmo, Helios, Helix, Iris, Kepler, Liora, Lumen, Luna, Lux, Naksh, Orion, Perseus, Rigel, Sirius, Ursa, Zagan, Zenith) | 20 in the docs (marketing says 25+). Every voice speaks every language. | not found | Yes: Custom Voices API, up to 120-second clip, 30 free clones | Inline tags like [pause] and [laugh], wrapping tags like <whisper>, speed 0.7–1.5x, pronunciation dictionary with phonetic spelling, timestamps |
Castor ("charismatic, down-to-earth, and easygoing"), Ara, Eve, Rex |
| Inworld | 95 | 14 native, 200+ cross-lingual | Yes, from a 30–250 character English description | Yes, instant (included pay-as-you-go) | TTS-2 takes bracketed plain-English direction (emotion, pace, volume, pitch, style) plus [laugh], [sigh], [breathe] and similar. Flash ignores direction but keeps the sound tags. |
Try several in the Inworld playground |
| Speechify Simba | 900+ catalog voices | Simba 3.2: English. Simba 3.0: 6 languages / 7 locales | not found | Yes, paid plans, consent-first | SSML (tags aren't billed), streaming | Browse the voice catalog |
| Google Gemini TTS / Chirp 3 HD | 30 (16 male, 14 female by Google's labels) | Gemini: 24 generally available plus about 70 in preview. Chirp: 53 locales. | Style by prompt ("say this warmly…") | Chirp instant custom voice ($60/1M) | Plain-English style prompts, tags like [sigh], [laughing], [whispering], pauses. Chirp pace 0.25–2x. |
Charon, Puck, Kore, Aoede, Achird, Zephyr |
| ElevenLabs | 10,000+ community voices. Older default voices expire Dec 31, 2026. | not re-checked in this pass | Yes: Voice Design, 20–1,000 character description | Yes, instant and professional | v4 audio tags like [whispers], [laughs], [excited], [sighs] |
Voice Library "conversational" filter |
| Cartesia | 500+ | 44 languages / 61 locales | not found | Instant (Pro), professional (Startup) | <speed>, <volume>, <emotion> (beta), [laughter] |
Its "Emotive" voices |
| OpenAI | 13 (alloy, ash, ballad, cedar, coral, echo, fable, marin, nova, onyx, sage, shimmer, verse) | Many (follows the input text) | Style via an "instructions" field | No | Free-text instructions on gpt-4o-mini-tts | marin, cedar (OpenAI's recommended) |
| Murf Falcon 2 | 150+ | not found exactly | not found | not found | not found in detail | Browse the Murf API page |
| Deepgram Aura-2 | 90 (41 English) | 7 | No | No | Limited | — |
| Azure | 400+ | 100+ | No | Custom/personal voice requires Limited Access approval | Full SSML, speaking styles | HD voices |
| Amazon Polly | 100+ (43 generative) | 40+ | No | No | SSML | Generative voices |
| Rime | 184 (Coda), 94 (Mist v3). FAQ says 600+ overall. | Coda 8, Mist v3 4 | No | Enterprise | Speed, pauses, spelling, pronunciation | — |
| Fish Audio | Large community library (count not found) | Multilingual | Yes: Voice Design API, $0.01 per request | Yes | Emotion cues | — |
| Luna TTS (VUI Labs) | 3 | not found | not found | not found | Expression tags, speed | — |
← swipe sideways to see all columns →
How many good conversational US-English male and female voices each provider has: no provider publishes this count, so it's not found. In practice:
- Google lists 16 male and 14 female voices.
- Deepgram has 41 English voices.
- OpenAI's two newest voices (marin and cedar) are its most natural.
- xAI's 28 aren't labelled by gender on the pages I read.
Audition before you decide.
Price and worked monthly costs
To make the numbers concrete, assume an average spoken reply of about 300 characters (roughly 20 seconds of audio):
- 100K characters/month is about 330 replies.
- 1M is about 3,300 replies (roughly 17 hours of audio).
- 10M is about 33,000 replies.
Costs below are what you'd pay with commercial rights for end users, after free allowances.
| Provider / model | List price per 1M chars | 100K / month | 1M / month | 10M / month | Free allowance and conditions |
|---|---|---|---|---|---|
| xAI grok-tts | $15 | $1.50 | $15 | $150 | Free credits not confirmed in official docs. Third-party guides mention a $25 sign-up credit and a $150/month data-sharing programme; the latter means sharing data, so avoid it if you want privacy. |
| Inworld TTS-2 (On-Demand) | $25 | $2.50 | $25 | $250 | About 70 minutes free, no subscription. Plans lower the rate to $20–$12.50, but whether the plan fee counts toward usage is not verified. |
| Inworld TTS-2 Flash | $15 | $1.50 | $15 | $150 | Same as above. Plans go down to $10–$7. |
| Speechify Simba 3.2 | $10 (Starter) / $8 (Pro) / $6 (Scale) | $0 (Free: 500K/month) | $10 (Starter plan includes 1.9M) | $91 (Starter $10 plus 8.1M top-up), or $99 on Pro (13.5M included) | Free tier allows commercial use but can't top up. Plans are month-to-month. |
| Google Gemini 3.8 Flash TTS | about $16.50 (Artificial Analysis estimate) | about $1.65 | about $16.50 | about $165 | Token pricing doubles on Jan 1, 2027 ($9 to $18 per 1M audio tokens), so expect roughly $33/1M after that (estimate). Preview, Enterprise API. |
| Google Chirp 3 HD | $30 | $0 | $0 | $270 | 1M characters free every month |
| ElevenLabs v4 Turbo | $40 ($11 during a 72%-off promo until Oct 12, 2026) | $299 | $299 | about $627 (estimate) | Scale plan ($299/month, 1.8M credits) required for end-user apps. 10M estimate = $299 + 8.2M × $0.04/1K, assuming 1 credit per character (not verified for v4). |
| ElevenLabs v4 | $80 | $299 | $299 | about $955 (estimate) | Same Scale requirement. Startup grant: 12 months, 33M characters (application needed). |
| Cartesia Sonic 3.6 | Overage $65 / $45 / $38 by plan | $5 (Pro, 100K included) | $49 (Startup, 1.25M included) | about $375 (Scale $299 with 8M included, plus 2M at $38) | Free tier (20K) is non-commercial |
| Murf Falcon 2 | $10 | $0 | $0 (covered by the $10 monthly credit) | $90 | $10 free credit every month. $10 minimum top-up, and credits don't expire. |
| OpenAI tts-1 / tts-1-hd | $15 / $30 | $1.50 / $3 | $15 / $30 | $150 / $300 | gpt-4o-mini-tts is billed by audio minute (about $0.015/min per OpenAI's estimate) |
| Deepgram Aura-2 | $30 | $3 | $30 | $300 | $200 one-time credit. Flux TTS ($45) has a 1:1 credit match up to $500 until Dec 31, 2026. |
| Azure Neural / HD | $15 / $22 | $0 (free F0 tier, 0.5M/month) | $15 / $22 | $150 / $220 | Custom voice $24 plus $4.03/hour hosting |
| Amazon Polly Neural / Generative | $16 / $30 | $1.60 / $3 | $16 / $30 | $160 / $300 | First 12 months: 1M Neural / 100K Generative free per month. New accounts can get up to $200 in credits. |
| Fish Audio S2.1 Pro | $15 per 1M bytes (≈ characters in English) | $1.50 | $15 | $150 | Pay-as-you-go, no minimum. The free model (s2.1-pro-free) is meant for testing. |
| Rime Coda / Mist v3 | $50 / $30 | $5 / $3 | $50 / $30 | $500 / $300 | Free minutes on Starter: the page says both "about 800 minutes" and "3,000 minutes" (conflicting) |
| Smallest Lightning v3.1 Pro | $17.50 on the pricing page ($19.50 on the model card) | about $1.75 | about $17.50 | about $175 | not found |
| Qwen 3.0 TTS Plus (Alibaba, international) | $20 | $2 | $20 | $200 | not found |
| MiniMax Speech 2.8 HD | $100 | $10 | $100 | $1,000 | not found |
| Hume Octave | Subscription. Overage $0.15–$0.05 per 1K by plan (secondary source). | not computed | not computed | not computed | Free and Starter can't buy extra usage |
| Resemble | about $0.0005/second of audio (secondary source) | not computed | about $30 (estimate) | about $300 (estimate) | Its official page now lists mainly detection products |
| PlayHT | — | — | — | — | Shut down: Meta acquired it in July 2025 and the platform closed Dec 31, 2025 |
← swipe sideways to see all columns →
Highlights:
- At 100K characters/month, Speechify, Murf, Google Chirp and Azure are all $0, and xAI is $1.50.
- At 1M, Murf is still $0, Speechify is $10, and xAI is $15. ElevenLabs is $299 because of the plan requirement.
- At 10M, Murf is $90, Speechify $91, xAI $150 and Inworld Flash $150. Cartesia is about $375 and ElevenLabs v4 Turbo about $627.
Scheduled changes and promos:
| Change | Ends or takes effect |
|---|---|
| ElevenLabs 72% off v4 / v4 Turbo, and 3x credits on Creator and above | Until Oct 12, 2026 |
| Deepgram Flux credit match | Through Dec 31, 2026 |
| Google Gemini 3.8 Flash TTS price doubles | Jan 1, 2027 |
| Old ElevenLabs default voices expire | Dec 31, 2026 |
| Inworld price increases | Inworld's terms promise 30 days' notice |
← swipe sideways to see all columns →
Commercial rights: can you play these voices to paying users?
| Provider | Embed in a paid app for end users? | You own the audio? | Must disclose it's AI? | Resale / other bans | Cloning consent |
|---|---|---|---|---|---|
| xAI | Yes, explicitly: you may "distribute or otherwise make the Bundled Service available to Customer's end users" (enterprise terms) | Yes: you own "all right, title, and interest in the Output in perpetuity" | You must not misrepresent output as human-generated | No leasing the service, no building a competing service, no using output to train models. The enterprise FAQ asks for attribution to Grok per brand guidelines. | Voice owner's rights needed |
| Inworld | Likely, but confirm. The general terms (June 11, 2025) license use "solely for your internal business purposes and as permitted by the documentation". | Outputs are assigned to you, subject to the terms. The service terms (May 14, 2026) say you keep rights in output. On termination you "must delete all Services, Models and Outputs". | not found as an explicit duty | No protected health info | Only voices you're authorised to use |
| Speechify | Yes: end users may "use, prompt, or download" output, but they can't upload their own voice files | Not stated in the API supplement; Studio terms say users own generated content | Yes: every use must carry "commercially reasonable disclosures" that it's AI, and your end-user terms must require the same | No reselling access, no using output to train voice models | Consent-first cloning, with warranties |
| ElevenLabs | Only on Scale ($299), Business, or Enterprise (OEM terms). Free/Starter/Creator/Pro are expressly excluded, and pay-as-you-go top-ups on a lower plan don't qualify. | Paid plans include a commercial licence. The free plan needs attribution. Beta-model output can't be used commercially. | End-user agreement minimum terms apply | "End User" wording is aimed at businesses. Whether merely playing audio to consumers counts as "making the service available" is a grey area, so ask them. | Strict verification for professional clones |
| Google Cloud | Yes (standard cloud terms; not re-read in this pass) | Yes, you keep your output (not re-read) | Not in the pages checked | Usual cloud acceptable-use rules | Custom voice needs consent |
| Azure | Yes (not re-read in this pass) | Yes (not re-read) | Yes: the Code of Conduct requires disclosure even for prebuilt voices | — | Limited Access approval |
| OpenAI | Yes (not re-read in this pass) | Yes (not re-read) | Yes: you must disclose the voice is AI-generated | — | No cloning offered |
| Murf | Yes, commercial and third-party use is allowed | Yes | not found | No reselling the voices themselves | not found |
| Cartesia | Yes on paid plans | not found | not found | Free tier is non-commercial | Pro and above |
| Fish Audio, Rime, Smallest, Hume, Qwen, Neuphonic, Luna | Terms not fully read (not found) | not found | not found | not found | not found |
← swipe sideways to see all columns →
This is a plain-English summary, not legal advice. Before launching a paid voice feature, ask your shortlisted providers to confirm in writing that consumer end-user playback is allowed on the plan you'll use.
Operations: hooking it into a Flask app on Replit
| Provider | Streaming | Official Python SDK | Formats that play on iPhone | Starter-tier limits | Data use |
|---|---|---|---|---|---|
| xAI | REST plus WebSocket (with interruption support) | not found (plain HTTPS works) | MP3 or WAV (also PCM and phone formats), 8–48 kHz | 50 concurrent WebSocket sessions per team. REST max 60,000 characters per request. | Not used for training. API data kept 30 days, then deleted (zero-retention option). Voice audio is never stored or used for training. SOC 2 and HIPAA BAA available. |
| Inworld | Yes | not found (plain HTTPS works) | MP3 and others | 5 concurrent requests on On-Demand | No training on your non-public materials |
| Speechify | Yes (streaming-native) | not found (plain HTTPS with a bearer key) | not checked | not found | not found |
| ElevenLabs | HTTP streaming plus WebSocket | Yes | MP3 and others | Concurrent requests: Starter 6/3, Creator 10/5, Pro 20/10, Scale 30/15 (Flash/Turbo vs other models) | Zero retention only on enterprise (not verified) |
| Cartesia | HTTP, SSE, WebSocket | Yes | MP3, WAV, PCM | Concurrent TTS: Pro 3, Startup 5, Scale 15 | not found |
| Yes | Yes | Chirp: MP3, OGG, WAV. The Gemini TTS API returns raw PCM, so wrap it as WAV or convert to MP3 on the server before sending it to the phone. | not found | Cloud data terms | |
| OpenAI | Yes | Yes | MP3, AAC, WAV, Opus | Tier-based | Not used for training by default (API) |
| Murf | Yes | not found | MP3, WAV | 5 concurrent (US-East), 2 elsewhere | not found |
| Deepgram | Yes | Yes | MP3, WAV and others | 45 concurrent on pay-as-you-go | not found |
| Fish Audio | WebSocket | not found | MP3, WAV and others | 5 concurrent (under $100 prepaid), 15 (≥$100), 50 (≥$1,000) | not found |
| Rime | HTTP plus WebSocket | not found | not found | 20 concurrent on Starter | SOC 2 Type II and HIPAA on Enterprise |
← swipe sideways to see all columns →
Reliability. All the main providers run public status pages. Uptime percentages were not published in a form I could read (not found). Recent incidents found:
- ElevenLabs: TTS and STT error spike on Aug 3, 2026.
- Cartesia: partial US outage July 29 and errors July 31, 2026.
- Inworld: TTS outages in March and May, and slow responses on Aug 24, 2026.
Company risk.
- PlayHT is the warning case. It was acquired, then shut down within about six months.
- Big platforms (Google, Microsoft, Amazon, OpenAI, xAI) are less likely to vanish, though they do retire models.
- Smaller specialists (Inworld, Cartesia, Rime, Speechify, Fish) move faster on quality but carry more change risk.
- Keep your code provider-agnostic so a switch takes an afternoon.
Drop-in pattern for Chance AI (Flask on Replit, iPhone):
- Keep the key server-side. Store the API key in Replit Secrets and never send it to the browser.
- Add a Flask route (e.g.
/speak) that takes the reply text, calls the TTS provider, and returnsaudio/mpeg(MP3). MP3 plays everywhere, including iPhone Safari. - Respect Safari's rules. iPhone Safari only allows sound after a tap, so create and "unlock" your audio player inside the user's first tap. Avoid raw PCM in the browser. A progress bar may show the duration as "Infinity" on blob URLs, so don't rely on it.
- Cache repeated phrases (greetings, errors) as files to save money.
- Make premium voices a setting that just switches the provider or voice ID in that one route.
Standouts beyond the big names
- Speechify Simba 3.2. #8 on Artificial Analysis, 96 on Vapi, 107 ms on Coval. Free 500K/month with commercial use, then $10/1M. English only. Disclosure rules apply.
- Qwen-Audio 3.1 TTS Plus (Alibaba). #3 on Artificial Analysis and #1 on the cloned-voice board, at about $19.30. Alibaba's realtime Qwen model measured 738 ms on Coval from the test location. Terms not reviewed (not found).
- Luna TTS (VUI Labs). #10 on Artificial Analysis with only 3 voices, and VUI's model was the fastest on Coval (43 ms). Price conflict: Artificial Analysis says $15, another site says $80 (unverified).
- Hume Octave 2. Emotion-aware and strong on the Hugging Face arena (1525), but #61 on Artificial Analysis and about $87.50 effective. Subscription-based.
- Rime (Coda / Mist v3). Very fast (Mist v3 at 37 ms by Rime's own figure), 20 concurrent on Starter. Coda had an 11.6% word error rate on Coval, which is high.
- Smallest.ai Lightning v3.1 Pro. $17.50–$19.50, #20 on Artificial Analysis, voice cloning from a 5-second sample.
- Fish Audio S2.1 Pro. $15 pay-as-you-go with no minimum, 83 on Vapi, 302 ms. Clear concurrency tiers.
- Neuphonic. #82 on Artificial Analysis. API price not published (not found). Its NeuTTS Air open model is Apache 2.0.
Open-source / self-hosting
| Model | Licence | Quality signal | Cost to run |
|---|---|---|---|
| Kokoro 82M | Apache 2.0 | Artificial Analysis 1064 (#54), HF arena 1477 | Hosted on DeepInfra: $0.62/1M (10M = $6.20). Small enough for CPU, but speed on Replit is untested. |
| Chatterbox (Resemble) | MIT | AA 1025; HD version 1102; HF arena 1480 | Self-host on a rented GPU: an RTX 4090 costs $0.34/hour (community) to $0.74/hour (secure) on RunPod, about $250–$540/month if always on. Serverless L4 is $0.69/hour, billed only while working. |
| Orpheus (Canopy) | Apache 2.0 code; the model is built on Llama 3.2, so Llama licence terms likely apply (check) | Vapi 89 | Same GPU costs as above |
| Breeze TTS 2 | Open weights | Top open model on Artificial Analysis (1216) | GPU required (size not checked) |
← swipe sideways to see all columns →
Verdict: self-hosting only beats xAI or Speechify on price at very high volume, and it adds work (servers, scaling, uptime). The exception is hosted Kokoro at $0.62/1M, as long as you can accept mid-table quality.
Voices to audition
| Provider | What to try | Official demo |
|---|---|---|
| xAI grok-tts | Castor, Ara, Eve, Rex, Leo | x.ai/api/voice |
| Inworld TTS-2 | Built-in voices plus voice design | inworld.ai/tts |
| Speechify Simba 3.2 | Filter the catalog to English | speechify.ai voice catalog |
| Google Gemini / Chirp 3 HD | Charon, Puck, Kore, Aoede, Achird, Zephyr | AI Studio speech · Chirp 3 HD voices |
| ElevenLabs | Conversational voices on v4 Turbo | Voice Library |
| OpenAI | marin, cedar | openai.fm |
| Cartesia | Emotive voices on Sonic 3.6 | Cartesia playground · Sonic |
| Murf Falcon 2 | Conversational voices | murf.ai/api |
| Fish Audio | Community voices | fish.audio |
| Blind-test yourself | Vote on pairs | Hugging Face TTS Arena |
← swipe sideways to see all columns →
Suggested test: paste three real Chance AI replies (a short answer, a long explanation, and an emotional or empathetic one). Play each candidate voice on the iPhone through the speaker and through earbuds, then pick the two you'd happily hear 50 times a day.
What changed since the first pass
- ElevenLabs pay-as-you-go doesn't unlock app use for your customers. Its OEM terms need the Scale plan ($299/month) or higher.
- Inworld's licence wording is tighter than it first looked. It says "internal business purposes" and requires deleting outputs on termination. Get written confirmation.
- xAI's terms are the most explicit about end-user distribution and output ownership, and it ranks #3 for humanness even though it sits at #26 on Artificial Analysis. Its docs list 20 languages, not "25+".
- Speechify Simba is a new contender: top-tier quality, the lowest price, and a free commercial tier.
- Cartesia dropped out of the top 5 because of the subscription requirement, low starter concurrency and a non-commercial free tier.
- Some third-party reviews are stale. One quotes grok-tts at $4.20 with 5 voices; the official figures are $15 and 28 voices.
What I couldn't verify
- Live Design Arena scores (only a Sept 2 snapshot)
- Artificial Analysis category and accent sub-boards
- Uptime percentages
- Reddit threads (blocked)
- xAI free credits
- Hume's exact plan prices
- Neuphonic API pricing
- Terms for Qwen, Fish, Rime, Smallest, Hume and Luna
- Resemble's current TTS price
- Whether Inworld plan fees count toward usage
- ElevenLabs v4 credits per character
- Independent MOS studies
Sources
- Artificial Analysis Speech Arena: provider voices · controlled voices · Luna TTS
- Vapi Humanness Index · Coval benchmarks · Hugging Face TTS Arena V2 · Design Arena TTS
- xAI: voice page · voice docs · enterprise terms · status
- Inworld: pricing · terms · TTS
- Speechify: pricing · API supplemental terms
- ElevenLabs: API pricing · OEM terms · concurrency
- Google: Gemini speech generation · Chirp 3 HD
- Cartesia pricing · Cartesia concurrency
- OpenAI TTS guide · OpenAI pricing
- Murf pay-as-you-go · Deepgram pricing · Azure Speech pricing · Amazon Polly pricing
- Fish Audio pricing and limits · Rime pricing · Smallest.ai pricing · Hume pricing · Qwen TTS (Alibaba)
- Open source: Kokoro · Kokoro on DeepInfra · Chatterbox · Orpheus · RunPod pricing