Chance AI: Text-to-Speech Voice Report (deep dive)

Text-to-speech providers compared for a paid AI chat app: quality, voices, price, rights and setup. October 6, 2026.

Research date: October 6, 2026. All prices are in US dollars per 1 million characters of text unless noted. Anything marked "not found" could not be confirmed from a public source. Prices and leaderboards change often, so re-check before you sign up.

The short version

Chance AI needs voices that sound natural on a phone, cost little, can be bought pay-as-you-go, and can legally be played to paying users, with a possible premium voice upgrade. With those needs in mind:

Pick Provider and model Why
Best overall xAI grok-tts (try the voice Castor) $15 flat, pay-as-you-go. Its terms explicitly allow giving the output to your end users, and you own the output "in perpetuity". It scores #3 of 21 on Vapi's blind "humanness" test and has the best pronunciation accuracy measured on Coval.
Best budget Speechify Simba 3.2 500,000 free characters every month with commercial use allowed. After that it's $10 for 1M, or $10/month for 1.9M. Blind-test results are near the top (#8 on Artificial Analysis, #2 on Vapi). English only.
Best raw quality ElevenLabs Eleven v4 Turbo #1 on the Artificial Analysis blind test. But putting it in a consumer app requires the $299/month Scale plan or higher; pay-as-you-go top-ups don't qualify.
Best premium upgrade Inworld Realtime TTS-2 Top-tier blind-test results (#7 on Artificial Analysis, #2 on the Hugging Face arena with its Max model, #4 on the cloned-voice board). Pay-as-you-go at $25, no monthly fee, very fast, and you can design voices from a text description. Read the licence caveats below.

← swipe sideways to see all columns →

Final top 5 for Chance AI

  1. xAI grok-tts. It has the cleanest commercial terms of any provider checked, a simple $15 rate, 28 voices including Castor, 50 simultaneous streams, and Safari playback guidance in its docs. Its weak spots are a middling Artificial Analysis score (#26) and slower first audio (354 ms).
  2. Inworld TTS-2 / TTS-2 Flash. It ranks near the top on every blind test that includes it, and its Flash model is one of the fastest measured (61 ms). It costs $25 (or $15 for Flash) pay-as-you-go. Its licence wording ("internal business purposes", and deleting outputs when the contract ends) needs a written OK from Inworld.
  3. Speechify Simba 3.2. It's the cheapest of the top-quality voices and has a free tier that allows commercial use. Its terms require you to tell users the voice is AI and to pass that duty on to them. English only.
  4. Google Gemini 3.8 Flash TTS. #5 on Artificial Analysis, and Google's family leads Design Arena. You can steer the style with a plain-English prompt, and it costs about $16.50. It's still a preview, and Google has scheduled the price to roughly double on Jan 1, 2027.
  5. ElevenLabs v4 Turbo, as the luxury option. It's the best-sounding voice today, but the $299/month floor only makes sense once premium subscribers cover it.

Runner-up: Murf Falcon 2 ($10, with $10 of free credit every month, about 1M characters, pay-as-you-go).

Dropped from the earlier top 5: Cartesia. It's still excellent, but its free tier is non-commercial, it needs a subscription, and the $5 plan allows only 3 simultaneous requests.

Jargon, in plain English

How the voices compare in blind tests

No single leaderboard is the truth. Each one tests different voices, scripts and listeners, so look for models that do well on several.

Model Artificial Analysis (provider voices, Elo / rank of 94) AA cloned-voice board (rank of 39) Vapi Humanness (human = 100) Hugging Face TTS Arena V2 (Elo) Design Arena (Sept 2 snapshot) Coval WER
Eleven v4 Turbo 1334 / #1 #2 (1173) not tested not listed not listed not found
Eleven v4 1321 / #2 #3 (1160) not tested not listed not listed not found
Eleven v3 1174 / #18 #9 (1072) 97 (#1) 1502 1290 (#4) not found
Qwen-Audio 3.1 TTS Plus 1292 / #3 #1 (1181) not tested not listed Qwen3 Flash 1040 not found
Cartesia Sonic 3.6 1278 / #4 #5 (1140) not found Sonic 2: 1514 Sonic 3.5: 1149 not found
Google Gemini 3.8 Flash TTS 1275 / #5 #14 (1050) not tested not listed Gemini 3.1 Flash 1459 (#1) Chirp 3 HD 1.6% (best)
Inworld Realtime TTS-2 1251 / #7 #4 (1145) TTS-1.5 Max: 79 TTS MAX 1557 (#2) 1.5 Max: 1149 4.5% (secondary source)
Speechify Simba 3.2 1242 / #8 Simba 3.0: #15 96 (#2) not listed not listed 4.3% (secondary source)
Luna TTS (VUI Labs) 1227 / #10 not listed not tested not listed not listed not found
Inworld TTS-2 Flash 1214 / #12 #10 (1071) not tested not listed not listed 5.4% (secondary source)
Murf Falcon 2 1155 / #21 not listed not found not listed Murf Gen2: 1103 not found
Fish Audio S2.1 Pro 1141 / #24 #19 (1015) 83 OpenAudio S2: 1519 not listed not found
xAI grok-tts (listed as "SpaceXAI TTS") 1136 / #26 #13 (1054) 93 (#3) not listed 1254 (#5) 1.7%
MiniMax Speech 2.8 HD 1173 / #19 #16 93 1539 02 HD: 1130 not found
Smallest Lightning v3.1 Pro 1171 / #20 not listed not found 1524 1149 not found
Azure HD 2.5 1131 / #28 not listed not found not listed not listed not found
OpenAI TTS-1 HD / gpt-4o-mini-tts 1103 / #36 not listed not found not listed GPT-4o Mini TTS 1225 (#6) not found
Polly Generative 1065 / #53 not listed not found not listed Polly 989 not found
Kokoro 82M (open source) 1064 / #54 not listed not found 1477 1040 not found
Rime Coda 1062 / #55 not listed not found not listed not listed 11.6%
Google Chirp 3 HD 1056 / #58 not listed not found not listed not listed 1.6%
Hume Octave 2 1051 / #61 not listed not found Octave: 1525 1162 not found
Chatterbox (open source) 1025 / #73 #35 not found 1480 955 not found
Deepgram Aura-2 not listed not listed not found not listed 1015 not found
Neuphonic 940 / #82 not listed not found NeuTTS Max 1439 not listed not found
Polly Neural 894 / #90 not listed not found not listed not listed not found

← swipe sideways to see all columns →

What the boards mean:

Forum and review sentiment (Reddit blocks automated reading, so this comes from search summaries and is only a rough signal):

Speed (time to first audio)

Model Coval median TTFA Vendor's own claim
VUI Labs (maker of Luna TTS) 43 ms not found
Inworld TTS-2 Flash 61 ms P90 about 20 ms (vendor)
Speechify Simba 3.2 107 ms not found
Inworld TTS-2 149 ms P90 about 100 ms
ElevenLabs Flash v2.5 / v4 Turbo 187 / 194 ms about 75 ms / about 100 ms
Deepgram Aura-2 256 ms not found
Murf Falcon 2 274 ms about 100 ms
Rime Mist v3 / Coda 279 / 319 ms 37 / 96 ms (P50)
Cartesia Sonic 3.5 / 3.6 286 / 304 ms under 90 ms
Fish Audio S2.1 Pro 302 ms not found
Smallest Lightning v3.1 Pro 342 ms under 100 ms
xAI grok-tts 354 ms not found. Vapi measured 460 ms (285 ms when streaming).
Google Chirp 3 HD 536 ms not found
Alibaba Qwen3 TTS Flash Realtime 738 ms not found
OpenAI gpt-4o-mini-tts 2,461 ms not found

← swipe sideways to see all columns →

Vendor numbers are usually measured in ideal conditions. Independent numbers include real network time. For a chat app that shows the text first, anything under about 400 ms feels fine.

Voice selection

Provider Built-in voices Languages Voice design from a prompt Cloning Delivery controls Standout voices to try
xAI grok-tts 28 (Ara, Eve, Leo, Rex, Sal, Altair, Atlas, Aurora, Carina, Castor, Celeste, Cosmo, Helios, Helix, Iris, Kepler, Liora, Lumen, Luna, Lux, Naksh, Orion, Perseus, Rigel, Sirius, Ursa, Zagan, Zenith) 20 in the docs (marketing says 25+). Every voice speaks every language. not found Yes: Custom Voices API, up to 120-second clip, 30 free clones Inline tags like [pause] and [laugh], wrapping tags like <whisper>, speed 0.7–1.5x, pronunciation dictionary with phonetic spelling, timestamps Castor ("charismatic, down-to-earth, and easygoing"), Ara, Eve, Rex
Inworld 95 14 native, 200+ cross-lingual Yes, from a 30–250 character English description Yes, instant (included pay-as-you-go) TTS-2 takes bracketed plain-English direction (emotion, pace, volume, pitch, style) plus [laugh], [sigh], [breathe] and similar. Flash ignores direction but keeps the sound tags. Try several in the Inworld playground
Speechify Simba 900+ catalog voices Simba 3.2: English. Simba 3.0: 6 languages / 7 locales not found Yes, paid plans, consent-first SSML (tags aren't billed), streaming Browse the voice catalog
Google Gemini TTS / Chirp 3 HD 30 (16 male, 14 female by Google's labels) Gemini: 24 generally available plus about 70 in preview. Chirp: 53 locales. Style by prompt ("say this warmly…") Chirp instant custom voice ($60/1M) Plain-English style prompts, tags like [sigh], [laughing], [whispering], pauses. Chirp pace 0.25–2x. Charon, Puck, Kore, Aoede, Achird, Zephyr
ElevenLabs 10,000+ community voices. Older default voices expire Dec 31, 2026. not re-checked in this pass Yes: Voice Design, 20–1,000 character description Yes, instant and professional v4 audio tags like [whispers], [laughs], [excited], [sighs] Voice Library "conversational" filter
Cartesia 500+ 44 languages / 61 locales not found Instant (Pro), professional (Startup) <speed>, <volume>, <emotion> (beta), [laughter] Its "Emotive" voices
OpenAI 13 (alloy, ash, ballad, cedar, coral, echo, fable, marin, nova, onyx, sage, shimmer, verse) Many (follows the input text) Style via an "instructions" field No Free-text instructions on gpt-4o-mini-tts marin, cedar (OpenAI's recommended)
Murf Falcon 2 150+ not found exactly not found not found not found in detail Browse the Murf API page
Deepgram Aura-2 90 (41 English) 7 No No Limited —
Azure 400+ 100+ No Custom/personal voice requires Limited Access approval Full SSML, speaking styles HD voices
Amazon Polly 100+ (43 generative) 40+ No No SSML Generative voices
Rime 184 (Coda), 94 (Mist v3). FAQ says 600+ overall. Coda 8, Mist v3 4 No Enterprise Speed, pauses, spelling, pronunciation —
Fish Audio Large community library (count not found) Multilingual Yes: Voice Design API, $0.01 per request Yes Emotion cues —
Luna TTS (VUI Labs) 3 not found not found not found Expression tags, speed —

← swipe sideways to see all columns →

How many good conversational US-English male and female voices each provider has: no provider publishes this count, so it's not found. In practice:

Audition before you decide.

Price and worked monthly costs

To make the numbers concrete, assume an average spoken reply of about 300 characters (roughly 20 seconds of audio):

Costs below are what you'd pay with commercial rights for end users, after free allowances.

Provider / model List price per 1M chars 100K / month 1M / month 10M / month Free allowance and conditions
xAI grok-tts $15 $1.50 $15 $150 Free credits not confirmed in official docs. Third-party guides mention a $25 sign-up credit and a $150/month data-sharing programme; the latter means sharing data, so avoid it if you want privacy.
Inworld TTS-2 (On-Demand) $25 $2.50 $25 $250 About 70 minutes free, no subscription. Plans lower the rate to $20–$12.50, but whether the plan fee counts toward usage is not verified.
Inworld TTS-2 Flash $15 $1.50 $15 $150 Same as above. Plans go down to $10–$7.
Speechify Simba 3.2 $10 (Starter) / $8 (Pro) / $6 (Scale) $0 (Free: 500K/month) $10 (Starter plan includes 1.9M) $91 (Starter $10 plus 8.1M top-up), or $99 on Pro (13.5M included) Free tier allows commercial use but can't top up. Plans are month-to-month.
Google Gemini 3.8 Flash TTS about $16.50 (Artificial Analysis estimate) about $1.65 about $16.50 about $165 Token pricing doubles on Jan 1, 2027 ($9 to $18 per 1M audio tokens), so expect roughly $33/1M after that (estimate). Preview, Enterprise API.
Google Chirp 3 HD $30 $0 $0 $270 1M characters free every month
ElevenLabs v4 Turbo $40 ($11 during a 72%-off promo until Oct 12, 2026) $299 $299 about $627 (estimate) Scale plan ($299/month, 1.8M credits) required for end-user apps. 10M estimate = $299 + 8.2M × $0.04/1K, assuming 1 credit per character (not verified for v4).
ElevenLabs v4 $80 $299 $299 about $955 (estimate) Same Scale requirement. Startup grant: 12 months, 33M characters (application needed).
Cartesia Sonic 3.6 Overage $65 / $45 / $38 by plan $5 (Pro, 100K included) $49 (Startup, 1.25M included) about $375 (Scale $299 with 8M included, plus 2M at $38) Free tier (20K) is non-commercial
Murf Falcon 2 $10 $0 $0 (covered by the $10 monthly credit) $90 $10 free credit every month. $10 minimum top-up, and credits don't expire.
OpenAI tts-1 / tts-1-hd $15 / $30 $1.50 / $3 $15 / $30 $150 / $300 gpt-4o-mini-tts is billed by audio minute (about $0.015/min per OpenAI's estimate)
Deepgram Aura-2 $30 $3 $30 $300 $200 one-time credit. Flux TTS ($45) has a 1:1 credit match up to $500 until Dec 31, 2026.
Azure Neural / HD $15 / $22 $0 (free F0 tier, 0.5M/month) $15 / $22 $150 / $220 Custom voice $24 plus $4.03/hour hosting
Amazon Polly Neural / Generative $16 / $30 $1.60 / $3 $16 / $30 $160 / $300 First 12 months: 1M Neural / 100K Generative free per month. New accounts can get up to $200 in credits.
Fish Audio S2.1 Pro $15 per 1M bytes (≈ characters in English) $1.50 $15 $150 Pay-as-you-go, no minimum. The free model (s2.1-pro-free) is meant for testing.
Rime Coda / Mist v3 $50 / $30 $5 / $3 $50 / $30 $500 / $300 Free minutes on Starter: the page says both "about 800 minutes" and "3,000 minutes" (conflicting)
Smallest Lightning v3.1 Pro $17.50 on the pricing page ($19.50 on the model card) about $1.75 about $17.50 about $175 not found
Qwen 3.0 TTS Plus (Alibaba, international) $20 $2 $20 $200 not found
MiniMax Speech 2.8 HD $100 $10 $100 $1,000 not found
Hume Octave Subscription. Overage $0.15–$0.05 per 1K by plan (secondary source). not computed not computed not computed Free and Starter can't buy extra usage
Resemble about $0.0005/second of audio (secondary source) not computed about $30 (estimate) about $300 (estimate) Its official page now lists mainly detection products
PlayHT — — — — Shut down: Meta acquired it in July 2025 and the platform closed Dec 31, 2025

← swipe sideways to see all columns →

Highlights:

Scheduled changes and promos:

Change Ends or takes effect
ElevenLabs 72% off v4 / v4 Turbo, and 3x credits on Creator and above Until Oct 12, 2026
Deepgram Flux credit match Through Dec 31, 2026
Google Gemini 3.8 Flash TTS price doubles Jan 1, 2027
Old ElevenLabs default voices expire Dec 31, 2026
Inworld price increases Inworld's terms promise 30 days' notice

← swipe sideways to see all columns →

Commercial rights: can you play these voices to paying users?

Provider Embed in a paid app for end users? You own the audio? Must disclose it's AI? Resale / other bans Cloning consent
xAI Yes, explicitly: you may "distribute or otherwise make the Bundled Service available to Customer's end users" (enterprise terms) Yes: you own "all right, title, and interest in the Output in perpetuity" You must not misrepresent output as human-generated No leasing the service, no building a competing service, no using output to train models. The enterprise FAQ asks for attribution to Grok per brand guidelines. Voice owner's rights needed
Inworld Likely, but confirm. The general terms (June 11, 2025) license use "solely for your internal business purposes and as permitted by the documentation". Outputs are assigned to you, subject to the terms. The service terms (May 14, 2026) say you keep rights in output. On termination you "must delete all Services, Models and Outputs". not found as an explicit duty No protected health info Only voices you're authorised to use
Speechify Yes: end users may "use, prompt, or download" output, but they can't upload their own voice files Not stated in the API supplement; Studio terms say users own generated content Yes: every use must carry "commercially reasonable disclosures" that it's AI, and your end-user terms must require the same No reselling access, no using output to train voice models Consent-first cloning, with warranties
ElevenLabs Only on Scale ($299), Business, or Enterprise (OEM terms). Free/Starter/Creator/Pro are expressly excluded, and pay-as-you-go top-ups on a lower plan don't qualify. Paid plans include a commercial licence. The free plan needs attribution. Beta-model output can't be used commercially. End-user agreement minimum terms apply "End User" wording is aimed at businesses. Whether merely playing audio to consumers counts as "making the service available" is a grey area, so ask them. Strict verification for professional clones
Google Cloud Yes (standard cloud terms; not re-read in this pass) Yes, you keep your output (not re-read) Not in the pages checked Usual cloud acceptable-use rules Custom voice needs consent
Azure Yes (not re-read in this pass) Yes (not re-read) Yes: the Code of Conduct requires disclosure even for prebuilt voices — Limited Access approval
OpenAI Yes (not re-read in this pass) Yes (not re-read) Yes: you must disclose the voice is AI-generated — No cloning offered
Murf Yes, commercial and third-party use is allowed Yes not found No reselling the voices themselves not found
Cartesia Yes on paid plans not found not found Free tier is non-commercial Pro and above
Fish Audio, Rime, Smallest, Hume, Qwen, Neuphonic, Luna Terms not fully read (not found) not found not found not found not found

← swipe sideways to see all columns →

This is a plain-English summary, not legal advice. Before launching a paid voice feature, ask your shortlisted providers to confirm in writing that consumer end-user playback is allowed on the plan you'll use.

Operations: hooking it into a Flask app on Replit

Provider Streaming Official Python SDK Formats that play on iPhone Starter-tier limits Data use
xAI REST plus WebSocket (with interruption support) not found (plain HTTPS works) MP3 or WAV (also PCM and phone formats), 8–48 kHz 50 concurrent WebSocket sessions per team. REST max 60,000 characters per request. Not used for training. API data kept 30 days, then deleted (zero-retention option). Voice audio is never stored or used for training. SOC 2 and HIPAA BAA available.
Inworld Yes not found (plain HTTPS works) MP3 and others 5 concurrent requests on On-Demand No training on your non-public materials
Speechify Yes (streaming-native) not found (plain HTTPS with a bearer key) not checked not found not found
ElevenLabs HTTP streaming plus WebSocket Yes MP3 and others Concurrent requests: Starter 6/3, Creator 10/5, Pro 20/10, Scale 30/15 (Flash/Turbo vs other models) Zero retention only on enterprise (not verified)
Cartesia HTTP, SSE, WebSocket Yes MP3, WAV, PCM Concurrent TTS: Pro 3, Startup 5, Scale 15 not found
Google Yes Yes Chirp: MP3, OGG, WAV. The Gemini TTS API returns raw PCM, so wrap it as WAV or convert to MP3 on the server before sending it to the phone. not found Cloud data terms
OpenAI Yes Yes MP3, AAC, WAV, Opus Tier-based Not used for training by default (API)
Murf Yes not found MP3, WAV 5 concurrent (US-East), 2 elsewhere not found
Deepgram Yes Yes MP3, WAV and others 45 concurrent on pay-as-you-go not found
Fish Audio WebSocket not found MP3, WAV and others 5 concurrent (under $100 prepaid), 15 (≥$100), 50 (≥$1,000) not found
Rime HTTP plus WebSocket not found not found 20 concurrent on Starter SOC 2 Type II and HIPAA on Enterprise

← swipe sideways to see all columns →

Reliability. All the main providers run public status pages. Uptime percentages were not published in a form I could read (not found). Recent incidents found:

Company risk.

Drop-in pattern for Chance AI (Flask on Replit, iPhone):

  1. Keep the key server-side. Store the API key in Replit Secrets and never send it to the browser.
  2. Add a Flask route (e.g. /speak) that takes the reply text, calls the TTS provider, and returns audio/mpeg (MP3). MP3 plays everywhere, including iPhone Safari.
  3. Respect Safari's rules. iPhone Safari only allows sound after a tap, so create and "unlock" your audio player inside the user's first tap. Avoid raw PCM in the browser. A progress bar may show the duration as "Infinity" on blob URLs, so don't rely on it.
  4. Cache repeated phrases (greetings, errors) as files to save money.
  5. Make premium voices a setting that just switches the provider or voice ID in that one route.

Standouts beyond the big names

Open-source / self-hosting

Model Licence Quality signal Cost to run
Kokoro 82M Apache 2.0 Artificial Analysis 1064 (#54), HF arena 1477 Hosted on DeepInfra: $0.62/1M (10M = $6.20). Small enough for CPU, but speed on Replit is untested.
Chatterbox (Resemble) MIT AA 1025; HD version 1102; HF arena 1480 Self-host on a rented GPU: an RTX 4090 costs $0.34/hour (community) to $0.74/hour (secure) on RunPod, about $250–$540/month if always on. Serverless L4 is $0.69/hour, billed only while working.
Orpheus (Canopy) Apache 2.0 code; the model is built on Llama 3.2, so Llama licence terms likely apply (check) Vapi 89 Same GPU costs as above
Breeze TTS 2 Open weights Top open model on Artificial Analysis (1216) GPU required (size not checked)

← swipe sideways to see all columns →

Verdict: self-hosting only beats xAI or Speechify on price at very high volume, and it adds work (servers, scaling, uptime). The exception is hosted Kokoro at $0.62/1M, as long as you can accept mid-table quality.

Voices to audition

Provider What to try Official demo
xAI grok-tts Castor, Ara, Eve, Rex, Leo x.ai/api/voice
Inworld TTS-2 Built-in voices plus voice design inworld.ai/tts
Speechify Simba 3.2 Filter the catalog to English speechify.ai voice catalog
Google Gemini / Chirp 3 HD Charon, Puck, Kore, Aoede, Achird, Zephyr AI Studio speech · Chirp 3 HD voices
ElevenLabs Conversational voices on v4 Turbo Voice Library
OpenAI marin, cedar openai.fm
Cartesia Emotive voices on Sonic 3.6 Cartesia playground · Sonic
Murf Falcon 2 Conversational voices murf.ai/api
Fish Audio Community voices fish.audio
Blind-test yourself Vote on pairs Hugging Face TTS Arena

← swipe sideways to see all columns →

Suggested test: paste three real Chance AI replies (a short answer, a long explanation, and an emotional or empathetic one). Play each candidate voice on the iPhone through the speaker and through earbuds, then pick the two you'd happily hear 50 times a day.

What changed since the first pass

What I couldn't verify

Sources