ElevenLabs released Eleven v4 and Eleven v4 Turbo on 28 September 2026, and on launch day Eleven v4 sat first on the Artificial Analysis Speech Arena at 1319 Elo, 43 points ahead of Cartesia Sonic 3.6. The model reads more than 90 languages, takes 10,000 characters per request, and clones a voice from 10 seconds of audio. The number that decides whether you should switch is on the API pricing page: $0.022 per 1,000 characters, marked "72% off until Oct 12". After that it is $0.08, the same as Eleven v3.

So for two weeks the best-rated voice on the board is cheaper than the second-best one. On 13 October it becomes 1.6 times the price of Cartesia and nearly five times the price of Gemini 3.8 Flash TTS. This piece compares v4 against the models it just passed, works out what an hour of narration costs on each, and says who should move a project now and who should wait.

What ElevenLabs shipped on September 28

Two models, aimed at two different jobs. Eleven v4 (model id eleven_v4) is the expressive model for narration, dialogue and character work. Eleven v4 Turbo (eleven_v4_turbo) is the real-time variant for voice agents, with a median inference latency of about 100 ms. Both are live in ElevenCreative, ElevenAgents and the ElevenAPI, according to the Eleven v4 product page.

Against Eleven v3, the concrete changes are:

  • Languages: 90+ against v3's 70+. TechCrunch reports the biggest quality jump is in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
  • Request size: 10,000 characters per generation (about 10 minutes of audio) against v3's 5,000, per the ElevenLabs models documentation.
  • Voice cloning: Instant Voice Clones now need 10 seconds of audio, and Professional Voice Clones are supported.
  • Audio tags: the inline direction syntax from v3 carries over and now stacks. You can write [excited, happy] or [said angrily in French accent], add sound cues like [light rain] or [phone buzzing], and the model follows the tags in sequence.
  • Consistency: ElevenLabs says speaker identity is now stable across regenerations, which matters most on long projects where a voice has to sound the same in chapter one and chapter twenty.

ElevenLabs describes v4 as an "entirely new architecture" rather than a v3 update, so treat old v3 prompts as a starting point to re-test, not a guaranteed drop-in.

3D slabs comparing the 5,000-character Eleven v3 request limit with the 10,000-character Eleven v4 limit
Characters per request: Eleven v3 5,000, Eleven v4 10,000.

The leaderboard lead is outside the error bars

Vendor-run preference tests are easy to wave away, so start with the independent number. On the Artificial Analysis Speech Arena, where listeners vote blind between two clips, the top of the table on launch day reads:

RankModelElo95% CISamplesPrice per 1M chars (list)
1ElevenLabs Eleven v41319±191,674$80.0
2Cartesia Sonic 3.61276±161,946$49.0
3Google Gemini 3.8 Flash TTS1267±162,198$16.5
4Alibaba Qwen-Audio-3.0-TTS-Plus1258±161,593$19.3
5Inworld Realtime TTS-21246±171,365$20.8
18ElevenLabs Eleven v31169

Two things stand out. First, the gap is real in statistical terms. Eleven v4's lower bound is 1300 and Sonic 3.6's upper bound is 1292, so the intervals do not overlap. A 43-point Elo lead means listeners picked v4 roughly 56% of the time in a direct pairing, which is a clear but not crushing margin. Second, v4 is 150 points above Eleven v3, which sits 18th. This is a generational jump for ElevenLabs, not a refresh.

ElevenLabs' own blind tests, which put v4 against Sonic 3.6, Inworld TTS-2 and both Gemini 3.8 TTS models, report v4 winning 65% to 81% of head-to-heads. That range is consistent with the arena, but it is the vendor's test, so the arena is the number to quote.

One caveat: 1,674 samples is the smallest count in the top three. New models often drift down a little as votes accumulate. Check the board again before the launch price ends.

The price: 72% off for 14 days, then 4.8 times Gemini

Here is the list that matters, from the ElevenAPI pricing page and the Artificial Analysis price column for the competitors:

ModelPer 1M chars, until 12 OctPer 1M chars, from 13 OctArena Elo
Eleven v4$22$801319
Eleven v4 Turbo$11$40not ranked
Eleven v3$80$801169
Cartesia Sonic 3.6$49$491276
Gemini 3.8 Flash TTS$16.5$16.51267
Gemini 3.8 Flash-Lite TTS$11$111241

During the offer, v4 is 55% cheaper than Sonic 3.6 and only a third more than Gemini 3.8 Flash TTS, while beating both on the arena. That is an unusually good deal for a model at the top of the board.

From 13 October, the picture inverts. At $80 per million characters v4 costs 1.6 times Sonic 3.6 and 4.8 times Gemini 3.8 Flash TTS, for a lead of 43 and 52 Elo points. Whether that premium is worth it depends entirely on your material. For a character-driven audio drama, maybe. For a product walkthrough read in a neutral voice, almost certainly not.

Google has the mirror-image problem. As we covered when Gemini 3.8 Flash TTS launched on 23 September, its audio output price doubles on 1 January 2027. Both vendors are using launch pricing to win workloads. The difference is that ElevenLabs' window is 14 days and Google's is three months.

3D stepped platforms showing $22, $49 and $80 per million characters
Per 1M characters: Eleven v4 at launch $22, Cartesia Sonic 3.6 $49, Eleven v4 from 13 October $80.

What an hour of narration costs

ElevenLabs prices v4 at "~$0.02/minute" at the launch rate on its own pricing page, which works out to roughly 1,000 characters per minute of speech. Using that same ratio for every model gives a like-for-like hourly figure (about 60,000 characters). Real speech rates vary with voice and pacing, so treat these as estimates for comparison, not quotes:

ModelPer hour, nowPer hour, from 13 Oct10-hour audiobook, now
Eleven v4$1.32$4.80$13.20
Eleven v4 Turbo$0.66$2.40$6.60
Cartesia Sonic 3.6$2.94$2.94$29.40
Gemini 3.8 Flash TTS$0.99$0.99$9.90
Qwen-Audio-3.0-TTS-Plus$1.16$1.16$11.60
Inworld Realtime TTS-2$1.25$1.25$12.50

The hourly spread is small in absolute terms. A 10-hour audiobook on v4 at list price is $48, against $29.40 on Sonic 3.6 and $9.90 on Gemini 3.8 Flash TTS. For a single creator producing a few hours a month, the difference is lunch money, and the quality gap may be worth it. For a studio localizing a course catalog into 20 languages, the same gap is a line item someone will ask about.

Retakes are the hidden multiplier. If v4's identity stability means you regenerate fewer lines, its effective cost per finished minute drops. That is the claim to test in your own project, because no leaderboard measures it.

The subscription math creators should check first

Most creators do not pay per character on the API. They pay a monthly plan and spend included characters. The ElevenLabs pricing page converts each plan's allowance into v4 characters at the launch rate, and the difference from v3 is large:

PlanMonthly priceEleven v4 included (launch rate)Eleven v3 included
Free$010,000 chars, ~10 min10,000 chars, ~10 min
Starter$6~273,000 chars, ~273 min75,000 chars, ~75 min
Creator$22 ($11 first month)1,000,000 chars, ~1,000 min275,000 chars, ~275 min
Pro$994,500,000 chars, ~4,500 min1,238,000 chars, ~1,238 min

On the $6 Starter plan, that is about 273 minutes of v4 against 75 minutes of v3, roughly 3.6 times the output, on a better model. ElevenLabs has not said what the included v4 amounts become after 12 October. Since the figures are computed from the $0.022 launch rate, expect them to shrink toward the v3 numbers once the rate returns to $0.08. If you have a backlog of narration, the next two weeks are when your plan stretches furthest.

3D stacks comparing about 273 minutes of Eleven v4 with 75 minutes of Eleven v3 on the $6 Starter plan
$6 Starter plan: about 273 minutes of Eleven v4 at the launch rate against 75 minutes of Eleven v3.

How to move a narration project to Eleven v4

A practical switch-over for a voiceover, audiobook or course project, completable in an afternoon:

  1. Pick three representative passages from your script: one neutral, one emotional, one with names or jargon. Keep each under 1,000 characters so tests stay cheap.
  2. Generate each on v3 and v4 with the same voice. In the app, choose the model in the Text to Speech panel. On the API, set model_id to eleven_v4; the API reference covers the request format.
  3. Rewrite your directions as stacked tags. Where v3 needed one tag per line, try combinations such as [sighs] [whispers] before a sentence and a sound cue like [door slams] at a scene break. Listen for whether the order is followed.
  4. Test identity stability. Regenerate the emotional passage three times. If the voice drifts between takes on v3 and holds on v4, that is the retake saving you are paying for.
  5. Re-chunk long scripts to 10,000 characters. v4 doubles v3's limit, so a 40,000-character chapter drops from eight requests to four, with fewer seams to edit.
  6. Re-clone if you use an Instant Voice Clone. A fresh clone from a clean 10-second sample is worth comparing against your old clone on v4 before you commit.
  7. Batch your backlog before 12 October if the tests pass. Generate the finished audio at the launch rate, and keep v3 or a cheaper model for drafts after the offer ends.

If you are choosing voices from scratch, the ElevenLabs Voice Library lists more than 17,500, filterable by narration, conversational, character and advertisement styles.

Eleven v4 Turbo for voice agents: fast, and unranked

Eleven v4 Turbo is the model for anyone building a talking assistant, a game NPC or a phone agent. ElevenLabs quotes a median time to first speech of 150 ms, against 262 ms for Cartesia Sonic 3.6 and 814 ms for OpenAI's GPT-4o mini TTS in its own September 2026 tests. For context, ElevenLabs lists the older v3 Conversational model at about 280 ms, and Flash v2.5 at about 75 ms of model latency. TechCrunch adds that v4 can start generating audio as soon as the LLM behind an agent starts producing its answer.

Two cautions before you swap it into production. The latency numbers are vendor-measured and exclude network overhead, so measure from your own region. And Turbo does not yet appear on the Artificial Analysis board, so its quality is not independently rated. The 1319 Elo belongs to the full v4 model, not Turbo. When Inworld TTS-2 launched with sub-200 ms latency, the real-world gap to competitors was smaller than the launch chart. Expect the same here.

On price, Turbo at $11 per million characters during the offer matches Gemini 3.8 Flash-Lite TTS. After the offer, $40 puts it at the same list price as v3 Conversational and Flash.

3D rods showing time to first speech of 150ms, 262ms and 814ms
Median time to first speech in ElevenLabs' tests: Eleven v4 Turbo 150ms, Cartesia Sonic 3.6 262ms, GPT-4o mini TTS 814ms.

Verdict: who should switch, and who should wait

Switch now if you produce character-driven audio: audio drama, games, animated shorts, expressive narration. v4 has the clearest quality lead on the independent board, stacked tags solve the direction problem v3 left open, and the launch rate makes it cheaper than Sonic 3.6 for the next two weeks.

Switch now, but budget for October if you are already on an ElevenLabs plan. The included minutes more than triple at the launch rate. Use them on finished audio, not experiments.

Wait if you produce high-volume, neutral narration. Gemini 3.8 Flash TTS sits 52 Elo points behind at under a quarter of v4's list price, and open-source voice cloning is free if you can run it locally. The v4 premium buys expressiveness you may not use.

Measure first if you build voice agents. Turbo's latency claims are strong, but they are the vendor's, and the model has no independent quality score yet.

Frequently asked questions

What is the difference between Eleven v4 and Eleven v4 Turbo?

Eleven v4 is the expressive model for narration and character work, with a 10,000-character request limit. Eleven v4 Turbo is tuned for real-time voice agents, with about 100 ms median inference latency and 150 ms to first speech. Turbo costs half as much per character.

How much does Eleven v4 cost?

On the API, $0.022 per 1,000 characters until 12 October 2026, then $0.08. Eleven v4 Turbo is $0.011 during the offer and $0.04 after. On subscription plans, v4 uses included characters, and the $6 Starter plan currently covers about 273 minutes.

Is Eleven v4 better than Gemini 3.8 Flash TTS?

On the Artificial Analysis Speech Arena, yes: 1319 Elo against 1267, with confidence intervals that do not overlap. Gemini 3.8 Flash TTS costs $16.5 per million characters against v4's $80 list price, so the answer for your budget depends on how much expressiveness your material needs.

Do my Eleven v3 audio tags work in Eleven v4?

The tag syntax carries over, and v4 adds the ability to stack several tags that the model follows in order. Because v4 is a new architecture, re-test your existing prompts rather than assuming identical output.

How many languages does Eleven v4 support?

More than 90, up from 70+ on Eleven v3. ElevenLabs says a voice recorded in one language can speak the others with a native accent.

How much audio do I need to clone a voice with Eleven v4?

Instant Voice Clones need 10 seconds of audio. Professional Voice Clones, trained on a longer recording for the highest fidelity, are also supported. Both require verified consent from the voice owner.