Modern Creator
ElevenLabs · YouTube

Eleven v4 Is Here: Everything You Need to Know

ElevenLabs introduces Eleven v4, a text to speech model that reads a script for context like a voice actor, and Eleven v4 Turbo, the same model tuned for real time voice agents.

Posted
today
Duration
Format
Demo
educational
Views
4.2K
279 likes
Big Idea

The argument in one line.

Eleven v4 reads a script for context the way a voice actor does, while Eleven v4 Turbo trades a little of that nuance for about 100ms latency so the same expressive model can run inside live voice agents.

Who This Is For

Read if. Skip if.

READ IF YOU ARE…
  • Anyone building with ElevenLabs who needs to choose between the standard v4 model and the low latency Turbo version for a specific use case.
  • Developers building voice agents, audiobooks, or long narrated content who have run into voices drifting or losing consistency across a project.
  • Creators working in Japanese, Brazilian Portuguese, Mandarin, or Cantonese who need better pronunciation and delivery than earlier ElevenLabs models offered.
SKIP IF…
  • You don't use ElevenLabs or a comparable text to speech API. This is a feature rundown of one company's new model, not a general voice AI explainer.
  • You're looking for pricing, benchmarks, or a side by side comparison against competitors. None of that is covered here.
TL;DR

The full version, fast.

ElevenLabs releases Eleven v4, a text to speech model built on a new architecture that reads a script's context the way a voice actor would, tracking who's speaking, what just happened, and how a line should land, so multi-speaker dialogue gets real overlap and timing even without audio tags. Eleven v4 Turbo is the same research tuned for about 100ms latency so live voice agents can use it in real time. Voice cloning holds up better across long-form regenerations, context stitching keeps a regenerated line consistent with the chapter around it, and language coverage improves most in Japanese, Brazilian Portuguese, Mandarin, and Cantonese. It's live now, free to try for 11 days.

Free for members

Chat with this breakdown — free.

Sign in and you get 23 free chat messages on us — ask for the hook, quote a framework, find the exact transcript moment, generate a markdown action plan. Bring your own key when you want unlimited.

Create a free account →
Chapters

Where the time goes.

00:00 – 00:36

01 · Cold open: what's new in v4

ElevenLabs states the pitch directly: v4 is the most expressive TTS model yet, built on a new architecture that reads tone, pacing, emotion, character, and context, with a low-latency Turbo sibling alongside it.

00:37 – 01:06

02 · Two models: v4 vs v4 Turbo

Both models share the same research; the only difference is latency. v4 targets produced content (narration, character work, dubbing) across ElevenCreative, ElevenAgents, and the API. Turbo targets real-time voice agents at about 100ms median latency.

01:07 – 02:12

03 · Script-aware delivery and audio tags

v4 reads a script the way a voice actor would, tracking who's speaking and what just happened. A scripted two-person exchange demonstrates overlap and timing without relying on audio tags, then the video shows how stacked audio tags direct delivery on top of that.

02:13 – 03:22

04 · Voice cloning stability and context stitching

Instant and Professional Voice Clones work identically across v4 and Turbo, but v4 holds speaker identity far more reliably across regenerations and long-form work. Existing PVCs need to be manually retrained on v4. Context stitching then keeps a regenerated line consistent with the rest of its audiobook chapter.

03:23 – 04:26

05 · Multilingual gains, Turbo for agents, and how to try it

v4 covers 90+ languages with the biggest gains in Japanese, Brazilian Portuguese, Mandarin, and Cantonese, plus better IPA support. The video closes by pointing agent builders back to Turbo for its bidirectional streaming and real-time latency, then directs viewers to sign up and try v4 free for 11 days.

Atomic Insights

Lines worth screenshotting.

  • Eleven v4 and Eleven v4 Turbo share the same underlying research and differ only in latency, not in expressive range.
  • Turbo runs at roughly 100ms median latency, fast enough for live voice agents while keeping v4's full expressiveness.
  • Eleven v4 infers tone and pacing from punctuation and sentence length alone, so audio tags only need to override what the text already implies.
  • Multi-speaker dialogue in v4 is generated with awareness of the full exchange, not as isolated lines stitched together afterward, so overlap and timing feel like a real conversation.
  • Voice identity in v4 holds stable across regenerations and long-form work, a known weak point in earlier ElevenLabs models.
  • A Professional Voice Clone trained on an older model has to be manually retrained on v4 before it gets the new model's stability gains.
  • Context stitching means regenerating one line inside an audiobook chapter now accounts for the lines before and after it, not just that line in isolation.
  • The on-camera host in this video is itself an Eleven v4 voice clone driving an AI avatar built inside ElevenCreative, making the demo its own proof.
Takeaway

What actually changed under the hood in v4

WHAT TO LEARN

Eleven v4's gains come from reading a script for context before generating it, not from adding more manual control, and that shift is worth understanding even outside ElevenLabs.

  • Context-aware generation beats manual tagging: a model that infers tone from punctuation and prior lines needs less hand-holding than one that requires explicit direction for everything.
  • Multi-speaker dialogue sounds real when it's generated as a full exchange rather than individual lines assembled after the fact, since timing and overlap depend on both sides.
  • Voice consistency over long content is a harder problem than a single good sample. It shows up specifically in regenerations, dialogue turns, and full-length projects, not in short clips.
  • Splitting a model into a produced tier and a real-time tier is a practical way to serve two very different latency requirements without maintaining two separate research efforts.
  • Migrating between model versions isn't automatic. Existing voice clones needed manual retraining to inherit v4's stability gains, which is a cost worth planning for before switching.
  • Regenerating one piece of a larger work, a line or a paragraph, needs awareness of what surrounds it or the seam shows. That's a general lesson for any AI content workflow, not just audio.
  • Time-boxed free access, 11 days here, is a low-friction way to get people to actually test a new model instead of just reading about it.
Glossary

Terms worth knowing.

Audio tags
Bracketed delivery directions such as [whispers] or [laughs] inserted into a script to steer how a line is spoken.
Instant Voice Clone
A voice clone built from about 10 seconds of clean audio with immediate setup, used for quick, lower-fidelity results.
Professional Voice Clone (PVC)
ElevenLabs' higher-fidelity voice cloning option, most accurate when cloning a voice in its own native language.
Context stitching
The model's ability to account for the surrounding lines or chapter when regenerating a single line of narration.
Bidirectional input streaming
Sending text to the model as it's generated, for example by an LLM, while audio starts streaming back before the full text arrives.
IPA (International Phonetic Alphabet)
A phonetic notation used to specify exactly how a word should be pronounced, useful for names and tricky terms.
Resources

Things they pointed at.

Quotables

Lines you could clip.

00:08
“Don't tell anyone, but we're breaking into the world's biggest candy store.”
playful cold-open line delivered as a live audio-tag demo→ TikTok hook↗ Tweet quote
01:07
“11v4 was designed to read a script the way a voice actor reads one.”
the thesis line, quotable on its own with zero setup→ newsletter pull-quote↗ Tweet quote
03:25
“If you hadn't noticed yet, this entire video is my 11v4 voice, made with an avatar entirely inside 11creative.”
the reveal that the host itself is the product demo→ IG reel cold open↗ Tweet quote
The Script

Word for word.

Read-along

Don't just watch it. Burn it in.

See every word as it's spoken — crank it to 2× and still catch all of it. The same dual-channel trick behind Amazon's Kindle + Audible.

metaphor
11v4 is here, and it's the most expressive text -to -speech model. With better voice control, long -form consistency, more reliable audio tag following. Don't tell anyone.
But we're breaking into the world's biggest candy store. And improved multilingual speech. It's built on an entirely new architecture, so it interprets tone, pacing, emotion, character and context to generate speech that feels performed.
Alongside it, we've also released 11v4 Turbo, a low latency version of the model. Here's everything you need to know about 11v4 and how to use it.
There are two models, and the same research sits behind both. The difference is latency. 11v4 is for produced content, narration, character work, voiceover, and dubbing.
You'll find it in 11 Creative, in 11 Agents, and through the API. 11v4 Turbo is for real -time, voice agents, live assistants, and interactive experiences. It runs at around 100 milliseconds median latency with high audio quality and the same full expressive range as V4.
You'll see Turbo in 11 Creative 2, but it was built for real time for 11 agents and the API. If you're making produced content, Use 11v4.
11v4 was designed to read a script the way a voice actor reads one. It knows who is speaking, what just happened, and how the lines should land while keeping the speaker sounding like themselves. Speakers respond to the context of the conversation, so lines aren't generated in isolation and assembled afterwards.
Multi -speaker exchanges get the overlap, timing, and reactive energy of a real conversation. And even without audio tags, the delivery follows what the script implies. I finished the presentation.
All that's left is adding the latest numbers. Wait. Wasn't Siren supposed to send you the figures?
Oh, right. I was just about to send them to you after lunch.
And speaking of audio tags, they've had a big upgrade too. Tags aren't new, but V4 follows tag sequences far more reliably than V3, including several stacked in a single line. The rule of thumb is to write the script as prose first, then direct it.
V4 already infers a lot from punctuation, sentence length, and context, so tags are there to shape the delivery when the text alone doesn't give enough direction. I thought you were coming. Honestly, I don't know what happened.
Wait, what?
Both cloning options work identically across V4 and V4 Turbo. An instant voice clone captures a voice with high fidelity from around 10 seconds of clean audio and setup is immediate. A professional voice clone is our best option and it's at its best when cloning a voice in its native language.
What's changed in V4 is stability. Speaker identity holds over an infinite texturation and it's far more reliable across regenerations, dialogue turns and long form work. In the past, the default for consistency was multilingual V2 over V3.
Now you can get high quality express... generations from your clone, but first you need to retrain your PVC on the new model. To do this, go to voice, click my voices, hover over the settings cog for your PVC, then click the 11v4 model to start fine -tuning.
And to show you how good this model is, if you hadn't noticed yet, this entire video is my 11v4 voice, made with an avatar entirely inside 11creative.
Context stitching has also improved a lot which matters for audiobooks. On longer projects you often need to regenerate a single line within a chapter and the whole chapter shapes how that line should be delivered. 11v4 is much better at taking the preceding and following context into account so the line fits your project no matter how many times you regenerate it.
V4 covers over 90 languages, with the strongest gains over V3 in Japanese, Brazilian Portuguese, Mandarin and Cantonese. Those gains go beyond pronunciation into rhythm, emotion and delivery. Support for the International Phonetic Alphabet is also much better, so complicated words get pronounced the way you intend.
So using my professional voice clone, Back to Turbo for anything real -time. Agent builders have always traded expressiveness against latency and Turbo closes most of that gap around 100 milliseconds with the expressive range of 11v4. It also supports bidirectional input streaming so text can be pushed as your LLM generates it and audio starts coming back early.
That gets you an agent whose confirmations, escalations and holds stop reading identically to one another and a single brand voice that stays recognizably itself across a long call. 11v4 is live now. Click the first link in the description, create a free account and pick 11v4.
model pick up and for the next 11 days you can use 11v4 for free thanks for watching
The Hook

The bait, then the rug-pull.

ElevenLabs opens with a direct claim: this is its most expressive text to speech model yet, and the on-camera host saying it is an AI avatar running on the new voice itself.

Frameworks

Named ideas worth stealing.

00:37concept

v4 vs v4 Turbo

  1. Eleven v4: produced content (narration, character work, voiceover, dubbing)
  2. Eleven v4 Turbo: real-time (voice agents, live assistants, interactive experiences)

Same underlying research split into two latency tiers so the model fits either a production pipeline or a live conversation.

Steal fordeciding which TTS tier to wire into a new voice product
CTA Breakdown

How they asked for the click.

VERBAL ASK
04:19product
“Click the first link in the description, create a free account and pick 11v4. For the next 11 days you can use 11v4 for free.”

Direct, time-boxed free-trial CTA delivered by the same AI avatar host, paired with an on-screen 'Eleven days' title card and a description link.

Storyboard

Visual structure at a glance.

open
hookopen00:00
dialogue demo
valuedialogue demo01:07
voice cloning
valuevoice cloning03:33
CTA: free for 11 days
ctaCTA: free for 11 days04:19
Frame Gallery

Visual moments.

One-click upgrade to your Google

Get more breakdowns in your search results

Add Modern Creator as a preferred source and Google shows you more of our breakdowns in Search, Top Stories, and AI Overviews. It only changes what you see, and you can undo it in your Google settings anytime.

Add to Preferred SourcesOpens your Google source preferences with us pre-loaded. Tick the box and you're done.
Watch next

More from this channel + related breakdowns.