Eleven v4 Audio Tags: Write Your Voice Script Like a Director
The most important control in ElevenLabs' new voice model is a pair of square brackets.
Eleven v4, released on September 28, takes a script with notes written straight into it. [whispers]. [nervous laugh]. [long pause]. [Gong sounds]. The model reads the note, then performs the line that follows. On its launch page, ElevenLabs shows a quiz-show scene with two voices, a host and a nervous contestant, plus a gong and crowd applause, all written as one block of text and generated as one piece of audio.
Here is the part the launch coverage skips. ElevenLabs' own documentation warns that because v4 was trained to produce both vocal delivery and sound effects, "a tag can occasionally be interpreted as a request for a sound effect rather than a delivery instruction (or vice versa)." Write [thunder] meaning a booming voice and you might get actual thunder.
So the skill that decides whether v4 sounds good is no longer picking a voice or nudging a slider. It is writing direction a machine cannot misread. That is a writer's skill, and a director's. It is worth learning properly.
What changed, in plain terms
Older text-to-speech read your words and guessed how they should sound. You had a few settings for how stable or varied the voice should be, and you rewrote sentences, added commas or swapped words until the reading came out right.
Eleven v3, the previous model, already understood these bracketed audio tags. Eleven v4 keeps the idea and makes it dependable. According to ElevenLabs, v4 "follows tag sequences more reliably than v3, sound effects included," and it reads a whole scene at once, so the second speaker can react to what the first one just said instead of sounding like a separate recording pasted in afterward.
Three other changes matter if you make long pieces like audiobooks, podcast series or game dialogue.
The voice holds steady when you redo a line. The v4 page says you can regenerate a line "once or fifty times and it's still the same person speaking." Anyone who has patched chapter nine of an audiobook and heard a subtly different narrator knows why that matters.
Long scripts stitch together cleanly. ElevenLabs says its context stitching keeps pacing and delivery steady across a long script, so the whole thing sounds like a single take.
Professional Voice Clones are back. These are the high-quality clones you train on a longer recording of a real voice. v3 did not support them. v4 does. Quick clones now need about ten seconds of audio.
It also covers more than 90 languages, and a cloned voice can speak another language with a native accent while still sounding like the same person.
ElevenLabs calls v4 its most emotive model and cites a #1 ranking on Artificial Analysis's voice leaderboard, plus a blind listening test where people preferred v4 about 75% of the time against four named competitors. That test is the company's own, so treat it as a claim to check with your own ears, not a settled fact.
Why the brackets change your job
If you have ever directed a voice actor, you know the useful notes are short and concrete. "Warmer." "Slower on the last line." "You just got bad news." Nobody says "be emotive." Actors need something they can play.
v4 works the same way, and ElevenLabs' docs say so directly. Their advice on avoiding the sound-effect mix-up is to write tags that "clearly describe the voice quality you want (e.g. [low, gravelly voice] rather than something that could be read as a sound cue)."
That one sentence is the whole craft in miniature. Compare these two versions of the same note:
[rough] I told you not to come back here.
[low, gravelly voice, barely holding his temper] I told you not to come back here.
The first could mean a rough sound, a rough edit, rough weather. The second can only mean one thing. It is longer, and it works better.
The docs make a second point that is easy to miss. A voice performs more reliably when the delivery you ask for is something that voice has already done. v4 can follow [whispering] on a voice that never whispered in its source recording, but ElevenLabs says "the result may not be optimal." In casting terms: pick the actor who can already play the part, then direct them. Do not ask a calm meditation voice to scream and blame the model when it strains.
Put this into practice
You can try all of this on the free plan, which gives you 10,000 credits a month (about ten minutes of speech, per the pricing table) for personal use. You need a paid plan, from $6 a month, to use the audio commercially. ElevenLabs is also running a launch promotion on Creator plans and up that stretches your credits for v4 until mid-October, which makes the next couple of weeks a cheap time to experiment.
Here is a way to build a short scene that works.
1. Write the scene clean first. Just the words each character says, as if for a table read. Get the dialogue right before you touch delivery. Tags cannot rescue a flat line.
2. Add one direction per line, in front of it. Keep each one to a few plain words describing the voice, not the action. [relieved, laughing a little] works. [walks across the room] does not, because there is nothing to hear.
3. Put sound effects on their own and name the sound. [door slams], [light rain], [phone buzzing]. Write them as a noise, not a feeling, so the model does not confuse them with delivery.
4. Use punctuation for pacing. v4 does not support the old pause codes some people used with earlier models. ElevenLabs' docs say to control pauses with audio tags, ellipses and the way you structure the text, and to use capital letters for emphasis. Their example, "It was a VERY long day [sigh] … nobody listens anymore," reads very differently from the same words typed plainly.
5. Set up two speakers in ElevenCreative. Choose Eleven v4 as the model, add a second speaker, and give each line to the right voice. Here is a small scene in the style of ElevenLabs' own examples:
Speaker 1: [tired, trying to sound cheerful] Morning. Coffee's on.
Speaker 2: [flat, not looking up] It's four in the afternoon.
Speaker 1: [long pause] [small, embarrassed laugh] Then it's very strong coffee.
[phone buzzing]
Speaker 2: [suddenly alert, quiet] Don't answer that.
6. Regenerate single lines, not the whole scene. When line three lands wrong, redo line three. This is where v4's steadier voices pay off.
7. Fix names with pronunciation marks. If a character's name comes out wrong, v4 accepts International Phonetic Alphabet spelling between forward slashes, and the docs say this is more consistent than in older models. For names you use a lot, ElevenLabs' pronunciation dictionary saves you from doing it every time.
If you are stuck, the text-to-speech screen has an "Enhance" button that adds tags for you using an AI writing assistant. ElevenLabs publishes the instructions it gives that assistant, and they are a useful lesson in themselves: never change the words, only add tags that describe something you can hear, and place each tag right before or after the part it changes. Use Enhance to see what it suggests, then rewrite the tags in your own words. Its guesses are generic. Your direction should not be.
Where it falls short
It still misreads tags sometimes. ElevenLabs says so plainly in its own docs: tags are "not perfect yet," and it describes reliability as an area it is still working on. Budget for retakes. A two-minute scene is not a one-click job.
Word-level timing is still loose. You can ask for a long pause. You cannot ask for a laugh to land exactly 1.4 seconds in. ElevenLabs says a fuller "Director's Mode" is in development, which tells you the precise-timing version does not exist yet.
Sound effects arrive baked in. Nothing in the launch material says you can export the gong or the rain as a separate track. If your mixer wants the effects on their own channel for the final edit, generate them separately.
It is cloud only and paid for real work. Every generation runs on ElevenLabs' servers and spends credits. The launch post gives no separate price for v4, so run a short script first and watch how many credits it takes before you commit an entire audiobook.
Some tags are unreliable across voices. The docs flag accent tags and a handful of playful ones as "experimental" and inconsistent. Test your exact tags with your exact voice before a deadline.
Ten seconds is enough to clone someone. That is great for your own voice and a real risk for anyone else's. Get consent in writing for any voice you did not record yourself, and tell listeners when a voice is generated.
The script is the session now
The old workflow split voice work in two. A writer handed over words, and someone else, an actor, a director, an engineer, turned them into a performance. Eleven v4 folds the second half back into the first. The script now carries the notes a director would have given in the booth, and the model plays them.
That favours people who already think in scenes: podcasters who write their own intros, game writers, audiobook producers, anyone who has sat in a session and said "one more, a bit warmer." It does not favour people hoping for magic from a bare paragraph.
Try it on something small this week. Take a short scene you have already written, add one clear direction per line, and generate it twice: once with vague tags, once with the specific kind. Listen to both back to back. That comparison will teach you more about this model than any benchmark, and you will come away knowing whether it has earned a place in your next project.
Medium metadata
Title: Eleven v4 Audio Tags: Write Your Voice Script Like a Director
Subtitle: ElevenLabs' new voice model performs the notes you write in square brackets. Its own docs admit a vague note can come out as a sound effect, so clear direction is now the skill.
Tags: AI Voice, Text To Speech, Podcasting, Audiobooks, ElevenLabs
Estimated read time: 7 minutes