Creative AI Briefing: Wednesday, September 30, 2026
Type "[whispered in a British accent] Beethoven. It's Beethoven, isn't it? [nervous laugh] Final answer." into a script, put a quiz host's line with "[Gong sounds]" above it, and press generate: you get a two-person scene with a sound effect in it, in one take. That is Eleven v4, which ElevenLabs shipped on Monday. The same week, LAION released an open voice-acting model that reads a script with a stage direction in brackets before every line. The pattern is plain. Voice generation stopped being "read this text aloud" and became "perform this scene the way I describe it," which means the skill that matters now is the one a radio director has, not the one a sound engineer has.
New models
Eleven v4 reads a script the way a voice actor would. ElevenLabs published Eleven v4 and v4 Turbo on September 28 (the post shows a September 29 update). You write direction into the text itself: inline tags like [laughs], [said angrily in French accent], [light rain] or [phone buzzing], or a plain-language note on how a line should land. The company says v4 follows these tags more accurately than v3, sound effects included, and that speakers in a multi-voice scene respond to what was just said instead of sounding like separately recorded lines. For long projects, three changes matter: the v4 page says regenerating a line "fifty times" keeps the same voice, long scripts stitch together without audible seams, and Professional Voice Clones (the high-fidelity, trained kind) are supported again after v3 dropped them. Instant clones now need about ten seconds of audio. It covers 90+ languages, and a cloned voice speaks other languages with a native accent. The "#1" ranking comes from Artificial Analysis's voice leaderboard, and the ~75% listener preference is ElevenLabs' own blind test against four named competitors (Cartesia, Inworld and two Gemini TTS models), so treat both as the vendor's framing. Price and access: available now in ElevenCreative, ElevenAgents and the API as eleven_v4. The free plan gives 10,000 credits a month but is for personal use; paid plans start at $6 a month. The catch: it is closed and cloud-only, and the post gives no per-character price change for v4, so check your credit burn on a short script before you commit a whole audiobook.
Humaneness Voice Small is an open voice-acting model you can direct line by line. laion/Humaneness-Voice-Small was created on Hugging Face September 29 at 11:39 UTC under CC BY 4.0, which allows commercial use of the weights with credit to LAION; the Qwen3 backbone and the separately downloaded MOSS audio codec keep their own licences. It speaks English and German. Its most interesting input format is a script: a GENERAL: line for the overall character ("frightened, breath-held, intimate, low-register"), then a SCRIPT: with a bracketed direction and optional duration before every sentence. The honesty on the card is unusual. LAION ships ten checkpoints from one training run, suggests S3 as a starting point, then says its own benchmark "does not establish S3 as significantly better than S10." The scripted format is the one that gives you per-line direction, and it is also the least reliable for getting the words right: the card measures about a third of words wrong in that mode, against about 4 to 5 percent on the two best caption-plus-transcript formats (the card warns these are not clean format rankings). Where it runs: research code on an NVIDIA card with PyTorch; it is not a one-line install. Free way to hear it: the listening Space holds 18,720 generated takes across all ten checkpoints. Four experimental preference-tuned add-ons followed this morning in laion/Humaneness-Voice-Small-DPO-LoRAs (September 30, 06:10 UTC), with LAION's own warning that better training scores are "not a claim that DPO improved listening quality."
Video
Generate a 360-degree video you can look around in a headset. rehan-fal/minimax-h3-360-equirect-lora (created September 29, 01:24 UTC) is an add-on for MiniMax H3 that makes the model output the whole sphere around the viewer, unwrapped into one wide frame, with H3's native sound. The author, who posts as rehan-fal, publishes measured geometry: the left and right edges of the frame (which must match, because they meet behind the viewer) correlate at 0.95, against 0.55 for plain H3. On fal's hosted H3 endpoint, a 10-second 4K clip "takes roughly 4–9 min and costs $2.00," per the card. The file sits under the MiniMax Community License. The honest limits are listed: a faint seam can remain behind you on busy scenes, footage skews handheld, and it is mono, not stereo. Full recipe in Workflows below. A sibling VR180 stereo add-on from the same author dates from September 3 and has a free demo Space.
Two research teams made H3 faster. NVIDIA's Efficient-Large-Model group posted the LongLive-Plug collection on September 29 (H3 files created 13:58 UTC, Wan files 16:13 to 16:18 UTC): speed-up add-ons for MiniMax H3, Wan2.1 14B and Wan2.2 5B. The H3 card says to use the few-step file and the companion file separately, not together, and says little else yet. Separately, pdmd2026/pdmd_4NFE_lora (September 29, 08:01 UTC) is a 4-step H3 add-on from a new distillation paper. It is a 1.38 GB file. The card is tagged Apache 2.0, but it modifies H3, so read the H3 licence before you ship with it. Neither card shows a side-by-side quality comparison (the PDMD paper reports the authors' own scores and user study), so test these before you trust them.
Image
Quiet on the image side overnight: the new Hugging Face uploads were mostly Qwen-Image 2.1 and Krea-2 style add-ons. The most useful image news is a week old and easy to miss. Midjourney's September 24 update (read via the Releasebot mirror, because midjourney.com blocks fetches) says inpainting and outpainting with the V8.2 edit model "will change only the pixels that you've selected," so repeated edits no longer degrade the rest of the image. It also fixed visible seams in --tile patterns for V8.1 and V8.2, which matters to anyone making textures or wallpaper repeats.
Audio and music
Beyond the two voice models above, the day's audio uploads were small conversions of existing speech models into faster or lighter formats (ONNX builds of OmniVoice and Qwen3-TTS, among others). Nothing new in music generation landed in the window. Eleven v4's inline sound effects ([crowd applause], [Gong sounds] in ElevenLabs' own demo script) are the closest thing to a music-and-sound story today: v3 could already do this, but ElevenLabs says v4 follows these tags more reliably, which is what makes it practical for a podcast producer to get the ambient bed and the stings in the same pass as the dialogue.
Open and local
The open story today is voice. Humaneness Voice Small is one of the few open models that treats a script with per-line acting notes as a first-class input, and it ships with the evaluation audio so you can judge it before installing anything. On the video side, the community keeps adding speed and new formats to MiniMax H3 rather than new models.
- laion/Humaneness-Voice-Small: direct an English or German voice line by line from a script, CC BY 4.0. Created September 29.
- laion/Humaneness-Voice-Small-DPO-LoRAs: four experimental preference-tuned add-ons for it, with their own listening Space. Created this morning.
- rehan-fal/minimax-h3-360-equirect-lora: 360-degree video with sound from a text prompt, ready for a Quest after one ffmpeg step. Created September 29.
- Efficient-Large-Model/LongLive-Plug-MiniMax-H3-few-step: NVIDIA research speed-up for H3, with Wan siblings in the same collection. Created September 29.
- pdmd2026/pdmd_4NFE_lora: a 4-step H3 add-on from a two-day-old paper. Created September 29.
Creative workflows
1. Make a 360 video and play it in a headset. Files: h3-360-equirect-lora-v1.safetensors on fal's minimax/h3/text-to-video/lora endpoint.
The steps. Send a request with the add-on's direct file URL at strength 0.75, aspect ratio 21:9, and prompt expansion disabled. Start the prompt with the trigger word eqr360, then the layout sentence from the card, then your scene, describing what is ahead, to the sides, behind and overhead. Preview at 768P (about a minute), then rerun the same seed at 2K or 4K. Squash the 21:9 output to exactly 2:1 with the card's ffmpeg line, then tag it as 360 with Google's spatialmedia tool. Copy it to a Quest, DeoVR or YouTube 360.
How it works. The add-on was trained on 181 real 360 clips that passed seam and pole checks, so H3 learns to draw the curved horizon and stretched top and bottom of a sphere laid flat.
Why it is good. A 360 establishing shot with sound for $2 and a few minutes, and the card tells you the one setting that trips people up: pass the direct file URL, because the repo-name form loaded differently and gave a different video on the same seed (tested September 29).
Where it breaks. A faint seam can appear directly behind the viewer on busy scenes. Longer renders can darken in the last few seconds, so keep the first ~7.5 s of a 10 s clip. Mono only, so no depth.
2. Direct a voice line by line with an open model. Files: checkpoints/S3/model_bf16.pt and code/infer.py.
The steps. Start with the caption format, because it gets the words right: a CAPTION: describing the voice, then an exact TRANSCRIPT: in quotes. Run infer.py --stage S3 with a frame budget (--frames, in 80 ms frames). Add --reference-wav with a clean clip under about three seconds for voice identity, never the target recording itself. Switch to the GENERAL:/SCRIPT: format only for lines where the per-sentence direction matters more than word accuracy, and listen to S3 and S10 on the same prompt before choosing.
How it works. The model was trained on every one of these prompt formats, so it treats a bracketed note like "(whispered panic, sharp inhale)" as instructions for delivery, not words to speak.
Why it is good. CC BY 4.0 weights (check the codec and backbone licences too), and 18,720 sample takes to audition first.
Where it breaks. Word errors in script mode, and the card says it has not shown reliable control over exactly when a laugh or gasp lands. It needs a CUDA machine and some patience.
3. Write a two-voice scene with sound effects in one pass. Source: the Eleven v4 page and its demo script.
The steps. In ElevenCreative's text-to-speech, choose Eleven v4, add a second speaker, and write each line with its delivery in brackets in front of it and sound effects as their own tags, as in ElevenLabs' quiz-show example ([warm], [long pause], [Gong sounds], [crowd applause]). Regenerate single lines until each lands.
How it works. v4 reads the whole scene, so a line's delivery can react to the previous one.
Why it is good. A radio-play draft, ad read or game bark set without a booth, a mixer or a separate sound-effects pass.
Where it breaks. You cannot export the effects as separate stems from the same generation as far as the launch material says, so a mixer who wants the gong on its own track will still need a separate pass. Free-plan output is personal use only.
Worth testing
- Eleven v4 in ElevenCreative: free tier, 10,000 credits a month. Tradeoff: personal use only on free, and the "#1" and "75%" claims are the vendor's.
- Humaneness Voice Small listening Space: audition the open model across checkpoints without installing it. Tradeoff: these are the authors' benchmark prompts, not yours.
- H3 360 add-on on fal: $2 per 10-second 4K render, per the card. Tradeoff: paid, and MiniMax Community License terms apply.
- VR180 stereo demo Space: a way to see H3 depth video in front of you without paying. Tradeoff: it runs on Hugging Face's shared free GPUs, so you may hit a daily quota; it covers the front half only, and it is a month-old release.
What actually matters from today's signal
The script is becoming the control surface for audio. Eleven v4 and Humaneness Voice Small arrived from opposite ends of the industry, one closed and polished, one open and candid about its flaws, and both want the same input: words plus stage directions. That changes who is good at this. A writer who can mark up a scene the way a director would ("building", "long pause", "nervous laugh") will now get better results than someone who knows every slider in the settings panel. For a solo podcaster, audiobook producer or game writer, the time goes into the script and the takes, not the edit.
The money split is sharp. Eleven v4 is the finished product, and you rent it with credits (a launch credit promotion runs on Creator plans and up until October 12). LAION's model is free to use commercially with credit, but it gets roughly a third of words wrong in the very mode that makes it interesting. For paid work this month, the closed model wins. The open one is worth your hour because it shows where free tools will be in a year, and because you can hear every one of its failures before you install anything.
The counter-signal is consent. Ten seconds of audio now clones a voice that holds steady across fifty regenerations and ninety languages. LAION's card asks you to disclose generated audio and avoid impersonation. That is the right default for anyone publishing this work, and it will not stay optional for long.
Source access notes: Hugging Face listings fetched through the API sorted by createdAt with cache busting; every Hugging Face date above is createdAt. ElevenLabs blog and v4 page fetched directly. Runway's changelog showed nothing new after September 23 (DaVinci Resolve plugin). fal's blog had nothing after September 17. Midjourney read through the Releasebot mirror (secondary). OpenAI, Google, Adobe, BFL, Stability, Luma, Suno, Comfy blog and civitai were not fetched directly this run; a WebSearch sweep found no creator launches from them in the window and stood in for them. No GitHub star counts are cited today because none of the named releases has a public code repo with stars worth quoting. Speed, quality and preference figures for Eleven v4, the 360 add-on and Humaneness Voice Small are the authors' own. The adversarial fact-check pass ran (Sonnet subagent). It confirmed every date, price, file size and figure. A second check could not confirm the author's fal affiliation from the model page, so the briefing now names the username only. It caught an overstated caption-mode word-error figure for Humaneness, an unsupported "first" claim, a licence line that ignored the model's codec and backbone licences, an implication that inline sound effects were new in v4 (v3 had them; v4 is more reliable), a missing note on the PDMD paper's own scores, and a missing credit promotion and GPU-quota caveat. All are fixed above.