LAION's Humaneness Voice Add-ons: 40 Emotions and 500 Voices, Each a One-Megabyte File
An open English and German voice model now comes with 557 tiny attachments: moods like Affection and Bitterness, delivery traits, and 500 ready-made voices. Here is which part is worth your time, how to hear it, and where the paperwork gets thin.
A voice used to be a big thing to move around. To get a new one out of an AI speech model, you recorded a reference clip or trained a whole custom version, and either way you were handling minutes of audio and gigabytes of files.
On October 3, LAION posted 557 voice add-ons for its open speech model, and each one weighs about 1.2 megabytes. That is smaller than most photos on your phone. Forty of them are emotions. Seventeen are delivery traits. Five hundred are individual voices.
So a mood is now a file you attach, the way you would drop a reverb plug-in on a track. That changes how you can direct a synthetic read. It also raises a question the files do not answer: whose voices are those 500?
What LAION actually released
There are two pieces.
The base model is Humaneness Voice Small, posted September 29. It is a voice-acting model for English and German. You direct it with two lines: a short caption describing how the line should sound, then the exact words in quotes. LAION's own example reads like a note to an actor:
CAPTION: Warm, softly amused.
TRANSCRIPT: "I cannot believe we made it."
The add-ons are Humaneness Voice Small Rank-1 LoRAs, posted October 3. Each one nudges the base model toward one thing. The 40 emotions run from Affection, Awe and Bitterness to Relief, Teasing and Triumph. The 17 traits cover things like high or low arousal, tension, emphasis and an ASMR delivery. The 500 voices carry codenames, not names.
Everything ships under CC BY 4.0. In plain terms, you can use it in paid work, including client jobs, as long as you credit LAION. For an open voice model, that is about as permissive as it gets.
Why an emotion as a file beats an emotion as a word
Anyone who has directed an AI voice knows the problem with typing "sad." The model hears the word, guesses, and gives you its average idea of sadness. Change the word to "grieving" and you might get something different, or the same thing with a slower pace. You are negotiating with a vocabulary you cannot see.
An emotion add-on works differently. LAION trained each one separately, on examples of that single emotion. When you attach the Bitterness file, you are not asking the model to interpret a word. You are tilting it toward a delivery it practiced.
The practical win shows up when you keep everything else fixed. Same line, same caption, same random seed (the number that makes a run repeatable), and only the add-on changes. Now you have a warm take, a bitter take and a tired take of one line, and they are comparable. Each sounds like the same performance pushed in one direction rather than three unrelated attempts.
That is how a voice director already works. You do not recast the actor when you want the line warmer. You give one note and run it again.
There is a rule you have to respect. The card says to attach to a fresh model and "do not stack multiple adapters unless you have independently evaluated that mixture." In practice, that means one mood per take unless you want to run your own tests. If you want "tired but hopeful," you are back to writing it in the caption.
Put this into practice
Start by listening, not installing.
1. Hear the base model first. LAION published a listening page with 18,720 takes from the base model across its ten training stages. The add-ons are not on it, but this tells you in ten minutes whether the underlying voice quality fits your project. The base card calls stage three "a reasonable provisional default" and suggests comparing it with stage ten before choosing a production voice. If the base model does not sound usable to you, no add-on will rescue it.
2. Decide which part you need. For most creators, that is the 40 emotions. They solve a real directing problem and carry no question about whose voice you are using, because they shape delivery rather than identity. The 17 traits are worth a look for specific jobs: the ASMR file for a sleep-story narrator, the high-tension file for a thriller trailer.
3. Run it locally. There is no hosted generator and no ComfyUI node (the drag-and-drop graph tool many creators use) yet, so this step needs someone comfortable with Python. Download the base model and the add-on repo into side-by-side folders. The card's script, infer_rank1.py, takes an add-on name, your caption and transcript, a language, a seed and an output filename. The example call uses --adapter-id emotion/Affection with the line "I am so glad you are here."
4. Build a take sheet. Pick one line from your script. Run it with no add-on, then with three or four emotions that could plausibly fit, keeping the seed fixed. Name the files by emotion. Drop them into your editor next to the picture and choose with your ears, the same way you would pick between takes from a session.
5. Write captions like stage directions. The base card recommends a short caption and the exact transcript in quotes. Short and specific beats long and poetic. "Warm, softly amused" gave LAION its example. "A voice dripping with the bittersweet memory of summers lost" gives the model more words to misread.
6. Credit LAION. CC BY 4.0 requires attribution. Put it in the end credits or the episode notes and you are covered for the license itself.
One more tip from the base card: if you do give the model a reference clip of a voice, keep it clean and under three seconds, and never use the recording you are trying to match as its own reference.
Where it breaks
LAION is unusually frank about this release, and the frankness is the best guide.
Nobody has proven the add-ons work at scale. The card's own words: "No broad efficacy claim yet." The voice-likeness numbers come from a two-voice pilot. The emotion and trait files have training logs, not listening tests. LAION also warns that training loss, its internal measure of progress, "is not an audio-quality or controllability score," which is a polite way of saying the numbers might not match what you hear.
The base model is shaky. The base card calls it "a research release with known shakiness, prompt sensitivity and imperfect word/timing/burst control." In practice that means words can slip, pauses can land in the wrong place, and a laugh or gasp you asked for may not arrive. Budget for several takes per line, and plan to fix timing in the edit.
Two languages. English and German only. If your project is in Spanish or Japanese, this is not your tool.
The 500 voices have no paper trail you can show a client. The add-ons carry codenames grouped into families such as anime, emolia, mediathek and refvoice, plus a numbered k series. The card does not say which recordings or which speakers they came from, and the index file lists training statistics, not sources. LAION's guidance is to use voice identities "only with appropriate rights and consent" and adds that a voice-matching score "does not verify speaker consent." That puts the burden on you, and you have no way to carry it, because you cannot ask permission from a codename.
My position is simple. Use the emotions and traits freely. Treat the 500 voice files as a research exhibit, not a casting catalog. If a client asks who the voice belongs to, "a file called mediathek_0047" is not an answer you want to give.
One mood at a time. Stacking is untested unless you test it yourself, so blended emotions mostly go back into the caption, where the old guessing problem returns.
The direction this points
The interesting part is not this one model. It is the shape of the tool.
When a mood is a one-megabyte file, it becomes something people can make, trade and swap. A studio could train its own "dry deadpan" for a recurring narrator and hand it to every editor on the show. A game team could keep a folder of emotions for one character and attach the right file per line of dialogue. That is closer to how sound designers already work with presets than anything prompt-based voice tools have offered.
Humaneness Voice Small is rough, bilingual and openly unfinished. Spend ten minutes on the listening page. If the base voice fits your work, pick one line and run it through three emotions. Then decide whether a mood you can attach beats a mood you have to describe.
Medium metadata
Title: LAION's Humaneness Voice Add-ons: 40 Emotions and 500 Voices, Each a One-Megabyte File
Subtitle: An open English and German voice model now comes with 557 tiny attachments for mood, delivery and identity. Which part is worth using, how to hear it first, and where the consent paperwork runs out.
Tags: AI Voice, Text To Speech, Voice Acting, Open Source, Podcasting
Estimated read time: 8 minutes