Audio Sliders: Change the Mood of an AI Music Cue Without Rerolling It
A free, open set of 35 dials for the ACE-Step music model moves one quality of the same piece (sad to happy, solo to full band) while everything else stays put. Here is what it is good for, how to try it, and where it falls apart.
Every AI music tool has the same blind spot. Ask for "a melancholy jazz piano piece," get something lovely, then ask for "the same thing, a bit happier," and you get a different song. New chords, new melody, new tempo. The words changed, so the whole piece changed with them.
That is fine if you want a song. It is useless if you are an editor who needs three versions of one cue: the quiet one under the interview, the one that lifts as the montage starts, and the one that hits at the reveal. Those three have to sound like the same piece of music, or the cut falls apart.
A small open-source project posted on October 2 takes a direct swing at this. It is called Audio Sliders, and the idea fits in one sentence from its README: drag a slider and "the same piece, same prompt and same seed, moves along one axis: sad to happy, solo to full ensemble, stiff to groovy."
My position, after reading the code, the model card and the author's own caveats: this is the right idea for AI music, a far better one than chasing longer prompts, and it is a research release you should treat as a sketchbook, not a finished instrument.
What a slider actually is
Start with the problem it solves. When you generate music with AI, the seed is the random starting point. Keep the prompt and the seed the same and you get the same clip back. Change one word in the prompt and the model takes a different path from that starting point, so the result drifts everywhere at once.
A slider is a tiny add-on file that sits on top of the music model and nudges it in one direction while the prompt and seed stay fixed. Push the "mood" slider toward positive and the model leans happier at every moment of the generation. The structure of the piece, set by the seed, mostly survives the trip.
The base model here is ACE-Step 1.5, an MIT-licensed open music generator that makes anything from ten seconds to ten minutes of audio. The sliders were built for its fast "XL turbo" version. The project also includes sliders for Stability AI's Stable Audio Open, which has its own license you must accept on Hugging Face before it will download.
There are 35 sliders in the published set. Twelve were trained from plain-language opposites: mood, ensemble, groove, harmony, melody, tension, brightness, density, energy, tempo, electronic and vintage. The rest are stranger and more interesting.
The sliders nobody named
The second group was not designed by anyone. The author analyzed more than 14,985 real recordings, looked for the main ways those tracks differ from each other, and trained sliders along those directions. Some of the results map to ideas musicians already use, like arousal (how keyed-up a piece feels) and valence (how positive it feels). Others are concrete contrasts such as strings to synth, or jazz to electronic.
The project's demo page puts it nicely: these are axes "found in 14,985 real recordings and nobody chose them."
Why does that matter to you? Because the hard part of asking for music in words is that you rarely know the right words. "Warmer" means something different to every composer. A slider learned from real music sidesteps the vocabulary problem. You do not describe the change. You hear it and decide how far to go.
Why this beats a better prompt
The obvious objection: just write a more careful prompt. Say "same melody, happier." Most current models do not hold onto a melody across prompts that way, so the request turns into a reroll with extra words.
Sliders also stack. The author measured that "effects roughly add, with interaction terms a fifth to a third of the main effects." In plain terms, pushing energy and ensemble together gives you something close to the sum of each push alone, with some overlap. That is predictable enough to work with. You can plan a build: start sparse, add players, then add energy.
And it is fast. The author's own timing is 0.4 seconds per ten-second clip on ACE-Step, and the slider adds no extra generation passes. That number comes from the author's setup, not an independent test, but the point holds: auditioning ten positions of a slider costs you seconds, not an afternoon.
The license is MIT, which means you can use the sliders in paid work. Check the license of whichever base model you pair them with too; ACE-Step is also MIT.
Put this into practice
There are two ways in: the browser demo, and running it yourself.
Try the demo first. The live demo page lets you drag 18 of the sliders across five example prompts, with every clip matched for loudness so you are hearing the change and not a volume jump. The page says that when it is served from the author's GPU server, you can also type your own prompt. Treat that part as a bonus that may not be up when you visit.
Spend ten minutes there with headphones before installing anything. Listen for two things: whether the piece still feels like the same piece at the extremes, and which sliders change more than their name suggests.
Run it on your own machine. You need an NVIDIA graphics card with real memory. ACE-Step's README says the XL turbo model needs at least 12 GB with offloading and compression turned on, or 20 GB without. Then:
- Install the package:
pip install "audiosliders[model,demo] @ git+https://github.com/takakhoo/audio-diffusion-control" audiobox_aesthetics - Start the slider app:
python -m audiosliders.server --sliders hf:ace-step-1.5-xl-turbo/text --backbone ace-turbo - Open the page it serves, write one instrumental prompt, and lock the seed.
Then build an editor's set of cues. Here is a recipe that turns the tool into something you can drop on a timeline:
- Write a short instrumental prompt that describes genre and instruments, not emotion. "Upright piano, brushed drums, double bass, small club." Leave the feeling to the sliders.
- Render the neutral version with every slider at zero. That is your bed.
- Render energy at -1 for the quiet version under dialogue.
- Render energy at +1 with ensemble pushed up for the peak.
- Line all three up in your editor. Because they share a seed, they tend to share a backbone, so crossfades between them land far more naturally than between three unrelated generations.
Train your own slider if you have a specific contrast in mind. The repo includes a training command, and the author says each slider takes "about 20 minutes" on one GPU. You can train from a pair of opposite prompts, or from two sets of clips with no text at all. If you have a library of your own "before" and "after" sounds, that second option is the interesting one. It is also the part with the least documentation for non-programmers, so budget an evening.
Where it breaks
The author is unusually frank, and you should take every caveat seriously.
Nobody has done a proper listening test yet. The model card says it plainly: "No listening study has been run yet." Quality was judged by two automated scoring models, not human panels. Your ears are the test.
The sliders leak. Move mood or melody and brightness changes too. The README says so directly. If you need surgical control over one thing, you will not get it here. You get a strong push in one direction with side effects.
Each slider has a sweet spot. Some are only meant to be used between -1 and +1. Others hold up to +2. Push past the range and the output stops sounding like music. The author lists usable spans per slider in the repo; read them before you crank anything.
Short instrumentals only. Testing was on ten-second clips, with some text sliders checked at thirty seconds. No vocals. If you need a three-minute song with a singer, this is not your tool this week.
Generated music is narrower than real music. The author measured that generated clips cover only part of the range real recordings do on each axis. The slider moves you along a line, but the line is shorter than you might hope. Do not expect the "full ensemble" end to sound like an orchestra.
It is a research project. One star on GitHub at the time of writing, a paper draft aimed at a 2027 conference, and a demo running on one person's server. It could change, stall or vanish. Download the weights you like.
The shape of the next instrument
The interesting thing here is not any single slider. It is the change in what you are doing when you use it.
With a prompt box, you are a client writing a brief and hoping the vendor gets it. With a slider, you are closer to a mixing engineer riding a fader: you hear the result, decide it needs a little more, and move your hand. That is how musicians already think. Nobody re-records a song because the chorus should feel brighter.
Audio Sliders is rough, short-form and unproven. It is also the clearest sign yet of where AI music tools need to go if they want working composers and editors to trust them. Spend ten minutes with the demo, pick the one slider that changes the music the way you would have, and notice how different that feels from typing "happier" and starting over.
Medium metadata
Title: Audio Sliders: Change the Mood of an AI Music Cue Without Rerolling It
Subtitle: A free, open set of 35 dials for the ACE-Step music model moves one quality of the same piece while everything else stays put. What it is good for, how to try it, and where it breaks.
Tags: AI Music, Music Production, Generative AI, Video Editing, Open Source
Estimated read time: 8 minutes