AI Storyboard From One Frame: The Free Next-Scene Add-On That Treats Your Prompt as a Cut
Most AI image editors are built to change the picture you hand them. Ask for "the camera pulls back" and they tend to give you the same frame, slightly zoomed out, with the same light and the same composition, because keeping your picture intact is what they were trained to do.
A small free add-on released on October 7 was trained to do the opposite. Give it a frame and a sentence starting with Next Scene:, and it makes the shot that comes after. Not a variation. The next shot.
That sounds like a minor distinction. It changes who the tool is for. An editor that preserves your image is a retouching tool. A model that moves to the next shot is a storyboard artist, and it rewards a very different kind of writing.
What it makes
The add-on is called Qwen-Image-2.1 Next-Scene. It plugs into Qwen-Image 2.1, the big open image model Alibaba's Qwen team released in September, and it was posted by akhaliq, who retrained an earlier version by lovis93 on the newer model.
You give it one frame and a direction. The card's own examples read like lines from a shot list:
- "Next Scene: The camera tracks forward and tilts down as rain begins to streak the lens."
- "Next Scene: Cut to a low-angle shot; sunlight breaks through the clouds behind her."
- "Next Scene: The camera pans right, revealing the fleet massing behind the ridge."
Out comes a new frame that carries the world of the first one (the place, the palette, the weather, the character) into a new angle. Feed that frame back in with the next direction, and the next, and you have a board. Five sentences can give a director five angles on one location before anyone books a scout.
You can try it right now in a hosted demo on Hugging Face, the main sharing site for open AI models. No install, no graphics card, though shared demos like this one can make you wait in a queue.
Why it moves instead of edits
The interesting part is how it learned. The training set is 1,144 pairs of consecutive frames pulled from two of Blender's open short films, released under the CC-BY license: Sintel, which is animated, and Tears of Steel, which is live action with effects. Each pair is "this frame" and "the frame that came next in the edit," with a caption describing the change.
So the model never learned "make this picture different." It learned what tends to happen across a cut in a real film. Wide shots give way to closer ones. A character's back becomes their face. Light shifts when the camera turns toward the sun. That is film grammar, and it is the thing general image models are weakest at, because they learned from single photographs that never had a "before" or "after."
This is also why it struggles with some things on purpose. The card says plainly that "static portraits are out of scope by design." A headshot has no next shot. There is nothing for the camera to discover.
The skill it asks for is writing shot directions
Here is my position: the add-on is less a picture tool than a writing tool, and the people who will get the most from it are the ones who already write good shot lists.
Look at those example prompts again. Every one leads with what the camera does (tracks, cuts low, pans right), then says what the move reveals. None of them describe the whole scene from scratch. The scene is already in the frame. Your sentence only needs to carry the change.
That is how a first assistant director or a storyboard artist writes. It is the opposite of how most people prompt image models, piling on adjectives about the subject and the style. Those adjectives are wasted here, because the frame already settled them. A vague direction like "make it more dramatic" gives the model nothing to cut to.
A few habits that follow from this:
- Name the camera first. Pull back, push in, pan, tilt, cut to a low angle, cut to an over-the-shoulder. The camera move is the instruction; everything else is detail.
- Say what the move reveals. "Revealing the fleet behind the ridge" gives the new frame a reason to exist. A pull-back with nothing to reveal tends to come out empty.
- Change one thing per hop. A move plus a reveal is plenty. Ask for a new angle, new weather and a new character at once, and you will lose the continuity that made the chain worth having.
- Think in coverage. Wide, medium, close, reverse. If the sequence would cut together in an edit, it will usually board well here.
The catch that comes with every chain: grain
Chaining is where the add-on earns its keep, and it is also where it breaks down.
Every time a generated frame goes back in as the next input, small artifacts come along and get amplified. The ComfyUI Wiki's write-up of the add-on reports that at strength 0.9 (how hard the add-on pushes the result), visible grain builds up "by hop 3 or 4." A board that starts crisp looks like a photocopy of a photocopy by the fifth panel.
The fix pairs two adjustments, and the author's own test grids (in the repo's mitigations folder) try two versions of it, each combining a lower strength with a light blur:
- Turn the add-on down for chains. Single shots work at strength 0.7 to 0.9. For chains, ComfyUI Wiki's write-up recommends 0.7 or lower.
- Blur every frame slightly before feeding it back. The same write-up gives a 0.5-pixel blur, too small to see by eye, to keep the noise from compounding.
The repo also keeps a softer version of the add-on, saved earlier in training (next_scene_step1500.safetensors), which the card describes as gentler on identity. Use the strongest version (next_scene_step2500.safetensors) when you want big moves, and the softer one when you need the same person to stay recognizable across six panels.
I would treat this as part of the instrument rather than a flaw to wait out. A cinematographer manages the noise floor of a camera. A board artist using this tool manages the noise floor of a chain.
Put this into practice
Start in the hosted demo before installing anything. Use a frame you own: a location still from a scout, a concept painting, a frame grab from your own footage.
- Pick a frame with space in it. Wide outdoor views, streets, interiors with depth. The training material was filmic scenes and establishing shots, and that is where the model is strongest.
- Write your shot list first, away from the tool. Five lines, each starting with a camera move and ending with a reveal. If you would not hand those lines to a camera operator, rewrite them.
- Run the first hop and judge it as a cut. Ask whether it would play after the first frame in an edit, not whether it is a pretty picture.
- Chain with the settings turned down. If you run it yourself in ComfyUI (a free app for running image models on your own machine), use strength 0.7 and add a 0.5-pixel blur to each frame before it goes back in. The card's single-image example uses 28 refinement passes; the author's own chain demo used 40.
- Keep the rejects. A hop that goes wrong often tells you your direction was ambiguous. Rewrite the line and run that hop again from the last good frame, not from the start.
- Lay the panels out in sequence and watch them as an animatic. Even a slideshow at two seconds a panel will show you whether the coverage cuts together.
If you want to run it on your own machine, the add-on loads like any other Qwen-Image 2.1 add-on in ComfyUI. Be ready for the base model's size; this is not a laptop job for most people, which is the best reason to start with the hosted demo.
Where it breaks
You cannot use the results in client work. The add-on itself is Apache 2.0, a permissive open license. That tag tells you nothing about what you can sell, because it rides on Qwen-Image 2.1, whose license (listed as qwen-research) says it is "for non-commercial purposes only." Treat everything this makes as previsualization: pitch boards, shot planning, internal references. Not deliverables, not posters, not frames in the final film.
It knows two films. The author calls this a first version, "1.1K pairs from two films," and says a second version "with broader footage is the planned upgrade." One animated fantasy and one sci-fi short teach a recognizable style of coverage. Ask it for a documentary handheld feel or a sitcom's flat lighting and you are outside what it learned.
It trained on small pictures. The card states training happened at 512 pixels square, and "larger canvases inherit the behavior but were not trained on." Expect softer detail at full size, which matters less for boards than for anything you want to print.
Continuity is approximate. A character's costume, a prop's position, the time of day can all drift between hops. A human storyboard artist would hold those fixed; this tool holds them loosely, and the softer version only helps so much.
People are the weak spot. Portraits are out of scope, and close-ups of faces push it toward the territory it was told to avoid. Sequences that live on performance (a conversation, a reaction shot) are better boarded by hand.
The real gift is the shot list
The output of this add-on is disposable by license and by design. What you keep is the five lines of direction you had to write to get it, and a quick check on whether your coverage actually cuts.
That is worth an afternoon. Take one frame from a project you care about, write the shot list you would give a crew, and see whether the model's version of your sequence plays. If it does not, the question worth asking is whether the problem is the model, or the list.
Medium metadata
Title: AI Storyboard From One Frame: The Free Next-Scene Add-On That Treats Your Prompt as a Cut
Subtitle: A small open add-on for Qwen-Image 2.1 learned from film cuts instead of photos, so it makes the next shot rather than editing yours. Good shot lists matter more than good prompts, and the license keeps it in previs.
Tags: Storyboarding, Generative AI, Filmmaking, Previsualization, AI Tools
Estimated read time: 8 minutes