Creative AI Briefing: Sunday, October 11, 2026
You can now take a five-second shot with a stranger walking through it, paint over them once in an image editor, and have an open video model rebuild the background behind them in short, rerollable chunks with your original soundtrack still attached. That release, and most of what shipped around it, works on material you already own: a frame that needs a next shot, a render that needs to finish sooner, a merged model that flickers less. The weekend's open work is repair and continuity, and that is the kind of AI a working editor can bill for without anyone noticing it was there.
New models
H3 Person Remover LoRA deletes a person from a video clip and rebuilds what was behind them. Akatz Labs posted it October 8 (Hugging Face createdAt 04:00 UTC, card refreshed October 10). The add-on does not find the person or invent the clean background on its own. SAM 3.1, a tracking model, follows the person from a text description such as "man in gray shirt"; you supply a clean version of the first frame, made in Photoshop or any image editor; the workflow paints the person green, and MiniMax H3 regenerates the covered area in overlapping windows, each window handing its last frames to the next. Window lengths follow 17n + 5 frames (22, 39, 56, 73), and the author recommends starting at 22 because longer windows cost more memory without reliably looking better. Output is silent by default; wire the original audio back into the final node to keep your sound. The workflow file is examples/Person-Remover-Window-Reroll-V1.json. License: the MiniMax H3 Community License, whose standard territorial grant, per the card, "excludes the US, EU, UK, and Republic of Korea," so most readers of this briefing need separate authorization before using it at all. Tested on an RTX 4090 with ComfyUI 0.37.0, which the author calls "a tested configuration, not a minimum hardware specification." No hosted demo. The honest catch, in the author's words: "H3 regenerates the full scene," so pixels outside the mask are not locked to your original, and a missed hand, shadow or reflection stays in the shot.
Qwen-Image-2.1 Next-Scene LoRA turns one still into the next shot of a sequence. Posted October 7 (createdAt 20:42 UTC) by akhaliq, retraining lovis93's earlier next-scene format on Qwen-Image 2.1. You give it a frame and a direction that starts with Next Scene: and leads with the camera ("Next Scene: The camera pulls back to a sweeping aerial view, revealing the fleet behind the cliffs"), and it renders the following shot rather than editing the input. Feed each output back in and you have a storyboard. It learned from 1,144 consecutive-frame pairs from the Blender open films Sintel and Tears of Steel, so it is strongest on filmic scenes, landscapes and establishing shots; static portraits are out of scope by design (per ComfyUI Wiki's writeup, secondary). The add-on is Apache 2.0, but it rides on Qwen-Image-2.1, which carries the non-commercial Qwen Research License, so treat the boards as previsualization, not deliverables. Free hosted demo (Space createdAt October 8, running).
Image
Kroma v0.3.1 makes the community's Krea 2 fine-tune fast without the usual quality loss. The Kroma repo dates to July 31; the new kroma-v0.3.1-turbo-opd.safetensors arrived with a card refresh on October 10 (ComfyUI Wiki dates the v0.3.1 post to October 7, secondary). This is second-day work: Lodestones trained the fast version on the pictures it actually produces mid-render, which the readme pitches as "Turbo speed without the usual distillation tax." In practice that means 8 to 12 steps instead of a long render, and add-ons trained on the base should load unchanged. The file is about 26 GB, the license is the Krea 2 Community License, and there are no benchmarks yet, so "no quality tax" is the author's intent until someone measures it. One trap: the readme's download table still lists the deleted v0.2 file.
Video
Veda makes MiniMax H3 reference-to-video renders roughly 1.4x to 2.2x faster end to end. The R2VA Preview went up October 7 (createdAt 04:49 UTC). It teaches the model which parts of the picture each part needs to look at, then skips the rest, while leaving text, prompt and audio handling untouched. By the authors' own measurements on one RTX PRO 6000 Blackwell card with the eight-step Turbo add-on, a 5-second 16:9 clip finishes 1.69x faster and a 14.4-second clip 2.20x faster; square clips gain least (1.37x). It is a preview trained for only 600 updates, under the MiniMax H3 Community License, and the card names no memory requirement.
LTX-2.5 OmniGen v12 folds five Lightricks add-ons into one model file. ApolloRaines posted it October 6 (createdAt 23:30 UTC): Refine Details, Restore, SDR-to-HDR, Alpha Gen and Layout-to-Render merged permanently, plus an edit aimed at flicker. The author's preliminary, single-seed tests show flicker down 12 to 27 percent on three image-to-video shots (a clock photo is the best at -27%) and a tie on the rest, and the author warns that part of the gain comes from the merge producing less motion. It is an 18.6 GB GGUF for ComfyUI-GGUF, eight steps at cfg 1.0, gated behind a contact-info form, and licensed under the LTX-2.x community license agreement, the same terms as the base model. Text in zooms still melts and fast fly-pasts still smear.
Self Gradient Forcing Plus is a research release for very long text-to-video. SGF+ (createdAt October 7, 15:49 UTC, Apache 2.0) builds on Wan 2.1 and claims "rollouts of up to 24 hours from only 5s training windows"; the default setting produces about four minutes at 16 fps. Its launcher reaches for eight GPUs. Research only for now (score 2).
Reka Rho-1 is a research preview, not a tool. Reka's October 5 post describes one model that reads and generates text, images and video and edits video by prompt, capped at 672x384 and, in Reka's words, "a functional proof-of-concept and an architectural direction, not a finished product." Reka offers it as a research preview by contact and does not say whether weights or an API will follow (score 2).
Audio and music
Steer-TTS is a 9-million-parameter English voice you direct with words, tags or dials. Whissle posted steer-tts-9m October 9 (createdAt 02:04 UTC). Type "furious, shouting," pick from 33 style tags such as whisper or low_pitch, or nudge 12 controls like rate and pause, and you can adjust pitch, loudness and length on single words. It is small enough to run on a CPU. The catches: 287 fixed voices and no cloning, English only, a mid-training checkpoint (step 116k of 150k) whose published scores come from step 58k, and a CC-BY-NC 4.0 license, so no paid work without a deal.
Antalia Mini now runs in a browser tab. Antalia Mini, the 31 MB Turkish voice released October 8, got a second-day ONNX port (createdAt October 10, 01:04 UTC, Apache 2.0) with a live web demo using WebGPU or WebAssembly. Nothing to install.
Open and local
The local story this weekend is plumbing for MiniMax H3: one tool to remove things from shots, one to render them faster, one to stitch long shots out of short windows without losing your accepted takes. Each depends on a base model whose license excludes the US, EU, UK and Korea by default, which is the single biggest fact about the whole H3 ecosystem and the one launch posts keep skipping.
- akatz-ai/h3-relay: ComfyUI nodes that build long MiniMax H3 shots from short windows and let you reroll one window without redoing the ones you liked. Why now: it is the reroll engine behind the Person Remover; 21 stars (repo, GPL-3.0 for the code only).
- ARahim3/DigUp: a free Mac app that finds photos, recordings and video in folders you choose by describing them, fully offline. Why now: trending on Trendshift today; 395 stars (repo, MIT, Apple Silicon, 865 MB model download).
- calesthio/OpenMontage: an agent-driven video production kit with eleven pipelines (explainer, documentary montage, talking head, dub) that can run free with local narration and open archives. Why now: on Trendshift's daily list; 66k stars (repo, AGPL-3.0).
- Jakubantalik/transitions.dev: copy-ready interface transitions (modal open, card resize, error shake) for web motion work. Why now: trending; 4.8k stars (repo, usable commercially, tooling MIT).
- wheresryan22/anatomy: a Claude skill that explains an idea by drawing it as an interactive isometric machine in SVG, with WebGL and 3D options. Why now: trending; 729 stars (repo, MIT).
Star totals are from shields.io this morning. Trendshift's own numbers are unlabeled, so no daily deltas are claimed.
Creative workflows
1. Remove a person from a shot and keep your soundtrack: H3 Person Remover (model card, workflow examples/Person-Remover-Window-Reroll-V1.json, H3 Relay)
The steps. Update ComfyUI until it has native MiniMax H3 and SAM 3.1 support. Put H3-Person-Remover-V1.safetensors in models/loras/. Conform your clip to 24 fps with width and height divisible by 32, ideally about five seconds and one continuous shot. Paint the person out of frame one in an image editor. Type a description for SAM 3.1, check the green mask, run with 22-frame windows, review each window and reroll the bad ones. Connect the original audio from Get Video Components to the final Create Video node.
How it works. The model never sees the person, only a green hole and your clean first frame. It fills the hole window by window, and each window inherits the last 18 frames of picture and sound from the one before, so the fill stays continuous.
Why it is good. Rerolling one window keeps every earlier window you approved, so a bad patch costs one window of render time, not the whole clip.
Where it breaks. The whole frame is regenerated, so check unmasked areas against the original. Shadows, reflections and stray hands the mask misses stay. A sloppy first-frame cleanup spreads through every later window, and hard cuts or big camera moves break continuity. Licensing excludes the US, EU, UK and Korea by default.
2. Board a sequence from one frame: Next-Scene chains (model card, recipe per ComfyUI Wiki)
The steps. Load next_scene_step2500.safetensors on Qwen-Image 2.1's edit workflow with your frame as the input. For a single next shot use strength 0.7 to 0.9. For a chain, drop strength to 0.7 or lower and add a 0.5-pixel blur to each frame before feeding it back. Start every prompt with Next Scene: and the camera move.
How it works. It learned what tends to happen between consecutive frames of two films, so it treats your direction as a cut, not a touch-up.
Why it is good. A director or storyboard artist gets five consistent angles on one location in the time it takes to write five sentences.
Where it breaks. At strength 0.9 grain builds up visibly by the third or fourth hop. It trained at 512 px, it avoids portraits, and the base model's research license keeps it out of client deliverables.
3. Carry-forward: check before you build on it, with SynthID Detector (Google's post)
The steps. Upload supplied assets at synthid.com before compositing them into client work, and log the result in the project file.
Why it is good. It turns a reputational risk into a two-minute check, and PetaPixel's October 9 report that Nikon disqualified a contest winner shows the stakes.
Where it breaks. It only detects marks from Google and named partners; a "no" proves little about open models.
Worth testing
- Antalia Mini in the browser (demo). Free, no install. Tradeoff: Turkish only, and browser speed depends on your GPU.
- Next-Scene on a location still you own (demo). Five hops, strength 0.7, blur on. Tradeoff: research-only base license, so it stays in previs.
- H3 Person Remover on one five-second plate. Tradeoff: needs a strong NVIDIA card, and the license territory may exclude you outright.
- DigUp on your archive drive (repo). Tradeoff: Mac only, and describing a shot cannot find "the good take."
- Steer-TTS for scratch narration. Tradeoff: fixed voices, English only, non-commercial.
What actually matters from today's signal
The work that shipped this weekend is invisible by design. Nobody watching a finished cut will see a removed passerby, a faster render, or a storyboard that helped a director pick an angle. That lines up with Creative Bloq's October 10 argument that AI's safest film role is the work that does "not influence what ends up on the screen." Repair tools are where the time savings are real and the audience backlash is absent, and a cleanup that used to take a compositor an afternoon now costs a few rerolls.
The money question is uglier than the craft question. Three of this weekend's most useful releases sit on MiniMax H3, whose default license excludes the US, EU, UK and Korea, and two more sit on Qwen-Image 2.1's research-only license. The open ecosystem is producing excellent tools that most professional readers cannot legally put in front of a client. Before you learn a workflow, read the base model's license, not the add-on's; the add-on's Apache tag tells you nothing about what you may sell.
The counter-signal: removal tools regenerate the whole frame, so "invisible" depends on you checking every pixel outside the mask. The release notes say so plainly. Believe them.
Source access notes: blog.adobe.com Firefly topic page returned empty; runway.com/news returned navigation only; blog.comfy.org JS wall; midjourney.com not attempted (blocked historically), WebSearch found no October update; Hugging Face API blocked from the shell, so every createdAt was read through WebFetch of the API; HN Algolia returned stale August results and GitHub trending returned a stale page, both discarded; civitai skipped. openai.com, deepmind, bfl.ai (latest post Sept 23), lumalabs (Oct 8 Claude Motion item already covered), stability.ai, elevenlabs.io (latest Sept 30) and suno.com (latest Oct 5) had nothing new and creator-facing in the window. Kroma v0.3.1's exact posting date is secondary (ComfyUI Wiki); the repo createdAt is July 31. Adversarial fact-check ran (Sonnet subagent). It confirmed every Hugging Face createdAt, the Person Remover settings and quotes, Veda's speedups, Steer-TTS figures, star counts and the Creative Bloq quote. It caught five errors, all fixed: OmniGen v12 is under the LTX-2.x community license, not Apache; the OmniGen flicker figures were rewritten from the card's actual table; the Kroma readme quote was misworded; Antalia Mini shipped October 8, not yesterday; and Reka never states that weights or an API are unavailable.