Creative AI Briefing: Thursday, October 1, 2026
A 15-second, single-take martial-arts fight scene, drafted at low resolution so you can watch the choreography and kill it before your graphics card spends the expensive pass polishing a bad punch, now runs on a 12 GB RTX 4070 Ti from one downloadable ComfyUI workflow. On the commercial side, Kling opened a limited early access to Kling 4.0 Flash and published a spec sheet for the full Kling 4.0, due in October, with up to 10 keyframes per clip and room for 15 reference images and videos. Both point the same way: AI video is turning from "roll the dice and wait" into something you steer while it renders. You pin the shot in advance, you look at a rough draft, and you only pay for the take worth finishing.
New models
Kling 4.0 lets you plan a 30-second shot like a storyboard instead of a slot machine. Kling AI's own Kling 4.0 vs 3.0 post, published September 30, says "Kling 4.0 Flash became available to a limited group of users for early access on September 28" and that "the all-new Kling 4.0 will officially launch in October." The changes that matter to someone making things: native clips of 3 to 30 seconds (Kling 3.0 topped out at 15), "up to 10 keyframes" where 3.0 had no multi-keyframe control, "up to 15 combined reference assets" including up to 10 images and up to 5 videos totalling 30 seconds, two-channel stereo audio, 10-bit HDR at 1080p and 4K, and 21:9 ultrawide output. In practice a product ad with four story beats becomes one generation with the beats pinned as keyframes, instead of four clips you stitch and color-match. The catch: the early access is to the Flash variant, the post does not say which of these specs Flash has, and Pandaily reports (secondary, not confirmed on Kling's page) that it opened to Ultra yearly subscribers only. Kling's post gives no pricing or credit cost, and every spec above is the vendor's description, not a test. Bloomberg covered the launch as part of the Kling spinoff's push toward a Hong Kong listing. Try it at kling.ai if your plan has Flash access, or plan shots now for the October release.
Video
A martial-arts fight workflow for MiniMax H3 that shows you a draft first. Work-Fisher/MiniMax-H3-Wuxia-Fight-Workflow was created on Hugging Face on September 30 at 17:39 UTC. It is a two-pass ComfyUI graph: pass one drafts motion and composition at low resolution, a learned upscaler enlarges that draft 1.5 times, and pass two adds detail. A preview node shows you pass one so you can cancel if the motion is wrong. It ships with a Chinese-language fight-choreography prompt template and was tested on an RTX 4070 Ti (12 GB) with 32 GB of system RAM. The honest catch is the download: 12 model files totalling 43.9 GB, six custom node packs pinned to specific versions, and a 15-second ceiling. MiniMax H3's Community License governs the base model, and the card warns that the other parts carry their own licenses.
MiniMax H3 now runs in two steps inside ComfyUI. The PDMD research team's 2-step H3 add-on (pdmd2026/pdmd_2NFE_lora, created September 23) only loaded in research code. On September 30, Iwannapose/minimax_h3_pdmd_2nfe_comfyui (created 21:45 UTC) converted it for ComfyUI's standard LoraLoaderModelOnly node, with a 4-step sibling created earlier that day. Fewer steps means less waiting per take, which is the difference between iterating and queuing. The converter reports every tensor matches the original. Settings are strict: strength 1.0, exactly 2 steps, no guidance. Apache 2.0 on the add-on; H3's own license still applies to the base. This is second-day work on a paper the team posted last week, and nobody has published side-by-side quality against the full step count yet.
Watch an H3 video form while it generates. OzzyGT posted minimax_h3_preview_blocks on September 30 (14:44 UTC): add-on pipeline blocks that hand you a picture of "what the model thinks the video will be" at every step. The rough color mode takes about 8 ms per preview at 960x544, a quarter-size detailed mode about 60 ms, and the full detailed one, using a 22 MB tiny decoder, about 280 ms. Raise an error from your preview function and the run stops, so you can abort a bad generation halfway. Apache 2.0. The catch: it is Python code for Diffusers users, not a ComfyUI node, and only the picture previews, not the soundtrack.
Two experimental H3 decoders. speach1sdef178/MiniMax-H3-X2-Detail-VAE (created September 30) is a ComfyUI decoder that outputs H3 video at twice the resolution, and its card openly documents a failed experiment: the extra-detail mode only works when you feed it a real reference image, because, in the author's words, "useful information disappears before the normal generated representation." corechan/MiniMax-H3-LightVAE (created October 1) is a slimmer decoder the author times at 5.2 s versus 7.1 s for the official one, both compiled with NVIDIA's TensorRT, for a 124-frame 1280x704 clip on an RTX PRO 6000 (author's own figures; without TensorRT the light decoder took 8.3 s). Both inherit H3's license, which grants no rights in the EU, UK, South Korea or the United States.
Image
Luma Variants turns one finished static ad into every placement and language. Luma launched Ad Variants on October 1: "Now any Luma user can upload one approved static ad, pick formats and languages, and Variants builds every version for you," keeping logo, headline and call to action while the layout and copy adapt. It covers five placements (Story 9:16, Portrait 4:5, Square 1:1, Medium Rectangle 6:5, Widescreen 16:9). It lives in Luma's Discover tab. Catch: static ads only, no published pricing, and no list of supported languages beyond Spanish in the example. A designer still needs to check every translated headline.
Photographic polish for Qwen-Image-2.1. SimpleTuner/Qwen-Image-2.1-LoRA-photo-aesthetics-v3 (created September 30) is a style add-on used at strength 1.0 with no trigger word. The trainers report the version trained at 512 px carries more fine detail than one trained at 1 MP, even when both generate at 1 MP. Qwen research license, so no paid work.
Krea 2 Turbo on AMD hardware. PuppetVision/krea-2-amd-rocm-optimized-comfy-triton (created September 30) packages an 8-bit Krea 2 build for ComfyUI on AMD's ROCm stack. On one Strix Halo machine the author measured a first generation dropping from 242.4 s to 85.0 s, and the model file shrinking from 26.28 GiB to 13.16 GiB. Tested on one chip only, and the author says so.
Audio and music
Raspy rock-soul vocals for the open YuE2 song model. becausereasons/yue2-grvl-raspy-rock-soul (created September 30, 13:23 UTC) is six add-ons that push YuE2-3B toward a raspy female rock-soul voice, trigger word grvl, with eleven demo MP3s you can play in the browser before downloading anything. Trained on recordings from the 1960s to 1990s; the card does not name the sources, which matters if you plan to release what you make. CC BY-NC 4.0 means no commercial use. English and female lead only, and songs can cut off abruptly at the length cap.
Dia2 dialogue voices without an NVIDIA card. markldn/Dia2-2B-GGUF (created October 1, 06:09 UTC) converts Nari Labs' Dia2 2B voice model for a custom native runtime, audio.cpp-dia2, that runs on AMD cards and plain CPUs. The smallest build is a 1.08 GB file that peaked at about 3.9 GB of graphics memory on the converter's AMD test machine. It needs that specific runtime (ordinary llama.cpp will not work), and the converter's own tests found the smallest build making more word errors. Apache 2.0 on the weights; the Mimi audio codec is CC BY 4.0.
Open and local
Today's local story is MiniMax H3 getting cheaper to iterate on from three directions at once: fewer steps (PDMD in ComfyUI), earlier exits (previews in Diffusers and the Wuxia draft pass), and faster decoding (LightVAE). None of it comes from MiniMax. All of it is community second-day work on a model released in July.
- Work-Fisher/MiniMax-H3-Wuxia-Fight-Workflow: a two-pass fight-scene workflow with a cancel-early draft, tested on a 12 GB card. Created September 30 (card).
- Iwannapose/minimax_h3_pdmd_2nfe_comfyui: two-step H3 video inside stock ComfyUI. Created September 30 (card).
- OzzyGT/minimax_h3_preview_blocks: see each step of an H3 generation and abort bad runs. Created September 30 (card).
- speach1sdef178/MiniMax-H3-X2-Detail-VAE: 2x-resolution H3 decoding in ComfyUI, with an honest write-up of what failed. Created September 30 (card).
- becausereasons/yue2-grvl-raspy-rock-soul: six vocal-style add-ons for open song generation, with demos. Created September 30 (card).
- markldn/Dia2-2B-GGUF: Dia2 voices on AMD and CPU. Created October 1 (card).
- kyutai/ovie-512: a 512-pixel version of Kyutai's single-photo new-camera-angle model, MIT license, from a March paper. Created September 29 (card).
Trendshift's front page showed no creative-media repos among its top movers today, so there are no star deltas to report.
Creative workflows
1. Draft a fight scene, check the choreography, then commit. Files: 【WORK-FISHER】26-9-28-MINIMAXH3-武戏高动态工作流.json and 打戏提示词模版_MiniMax15秒一镜到底_v1.4.md.
The steps. Install the six node packs at the versions listed on the card (the core one is comfyui-minimax-h3-audio-T8 v1.86.0). Download the 12 model files into the folders the card names, and rename the upscaler file after download as instructed. Load a character reference into the "参考图1" (Reference Image 1) group, write your fight in the "简单提示词写入" (simple prompt) box using the template, set the duration (the card suggests about 5 seconds for a first try, 15 maximum), and queue. When the pass-one preview appears, judge the motion. Cancel if it is wrong.
How it works. The expensive part of video generation is detail at full resolution. Motion and composition get decided early and cheaply, so the workflow renders those small first, enlarges the rough version, and spends the heavy pass only on drafts you approved.
Why it is good. Fast action is where AI video fails most, so being able to throw away a bad take after the cheap pass saves most of the wasted compute. It fits on a mid-range 12 GB card.
Where it breaks. 43.9 GB of downloads, pinned node versions that can conflict with what you already have, and Chinese-language UI labels and prompt template. The int8 model, six add-ons from four different authors and a bridge model all carry their own licenses, so check each before client work.
2. Two-step H3 in stock ComfyUI. File: minimax_h3_pdmd_2nfe_comfyui.safetensors.
The steps. Load your normal H3 graph, add LoraLoaderModelOnly, point it at the file at strength 1.0, set steps to 2, use H3's own scheduler, and turn guidance off. Use the 4-step file (the _v6 one; the converter says earlier uploads silently did nothing) when two steps looks too rough.
How it works. A slow model was used to teach a student to land close to the same result in far fewer passes, and the add-on carries that student's changes.
Why it is good. Short waits make it practical to try ten prompts instead of two.
Where it breaks. Any strength other than 1.0 changes the behavior, and there is no published quality comparison yet. Expect softer detail than a full-step render and judge it on your own shots.
3. Carry forward: the 360 headset workflow from September 30. The H3 equirect add-on recipe from yesterday's briefing remains the best documented end-to-end pipeline this week; see that briefing for steps and limits.
Worth testing
- Kling 4.0: try keyframed 30-second clips if your plan is in early access. Tradeoff: access is limited and pricing is unpublished.
- Luma Ad Variants: run one finished ad through all five placements. Tradeoff: static only, and machine-translated headlines need a human check.
- GRVL demo MP3s: hear the vocal add-ons before installing anything. Tradeoff: non-commercial license and undisclosed training recordings.
- PDMD 2-step for ComfyUI: free and drop-in if you already run H3. Tradeoff: no independent quality comparison, and H3's territory restrictions still apply.
What actually matters from today's signal
The work in AI video is moving from generation to direction. Kling 4.0's ten keyframes and fifteen references are a storyboard you hand the model up front. The Wuxia workflow's draft pass and OzzyGT's step previews are a director's monitor, letting you call cut before the expensive part. Both reward the same skill: knowing what the shot should look like before you press the button. A creator who storyboards will waste far less money than one who rerolls.
For local creators, the money story is H3. Three separate community releases in one day cut steps, add early exits and speed up decoding, and none of them came from MiniMax. That is what an ecosystem looks like when the base model is good enough to build on. The catch is the base. The add-ons themselves are mixed (the PDMD conversion and the preview code are Apache 2.0), but none of them runs without H3, and the H3 Community License text quoted on the LightVAE card grants no rights in the EU, UK, South Korea or the United States. Read that license before you build a paid workflow on any of this.
The counter-signal: the biggest feature claims today, Kling's 30 seconds, HDR and stereo, are a vendor comparison table with no independent tests and no price. Treat them as a plan, not a purchase, until October's full release shows what a 30-second clip costs and how often the tenth keyframe is actually honored.
Source access notes: Every Hugging Face date is createdAt from the API (/api/models/<id>) with cache busting. The cloud and local shells could not reach the Hugging Face API directly; all HF data came through WebFetch. Kling specs come from Kling AI's own blog; the Ultra-yearly access detail is Pandaily's report only. Bloomberg was robots-blocked, so only the headline and date are cited. Runway's news page, Stability's news page, Google's AI blog and Adobe's blog returned no dated items (JS shells or 404); blog.comfy.org and the ComfyUI releases page were not readable; the Krea release notes did not render. OpenAI's news page showed no creative-media posts for September 28 to October 1 (DevDay focused on agents). Suno's newest post is a Studio EQ tutorial (September 29), and ElevenLabs' newest remains Eleven v4 (covered yesterday). Midjourney, Ideogram, Recraft, Pika and HeyGen showed nothing dated in the window. The HN Algolia feed had no creative-AI stories. Civitai not attempted. Writing section omitted: nothing qualified. Adversarial fact-check pass ran (separate agent, primary sources): it caught that Kling's early access is for the Flash variant while the spec table describes the full October model (lede and New models item reworded), an altered Luma quote (now quoted as written), the Wuxia add-on count (six from four authors, not four), an overstated license claim in the closing section (PDMD conversion and preview code are Apache 2.0; only the base and two decoders carry H3's license), an invented "ten minutes" figure in the opening (cut), and missing context on LightVAE timings (TensorRT on both sides) and Dia2 memory (graphics memory on AMD). All Hugging Face dates, file sizes, node versions and Kling counts checked out.