FervorCreative AI
Live Latest 22.09.26 · morning 72 tools tracked 210 workflows indexed 187 topics Hot: MiniMax H3, ComfyUI, Krea 2

Three separate releases this weekend moved a camera through footage that was already shot, and none of them did it with a prompt sentence, which means the control surface for AI video is becoming a depth pass rather than a description.

WorldCrafterSupra2-IMGCrossView-WarpMoGe 3SpanSynth-Editvideo-genimage-gencomfyuilocal-creative-aiopen-weightscreative-workflowsmusic-genlicensing-provenance

Creative AI Briefing: Tuesday, September 22, 2026

Take a clip you shot last month, drag a camera marker around a sphere until the subject is three-quarters instead of head-on, and render the same performance from the new angle. That shipped Sunday, free, as one custom node and one adapter file. Look at what else landed in the same 48 hours and the shape of the week is obvious: three separate teams moved a camera through footage that already existed, and not one of them did it by writing a sentence about where the camera should go. Each one measured the scene instead. Two of them build an explicit depth map of your footage and hand that to the generator as the instruction; the third keeps its own internal three-dimensional record of what it has already drawn. ComfyUI sharpened the depth tool that feeds the first two the same weekend. The frontier labs published nothing creative-facing in the window. The community wrote the week.

New models

WorldCrafter-Fast went up Monday, September 21, from Tencent's ARC lab: a video world model you steer with camera moves instead of sentences. Feed it one image, or just a prompt, then give it a trajectory. forward1x2 yaw_left30x3 backward1 walks you forward twice, pans left thirty degrees three times, then backs up, thirty-three frames per chunk. What separates it from the adapters below is that it holds onto what it already rendered. The paper (arXiv 2609.24984) calls it a camera-queryable implicit 3D-aware memory, which in practice means that if you pan away from a bookshelf and pan back, it is still the same bookshelf.

The catch is the size. The Hugging Face repository reports 147 GB of storage: two full transformers at roughly 57 GB each, a 22.7 GB prompt-reading model and a 2.5 GB scene encoder, which is about 139 of it, with the remaining 8 spread across two 2.5 GB adapters, two 1.06 GB partial transformer files and a 0.51 GB autoencoder. "Fast" refers to the step count, not the footprint. Output is 384 by 640. Resuming a rollout works only on the Base checkpoint. Linux, an NVIDIA card, Python 3.11 and a CUDA 12.8 build of PyTorch 2.10 are required, and the interactive keyboard demo is, in the README's own words, currently being debugged. There is no Apache or MIT tag, only a LICENSE.txt carrying terms of use, so treat commercial use as an open question until you have read it. Repository at 150 stars. Free hosted demo: hugging-apps/worldcrafter-demo.

Supra2-IMG arrived twenty-five minutes earlier the same day and is interesting for the opposite reason. SupraLabs trained a 104.1 million parameter text-to-image model from scratch on a single rented H100 and released it Apache 2.0. The whole checkpoint is 417 MB. It makes 256 by 256 images, reads prompts through a frozen Flan-T5-Base at 128 tokens, and decodes through the old SD-VAE-FT-MSE, both pulled from their own repositories rather than shipped in the download.

The training run is the story: ten epochs across FLUX-Reason-6M, 5.6 million images after preparation, on one H100 SXM 80GB rented from RunPod. The ComfyUI Wiki write-up headlines ten hours and says nine in the body including data prep, so read it as roughly a day of one rented card. At spot rates that is a rounding error against anything else in this section.

Nobody should replace their image model with this. What it is good for is fine-tuning and pretraining experiments that are unaffordable at 8 billion parameters. No ComfyUI support, because it is a custom SupraDiT architecture rather than a diffusers pipeline; you run the shipped script at seed 0, CFG 3.0, 50 steps. Weights and code, 86 likes. Free demo: hugging-apps/supra2-img-demo.

Video

CrossView-Warp v1 for MiniMax H3 landed Sunday, September 20, and it is the clearest expression of the week's pattern. Cseti's adapter re-renders a clip you already have from a different camera angle. You set an azimuth and elevation offset numerically, and the result lands where the geometry says it should.

The mechanism deserves a plain sentence. The adapter reads two videos at once. The first is your clip reprojected into the new viewpoint by a depth warp, with magenta filling every region the original camera never saw. That carries the viewpoint change. The second is the untouched clip, which carries identity, lighting and texture. The warp goes into the guide path, the original goes into the reference path, and the trigger word crossview switches the adapter on. Because the offset is geometry rather than a prompt phrase, the camera goes where you put it.

The author publishes the boundaries, which is rare enough to name. In the companion node: azimuth holds to about 45 degrees in the green zone and 90 in yellow, and the author says the training set is spread evenly out to 90, so sideways is genuinely covered; elevation is the asymmetric one, green from plus 30 to minus 20 and yellow to plus 45 and minus 35, with looking up from below minus 20 covering only 3.1 percent of the training set. Near-zero angles misbehave because the warp is nearly identical to the source. And the distance control, in the README's own bolded words, "doesn't work as expected due to some dataset problems which will be solved in the next release." Angles are what this is good at. Dollies are not.

Published settings on the H3 model card: strength 0.8 to 1.0, with 0.8 recommended on the distilled path; 124 frames; a first pass at 0.5 megapixels and a second at 1.5, both 16:9; res_multistep at 8 steps with the DMD 8-step turbo adapter. The released checkpoint is step 3,500 of a planned 6,000, trained on 719 Blender scenes rendered across camera offsets, at 512 by 512 and 73 frames, in 33 hours 55 minutes on an RTX PRO 6000. Licence is the MiniMax H3 community licence, inherited from the base. Node at 59 stars, Apache 2.0. This is the third CrossView release after two on LTX-2.3, and the first built on a reference path rather than prompt-driven camera phrases. The author moved off text control deliberately.

Image

An AMD build of Qwen-Image-2.1 appeared Tuesday morning, September 22, from EliovpAI, compressed to MXFP4 for ROCm cards, about 9.4 GB total. What separates it from the pile of community conversions is the measurement trail it ships: an evaluation folder with matched pairs of compressed and full-precision renders across eight prompt categories, MI355 timing files, pinned lock files, integrity tests and a source manifest. You can reproduce the author's numbers rather than take them. The inherited licence is Qwen Research, which is research and evaluation only; selling work made with these weights requires a separate agreement.

The two image stories bracket the medium neatly: one team spends a day of one GPU making a model from nothing, another spends a weekend making a very large model fit on a card it was never built for.

Audio and music

SpanSynth-Edit shipped Tuesday morning, September 22, Apache 2.0 on both code and weights, and it does something a musician will understand instantly: it lets you change the notes in a recording by changing the MIDI.

Select a region, revise the score for it, and the model regenerates only that span, taking timbre and room from the surrounding audio so the patch sits in the same mix. Outside the generated region the output matches the input sample for sample, which is what makes it usable rather than a novelty. Two methods ship: the ordinary path takes the revised MIDI alone, the flowedit path takes original and revised so it can reason about what changed. The hardware ask is modest for once: 6 GB of graphics memory recommended, measured peak around 3.9 GiB at defaults, Apple Silicon supported through Metal and tested on an M1 Pro with 16 GB. The real limit is length. One run handles at most 20.48 seconds, output is 48 kHz mono, and the paper is still forthcoming. Take the results, not the claims, until it lands.

Also Monday, September 21: OpenStems published four music source separation bundles in GGUF form, each zipped with a matching SHA-256 file, around 870 MB total, including a karaoke model. The repository carries no licence tag at all, which for stem separation you are going to publish from is a real problem rather than a paperwork one. Useful, unattributable.

Open and local

Here is the connective tissue. On Sunday, September 20, ComfyUI v0.37.0 put MoGe 3 into core. MoGe 3 is Microsoft Research's monocular geometry model (paper, weights, MIT): one ordinary frame in, a depth map and a surface normal map out, with thin structures intact because the third generation refines its estimate on a sparse three-dimensional shell rather than in the flat image. Hair, railings and window frames survive.

That is the file a depth-warp adapter wants. Cseti's node README says it outright: MoGe "is the better path and needs no install," and its real metric geometry "roughly halves the warp error" against a relative depth map. Be precise about what changed, though. Run MoGe Inference was already in ComfyUI, and the node still loads MoGe 1 and MoGe 2. What v0.37.0 added is the third-generation checkpoint and its refiner: a sharper measurement, not a new capability. WorldCrafter touches none of it and runs entirely outside ComfyUI. This is convergence, not causation.

The same version brought Qwen-Image 2.1 into core with three templates and two editing nodes, faster prompt reading, automatic fast-disk loading on NVMe drives, and a default Empty Latent Image size raised from 512 to 1024. Two fixes worth knowing: MiniMax Music 3 no longer renders noise with the compiler graphs on, and ACE-Step decoding no longer crashes on cards without bfloat16.

  • TencentARC/WorldCrafter: camera-steered video world model that remembers what it already drew (link, 150 stars).
  • cseti007/ComfyUI-CrossViewWarp: the orbit picker and live preview that build the depth warp; Apache 2.0, no extra packages required (link, 59 stars).
  • mimbres/spansynth-edit: MIDI-guided synthesis and span editing of real recordings, Apache 2.0, runs on a laptop GPU (link).
  • SupraLabs/Supra2-IMG: a 417 MB image model and the script that made it, for anyone studying small-scale pretraining (link).
  • Comfy-Org/ComfyUI: v0.37.0, the release that made sharper geometry a default (link).

Creative workflows

1. Re-shoot a clip from a different angle without reshooting it. Adapter: CrossView-Warp v1, file MiniMax-H3_Ref2VA-LoRA-CrossView-Warp_v1_3500.safetensors. Node: ComfyUI-CrossViewWarp.

The steps. Update to ComfyUI v0.37.0 so Run MoGe Inference is present. Load your clip, run MoGe on it, and wire the moge_geometry output into the CrossView Warp node rather than a relative depth map. Run the graph once to cache the clip. Now orbit the live preview by dragging the picture; magenta shows everything the original camera never saw. Park the playhead, aim, press KEY to set a keyframe, repeat for a move. Wire warp to the guide path and the original clip to the reference path, put crossview in the prompt, set strength 0.8, render.

How it works. The node reprojects every pixel of your footage into the new viewpoint using measured geometry, then hands that reprojection to the model as an instruction. Identity and lighting come from the untouched clip on a separate input. You are not describing a camera move, you are showing one.

Why it is good. The live preview calls the same code as the final render: measured against a full run, at most 4 pixels out of 147,456 differ. What you audition is what you get, without queueing a job.

Where it breaks. Stay inside the green zone, about 45 degrees of azimuth and plus 30 to minus 20 of elevation. Yellow runs to 90 azimuth and plus 45 to minus 35, and sideways that is real coverage; vertically it is not, with below minus 20 at 3.1 percent of the training set. The distance control does not work in this release and the author says so. The step that fails silently on the H3 path is step 5 of the card's own list: resize the warp and the source to the same size before both nodes. Nothing warns you when they differ. (On the older LTX path the equivalent trap is latent_downscale_factor, which must be 1 on both reference guides.) Licence is the MiniMax H3 community licence, so read it before selling the result.

2. Get a usable depth and normal pass out of a single frame. Template: utility_moge3_geometry_estimation, shipped with ComfyUI v0.37.0. Weights from Comfy-Org/MoGe, file moge_3_vitg_fp16.safetensors in models/geometry_estimation/.

The steps. Load one image, run the template, set refine_steps on the Run MoGe Inference node (default 3, maximum 8, 0 turns the refiner off), and take the depth map and normal map out the other side. For a 360 image, the panorama node splits an equirectangular frame into twelve perspective views and merges the depth back.

How it works. Instead of guessing depth pixel by pixel in the flat image, it builds a rough three-dimensional shell and sharpens the estimate there, which is why thin things stop smearing.

Why it is good. No install beyond the weights, because the sparse kernels were already in ComfyUI for other work. The node also reports estimated field of view, which is what you need to match a virtual camera to a real plate.

Where it breaks. refine_steps does nothing on MoGe 1 or MoGe 2 checkpoints, which the same node still loads, so a wrong file in the dropdown gives you a silently worse result with no error. One image in, one geometry pass out; temporal consistency across a moving shot is not something this guarantees on its own.

3. Change a few notes inside a finished recording. Tool: SpanSynth-Edit, Apache 2.0.

The steps. Install with the sparse checkout in the README, which skips the demo audio. Align MIDI to your track, edit the notes you want changed, then run spansynth-edit edit --audio track.mp3 --midi revised.mid --edit-start 6.40 --edit-end 14.08 --cfg 2.0 --output results/. If you have the original MIDI too, add --method flowedit --source-midi original.mid so the model can see before and after. You get the full crop, the generated span alone, and a settings file.

How it works. The model generates only the selected span, conditioned on your MIDI and on the surrounding audio, which is what supplies instrument timbre and room so the patch matches.

Why it is good. Outside the edited region the output is bit-identical to the input. That is a guarantee, not a hope, and it makes the tool safe to put in a real session.

Where it breaks. One run covers at most 20.48 seconds, so anything longer is a stitching job you do yourself. Edit boundaries round outward to 40 milliseconds. Output is mono at 48 kHz, so this is a source for a track, not a master. Audio and MIDI must be aligned or offset correctly, and the instrument vocabulary merges MIDI programs into groups, so five piano programs land on one category. The paper is not out yet.

Worth testing

  • hugging-apps/worldcrafter-demo. Camera-steered scene exploration with nothing to install. Tradeoff: the local version is 147 GB at 384 by 640, so this is a look at the idea, not a preview of your own pipeline.
  • hugging-apps/supra2-img-demo. A 417 MB model that cost about a day of one GPU. Tradeoff: 256 by 256, and not competitive with anything you already use. The price of the run is the point, not the picture.
  • Ruicheng/MoGe-3. Drop in a frame and see whether your footage produces a clean depth pass before committing to a re-camera workflow. Tradeoff: a still tells you nothing about how the estimate holds across a moving shot.
  • SpanSynth-Edit on a laptop. Around 3.9 GiB peak, Apple Silicon supported. Tradeoff: 20 seconds a run, mono output, no published paper yet.
  • Qwen/Qwen-Image-2.1. Now native in ComfyUI core with three templates. Tradeoff unchanged: research and evaluation only.

What actually matters from today's signal

The control surface for AI video changed hands this weekend and hardly anyone said so out loud. For two years the interface has been prose: you wrote "slow dolly in, slight low angle" and hoped. What shipped between Sunday and Tuesday was three different answers to the same question, and all three replaced the sentence with a measurement. CrossView-Warp warps your footage with real metric geometry and hands the warp over as the instruction. WorldCrafter keeps an internal 3D memory so the camera can leave a thing and come back to it. ComfyUI sharpened the geometry estimator sitting under the first approach. Nobody coordinated this. A depth map is not a poetic description of a camera move. It is a camera move.

For a working creative the consequence is a shift in where your time goes. Prompt craft for camera language stops paying, because the model is no longer guessing at your adjectives. What starts paying is the pre-pass: whether your plate produces a clean depth estimate, where the occlusions are, what you will put in the magenta holes the original camera never saw. That is closer to matchmove work than to writing, and people who have done matchmove will be good at it immediately. It also means an angle you did not shoot is recoverable from one you did, within about 45 degrees, which changes what coverage means on a small production.

The counter-signal is about scale. WorldCrafter-Fast is 147 GB of weights to render at 384 by 640, Linux only, NVIDIA only, interactive demo admittedly broken. The adapter that actually works on your machine today is step 3,500 of a planned 6,000, with a dead dolly control and a licence inherited from a base model that restricts what you do with the output. These are not products. They are the community doing second-day work on somebody else's model, publishing the failures alongside the wins, and shipping faster than labs whose release pages sat empty all weekend. Good trade, usual price: read the licence file before you bill anyone for what comes out.


Source access notes: Hugging Face dates from the API createdAt field with cache busting, never listing "Updated" timestamps. GitHub file contents via raw.githubusercontent.com with cache busting; star counts from shields.io with cache busting (WorldCrafter 150, ComfyUI-CrossViewWarp 59). The GitHub Releases API returned an empty body for both ComfyUI repositories, so v0.37.0 contents are sourced from the ComfyUI Wiki write-up, which cites core pull requests #16400, #16381 and #15623; the version number and feature list are therefore secondary-sourced. The Hugging Face daily-papers endpoint and the ComfyUI docs changelog both returned responses too large to read within the fetch limit and were skipped. Midjourney, Civitai and blog.comfy.org were not reached. OpenAI, Google DeepMind, Runway, Black Forest Labs, Stability, ElevenLabs, Adobe, fal and Replicate were checked and published nothing creative-facing inside the window. One discrepancy is left visible: the Supra2-IMG coverage headlines a 10-hour training run and states 9 hours including data preparation in the body, and both are reported above.

An adversarial fact-check pass ran against primary sources and caught five problems, all corrected: two wrong weekdays (September 21 is a Monday, September 20 a Sunday, which also contradicted the draft's own later use of the same dates); a "four-step distilled path" phrase carried over from the secondary write-up when the H3 model card states only an eight-step res_multistep setting, so the step count has been dropped rather than guessed; a stale star count for TencentARC/WorldCrafter (148, now 150); and an overstated causal claim that ComfyUI's MoGe 3 release enabled the camera-control wave, when Run MoGe Inference already existed in core and WorldCrafter has no ComfyUI or MoGe dependency at all. The meta-pattern is now stated as convergence rather than causation. The pass confirmed the WorldCrafter licence position, the Apache 2.0 tags on Supra2-IMG, SpanSynth-Edit and the node, MoGe 3 weights as MIT and created August 18 rather than this week, the Qwen MXFP4 build as research-only, the missing licence on OpenStems, and the separation of CrossView settings between the H3 card and the node README.