FervorCreative AI
Live Latest 29.09.26 · morning 86 tools tracked 288 workflows indexed 236 topics Hot: MiniMax H3, Qwen-Image-2.1, ComfyUI

The day's releases let you describe where a piece should end up (the corrected sentence, the final frame layout, the path a sound travels) and hand the in-between to the model, which moves the creator's job from generating to directing.

EditVoiceLIFTPhysWaveMiniMax-H3 ORB360 CardSpinMiniMax-H3-QuantFunc-4bitviggle-animate-workflowInSpatio-World 1.5audio-genvoice-clonevideo-genlocal-creative-aicomfyuilora-finetuning

Creative AI Briefing: Tuesday, September 29, 2026

Record a voiceover, notice you said "Tuesday" where the script says "Thursday", and fix it by typing the corrected sentence: the new model changes the words in your recorded voice instead of making you book the booth again. That is EditVoice, and its weights went up yesterday. It fits a pattern running through the last 48 hours. The useful releases stopped asking you to describe a whole clip and started asking where it should end up. The corrected sentence. Where the car sits in the last frame. The path a siren takes around the listener. You name the destination, and the model fills in the journey. The big labs had nothing new for creators on their blogs this morning, so all of this came from researchers and the community.

New models

EditVoice rewrites the words of a recording and leaves the voice alone. DOOD02/EditVoice was created on Hugging Face September 28 at 06:17 UTC, four days after the paper (submitted September 24). You give it three things: the original audio, its transcript, and the transcript you wanted. It inserts, deletes or swaps the words that changed, and the output can come out longer or shorter than the original. That second part matters. The authors note that recent parallel speech generators typically need the output length fixed in advance, and they describe EditVoice as, to their knowledge, the first of its kind that does not. The same weights also do plain text-to-speech from a short reference clip, and voice conversion. The authors trained the editing path specifically for noisy recordings. Licence is plain: the code is Apache 2.0, and the weights are a file-by-file mix of Apache 2.0 and MIT (the editing voice renderer comes from Kimi-Audio under MIT, the rest from CosyVoice2 and the authors under Apache 2.0). Those are commercial-friendly licences, though the card says no single licence covers every file. The catch: it needs a text-processing package called ttsfrd, which the card says to "obtain them separately from an authorized source" and which the authors do not ship; its own licence terms are not covered. Until someone packages that step, installing it takes an afternoon rather than a click. It is English only, trained on 10,000 hours of GigaSpeech, and has no hosted demo. The audio samples page is the only way to hear it without installing anything. The GitHub repo shows 0 stars, so you would be among the first people through the door.

Video

LIFT lets you draw what should be in the last frame, then moves the camera there. Overdog/LIFT, created this morning at 03:10 UTC under Apache 2.0, starts from a single image. You set a camera path, then draw boxes on the final frame with a short label on each ("a red kiosk here", "a dog there"), and it generates the shot that travels from your photo to that layout. Camera-control models already let you move the camera; LIFT adds control over what you find when you arrive. It sits on the 1.3B Wan2.1-Fun camera-control model, so it is small by current standards; the card gives no memory figure, so check before you plan around it. The weights are research-grade: the code repo has 0 stars, and the project page is where the examples live.

MiniMax H3 at 4-bit, with one loader swap in ComfyUI. QuantFunc/Minimax-H3-Quantfunc-4bit (September 28, 13:55 UTC) compresses H3's weights to a quarter of their original precision, into two 12.37 GB files: a 4-step file for text-to-video and first/last-frame control, and an 8-step file for reference-driven generation. The company's own RTX 4090 figures put each step at 3.2 seconds against 10.2 for FP8. That is per step on the core model only (the INT8 ConvRot build measured 8.5 s), and those are QuantFunc's numbers, not an independent test. Audio generation is kept. A sister file for LTX-2.5 landed the same day. The catch: these use QuantFunc's own sealed file format and load only through ComfyUI-QuantFunc (393 stars), and the H3 file stays under the MiniMax H3 Community License, territory exclusions included (the LTX-2.5 file follows the LTX-2 Community License).

Viggle-Animate in one node. aireet/viggle-animate-workflow (September 28, 03:08 UTC) packages the Viggle-Animate character-animation model as a 19.6 GiB compressed checkpoint, a single ComfyUI node and a ready workflow file. You give it a driving clip and one still, and the character in the still performs the clip. Full walkthrough below.

Walk into a photograph, in a browser tab. blanchon/inspatio-world-v1.5 (September 28, 21:32 UTC) repackages the InSpatio-World 1.5 world model so a single command turns a photo into a video you steer with forward and turn moves. There is a free Space to try it. Each component keeps its own licence, and the depth component is CC BY-NC 4.0, which blocks commercial use of the package as shipped.

Audio and music

PhysWave places a sound in 3D space and moves it along a path you set. 10wind/PhysWave (September 28, 19:44 UTC, Apache 2.0, EMNLP 2026) takes a description of a sound plus waypoints, or a plain-language instruction that an outside language model turns into positions, and returns about ten seconds of first-order ambisonic audio. Ambisonics is the surround format that VR players and 360 video use. For game audio and immersive film sketches, that means a positioned, moving effect without hand-automating a panner. The catch is fidelity: output is 16 kHz, about the quality of a phone call, and one sound source at a time. The plain-language mode needs an API key: by default it sends your instruction to Gemini 2.5 Flash Lite through OpenRouter, and the README says other parsing models have not been tested. The code repo has 3 stars.

The YuE2 style add-on trend continued overnight with a quiet-storm R&B and slow-jam add-on (September 29, 01:32 UTC) from the same creator as Sunday night's folk one. It is CC BY-NC 4.0 again, so treat it as a sketchpad.

Image

Second-day tools, useful but not headline news. mlx-community/HEART-fp16 (September 28, 23:31 UTC, Apache 2.0) brings the HEART upscaler and restorer to Apple Silicon. AcademiaSD/TAE-Qwen-Image-2.1 (September 29, 04:59 UTC, Apache 2.0) is another tiny preview decoder, so you can watch a Qwen-Image-2.1 render form instead of waiting blind. ming0531/PhaSR (September 28) is a third-party mirror of the CVPR 2026 shadow-removal and lighting-evening model (the card calls itself a draft, not an author-endorsed release). No licence is listed for the weights, so do not ship with it yet.

Open and local

The local story today is about cost and packaging, not new capabilities. H3, among the most-downloaded open-weights video models on Hugging Face, yesterday it got a 4-bit build aimed at consumer cards plus a one-node character-animation package. The underlying models stayed the same. What changed is how many people can actually run them this week.

Creative workflows

1. Turn one lucky glitch into a reusable effect. Files: minimax_h3_orb360_cardspin_step50.safetensors, with the prompts in prompts/cardspin_caption.txt and prompts/orbit_realscene_caption.txt. The steps. In ComfyUI, load a MiniMax-H3 Ref2VA model and add the add-on with a standard Load LoRA node at strength 1.0 (model only). Give it one portrait as the reference and paste the card-spin prompt as the text. Keep 124 frames at 24 fps, because the prompt's timestamps (edge-on at 1.0 s, back of the card from 1.7 s, back on the photo at 5.125 s) assume that length. Try several seeds. Swap to the orbit prompt and the same file does a plain 360-degree camera orbit. How it works. The team was testing an orbit add-on on real photographs and fed it an 1867 Julia Margaret Cameron portrait. The model treated the photograph itself as the object and spun it like a card. They captioned that single generated clip with timestamps and trained 50 more steps on it, about half an hour on one card. The effect carried over to photos it had never seen. Why it is good. It is a documented recipe for keeping a happy accident. You do not need a new dataset for the effect. You need the clip, an honest caption with times in it, and a short top-up training run on an add-on that already works (here, a 750-step orbit add-on trained on four Blender renders). Where it breaks. Learned from one example, so the card always turns the same way at roughly the same speed, and some seeds commit less. The top-up run peaked at 73 GB of graphics memory, so training it needs a rented workstation card even though using it does not. MiniMax H3 Community License, and the authors state plainly: "It is not MIT."

2. Make any character perform a clip, in one node. Files: comfyui/workflows/viggle-animate-workflow.json, the ComfyUI-SlimDiT node pack, and minimax_h3_ref2va_slimdit_int8_convrot.safetensors. The steps. Clone the repo, copy ComfyUI-SlimDiT into custom_nodes, then copy the slimdit folder inside it, which is the step the authors say tripped them up. Download the three upstream files the card lists (a speed-up add-on, a text-conditioning file and the H3 video decoder), install Video Helper Suite, load the workflow and press Run. The first run reproduces the included dachshund example. How it works. Viggle-Animate reads motion from the driving footage and appearance from your still. The node also picks an attention method for each call, based on clip length and free memory: a fast path for short clips, the standard one for long clips. Why it is good. The authors measured their choices. On an RTX 5090, a 124-frame render took 33.2 s on the fast path against 38.3 s standard. On a 372-frame clip, an in-place quantized attention kernel used 31.8 GiB and showed ghosting from around frame 124, while standard attention finished clean in 26.1 GiB. So the node routes long clips, or tight memory, to the standard path automatically. Where it breaks. Their own tip is the limit: the still needs to match the driving shot's pose, framing and light, ideally as a repainted frame of that shot, or the character drifts off-model. The checkpoint is 19.6 GiB. H3 licence applies.

3. Swap one loader and run H3 at 4-bit. Files: minimax_h3_fl2va_4step_quantfunc_int4_r128.safetensors or the Ref2VA 8-step file, via ComfyUI-QuantFunc. The steps. Install or update ComfyUI-QuantFunc. Replace your model loader with the QuantFunc loader, pick the file that matches your mode, and use its matching step count and scheduler. Leave everything else in the workflow alone. How it works. The speed-up add-on and supporting parts are fused into the file, so there is nothing else to attach. Why it is good. A graph you already trust gets faster without being rebuilt. It supports every NVIDIA card from the RTX 20 series up. Where it breaks. The sealed format means no other loader will open it. The quality figure is QuantFunc's own internal comparison. Memory needs are not stated as a number, so test your own resolution before committing a job.

Worth testing

  • InSpatio-World Space: free, and you can steer through a photo in minutes. Tradeoff: the bundled depth model is non-commercial.
  • EditVoice audio samples: listen to the editing examples before you commit an afternoon to the install. Tradeoff: the samples are the authors' own picks, and the install needs a package they do not ship.
  • CardSpin on its model page: the card shows a hosted inference widget through WaveSpeed. Tradeoff: a paid provider, and results vary a lot by seed.
  • LIFT project page: see whether last-frame layout control holds up on your kind of shot. Tradeoff: no hosted demo yet, so trying it on your own images means a local install.

What actually matters from today's signal

Directing beats describing. For three years the creative loop has been: write a paragraph, roll the dice, write a better paragraph. EditVoice, LIFT and PhysWave each replace the paragraph with something an editor, a cinematographer or a sound designer already thinks in: the corrected line, the final framing, the path of a sound. That changes where your time goes. You spend it deciding what you want, not coaxing a model into guessing. EditVoice matters most for working creators this week, because a podcast or video essay fix that used to mean a re-record now means typing a sentence. Its licence lets you get paid for the result.

The counter-signal is the install. Every one of these arrived as research code with low star counts, and EditVoice depends on a package the authors will not ship. Last year's pattern will probably repeat: someone wraps these in a ComfyUI node or a hosted demo within a couple of weeks, and that is when they become tools rather than papers. If your time is expensive, wait for that wrapper. If you are the kind of person who builds wrappers, this is the week.

The H3 packaging wave is a quieter story about money. A 4-bit build and a one-node character-swap package do not add capabilities. They move H3 from "rent a card" to "maybe your card", and that decides what a solo creator can afford to try. Every H3 derivative still carries the Community License and its territory exclusions, which remains the biggest reason to read the licence before you publish.


Source access notes: Hugging Face listings fetched through the API with sort=createdAt and cache busting; every date above is createdAt. The big labs' blogs (Google, ElevenLabs, fal, Replicate, Hugging Face) showed no new creator releases in the last 48 hours. The newest items were the fal H3 Max write-up (Sept 17) and BFL's FLUX 3 Action post (about five days old), both outside the window. blog.adobe.com returned an empty page. The GitHub releases API returned nothing. The HN Algolia search returned only stale results (2014 to July 2026) and was discarded. The adversarial fact-check pass ran (Sonnet subagent). It confirmed every date, size, star count and speed figure. It caught a wrong weekday on the folk add-on, three unsupported superlatives (LIFT, CardSpin, H3), an overgeneralised EditVoice length claim, a misattributed vocoder and misquoted install line, an invented PhysWave example sentence, an H3 licence sentence that wrongly covered the LTX file, viggle ghosting attributed to the wrong kernel, and PhaSR's mirror status. All are fixed above. A WebSearch sweep turned up no lab launches from Runway, Midjourney, Luma, Kling or Suno in the window; the Suno item it found was a secondary report of a September 3 policy change, not a model launch. OpenAI, Runway, BFL, Stability, Luma, Midjourney, Comfy blog and civitai were not fetched directly this run, to save budget, and the WebSearch sweep stood in for them. Star counts come from shields.io with cache busting. Speed and quality numbers for QuantFunc and viggle-animate-workflow are the authors' own.