FervorCreative AI
Live Latest 20.09.26 · morning 63 tools tracked 178 workflows indexed 177 topics Hot: MiniMax H3, ComfyUI, LTX-2.5

Two independent publications yesterday did the thing the add-on boom skipped, which is measure whether the add-on is doing anything at all, and one of them found that five of six popular adapter formats load nothing and serve you the plain model without saying so.

dptlabFLUX.2 Klein 4BMiniMax H3 Multishot WorkflowMiniMax H3YuE2Z-Image Turboimage-genvideo-genmusic-genlora-finetuningcomfyuilocal-creative-aiopen-weightscreative-workflows

Creative AI Briefing: Sunday, September 20, 2026

Type a three-shot scene into one text box, queue it once, and get back a single video where the same face and the same voice carry across all three shots with no visible cut. That workflow published yesterday with its measurements attached, which is the day's actual story. Two people, working separately on different media, spent their release not on a new capability but on finding out whether the add-ons everyone is now downloading actually do anything. One ran six adapter formats head to head on the same image model, same data, same budget, and reported a result nobody wanted: five of the six load nothing through the standard loader and hand you the plain base model without an error. The other measured drift across chained video shots and found that the dials most people tune are operating on the wrong thing.

New models

Neither headline item today is a model, and both are more useful than most models.

dptlab's six-way adapter benchmark (OnePunchMonk101010, six repos, Hugging Face createdAt September 19, 17:34 to 17:36 UTC) takes FLUX.2 Klein 4B (Apache 2.0, createdAt January 14, 3.88 billion parameters, 408,799 downloads) and trains the same subject six different ways: LoRA, DoRA, LoHa, LoKr, OFT and BOFT. Identical data, identical budget, 500 steps at 512 pixels, seed 42, from the SynCD multi-view dataset. The job is subject personalization: hand the model a few photos of one object, get it back in scenes it has never seen. Every checkpoint is Apache 2.0 and the code is MIT.

Method Prompt (CLIP-T) Subject (DINO) Trainable params Latency
LoRA 0.9646 0.4320 13.8M 1558 ms
DoRA 0.9572 0.4424 14.5M 2289 ms
LoHa 0.9710 0.2907 13.8M 2017 ms
LoKr 0.9502 0.2946 0.9M 2104 ms
OFT 0.9176 0.2750 1.4M 1596 ms
BOFT 0.8727 0.2680 3.0M 2441 ms

Read the columns together, as the cards insist: an adapter that learned nothing scores well on the first, one that memorized its training shots scores well on the second. LoRA and DoRA are the only two that move both, and DoRA's small subject win costs 47 percent more time per image. The four exotic formats all land near the "learned nothing" end, for different reasons: LoKr, OFT and BOFT were sold on training far fewer parameters, while LoHa trains the same 13.8 million as LoRA and was sold on packing more into them.

Then the part that matters more than the table. Five of the six cards carry the same warning: this format is not something pipe.load_lora_weights() can read, and if you try, it logs "no LoRA keys found" and silently serves the base model. The repository extends that to deployment, where no inference server loads a non-LoRA adapter either. Caveats are on the page too: the LoRA card flags an unresolved warning about missing target modules and calls its own numbers provisional, the Status section still says no training run has happened (stale, given these checkpoints), and the repo has zero GitHub stars, which is to say nobody has checked this yet. Free demo of the base: black-forest-labs/FLUX.2-klein-4B.

MiniMax H3 Multishot Workflow (LaDruid/MiniMax-H3-Multishot-Workflow, createdAt September 19, 20:29 UTC, Apache 2.0, DOI 10.57967/hf/10523) is four ComfyUI graphs plus a node pack (jlucasmcrell/ComfyUI-H3-Multishot, 62 stars) that chain H3 shots into one continuous take with the voice and face held across the joins. What sets it apart from every other workflow drop is that its settings file is a measurement log: each dial carries what it does, the shipped default, and what breaks if you move it, and three dials are marked as doing nothing under the default continuity mode, so nobody wastes an afternoon tuning dead controls. The pack is 10 MB; the weights are the GGUF H3 files you probably already have, Q8_0 for 32 GB cards, Q5_1 for 24 to 32 GB, Q4_0 for 16 GB. Needs ComfyUI v0.30.0 or newer. The CORE graph runs on stock ComfyUI with nothing else installed; the full graph wants five third-party packs. The catch is time: on a 24 GB 3090 at the shipped 736x1280 and 192 frames, later shots run around 98 seconds a step, roughly 23 minutes per shot, because the reference payload pushes most of the weights off the card.

Image

OrionLLM/Photon-P2.1 (createdAt September 19, 18:07 UTC) is a style add-on for Z-Image Turbo aimed at photographic lighting: softer falloff, subjects sitting in their environment rather than floating on it. Apache 2.0, commercial use fine, and it keeps the base model's few-step speed. There are no numbers and no comparison grid, so every claim is the author's. Try the base free at mrfakename/Z-Image-Turbo.

The volume story yesterday was identity. One account, AiMamis, published more than thirty Krea 2 Turbo character add-ons between 18:17 and 20:47 UTC, most of them named people. Nothing technical distinguishes them, which is the point: training a likeness has become cheap enough that it happens in batches, and the licensing across the day's character files ranges from OpenRAIL++ to non-commercial to nothing declared at all.

Video

weihang44/Causal-Forcing-Memory-Checkpoints (createdAt September 20, 02:09 UTC) publishes four training checkpoints of a context compressor for long video, at a memory ratio of 0.1, meaning the model keeps a tenth of the history it would otherwise carry. Four steps of one run, 5.682 GB each, 22.73 GB total, with SHA256 checksums and a pinned commit of the inference repository needed to load them. Research plumbing rather than a creator tool: no license is declared and the loading path is three models deep. It signals where the effort is going, because long-video drift is what the workflow above spends its whole settings file fighting.

Audio and music

The YuE2 wardrobe grew a fifth family. becausereasons/yue2-qwwl-qawwali-sufi-tabla (createdAt September 19, 17:26 UTC) ships two add-ons, 177 MB each, one for live qawwali as a party performs it and one for the same voice in dark studio production. Both trained on a single RTX 5090, both CC BY-NC 4.0, so they are for your own work and not a client's. The card is a failure log as much as a release. The first attempt was one file trained on a 71-item mix and it came back as film pop. What fixed it, in order: one sound per file rather than three averaged together (planner loss fell 0.33 on the mixed set against 0.71 and 0.42 on the focused ones), building from the purest live recordings rather than studio crossover material, and switching the score mode off, because YuE2 writes itself a melody in notation first, and the base model has never heard qawwali, so it writes a tidy four-bar pop tune and the voice follows it no matter what your add-on learned. One finding saves anyone a week: real tabla passed through the model's audio codec alone, no generation, came back with every frequency band within 1 dB. A weak tabla is a training problem, not a codec problem.

Open and local

The local scene is producing measurement now, not just repacks. The compressed-file crowd is still busy, but the two items people will still be using in a month are both documentation:

Creative workflows

1. Chain three video shots into one take that holds the same face and voice. Direct link: LaDruid/MiniMax-H3-Multishot-Workflow. Start with workflows/H3_Seamless_Chain_CORE.json, which needs nothing but the pack and stock ComfyUI.

The steps. Copy ComfyUI-H3-Multishot into ComfyUI/custom_nodes/ and restart. Pull the H3 GGUF checkpoint at the quality your card takes, plus the text encoder and both VAE files (there are two, video and audio) from Comfy-Org/MiniMax-H3. Open the CORE graph. Type your shots into the sampler's script box, one per shot, with --- on its own line between them. Leave continuity on first_frame, seed_per_shot on, and the VRAM reserve at 0. Queue once.

How it works. Each shot conditions on the tail of the one before it, so the model is handed the previous shot's ending rather than starting fresh. Identity lives in that conditioning, not the seed, which is why varying the seed per shot holds the face while reusing one seed everywhere drifts both the face and the voice.

Why it is good. Every number is attached to a decision. Setting memory_frames to 0 drops brightness drift from 1.055 to 1.022 per join and colour drift from 1.086 to 1.039, and stops the drift accelerating. The memory reserve sizes itself: a hand-set value that suited shot 1 was wrong for shot 2, left the model 399 MB short, and collapsed a render from 18.8 seconds a step to 283. Capping it turned a 37-minute chain into 14.

Where it breaks. Texture ratchets about 1.3x at every join and keeps climbing, with no honest fix after the fact, because blur is the only lever and it destroys real detail. Set chain_gain_control to flatten past five shots and accept one house texture. Resolution cannot change mid-chain. Per-shot colour and brightness corrections do nothing under the default continuity mode, because the drift rides a raw latent stored before decode while those dials operate on decoded frames; correct the finished master instead. Long-take mode still measures about 13 percent extra fine texture per join, so keep extended takes to roughly four windows.

2. Find out whether your image add-on is actually loaded. Direct link: flux2-klein-peft/ in dptlab, specifically scripts/merge_and_export.py.

The steps. Load your base model, load the add-on, render one image. Then set the adapter strength to zero and render the same prompt at the same seed. Identical images mean the add-on was never loaded. For any format that is not LoRA, skip the test and run merge_and_export.py first, which folds the adapter into the weights and writes an ordinary model directory every tool can read. If you compress afterward, merge first and compress second, never the reverse: an adapter is a difference against the exact weights it was fitted to, so compressing the base shifts those weights out from under it.

How it works. The standard loader looks for a specific key naming scheme in the file. LoRA files have it. DoRA, LoHa, LoKr, OFT and BOFT do not, so the loader finds nothing, logs a line most people never read, and proceeds with the untouched model. You get an image. It looks fine. It is not your image.

Why it is good. Two renders, and it catches a class of failure that produces no error at all. It also explains a forum complaint you have seen: "my trained file does nothing."

Where it breaks. The zero-strength trick only works on formats the loader can read in the first place, which is the joke. And merging costs you the thing that made add-ons useful, because a merged model is a full copy per subject rather than a small file you swap per job.

Worth testing

  • black-forest-labs/FLUX.2-klein-4B: the base model from the six-way sweep, free, in a browser. Tradeoff: a hosted queue, and it does not run the add-on comparison that makes the sweep interesting.
  • prithivMLmods/FLUX.2-Klein-LoRA-Studio: community add-ons on Klein with no local install, and a working illustration of the LoRA-only constraint, since everything it offers is in that one format. Tradeoff: it runs the 9B Klein, not the 4B the benchmark used, so those six checkpoints will not load here.
  • mrfakename/Z-Image-Turbo: the base under Photon P2.1, free and fast. Tradeoff: you are judging the base, not the style add-on, which publishes no comparison of its own.
  • workflows/H3_Seamless_Chain_CORE.json from the multishot pack: three chained shots on stock ComfyUI, no third-party packs. Tradeoff: local only, and a 24 GB card spends about 23 minutes a shot at the shipped resolution.
  • becausereasons/yue2-qwwl-qawwali-sufi-tabla: qawwali and dark Sufi fusion for YuE2, with demos and prompt files included. Tradeoff: non-commercial only, and the author judged everything by ear, because speech recognition is useless on this material.

What actually matters from today's signal

The add-on boom has been running about a month and until yesterday nobody had published the obvious control experiment. The result is worse than "some formats underperform." Five of six formats fail open. You train a file, you load it, you get an image, and the image is from the plain model. No exception, no red text, one log line about missing keys that scrolls past. If you have ever trained a style or a likeness and concluded the training did not take, check the format before you blame the training.

That points at something structural. LoRA has an advantage the score table does not show: it is the only format the whole ecosystem reads. Diffusers reads it, ComfyUI reads it, hosted services read it, and it stays a small swappable file rather than a full model copy. Everything else must be merged down before it can be deployed, at which point it stops being an add-on. When two methods tie on quality, that asymmetry is the whole decision, and it is why a format being smaller during training does not matter if it is larger in practice at delivery.

The counter-signal is sample size. Six prompts, one subject, one seed, from a repository with zero stars and a stale status section. The author says so on every card, which is why the work is worth reading at all, and why you should run the two-render check yourself rather than take the ranking as settled. The useful thing here is not the leaderboard. It is the test.


Source access notes: Every Hugging Face date comes from the model API's createdAt, never a listing "Updated" timestamp; all API calls carried a cache-busting parameter. Star counts come from shields.io with cache busting. The GitHub API returned empty responses for OnePunchMonk/diffusion-post-training-lab and for raw RESULTS.md at both main and master, so the six-method comparison is built from the six model cards and the repository README fetched via raw.githubusercontent, not from RESULTS.md. blog.comfy.org served a JavaScript wall as usual. The multishot workflow's README exceeded the fetch size limit; INSTALL.md and SETTINGS.md were read in full instead. Nothing creative-adjacent in the window from OpenAI (latest was advertising, September 16), Adobe, Google DeepMind, Runway, Black Forest Labs, Stability, Luma, ElevenLabs, Suno, or fal (whose September 17 H3 Max post was covered two days ago). Web search surfaced no vendor launches dated September 19 or 20. Writing section omitted: nothing in the window affected content creation.

An adversarial fact-check pass ran before publication and caught seven problems, all corrected above: a download count stated as 2,398 when the API returns 4,019; a "most-downloaded" superlative the evidence contradicted; a hosted-demo link offered as a way to test the benchmark's checkpoints when that Space runs the 9B Klein and the benchmark used the 4B; a per-shot render time given as 23 minutes in one section and 20 in another; a caveat about the evaluator model attached to the score table when it belongs to an optional judge never run for this sweep, now removed; LoHa grouped with the small-parameter formats when it trains the same count as LoRA; and an unsourced count of character add-ons, now attributed with a link and a verified time window. Every benchmark figure, date, license, drift measurement, training-loss number, file size, star count and Space link verified against primary sources.