Creative AI Briefing: Wednesday, October 7, 2026
You can now type a line of dialogue and get back a five-second clip of a person saying it, with matching sound, from a model whose MIT license lets you sell the result. That is Kandinsky 6.0 Video, open-sourced yesterday, and it sits in a week where the open side keeps rebuilding entire paid jobs rather than single features: a Premiere-style editor, an After Effects-style compositor, a Lightroom-style raw developer, a CPU voice cloner. The bill for all of it is time. A clip that costs credits elsewhere can cost you twenty minutes on a mid-range card here at full quality, and the editors are honest that they are not ready for client work yet.
New models
Kandinsky 6.0 Video makes picture and sound together, and the license is MIT. Kandinsky Lab's README says the family was open-sourced on October 6; the Hugging Face repos were created on September 9 (API createdAt) and last updated on October 6, the day of the announcement. There are two sizes: Lite at 3 billion parameters and Pro at 29 billion (the paper says 29B; the Diffusers collection rounds to 30B). Both make five-second clips with 44 kHz audio, including lip-sync, from a text prompt or a starting image, and a separate upscaling add-on takes the result to 1920×1080. The Pro distilled card runs in 10 passes instead of 50, which is the version to start with. Every checkpoint is tagged MIT and ungated, which means commercial use with no revenue ceiling and no territory carve-out, unlike the LTX community license (free under a revenue threshold) or, as this briefing reported on September 18, MiniMax H3's (whose territory clause excludes the EU, UK, Korea and the USA). Where it runs: NVIDIA only, Python 3.13 or 3.14, plus ComfyUI nodes (kandinsky6, kandinsky6-sr) and Diffusers. The catch is the clock. The README's timing table covers the full 50-pass (non-distilled) models, measured after warmup: a Lite clip takes 437 seconds on an RTX 4090 and 1,310 seconds (about 22 minutes) on an RTX 5060 Ti; Pro at full HD is 1,247 seconds on a 4090. The README gives no timings for the 10-pass versions, which should be considerably faster, but nobody has published how much. The lip-sync and the 47% speech-error drop after reinforcement training are the authors' own claims. Try it first in the free Pro distill Space, which runs on shared ZeroGPU hardware, so expect a queue.
Image
ComfyUI v0.39.0 changes two defaults that matter to finishing work. The October 5 changelog makes 16-bit float the default for EXR saves and raises the default video encode quality (CRF 18 on h264, CRF 24 on AV1). If you grade AI frames in Resolve or Nuke, that is more latitude in the highlights without touching a setting. The same release removes the transparent-background option from the GPT Image 2 node, so any graph that relied on it for cutouts needs a new route. The open transparency models covered here in late September (Qwen-Image-2.1 and Ming-Image-Design-Layer) run on your own machine.
Custom nodes can now offer repeatable rows. The same release adds a "DynamicGroup" input so a node can show a stack of add-on rows with their strengths, the way a LoRA loader should have worked all along. It is a tool for node authors, so the benefit arrives as node packs adopt it.
Video
Pull a clean alpha matte from ordinary footage, no green screen. Lightricks' LTX-2.5-22b-IC-LoRA-Alpha-Gen was created on September 28 and updated October 1; it has not crossed this briefing until now, so treat it as a second look. Feed it an RGB clip, leave the prompt empty, and it returns a grayscale matte frame-aligned to your source, with the card naming hair, fur, smoke, fire, sheer fabric, glass and water as targets. The limits are firm and stated: 145 frames maximum before "RGB content starts leaking into the matte," 1920×1088 maximum, no way to tell it which subject you mean, and full-HD 145-frame clips need an H100 or B200-class card. License is the LTX-2 community license.
Luma Ray 2 leaves ComfyUI on October 24. The v0.39.0 changelog marks the Ray 2 partner nodes deprecated with removal on that date. If a saved graph calls Ray 2, swap the node now rather than finding out mid-deadline.
Kling 4.0 is still not public. fal's September 29 explainer says the full model launches in October with no date set, and Flash remains limited early access. Nothing from Kling's own channels confirmed a launch in the last 48 hours.
Audio and music
KittenTTS 2 clones a voice on a CPU. KittenML/kitten-tts-2 was created on September 30 and last updated October 2, so this is second-day coverage of a release that has been climbing the text-to-speech trending list. It is a 1.7-billion-parameter speech model that clones a voice from 5 to 30 seconds of one speaker, ships 47 built-in voices, including voices for nine languages beyond English, and runs on a CPU out of the box (pip install kittenml). The packed weights are about 954 MB (a 469 MB variant also exists), small enough for a laptop, though the full repo with every variant is several gigabytes. License, in plain words: free for research, non-commercial and small commercial use until you or your company pass $1 million in annual revenue or total funding, after which you need a paid license from Stellon Labs. No hosted demo exists yet. Clone only voices you have permission to use.
Eleven v4 lands inside ComfyUI. v0.39.0 adds Eleven v4 and v4 Turbo to the text-to-speech and text-to-dialogue partner nodes, so a voice line can now be generated in the same graph as the shot it plays over. Partner nodes are paid services, not local models.
Open and local
The storytold team's clean-room rebuilds of Adobe's suite keep stacking up on Trendshift. After PhotoCraft last week, three siblings started trending on October 3 and now hold four-figure star counts. All are dual MIT or Apache 2.0, all run on macOS, Windows and Linux, and all expose an MCP server, meaning an AI assistant can drive the real app rather than generate pixels. All are also openly unfinished.
- storytold/filmcraft: a Premiere-style editor written in Rust, with timeline trims, a Lumetri-style color panel and an MCP server. The README rates its feature checklist at about 87% and "ready for real work" at 50 to 60%, with H.264 as the only delivery codec (it also exports ProRes, DNxHR and image sequences). 2.2k stars per shields.io (repo).
- storytold/effectcraft: an After Effects-style compositor with 306 effects and Lottie support; first commit October 1, and the README says it "isn't yet a replacement for After Effects on client work." Cannot open .aep files. 1.1k stars per shields.io (repo).
- storytold/lightcraft: a Lightroom-style library and raw developer at about 79% by feature count, though the README puts it nearer 60 to 70% as a day-to-day Lightroom replacement; CR3 or compressed Fujifilm raws open only as embedded previews, and its "AI masks" are classical heuristics. 1.8k stars per shields.io (repo).
- storytold/photocraft: the Photoshop-style editor that started the run, now 11k stars per shields.io, up from 2.4k yesterday (repo).
- kandinskylab/kandinsky-6: code for today's lead, picture plus sound under MIT. 138 stars per shields.io, a day after release (repo).
- Lightricks/ComfyUI-LTXVideo: home of the IC-LoRA workflow the Alpha-Gen matte runs on. 4.2k stars per shields.io (repo).
- KittenML/KittenTTS: the library behind the CPU voice cloner above. 15k stars per shields.io (repo).
Creative workflows
1. Render a talking clip with sound on your own card, with Kandinsky 6.0 (just download pro-distill).
The steps. Clone the repo, run just setup (it picks the right PyTorch build for your card), then just download pro-distill and just generate "your prompt". Write the spoken line into the prompt. Outputs land in a timestamped folder alongside the expanded prompt and the settings used. For the full-HD pass, run the kandinsky6-sr upscaler, or install both kandinsky6 and kandinsky6-sr through ComfyUI Manager and work in a graph.
How it works. One stream of the model draws the frames and a second stream writes the audio, and the two keep checking each other while they generate, which is why the mouth and the sound land together instead of being dubbed afterwards.
Why it is good. Dialogue, room tone and effects in one pass, under a license that lets you deliver it to a paying client.
Where it breaks. Five seconds is the ceiling. At full quality a Lite clip takes over twenty minutes on a mid-range card; the 10-pass version is faster but unmeasured, so test it on your own card before planning a schedule, and draft ideas in the free Space. Lip-sync quality is the authors' claim, unmeasured by anyone else yet.
2. Pull an alpha matte without a green screen, with LTX-2.5 Alpha-Gen (ltx-2.5-22b-ic-lora-alpha-gen-0.9.safetensors).
The steps. Put the file in models/loras. Load the ltx-2.5-22b-distilled model and the LTX-2.5_V2V_ICLoRA_Single_Stage_Distilled.json workflow from ComfyUI-LTXVideo. Connect your clip as the reference video at strength 1.0, set the add-on to 1.0, leave the prompt empty, and run only the first stage at native size. Keep clips at 145 frames or fewer, with width and height divisible by 32; pad rather than stretch.
How it works. The add-on teaches the video model to answer "here is a shot, draw what belongs in front" as a black-and-white video, so the matte comes out as a second clip lined up frame for frame with yours.
Why it is good. It aims at the material keyers hate: flyaway hair, smoke, glass, water.
Where it breaks. The model picks the foreground itself, so a shot with two subjects may give you the wrong one. Long clips leak picture into the matte. Full HD needs data-center hardware, so most people will rent a card or work at lower resolution.
3. Clone a narrator's voice on a laptop with KittenTTS 2.
The steps. pip install kittenml. Record 5 to 30 seconds of clean speech from one person who has agreed to it. Call m.generate("Your line here.", reference="my_voice.wav") as the card shows. Try the built-in voices first to judge quality before cloning.
How it works. The model listens to your sample as part of the request and continues in that voice, so there is no training step and no upload.
Why it is good. Scratch narration for an animatic, offline, on a machine without a graphics card.
Where it breaks. The license caps commercial use at $1 million in revenue or funding. No published real-time figures on CPU (the card says its llama.cpp-based C++ build is the fastest route), and no hosted demo to hear it before installing.
Worth testing
- Kandinsky 6.0 Pro distill Space: type a line of dialogue and hear whether the mouth matches. Tradeoff: shared free hardware, so queues, and only five seconds per clip.
- LTX-2.5 Alpha-Gen: run your worst hair shot through it at half resolution. Tradeoff: no subject selection and a 145-frame cap.
- EffectCraft: rebuild one lower-third you already made in After Effects. Tradeoff: no .aep import, no plug-ins, a week old.
- ComfyUI
--offlineflag (v0.39.0): per the ComfyUI README, it disables the paid API nodes and keeps built-in functionality offline, which suits NDA jobs. Tradeoff: every partner node (Eleven v4, GPT Image 2 and the rest) stops working while it is on.
What actually matters from today's signal
The open side is no longer shipping parts. In one week it has shipped a sound-on video model, a matte generator, a voice cloner, and four editors that each copy a whole Adobe app. Between them they touch every stage from first frame to final export, with no subscription attached. That is new, and it changes the question from "can I do this without a subscription" to "how much of my week am I willing to spend waiting and patching."
The waiting is real. Twenty-two minutes for five seconds of full-quality Lite on a mid-range card is a render-farm habit, not an iteration habit. The patching is real too: every storytold README says some version of "not yet for client work," and LightCraft's own notes say its color calibration is neutral for non-DNG raws. Budget these tools for drafts, personal projects and the one job where the license matters more than the turnaround.
The license is the part to watch. Kandinsky's MIT tag carries none of the revenue caps or territory clauses attached to the LTX and MiniMax H3 licenses this briefing has tracked, and the counter-signal sits in the same changelog: GPT Image 2's transparency vanished from a node with one bullet point, and Luma Ray 2 goes in seventeen days. Hosted features disappear on someone else's schedule. Open weights you have downloaded do not.
Source access notes: Bucket A was quiet: OpenAI (no creative posts Oct 5–7), Google blog and DeepMind (no Veo/Imagen/Lyria/Flow posts in window), Adobe (latest Oct 2; MAX is November 10–12), ElevenLabs (latest Sep 30), BFL (latest Sep 23), fal blog (latest Sep 17), Replicate blog (latest April), Suno (Albums editorial Oct 5, already covered yesterday). Runway's news page redirected to runway.com and was not re-fetched. Stability's news page carries no dates. Luma's blog index is stale (2025). Midjourney not checked directly; WebSearch surfaced nothing dated this week. HN Algolia search_by_date returned August 28 stories (stale cache), so HN was not used. The GitHub API and gh were blocked (403); repo creation dates come from READMEs and Trendshift first-trending dates, and the shields.io created-at badge returned only "september". GitHub commit pages are blocked by robots.txt. The HF API sort=createdAt listing served stale (September) results, so discovery used trending pages with every date checked through the per-repo API. The ComfyUI blog and civitai were not attempted. Writing section omitted (nothing creator-facing). PetaPixel's Skylum Aperty shutdown story was dropped because its dates were internally inconsistent. Adversarial fact-check ran (Sonnet subagent) and caught: the Kandinsky timing table describes the 50-pass models, not the distilled ones (now stated, with no distill timings claimed); an inferred "private for a month" (removed); FilmCraft exports more than H.264 (now "only delivery codec"); LightCraft's 60 to 70% day-to-day figure (added); KittenTTS download size overstated as the default (corrected); an unsourced billing claim for Eleven v4 nodes (removed); the --offline behavior now cited to the ComfyUI README; and two superlatives softened. The KittenTTS cloning call was re-verified against the card.