Creative AI Briefing: Friday, October 9, 2026
You can now write an animated explainer in Claude, open it in Luma as a video, and have Luma restyle it or reframe it to 9:16 without re-animating a frame. That handoff shipped yesterday, and it rhymes with everything else this week. Magnific's new image model hands you a grid of 8 or 16 drafts for the price of one finished image. Vidu's new video preview starts at $0.014 a second. In one run of a new public benchmark, a model costing about a cent wrote a working motion-graphics video on the first try. The draft has become the cheap part. What you pay for now, in money or in hours, is the one version that ships.
New models
Vidu Q4 Preview puts 15 reference images and 3 voice clips into one shot. ShengShu released it October 7 as a public preview on the Vidu web app and API. What you can make: a clip of up to 16 seconds where the same character keeps their face and their voice, with sound generated alongside the picture. The API docs accept up to three MP3 voice references of 3 to 12 seconds each, output from 540p to 4K, in five aspect ratios (no 21:9). In the release's words, "Launch pricing starts at $0.014/sec"; it does not say which resolution that buys, and adds that final pricing and terms "may vary by plan and region." The catches: it is a preview whose final behavior may change, the docs set file limits on voice references but say nothing about consent, and the commercial-use terms are not stated on the product page. Read the plan terms before client work.
Magnific One adds an art-direction step before the image. Launched October 8 (VentureBeat; Magnific's own page confirms the draft grid, 4K, Brand Kit and MCP). You work in two gears: Draft gives "a grid of 8 or 16 options," then you take one to Final at up to 4K. A Brand Kit holds a client's palette, type, imagery and tone, and VentureBeat reports it checks outputs against the kit and flags mismatches. Up to five reference images, per VentureBeat. It also works from Claude and ChatGPT through MCP (a connector standard). The catches: it is subscription-only (VentureBeat lists Premium at $14.50 a month billed annually, and estimates roughly $0.12 to $0.73 per image assuming you use every credit), and VentureBeat flags a conflict it could not resolve: Magnific's CEO told it the model is "powered by GPT-Image-2" with some post-training, while the FAQ says "Nothing was trained." Through October 15, eligible higher tiers get unlimited High-quality 2K and 4K generations in the desktop app; Max quality still costs credits.
Iris-3B is an open image model whose authors published that their big idea did not pay off. Speridlabs put the weights up October 5 (Hugging Face createdAt 14:51 UTC) and the paper followed October 7. Iris draws pixels directly instead of working in the compressed stand-in format that most image models use, which in theory should keep fine detail. The same backbone also ships fine-tuned for depth maps and 4x photo restoration. License: Apache 2.0, so you can use it commercially. The honest catch is in the abstract: "We find no significant improvement from using a pixel-space generative prior" on depth or restoration. On the authors' own scores it ties Qwen-Image on the OneIG benchmark (0.540 vs 0.539) but trails it on lettering (0.857 vs 0.943). Each of the three weight sets is about 12 GB, it needs an NVIDIA card, and the default is 100 passes per image, so expect to wait. Try it free in the Space.
Image
Rembrandt is a free, open Lightroom-style raw developer with on-device AI masks. The repo was created September 29, has shipped ten releases since October 4 (v0.3.7 on October 7, v0.3.8 and v0.3.9 on October 8), and reached Show HN on October 8. It reads raw files from "1,000+ cameras," makes subject, sky, background and depth masks on your own machine, and adds AI denoise, 2x and 4x enlargement and relighting. The "ask in words" box is "a vocabulary, not a language model," so it matches set phrases rather than understanding requests. GPL-3.0, runs on Mac, Windows and Linux, imports Lightroom Classic catalogs. The catch: builds are not code-signed yet, so your OS will warn you, and it is about ten days old.
Video
Claude Motion videos now open directly in Luma. Luma announced it October 8. Claude Motion, in beta on Claude Team and Enterprise plans, writes code that animates text, charts, shapes and images. Add the Luma connector and the animation "will appear in Luma as a video," where Luma's Ray and Uni models can restyle the look while keeping the motion, reframe to 9:16, 1:1, 4:3 or 21:9, or spin out new versions. Luma's post gives no pricing, and individual Claude plans are not in the beta.
Prism now runs on a 24 GB laptop, slowly. Tencent's picture-plus-sound video model got a ComfyUI port on October 6 (createdAt 11:53 UTC), MIT, compressed to 8-bit numbers so its peak memory use was about 7.76 GiB in the author's run. The catch is time: one sample clip took about 47 minutes on an RTX 5090 laptop, the download is about 38 GiB, and the author says the sample's audio "has not been listened to yet." Second-day work, honestly labeled.
Audio and music
Stable Audio 3 Small SFX now serves from a Mac app with no Python. A mirror created October 8 lets the mlx-serve app pull Stability's sound-effects model and answer requests on your own machine. The page claims "up to two minutes of 44.1 kHz stereo audio in about a second on Apple Silicon," without naming a chip. The catch: the mirror exists to skip Stability's sign-up gate, but the Stability Community License still applies (free commercial use under US $1M annual revenue, with registration).
VoTSpeech designs a voice from a written description. Weights landed October 9 (createdAt 06:53 UTC): a 2B model making 48 kHz Chinese and English speech from text plus instructions like the voice you want. CC BY-NC 4.0, so no paid work. Three stars, no paper, no demo yet.
Open and local
The local story today is breadth rather than a headline model: a raw developer, a pixel-drawing image model, a sound-effects server and a laptop path for picture-plus-sound video, all free to run, each with an honest catch about time or polish.
- speridlabs/iris-3b: Apache image model plus depth and 4x restore from one backbone, with code. Paper out October 7; 60 stars (shields.io). (repo)
- thesnarkitecht/rembrandt: free raw developer with local AI masks and denoise. Three releases in two days, Show HN October 8; 70 stars. (repo)
- yi1108/printfilm: self-hosted pipeline that turns a script into storyboard, images, video and a final cut for short dramas, MIT, Chinese docs, needs a paid upstream key. On Trendshift's daily list, tagged new; 5.1k stars. (repo)
- ddalcu/mlx-serve: native Apple Silicon server that now pulls Stable Audio 3 Small SFX for local sound effects. New SFX mirror October 8; 1.8k stars. (repo)
- T8mars/Comfyui-Prism-T8: ComfyUI nodes that run Prism's joint video-and-audio pipeline. Paired weights posted October 6; 14 stars. (repo)
- ywbn/VoTSpeech: instruction-driven voice design code, non-commercial weights. Weights posted October 9; 3 stars. (repo)
Creative workflows
1. Turn a written explainer into a reframed social cut: Claude Motion to Luma (Luma's launch post)
The steps. Build the animated explainer in Claude Motion. Add the Luma connector in Claude. Open the animation in Luma, where it arrives as a video. Ask Luma's agent to restyle it while keeping the motion, then ask for 9:16 and 1:1 versions for social.
How it works. Claude writes the animation as code, so timing and layout are exact and editable. Luma treats the render as source footage and runs its video models over it.
Why it is good. You get precise motion design (charts that hit the right numbers, type that lands on the beat) and a generated finish, without animating twice.
Where it breaks. The beta is Team and Enterprise only. Once the piece is in Luma, it is a video: change a number in the chart and you go back to Claude and re-export. Pricing on the Luma side is not published in the post.
2. Picture-plus-sound video on a 24 GB laptop with Prism-Comfy (model card, nodes)
The steps. Install the Comfyui-Prism-T8 nodes. Download the seven model files (about 38 GiB) into models/diffusion_models, models/text_encoders and models/vae as the card lists. Load 01_native_i2va.json: one image in, 848x480, 49 frames at 24 fps (about two seconds), separate prompts for picture and sound. Turn on chunked CPU offloading and leave it running.
How it works. Offloading swaps pieces of the model between graphics memory and system memory, trading speed for fit.
Why it is good. It is a documented, MIT-licensed path to Prism's synced sound on a laptop.
Where it breaks. About 47 minutes for one two-second clip in the author's run. The author reports slight composition drift and soft texture and has not listened to the sample's audio. 720p and long clips are untested.
3. Depth maps and 4x restores from Iris-3B (README)
The steps. Clone the repo and pip install -e .. Download the depth/ or upscaler/ folder. Run python scripts/depth.py photo.jpg --out depth_out for a depth map, or python scripts/upscale.py photo.jpg --out upscaled for a 4x restore.
How it works. The same backbone that makes images was fine-tuned for each job.
Why it is good. A depth pass for fog, relighting or parallax, plus a restore, from Apache-licensed weights.
Where it breaks. Depth is relative, not measured distance. The authors found the pixel approach no better than a standard model for these jobs, and each folder is another 12 GB.
Worth testing
- Iris-3B Space: free text-to-image, depth and restore in a browser. Tradeoff: 100 passes per image by default, so it is slow, and lettering trails Qwen-Image on the authors' own test.
- Vidu Q4: test one character with a voice reference across three shots. Tradeoff: preview pricing and terms may change, and commercial rights are not stated on the page.
- Magnific One: try Draft mode on a client brief. Tradeoff: subscription only, the October 8 to 15 unlimited offer covers only higher tiers in the desktop app, and the underlying model story is unresolved.
- Rembrandt: run one shoot through it next to your usual raw developer. Tradeoff: unsigned builds and about ten days of history.
- cliphou.se: free motion-graphics templates built for coding agents, editable in the browser with no watermark. Tradeoff: no license terms beyond "free to remix," and export formats are not specified.
What actually matters from today's signal
The price of a first attempt is collapsing on every front at once. Magnific sells drafts in grids of 16. Vidu launches at a cent and a half a second. Cliphouse Research's benchmark, dated October 8, had ten language models write the same motion-graphics briefs as code, and in the product-launch run shown on the page Claude Haiku 5.5 passed first try for $0.011, against $0.374 for Opus 5.5. The page calls two runs "exploratory evidence," so read that as direction, not ranking. Still, a working animated draft for about a cent changes what a designer does with an afternoon.
That moves your budget. When drafts are nearly free, the scarce thing is a finishing pass that does not throw the draft away. Claude-to-Luma is that idea as a product: precise motion from code, then a generated look on top. Magnific's Draft-to-Final with a Brand Kit is the same idea for stills. The tools to bet on are the ones with a clean handoff, because rebuilding a draft by hand eats exactly the time the cheap draft saved.
The counter-signal sits on PetaPixel today under the headline "No, I Don't Want to Restyle My Photo With AI". Cheap drafts produce a lot of things nobody asked for. And the open side is honest about its bill: Prism costs 47 minutes a clip on a laptop, and Iris's own authors say their headline approach bought nothing on depth or restoration. Cheap to start is not the same as cheap to finish.
Source access notes: blog.adobe.com Firefly topic page returned empty; Releasebot showed no Adobe creative updates after September 30. blog.google listing carried no dates. runway.com/news returned no listing (changelog used: nothing new since October 2). blog.comfy.org JS wall. ComfyUI GitHub releases API and github.com/trending not fetched (API errors); Trendshift showed values without labeled deltas, so only shields.io totals are cited. fal.ai/models blocked by robots. Replicate blog newest post April 15. openai.com, elevenlabs.io, bfl.ai, stability.ai, suno.com, HeyGen and Krea had nothing creator-facing in the window. Midjourney checked via Releasebot (secondary): October 1 alpha changelog only. Kling 4.0 full release still unconfirmed by a primary source. civitai not attempted. HN Algolia dated search used for discovery. Magnific's model provenance is reported by VentureBeat and not resolved by Magnific's page.
Adversarial fact-check ran (Sonnet subagent). It confirmed the Vidu, Luma, Iris, Prism-Comfy, audio and star figures and caught six problems, all fixed: Rembrandt's age (about ten days, not three weeks) and release count; an unsupported "first" claim on Prism-Comfy; a reworded Vidu price quote; the benchmark generalized beyond one run in the opening; Magnific's GPT-Image-2 claim attributed to the CEO interview and the unlimited promo's tier, app and quality limits; and "block" vs "chunked CPU" offloading. The PetaPixel headline and Rembrandt Show HN were confirmed from this run's own fetches of petapixel.com and HN Algolia (Show HN, 2026-10-08 21:00 UTC).