FervorCreative AI
Live Latest 07.10.26 · morning 117 tools tracked 403 workflows indexed 302 topics Hot: MiniMax H3, ComfyUI, ArtCraft

The weekend's open releases were finishing tools rather than generators, so cutting out a subject, finding a shot in a drive full of footage and upscaling a still now run on your own machine, mostly under licenses you can work with.

PlateExtractSCM Screen MemoriesFastH3NacreArtCraftMadStickArtPrismlocal-creative-aiai-editingupscalingvideo-genopen-weightscreative-workflows

Creative AI Briefing: Tuesday, October 6, 2026

Shoot a frame with your subject in it, shoot the same frame empty, and a free add-on will now hand you the subject as a transparent PNG with soft hair edges intact. That is PlateExtract, and it sets the tone for the last 48 hours. The big labs posted nothing new for image, video or sound, while the community shipped the dull, valuable middle of a creative job: cutting things out, finding the shot you know you filmed, sharpening a small image, sketching a storyboard. Almost all of it runs on hardware you already own.

New models

PlateExtract turns a clean plate into a transparent cutout. trmz/plate-extract-qwen-image-2.1 was created on October 4 at 10:10 UTC by Xavier Jara. You give it two images: the composite (subject in the scene) and the clean background (same scene, no subject). It returns the foreground as an RGBA PNG with "soft edges and partial transparency," which is the part a compositor cares about, because hair, smoke and glass are where hand masks fall apart. The example files in the repo cover an anime frame and a watercolor, so it is not only for photos. It is an add-on for Viggle's fast four-pass build of Qwen Image 2.1, runs in four passes, and the author tested it on a 16 GB NVIDIA card with the model partly parked in system memory and 64 GB of RAM. One command runs it: python inference/qwen_extract.py --composite picture.png --background clean_background.png --output extracted.png. The catch is the license. The add-on weights sit under the Qwen Research License, so client and paid work are off the table; only the inference script is MIT. No hosted demo and no ComfyUI node yet.

FastH3 V2 puts MiniMax H3's picture-plus-sound video on a single RTX 5090. Hao AI Lab's FastVideo-FastH3-8-Step-V2-NVFP4-Consumer was created on October 5 at 19:01 UTC. It compresses the team's eight-pass FastH3-8-Step-V2 (created September 4) to four-bit numbers. H3 makes video with synced sound from a text prompt; V2 does it in eight passes instead of the base model's long schedule, and the consumer build targets "a single Blackwell GPU (RTX 5090, RTX PRO 6000) or a DGX Spark." If your machine has under 64 GB of system memory, the card tells you to grab a 155 MB precomputed table (adaln_tables.pt) so you skip loading 26 GB of timing data. The honest catches: the V2 card warns that "Difficult motion, fine detail, and some audio may remain below the base MiniMax H3 model" and that the first-and-last-frame and reference modes were not sped up; the consumer card publishes no speed figures at all; and V2 inherits the MiniMax H3 Community License, whose territory restrictions earlier briefings have flagged. No hosted demo.

Image

Nacre is a 4× photo upscaler that is open all the way down. xocialize/nacre-v1 (created October 5, 19:38 UTC) is a small model, 118.6 million numbers in its main network, that blows a degraded photo up four times in four passes. What makes it different is the paper trail: it was trained on 122,021 tiles cut from 51,759 Wikimedia Commons photos under CC0, public-domain or CC BY terms, with an ATTRIBUTION.csv crediting every one, and the weights are Apache 2.0 with an MIT-licensed image compressor. You can use it on client work. The author's own benchmark puts it level with ResShift v3 (23.98 against 23.93 on a fidelity score for the RealSR test set), which is a modest claim made honestly. Mac users get an MLX port. The card's own warnings: "Generative detail is invented, not recovered," it "damages text that was already legible," and it was never trained on motion blur.

UltraSharp V2 Lite now runs on Apple silicon, for personal projects. swiftail/UltraSharpV2-Lite-onnx-coreml (created October 6, 09:01 UTC) re-exports Kim2091's popular 4× upscaler so it runs through Core ML at about one second per 544-pixel tile on an M3 Pro, per the card. The license is CC BY-NC-SA 4.0, so it is non-commercial. Pair it with Nacre when you need the commercial option.

Video

Prism makes 2K video with dialogue, effects and score, if you own a render farm. Tencent Hunyuan, Fudan and Zhejiang University published Prism code and preview model files (FrancisRing/Prism, MIT). The Hugging Face repo was created on September 24 and its card refreshed today; mirrors appeared on October 5, and the README dates its release only as "2026-x-xx," so we cannot pin the ship day. It turns an image and a prompt into video with matched sound at up to 2560 × 1440. The catch is the hardware: 720p needs one 80 GB data-center card, and 1080p or 2K need four or more. Research to watch, not a tool for this week.

Audio and music

Two music models get Mac builds. Avdpro/Stable-Audio-3-Small-MLX and Avdpro/ACE-Step-1.5-Turbo-MLX, both created October 5, convert Stability's music-and-effects model and ACE-Step's fast song model for Apple's MLX framework. The ACE-Step build is about 10 GB and MIT licensed; the Stable Audio build carries the Stability AI Community License plus Gemma terms, which requires registration for commercial use. Neither card publishes speed figures, so time a cue yourself before planning a session around it.

Suno says Albums has launched. An October 5 editorial post describes Albums as "a way for you to turn your best music into full-length projects," but gives no launch date, plan details or how-to. Treat it as a pointer, not documentation.

Open and local

The local story today is post-production. SCM searches your own footage by description on a Mac. PlateExtract and Nacre handle the cutout and the upscale. Hao AI Lab's FastH3 V2 brings audio-video generation to one consumer card. And the team behind PhotoCraft now lists seven open-source desktop apps, which drew a long and skeptical Hacker News thread.

  • allenv0/SCM: search every photo and every frame of video in a folder by describing it, with jumps to the timecode, plus search by on-screen text and spoken dialogue. Repo created October 3, Show HN on October 4 (166 points at last check), MIT, 389 stars per shields.io (repo).
  • storytold/photocraft: the free, Photoshop-compatible editor an agent can drive. Now 2.4k stars per shields.io, up from 1.3k yesterday (repo).
  • storytold/artcraft: the same team's AI filmmaking desktop app. On October 4 the team posted ArtCraft Apps, a lineup of seven apps (PhotoCraft, VectorCraft, FilmCraft, LightCraft, PrintCraft, EffectCraft, DesignCraft). Only PhotoCraft and PrintCraft are labeled "early alpha"; the rest say "in development." 2.6k stars per shields.io (repo).
  • hao-ai-lab/FastVideo: the toolkit that runs FastH3 V2, now with a single-card consumer path. 4.6k stars per shields.io (repo).
  • Tencent-Hunyuan/Prism: code for the 2K picture-plus-sound model above. 22 stars per shields.io, data-center hardware only (repo).
  • OpenCut-app/OpenCut: the open-source CapCut alternative, still the obvious place to cut everything above. Not new (repo).

Creative workflows

1. Pull a subject off a clean plate with PlateExtract (loras/extract.safetensors, inference/qwen_extract.py).

The steps. Lock the camera. Shoot the frame with the subject, then the same frame empty. Install the repo's inference/requirements.txt, then run python inference/qwen_extract.py --composite picture.png --background clean_background.png --output extracted.png. The script fetches the add-on if you have not downloaded it. Drop the RGBA result over a new background and check the edges at 200%.

How it works. The model sees both images and learned, over 15,000 training steps, which pixels belong to the scene and which to the thing in front of it. Because it paints the alpha channel directly, it can keep half-transparent edges instead of cutting a hard line.

Why it is good. It is the old difference key with judgment added: it can keep hair and soft shadow that a pixel-subtraction key throws away, and it works on drawn frames as well as photos.

Where it breaks. Research license, so no paid work. Stills only; running it frame by frame on video will likely flicker (our expectation, not the author's claim). Needs a 16 GB NVIDIA card and 64 GB of RAM. And it depends on a true clean plate: if the light or framing shifts between shots, expect it to struggle.

2. Find the shot you know you filmed, with SCM.

The steps. brew tap allenv0/scm, brew trust allenv0/scm, then brew install --cask allenv0/scm/scm on an Apple-silicon Mac running macOS 12 or later. Add your footage folder as a watched folder and let it index. Use Scenes mode for a described moment ("wide shot, rain on the window, someone at the desk"), Dialogue for a line you remember, OCR for a sign or slate. Click a result to jump to the timecode.

How it works. It cuts each video into scenes and turns each one into a searchable fingerprint of what it looks like, using a local vision model you pick from a short list (the default is about 435 MB, downloaded once). Whisper transcribes speech and Tesseract reads visible text, so three different kinds of memory point at the same frame.

Why it is good. Nothing leaves the machine, which matters for client footage under NDA, and renamed duplicates are caught by content rather than filename.

Where it breaks. Mac only. Indexing a large drive will take time, and the README gives no speed figures. Search by description is only as good as the vision model's sense of your shots, and the optional chat mode needs a separate local language model.

3. Generate picture-plus-sound shots on one card with FastH3 V2 NVFP4 Consumer (adaln_tables.pt, fastvideo_inference.json).

The steps. Install FastVideo with uv and the CUDA 13 build of PyTorch, as the V2 card shows. The V2 card's eight-pass example is python examples/inference/basic/basic_fasth3_8step.py --prompt "your prompt" --no-warmup --repeats 1. For the consumer build, the card's only code downloads the 155 MB adaln_tables.pt and sets FASTVIDEO_H3_ADALN_TABLE to it, which you need if your PC has under 64 GB of RAM. The card gives no full run command, so expect some setup.

How it works. The team trained a fast student to match H3's results in eight passes, then shrank the model and its prompt reader to four-bit numbers so they fit on one card.

Why it is good. Synced dialogue and sound effects on a desk machine, with no per-second fees.

Where it breaks. Blackwell cards only. The precomputed table works for the included eight-pass schedule only and errors on anything else. No published speed on a 5090, weaker fine detail and hard motion than the full model, and the H3 license.

Worth testing

  • MadStickArt Playground: describe a moment, pick a camera angle and shot type, get a 2.39:1 stick-figure storyboard frame in the browser. The model (created October 6, Apache 2.0) has 12,192 learned numbers and was trained on 66 sci-fi story beats. Tradeoff: fixed vocabularies, no character identity, no continuity between frames.
  • Nacre v1: run your worst archive scans through it. Tradeoff: invented detail and mangled small text, so keep it away from documents and logos.
  • SCM: point it at last month's B-roll. Tradeoff: Mac only, and a big library means a long first index.
  • Flattener (created October 5, Apache 2.0): uvx flattener-scan -i photo.jpg scan.jpg turns a phone photo of a sketchbook page into a flat, upright scan. Tradeoff: pages cut off by the photo, heavy occlusion and low-resolution text give "incomplete scans or uncertain orientation."

What actually matters from today's signal

The headline releases have gone quiet for a few days, and the useful work moved to the bench. Look at what shipped: a cutout tool, a footage search, two upscalers, a page flattener, a storyboard sketcher. None of these make a viral clip. All of them take an hour out of a real job. For a working editor or compositor, that is the better trade. Generators get you a shot; finishing tools decide whether you can use it.

Money and licensing split the list cleanly. SCM, Nacre, Flattener and MadStickArt are MIT or Apache, so they can sit in a paid pipeline today. PlateExtract and UltraSharp V2 Lite are research or non-commercial, which makes them great for testing a technique and useless for invoicing it. Read the license line before you build a habit around a tool, because swapping a cutout step out of a finished pipeline is more painful than choosing well up front.

The counter-signal is ArtCraft. Seven open desktop apps aimed at Adobe's lineup, five of them still "in development," drew 154 Hacker News comments, many asking whether code written this fast by AI can be maintained, and one PDF-library maintainer warning that quality takes testing, not just tokens. A team member told HN they are "still super early alpha." Treat the promise as real and the timeline as unknown. PhotoCraft's star count roughly doubled since yesterday's briefing (1.3k to 2.4k on shields.io), so people want it to work. Keep your Adobe subscription until it does.


Source access notes: Adversarial fact-check ran (Sonnet subagent) and caught: a wrong creation date for the base FastH3 V2 repo (September 4, not October 5), limitation quotes attributed to the consumer card that live on the V2 card, an incomplete run command, Nacre's PSNR-Y figure mislabeled as sharpness and generalized to several test sets, SCM described as storing text descriptions (it stores visual fingerprints), a stale HN point count, an HN comment presented as an official statement, and a missing Gemma clause on Stable Audio's license. Every Hugging Face date is createdAt from the API (/api/models/<id> or listings sorted by createdAt, cache-busted), fetched through WebFetch. Prism's ship date could not be pinned: the HF repo was created September 24, the card refreshed October 6, and the GitHub README lists "2026-x-xx". SCM's creation date (October 3) comes from shields.io's created-at badge because the GitHub API returned 403; star totals are shields.io. GitHub's trending page served stale, unrelated content and Trendshift's daily list held no creative repos today, so no daily deltas are given. The HN Algolia default feed returned September items; a date-filtered query worked. Adobe's Firefly blog returned an empty page and WebSearch surfaced nothing dated in the window. OpenAI (advertising and EU provenance posts), Google, ElevenLabs (Eleven v4 ran September 28 and was covered), BFL, fal, Replicate, Luma and Runway had nothing new for creators inside the window; Runway's October 2 changelog items (Ideogram 4.5 in MCP, Seedance 2.5 Draft) ran in earlier briefings. Kling 4.0 full release is still unannounced beyond "October"; only 4.0 Flash early access (from September 28) exists, per secondary reporting. Stability's news page shows no dates. Midjourney, Krea, Ideogram, Pika, HeyGen, Udio and Recraft surfaced nothing dated via WebSearch. ComfyUI's releases page served a 2024 snapshot and was skipped. DMAD (October 2) and the FastH3 Trim builds (October 3) were skipped as covered or outside the window.