FervorCreative AI
Live Latest 08.10.26 · morning 120 tools tracked 417 workflows indexed 309 topics Hot: ComfyUI, MiniMax H3, ArtCraft

Control got cheaper this week on both sides of the fence, with Google halving the price of a reference-heavy image model while open add-ons let a single photo or a 24 GB card do work that needed a reshoot or a rented server a week ago.

Nano Banana 2.1Qwen-Image-2.1 Multiple Angles LoRADMADVectorCraftCompositorComfyUIimage-genvideo-gendesign-toolsopen-weightslocal-creative-ai

Creative AI Briefing: Thursday, October 8, 2026

Hand Google's newest image model a stack of up to 14 reference photos, four characters and ten props, and it will build a scene that keeps them all recognizable, for about three and a half cents an image at 1K. That is Nano Banana 2.1, launched Tuesday, and it arrived in ComfyUI the next day. The same 48 hours brought an open add-on that re-photographs a single picture from any of 72 camera positions and a set of ComfyUI workflows that get MiniMax H3 clips with sound out of a 24 GB gaming card. The pattern: control is getting cheaper. Consistency, camera angle and render length used to cost a reshoot or a rented server. This week each one costs a few cents or a spare evening.

New models

Nano Banana 2.1 is Google's new everyday image model, and at 1K it costs half what Nano Banana 2 did. Google released it on October 6; the DeepMind model card is dated that day and the Gemini API page lists it as stable under the id gemini-nano-banana-2.1. What you can make: composites built from up to 14 reference images, with consistency for up to four characters and ten objects, at 1K, 2K or 4K, plus panoramic frames up to 8:1 that no longer show tiling seams at 2K and 4K. You pick a "thinking" level (minimal, medium, high) that trades speed for planning, which matters most for infographics and poster text. Per Google's pricing page, a standard 1K image costs $0.0336 against $0.067 for Nano Banana 2, 2K is $0.0504, and 4K is $0.113 (Decrypt reported $0.0756 for 4K; Google's own page says $0.113, and $0.0567 in batch). Batch jobs halve every price. There is no free API tier. It is rolling out in the Gemini app, AI Mode, AI Studio, Flow and Stitch, and Google's deprecations page names it the replacement for Nano Banana 2 and the retired Imagen 4 models, though it lists no shutdown date for Nano Banana 2 yet (one outlet reported October 29; Google's page does not say that). The catch, in Google's own words on the model card: small text and long paragraphs often render poorly, character consistency "is not always perfect," masked and doodle edits follow instructions only partially, and it sometimes confuses left and right. The headline preference score (1,050 Elo against 990 for Nano Banana 2) is Google's own test of the Thinking variant. Try it in the Gemini app or AI Studio before wiring it into anything that bills (Google's pages list no free API tier; one secondary outlet says the Gemini app offers free use under daily limits).

Image

Re-photograph a single picture from another angle with the Qwen-Image-2.1 Multiple Angles add-on. akhaliq/Qwen-Image-2.1-Multiple-Angles-LoRA was created on October 6 at 22:06 UTC. Give it a photo and a line like <mva> right side view, high-angle shot close-up and it redraws the subject from that camera position: 12 directions around the subject in 30-degree steps, four heights from eye level to straight down, each with an optional close-up, 72 framings in all. The author recommends the v2 file at step 1,500 (319 MB) at strength 0.8 to 1.0, and keeps the smaller v1 file as the fully checked fallback. The add-on itself is Apache 2.0, but it only runs on top of Qwen-Image-2.1, whose license is the Qwen research license, so treat outputs as non-commercial until Qwen says otherwise. Honest limits from the card: trained at 512 pixels square, so larger canvases are untested; typed degree numbers are ignored, only the word labels work; fine fur softens on real photos. The author's own similarity benchmark actually scored the add-on slightly below the base model (0.832 vs 0.857), and the card explains why that metric cannot tell whether the camera moved. Judge it by eye.

Nano Banana 2.1 is already a ComfyUI node. ComfyUI v0.39.1, dated October 7, adds it as a single partner node that does both generation and editing, with 1K to 4K output, the three thinking levels and up to 14 references. Partner nodes are paid API calls; v0.39.0's new --disable-partner-nodes launch flag switches them off.

Video

DMAD's four-step MiniMax H3 now fits a 24 GB card and runs 15 seconds in ComfyUI. This briefing covered the DMAD weights on October 3, when the README listed a single 80 GB data-center card. Since then the GitHub README has logged a --low-vram mode for 24 GB consumer cards (October 4), ComfyUI support (October 5) and, on October 7, ready-to-run 4-step and 8-step ComfyUI workflows that produce 15 seconds of video with sound. The README says the low-memory path peaks at 12.9 GiB of graphics memory while sampling and 14.2 GiB while decoding and produces output identical to the full path. The time cost is the authors' estimate rather than a measurement: about a minute per step on an RTX 4090 or 5090, so four to eight minutes of sampling per default five-second clip, plus decoding, and about twice that on a 3090. You still need the roughly 170 GB MiniMax H3 base download and must accept the MiniMax H3 Community License, whose territory clause this briefing has flagged before. The workflow files sit in the weights repo under minimax_h3/workflows/, and a free Space makes five-second drafts.

Audio and music

A 34 MB Turkish voice that runs faster than real time on a CPU. canberkkkkkk/ema-lightning, created October 1 and the top trending text-to-speech repo this morning, is an 8.6-million-parameter Turkish speech model under Apache 2.0, so commercial use is fine. The author reports about 6 times real time on a cloud-server CPU and roughly 440 times on an RTX 4090, plus 0.92% word error on their own benchmark. One voice, Turkish only, no cloning and no emotion control, and the author calls its naturalness "good, not the best." If you narrate for Turkish audiences, the free Space lets you judge it in a minute. The card asks that published audio be labeled AI-generated.

Open and local

Two free desktop editors are trending at once, from opposite directions. Compositor is one developer's Mac-only Photoshop alternative that shipped 1.0 on September 16 (per its GitHub releases) and now sits at 12k stars. VectorCraft is the storytold team's Illustrator rebuild, the latest of their clean-room Adobe rebuilds to post a self-rated status report. Both are built so an AI assistant can drive them: Compositor projects are folders of PNG layers plus a manifest that update live when a script writes to them, and VectorCraft ships an MCP server, a standard plug-in port that lets an assistant like Claude operate the real app.

  • robbietilton/Compositor: a free MIT image editor for Apple-silicon Macs with layers, masks, adjustment layers, Select Subject and Content-Aware Fill; opens PSD and PSB (8-bit RGB only, no CMYK) but exports JPEG only. Released September 16; 12k stars per shields.io (repo).
  • storytold/vectorcraft: an Illustrator-style vector app in Rust with SVG, PDF and PDF-compatible .ai import and export. Its October 7 status note puts it at 69 to 75% of Illustrator's features, 40 to 55% by a power user's standard, with no releases published yet. 2.7k stars per shields.io (repo).
  • storytold/photocraft: the Photoshop rebuild that started the storytold run, now 21k stars per shields.io, up from 11k yesterday (repo).
  • storytold/artcraft: the team's AI image and video workspace that routes to hosted models such as Nano Banana, Seedance and Kling. 5.6k stars per shields.io (repo).
  • shader-effects-inc/shaders: 200+ animated web effects (glass, metal, noise, gradients) you design on a free canvas and export as code for React, Vue, Svelte or plain JavaScript; the library is MIT, the Framer plugin is in the paid tier. 3.1k stars per shields.io (repo).
  • Yzmblog/DMAD: code and the new 24 GB ComfyUI workflows for four-step MiniMax H3 (repo).

Creative workflows

1. Turn one product or character photo into a turnaround set with the Multiple Angles add-on (checkpoints_v2/angles_v2_full_qwen21_multiple_angles_v2_qwen21_multiple_angles_v2_000001500.safetensors).

The steps. Load Qwen-Image-2.1 in edit mode (the card's Diffusers example uses QwenImage21Pipeline with 40 steps and true_cfg_scale=1.0), add the v2 step-1,500 file at strength 1.0 as the card suggests, and feed your photo. Run one prompt per view, always starting with <mva>: <mva> front view, eye-level shot, <mva> right side view, eye-level shot, <mva> back view, eye-level shot, <mva> left side view, eye-level shot. Keeping the seed fixed across the set is our suggestion, not the card's. Work at 512 pixels square, then upscale.

How it works. The add-on learned from 13,328 pairs of rendered 3D objects and rigged characters photographed from known positions, so the angle label works like a camera instruction rather than a style word.

Why it is good. A front, side and back set from one reference is what a character designer, a product shooter or a 3D modeler needs before anything else, and it used to mean a reshoot or a sketch session.

Where it breaks. If the result starts copying your original view, the card says lower the strength toward 0.8 rather than switching files. Fine detail softens on real-world subjects. The base model's research license rules out client work for now.

2. Render 15 seconds of MiniMax H3 with sound on a 24 GB card, with DMAD (dmad_minimax_h3_4step_lora_critic_comfyui.safetensors).

The steps. Accept the MiniMax H3 license on Hugging Face and download the base model, skipping the FL2VA, Ref2VA and transformer_ref folders as the README shows. Put the ComfyUI-format DMAD file in your add-on folder. Open minimax_h3/workflows/dmad_h3_4step_15s_podcast.json for speed or the 8-step dmad_h3_8step_15s_wok.json for quality. Sample with the stock lcm sampler and simple scheduler, or the ComfyUI-DMAD nodes.

How it works. DMAD trains a student copy of MiniMax H3 to land in four steps what the original needs fifty for, so each take costs a fraction of the passes.

Why it is good. It moves a 33-billion-parameter picture-plus-sound model from rented hardware to a card many editors already own.

Where it breaks. The 24 GB timings are estimates, not measurements. The model's default output is about five seconds (124 frames), so watch 15-second takes for drift late in the shot. The 170 GB base download and the territory clause in the MiniMax license do not go away.

3. Build a consistent cast scene with Nano Banana 2.1 in ComfyUI (v0.39.1 partner node).

The steps. Update ComfyUI to v0.39.1. Load up to four character reference images and your prop references into the Nano Banana 2.1 node. Write the scene in plain sentences naming each character the way you labeled the references. Set thinking to high for posters and diagrams, minimal for quick drafts. Draft at 1K ($0.0336 each per Google's API pricing), then rerun the keeper at 2K or 4K.

How it works. The model reads every reference alongside your prompt and plans the layout before drawing, which is what the thinking level controls.

Why it is good. Storyboards and comic panels with a recurring cast, at half the old per-image price.

Where it breaks. Google itself says character consistency is not always perfect and small text is often blurry at 1K, so judge layout, not lettering, in drafts. Every run is a paid call, and the node goes away if you launch ComfyUI with --disable-partner-nodes.

Worth testing

  • Nano Banana 2.1 in the Gemini app: give it a 4:1 panoramic banner brief with real copy on it. Tradeoff: Google warns small text and long paragraphs often render poorly, so check every letter.
  • DMAD ZeroGPU Space: write a scene plus a soundscape line and listen for sync. Tradeoff: five seconds only, shared free hardware with queues.
  • EMA Lightning Space: paste a Turkish script and time it. Tradeoff: one voice, no emotion control.
  • Compositor (brew install --cask robbietilton-compositor): open a layered PSD you already have. Tradeoff: Apple silicon and macOS 26 only, JPEG is the only export, CMYK files will not open.

What actually matters from today's signal

The price of consistency just fell by half, and that changes how you should work. At three and a half cents a 1K image, the sensible habit is to draft every frame of a storyboard at 1K with minimal thinking, throw most away, and pay for 4K only on the keepers. Batch pricing halves it again if you can wait. Nobody should be iterating at 4K anymore.

The open side answered with control, not quality. A free add-on that moves the camera around a single photo and a 24 GB route to fifteen seconds of picture with sound both attack the expensive part of a shoot: going back for another angle or another take. Neither is finished, and both sit on licenses (Qwen research, MiniMax H3 territory terms) that make client work a legal question before it is a technical one. Read the base model's license before you fall in love with the add-on.

The risk is dependency. Google's deprecations page already lists Nano Banana 2.1 as the recommended replacement for seven older image models, and ComfyUI's own changelog is removing Luma Ray 2 on October 24. Every hosted model you build a pipeline around will be swapped out on someone else's calendar. Keep your prompts and reference sets in your own folders, and keep one open route working for each job you cannot afford to lose.


Source access notes: Bucket A: OpenAI (GPT-6 posts Oct 7, nothing creative), ElevenLabs (latest Sep 30), BFL (latest Sep 23), fal blog (latest Sep 17), Adobe (latest Oct 2, Substance), Runway changelog (latest Oct 2, Seedance 2.5 draft mode already covered), Suno (Albums Oct 5, Speech beta Oct 1 already covered), Luma (marketing posts only). blog.google returned no dated Nano Banana 2.1 post, so the launch is sourced to Google's API docs, pricing page, deprecations page and DeepMind model card, with Android Authority and Decrypt for the date and rollout list. Google's announcement posts on X were not fetched. fal.ai/models is blocked by robots.txt. Midjourney, Kling and Stability: WebSearch showed nothing dated in the window; Kling 4.0 is still unreleased per secondary sources. GitHub's releases page omits the year, so Compositor's 2026 date is inferred from the macOS 26 requirement. The HF API, GitHub raw and shields.io were reached through WebFetch; direct curl from both shells is blocked by egress policy. HN, the ComfyUI blog and civitai were not attempted. Writing section omitted (nothing creator-facing). Aftershoot Photon (debuted Sep 18) was dropped as old news. Adversarial fact-check ran (Sonnet subagent) and caught: an unsupported "default model" label for Nano Banana 2.1 (removed); an unsourced "free" claim for the Gemini app (now attributed to a secondary source); EMA's CPU figure measured on a cloud-server CPU, not a laptop; DMAD's peak memory (now 12.9 and 14.2 GiB exactly); an unsupported claim that DMAD students were trained on five-second clips (reworded); and the Multiple Angles starting strength, which the card puts at 1.0 (fixed, with the fixed-seed tip labeled as ours). Prices, deprecation counts, file names, dates and star counts all held. The article fact-check later corrected the --offline flag wording here: the changelog ties disabling partner nodes to --disable-partner-nodes.