FervorCreative AI
Live Latest 03.10.26 · morning 102 tools tracked 345 workflows indexed 267 topics Hot: MiniMax H3, ComfyUI, ACE-Step 1.5

The community spent a quiet lab week turning generations into things you adjust, so a music cue gets a happier version without a reroll, a LEGO idea comes back as an editable parts file, and a talking video clip costs four steps instead of fifty.

Audio SlidersACE-Step 1.5ldraw-novaDMADMiniMax H3Ming-Image 0.1 Designmusic-gen3d-genvideo-genlocal-creative-aiopen-weightscreative-workflows

Creative AI Briefing: Saturday, October 3, 2026

Generate a ten-second jazz piano cue, then drag one slider from sad to happy and hear the same piece, same notes and same seed, brighten in front of you. That is Audio Sliders, an open add-on set for the ACE-Step music model posted on Friday night. It landed in the same 24 hours as an agent that turns a sentence into a buildable LEGO model you can open in Blender, and a research team's fast version of MiniMax H3 that makes a five-second clip with sound in four steps. The big labs were busy with language models and fundraising. The independent builders shipped control instead: ways to change one thing about a result without starting over.

New models

Audio Sliders gives an AI music model 35 knobs. Taka Khoo published the weights on Hugging Face on October 2 at 21:37 UTC, with code in takakhoo/audio-diffusion-control. Each slider is a tiny add-on for ACE-Step 1.5 XL turbo that moves one musical quality while the prompt and seed stay fixed: "sad to happy, solo to full ensemble, stiff to groovy," in the README's words. Twelve sliders come from text concepts (mood, ensemble, groove, harmony, melody, tension, brightness, density, energy, tempo, electronic, vintage); others were found by analyzing real recordings (14,985 of them, per the demo page), such as strings to synth. Sliders stack: the author reports that "effects roughly add, with interaction terms a fifth to a third of the main effects." Speed is the author's own figure, 0.4 seconds per ten-second clip on ACE-Step, with no extra generation passes. The license is MIT, so you can use it in paid work. Where it runs: ACE-Step's README says the XL turbo model needs at least 12 GB of graphics memory with offloading plus compression, or 20 GB without offloading. The catch, from the author: "No listening study has been run yet," sliders bleed into each other (mood and melody also change brightness), some only work between -1 and +1, and testing was on short instrumentals (ten seconds, with some text sliders also checked at thirty). Try it free in the live demo, which generates live on the author's hardware and may not stay up; the repo says the demo exposes 18 of the sliders.

ldraw-nova turns a LEGO idea into a model file you can build, render and edit. It hit Hacker News on October 2 as a Show HN (113 points at crawl time; 136 GitHub stars per shields.io). You give an AI agent an idea, guide it, and it writes a plan and then a Python script that outputs the model in LDraw, the long-standing open file format for LEGO designs. You get the LDraw source, a 3D viewer, a VR mode, a Blender-ready .glb file, and the full chat log. The README names Claude Opus 5.5 and GPT-6 Astra as the models that made it work and warns that "only expensive, high-end models, are currently capable of generating large-sized and correct models," so plan on a real API bill for anything big. It runs locally in Docker (two repos at tag v0.6.0, about 5 GB to build). The code is AGPL-3.0, which matters only if you host it as a service for others. No hosted demo; example builds such as examples/atlas-crane/ ship in the repo. LEGO is a trademark of the LEGO Group, and the project is unaffiliated.

Video

DMAD makes MiniMax H3 clips with sound in four steps. The paper went to arXiv on October 1 and the weights on October 2 at 00:48 UTC. Two 1.4 GB add-on files turn the 50-step MiniMax H3 into a four-step model that outputs 1344 × 768 video with stereo audio, 124 frames at 24 fps, about five seconds. Fewer steps means you wait less per take, which is the difference between previewing ten ideas and committing to one. Quality claims are the authors' own: the abstract reports 79.1% human preference over an older fast method (DMD2), and the project page shows 84.6% against another (rCM). It runs through the authors' inference.py or Diffusers, with no ComfyUI support listed. The GitHub README says the base download is about 170 GB and that it runs on a single 80 GB GPU, which puts it on rented hardware for most people. Same MiniMax H3 Community License as the base.

A lighter decoder shaves H3's final step. corechan/MiniMax-H3-LightVAE (created October 1, 03:37 UTC) is an unofficial replacement for the part that turns H3's internal result into frames. On an RTX PRO 6000 at 1280 × 704 and 124 frames, the uploader measured 5.2 seconds with TensorRT against 7.1 seconds for the official decoder. Two seconds per clip is small alone; next to a four-step model it is a noticeable share of the wait. 3.3 GB per format (safetensors or ONNX), diffusers format, no ComfyUI files, and the card points readers to the MiniMax H3 license's territory restrictions, so read them first.

Kling 4.0 is still rolling out. Kling's own guide still says "the full Kling 4.0 model is scheduled to launch in October 2026," with early access since September 28. An October 2 creator interview on the Kling blog is marketing, not a release.

Image

Ming-Image 0.1 Design runs on AMD cards now. voltaire321/Ming-Image-0.1-Design-GGUF (created October 3, 00:58 UTC) shrinks inclusionAI's MIT-licensed design model, which is good at transparent backgrounds and pixel art, for stable-diffusion.cpp. The uploader reports that on AMD graphics without bf16 support, the original took "about 120 s a step," while these files make "a 1024×1024 image in about 35 s end to end." You need 6.54 GB for the image model plus 10.52 GB for the text encoder. Second-day work: a conversion, not a new model, and MIT like the original.

Audio and music

Audio Sliders leads above. On the research side, today's top Hugging Face paper is Tencent's Adaptive Reward Routing for joint audio-video generation (55 upvotes), and "Align Then Reason" proposes an automatic judge for lip-sync in dubbing. Neither is something you can use yet.

Open and local

The local story is about cost per idea. One person made music editable by dial on a 12 GB card (with offloading and compression), a lab cut H3 drafts from fifty steps to four, a hobbyist cut H3's decode time, and another made a design-focused image model usable on AMD hardware. Nothing here needs a subscription, though ldraw-nova needs an API key.

  • takakhoo/audio-diffusion-control: 35 sliders that change one quality of an AI music clip without changing the piece. Released October 2, 1 star per shields.io, so you are early (repo).
  • anteloc/ldraw-nova: describe a LEGO model, get LDraw and Blender files back. Show HN on October 2; 136 stars per shields.io (repo).
  • Yzmblog/DMAD: four-step MiniMax H3 audio-video generation. Code for the October 1 paper; 38 stars per shields.io (repo).
  • ace-step/ACE-Step-1.5: the MIT music model the sliders ride on, 10 seconds to 10 minutes per song. 13k stars per shields.io (repo).
  • leejet/stable-diffusion.cpp: the no-Python image runtime behind today's Ming-Image conversion. 7.5k stars per shields.io (repo).
  • Comfy-Org/ComfyUI: 136k stars per shields.io; none of today's releases ship ComfyUI files yet, so expect community wrappers (repo).

Trendshift's top movers today were agent and inference tools; ldraw-nova was the only creative repo listed, so there are no other star deltas to report.

Creative workflows

1. Make calm, build and peak versions of one music cue with Audio Sliders (weights: takakhoo/audio-sliders, backbone ace-step-1.5-xl-turbo).

The steps. Install with pip install "audiosliders[model,demo] @ git+https://github.com/takakhoo/audio-diffusion-control" audiobox_aesthetics, then start the local slider app with python -m audiosliders.server --sliders hf:ace-step-1.5-xl-turbo/text --backbone ace-turbo. Write one instrumental prompt and keep the seed fixed. Render at energy -1, 0 and +1. Then try energy plus ensemble together for the peak.

How it works. Each slider is a small add-on that nudges the model along one direction it learned, so the same seed produces the same piece moved along one axis.

Why it is good. An editor gets three related cues that cut together, instead of three unrelated rerolls that share a genre and nothing else.

Where it breaks. Short instrumentals only, sliders leak into each other, and quality drops past each slider's usable range (-1 to +1 for some, up to +2 for others). No vocals tested.

2. Describe a LEGO build and take it into Blender with ldraw-nova (Docker setup in anteloc/ldraw-nova-docker).

The steps. Clone both repos at --branch v0.6.0, run docker compose build and docker compose up -d inside ldraw-nova-docker, and open https://localhost:8443. Describe the model, review the plan, steer as it renders and corrects. Export the .glb into Blender for lighting, or keep the LDraw file for building.

How it works. The agent never draws pixels. It writes a script that places real parts by coordinates, renders the result, looks at it and fixes mistakes.

Why it is good. The output is an editable parts list, not a picture of bricks, so you can change a color or swap a piece by hand afterward.

Where it breaks. The author says only expensive top models produce large, correct builds, setup needs Docker and a self-signed certificate for VR, and the README does not say whether builds hold together physically.

3. Draft a five-second talking clip in four steps with DMAD (dmad_minimax_h3_4step_full_critic.safetensors).

The steps. Download MiniMax H3 with hf download MiniMaxAI/MiniMax-H3 --local-dir models/MiniMax-H3 --exclude "FL2VA/*", (the README excludes a few more folders), then run inference.py from the DMAD repo with the base folder, the add-on file and a --prompt such as its own "A polar bear is playing the violin in the snow." The card recommends four steps and no guidance setting.

How it works. A training method taught a small add-on to jump to the finished clip in four moves instead of fifty.

Why it is good. Picture and sound arrive together, fast enough to try several takes of a line.

Where it breaks. No ComfyUI node, a roughly 170 GB base download, a single 80 GB GPU per the README, no published timing, and fast versions usually trade some detail for speed.

Worth testing

  • Audio Sliders live demo: type a prompt, stack two sliders, listen. Tradeoff: no listening study yet, and the demo runs on one person's server.
  • ldraw-nova examples: open the shipped example builds before paying for your own. Tradeoff: real builds need a top-tier model and the API bill that comes with it.
  • DMAD: for anyone already running H3 locally, the "full critic" file scored higher on the authors' benchmark. Tradeoff: the gains are self-reported and there is no ComfyUI route.
  • Ming-Image Design GGUF: pixel art and transparent assets on an AMD card. Tradeoff: about 17 GB of files and about 35 seconds per image on the uploader's machine.

What actually matters from today's signal

The prompt is losing its monopoly on control. For two years, the only way to change an AI result was to rewrite the words and roll again, which throws away everything you liked about the last take. A slider that moves one quality of the same clip, and an agent that hands back an editable parts file, both keep the take and let you adjust it. That is how musicians and designers already work: you do not re-record a song to make the chorus brighter.

The time math changes with it. Four steps instead of fifty and a decoder that is two seconds faster both point at the same goal, making the next try cheap enough that you stop guarding your rerolls. For a creator on one graphics card, that means an afternoon of tries instead of a handful.

The counter-signal is permanence. Solo, Mozilla's AI website builder, is shutting down (per a Hacker News post on October 2), and its shutdown FAQ says all sites and data are deleted on November 30, 2026, with an export (an HTML copy of the site, plus a CSV of image links and form submissions) that leaves out the image files themselves. Hosted AI tools come and go; MIT-licensed weights on your own drive do not. If you built anything on Solo, export it and download your images now.


Source access notes: All Hugging Face dates are createdAt from the API (cache-busted): audio-sliders 2026-10-02 21:37 UTC, DMAD 2026-10-02 00:48 UTC, MiniMax-H3-LightVAE 2026-10-01 03:37 UTC, Ming-Image-0.1-Design-GGUF 2026-10-03 00:58 UTC. Star totals from shields.io with cache busting; the GitHub API returned 403 for ldraw-nova, so its repo creation date is not stated. The Show HN point count is from the HN Algolia API at crawl time. Bucket A was quiet: OpenAI (GPT-6 and dots, no creative media), ElevenLabs (newest product post still Eleven v4), Suno (Speech beta, covered yesterday), BFL (FLUX 3 Image covered yesterday), Adobe (newest Sept 29), DeepMind (nothing creative in window), Stability (undated index), Luma (SEO roundups only), fal blog (newest Sept 17), Replicate blog (nothing recent). Runway's news page and Google's AI blog did not render dated items; Midjourney updates and ByteDance Seed did not render; fal's model index is blocked by robots.txt. ComfyUI and Diffusers GitHub release pages served stale content (2024 dates) and were not used. Civitai not attempted. Writing section omitted: nothing qualified. Adversarial fact-check pass ran (separate agent, primary sources): it caught a dropped "quantization" condition on ACE-Step's 12 GB figure, an unconfirmed DMAD --prompt-file flag (replaced with the README's --prompt example) plus the missing 170 GB download and 80 GB GPU requirement, an over-broad "ten-second only" testing claim (some sliders were tested at thirty seconds), an unverified "jazz to electronic" discovered axis (cut), a loose usable-range figure, the demo exposing 18 rather than 35 sliders, LightVAE's two-format size and license territory note, and unsupported "Mozilla said this week" wording on Solo (now attributed to the HN headline; the FAQ itself does not name Mozilla). All HF createdAt dates, star counts, quotes, file sizes and speed figures checked out. The article fact-check pass later flagged that a Hacker News front-page claim could not be confirmed, so Solo and ldraw-nova are now cited to their HN posts only. Not independently rechecked: the HF papers upvote counts.