Creative AI Briefing: Sunday, October 4, 2026
Type "I am so glad you are here," attach a file the size of a phone photo, and an open voice model reads the line with affection instead of flat narration. LAION posted 557 of those files on Saturday: 500 voices, 40 emotions and 17 delivery traits, each one a separate add-on you swap in. The same day brought two add-ons that teach the MiniMax H3 video model to draw on a real pixel grid, and one that dresses a person in an outfit from a second photo without nudging the rest of the frame. The labs were quiet this weekend. What arrived instead is customization as a download: one small file per job, attached to a model you already have.
New models
LAION's Humaneness Voice add-ons turn one open voice model into 557 voices and moods. laion/Humaneness-Voice-Small-Rank1-LoRAs was created on October 3 at 15:15 UTC. It holds 500 voice-identity add-ons, 40 emotion add-ons (the card's example is emotion/Affection) and 17 delivery-trait add-ons such as a tempo shift. Each add-on trains only 286,720 numbers, which is why it is tiny next to the base model. The base, laion/Humaneness-Voice-Small (created September 29), is an English and German voice-acting model that you direct with a short CAPTION: describing the delivery and a quoted TRANSCRIPT: of the exact words. Everything is CC BY 4.0, so you can use it in paid work as long as you credit LAION. Where it runs: Python scripts from the repo (infer_rank1.py), no ComfyUI node and no hosted generator, and the card gives no memory figure. The honest catch is large. The card says "No broad efficacy claim yet," the identity numbers come from a two-voice pilot, and the base card calls the model "a research release with known shakiness, prompt sensitivity and imperfect word/timing/burst control." The card also tells you to attach one add-on at a time ("do not stack multiple adapters unless you have independently evaluated that mixture") and to use voice identities only "with appropriate rights and consent." Before installing anything, listen: LAION's public listening page holds 18,720 takes across the base model's ten training stages.
Video
Two add-ons teach MiniMax H3 real pixel art, posted about half an hour apart. JOYUNGI/pixelart8x-h3-lora (created October 3, 12:48 UTC) aims for "one colour per cell on a fixed 8 px grid, a limited palette, and frame-by-frame sprite-style motion." Start the prompt with "pixel art animation, crisp 8x pixel grid, limited palette." followed by your scene, and the card ends its example with "Sound: none." Keep sizes in steps of 32 pixels. The author trained it on 20 stills and 60 clips and warns the set is mostly "full-body anime characters in military-style outfits on flat backgrounds," so expect that look to leak in. ij/PixelTune-MiniMax-H3-LoRA (created October 3, 13:24 UTC) takes a stricter route: you declare a native canvas such as 128 × 128, generate at four times that size, then snap the result back to the grid with one shared palette. Its author calls it an early 1,000-step checkpoint of a planned 5,000, trained on 153 clips, and says exact pixel boundaries "may be inconsistent." Both carry the MiniMax H3 Community License, which has territory restrictions worth reading. Neither ships a ComfyUI workflow, but fal's H3 text-to-video LoRA endpoint accepts an add-on by file URL at $0.0625 per second at 480p.
Wan2.2 image-to-video on an iPhone, for app builders. aether-models/wan22-ti2v-5b-fast-int8 (created October 4, 01:20 UTC) converts FastVideo's fast Wan2.2 5B model into a bundle for Apple's on-device runtime on iOS and macOS 27. It turns a still plus a prompt into a two-second clip (49 frames at 24 fps) in three steps, at sizes from 320 × 576 to 672 × 384. Apache 2.0. The catch: this is a developer bundle, not an app. It needs a special Apple memory entitlement, a 13.02 GB download and about 37.75 GB of storage after first-run setup.
Image
Dress a person in an outfit from a second photo, and keep everything else where it was. ausboss/Qwen-Image-2.1-Outfit-Swap-Consistency-LoRA (created October 4, 04:03 UTC) is an add-on for Qwen Image 2.1. Person in image 1, outfit in image 2, and a sentence that names what must not change. The author tested 16 swaps the add-on never saw, with the same seeds and settings, and reports that the median shift of the picture dropped from 3.6 pixels to 0.1, changed face pixels from 14.1% to 1.7%, and repainted background from 10.8% to 0.9%. Those are the author's own measurements. A ComfyUI workflow ships in the repo. The catches: shoes "are the stubborn part," poses never change to suit the new clothes, and the base model's Qwen research license limits what you can do commercially, so read it before using this for a client catalog.
Audio and music
The voice add-ons above lead. One more for Mac users: mlx-community/Irodori-TTS-v4-Large-8bit (created October 4, 04:55 UTC) brings Aratako's Japanese voice-cloning model (created September 27, Gemma license) to Apple Silicon in about 4.5 GB through mlx-audio, with a short reference clip setting the voice. Second-day work: a conversion, not a new model.
Open and local
The local story is a pile of attachable parts. A voice is a small file, a pixel grid is a small file, an outfit swap is a small file, and a phone can run a two-second video model if a developer packages it. None of it needs a subscription, though none of it is a one-click app yet either.
- neilsonnn/image-blaster: one photo in, an explorable 3D scene with object models and sound effects out, run through Claude Code. Not new (Trendshift ranked it #1 in TypeScript on May 15) but back on Trendshift's daily list with +359 stars; 6.3k stars per shields.io, MIT (repo).
- zcbacxc/movie-narrator: give it a film and a style, get a narrated recap video with subtitles. +32 on Trendshift today, 642 stars per shields.io, AGPL-3.0, and its default voice is for personal non-commercial testing only (repo).
- OpenCut-app/OpenCut: an open-source CapCut alternative for editing the clips all these models make. +49 on Trendshift, 92k stars per shields.io (repo).
- meituan-longcat/LongCat-Video: Meituan's open video generation model. +40 on Trendshift, 8.9k stars per shields.io (repo).
- modelscope/DiffSynth-Studio: the Python toolkit PixelTune is tested with, and the route to run it locally. 13k stars per shields.io (repo).
Creative workflows
1. Make a pixel-art animation that sits on a true grid, with PixelTune (pixeltune_h3_full_canvas4_v3_step_01000.safetensors).
The steps. Pick a native canvas in multiples of 8, say 128 × 128. Generate at four times that, 512 × 512, with the card's prompt pattern: "[Shot 1] pxgrid4 native_res_128x128 source_fps_10. Native pixel canvas: 128 by 128 pixels. Each native pixel is a 4 by 4 solid-color square… 2D pixel-art animation." then your scene. The tested settings are 50 steps and strength 1.0. Then clean up: snap every 4 × 4 block to one color, build one palette for the whole clip with no dithering, and enlarge with nearest-neighbor scaling.
How it works. The model learned to paint in 4 × 4 blocks, so the cleanup step only has to tidy cells that are already close, instead of guessing a grid from soft video.
Why it is good. One palette per clip stops the flicker you get when every frame picks its own colors, which is what usually gives AI "pixel art" away.
Where it breaks. It is an early checkpoint, the card warns that "static-region stability and exact pixel boundaries may be inconsistent," local use needs DiffSynth-Studio and the full H3 download, and long clips are untested.
2. Swap an outfit in ComfyUI without the frame drifting, with the Outfit Swap Consistency add-on (qwen-image-2.1-outfit-swap.safetensors, workflow workflows/qwen-image-2.1-outfit-swap.json).
The steps. Install the AusBoss node pack (2.5.1 or newer) on ComfyUI 0.38 or newer and load the workflow. Load the add-on at strength 1.0, with no trigger word. Feed the person as image 1 and the outfit as image 2. Prompt: "Dress the person in image 1 in the [outfit description] shown in image 2. Keep their face, hair, hands, pose and the background exactly the same." The author used 25 steps, CFG 1 and euler/simple.
How it works. The add-on learned from 66 hand-picked before-and-after pairs, each one chosen because the person and background stayed put, so it learned that a swap should touch only the clothes.
Why it is good. A product shot that stays registered to the original can be layered over it in Photoshop and masked, instead of redone.
Where it breaks. Shoes, accessories that cross the garment, poses that suit the old clothes, untested sizes (the author sampled at about 2 megapixels), and the research license.
3. Read a line with a chosen emotion, using a Humaneness Voice add-on.
The steps. Download the base model and the add-on repo side by side. Run the card's infer_rank1.py with --adapter-id emotion/Affection, a prompt of CAPTION: warmly affectionate... plus the quoted TRANSCRIPT:, --language en and a fixed --seed. Change only the add-on and compare takes.
How it works. The base model already knows how to act; each add-on nudges it toward one voice, one mood or one trait.
Why it is good. A fixed seed and a swapped add-on give you comparable takes of the same line, like asking an actor for a warmer read.
Where it breaks. One add-on at a time, research-grade timing control, English and German only, and no consent check built into the voice files.
Worth testing
- Humaneness listening page: hear the base model before installing anything. Tradeoff: these are base-model takes, not the new add-ons.
- H3 LoRA on fal: paste the JOYUNGI or PixelTune file URL and render a 480p test. Tradeoff: you pay per second, and both add-ons are experimental.
- Outfit Swap workflow: try it on a product shot you already own. Tradeoff: research license and stubborn shoes.
- image-blaster: turn one concept painting into a walkable scene. Tradeoff: it needs World Labs and fal API keys (World Labs generation is paid), its sound effects come from ElevenLabs, and it was built by a World Labs designer as an open project, not an official World Labs product.
What actually matters from today's signal
The unit of customization keeps shrinking. A year ago, teaching a model your look meant a training run. Today a voice, an emotion, a pixel grid or an outfit rule is a file you download and attach, and the people making those files are hobbyists and research groups, not the labs. For a creator, that changes where the money goes: you stop paying for a better general model and start collecting the specific add-ons that fit your work.
It also changes what you have to check. Every one of today's add-ons arrives with its author's own numbers and its author's own warnings, and the warnings are the useful part. The pixel-art add-ons admit their training sets are tiny. The outfit add-on admits shoes defeat it. The voice add-ons admit nobody has run a listening study across 500 voices.
The risk sits in the licenses and the consent. The voice files are CC BY 4.0, which is generous, but a voice-identity file is a person's voice, and the card's consent line is a request, not a lock. The outfit swap rides on a research license, and the pixel add-ons on a license with territory limits. Collect add-ons, but read the paperwork under each one before it reaches a client.
Source access notes: Every Hugging Face date is createdAt from the API (/api/models/<id>, cache-busted) or from API listings sorted by createdAt. The direct Hugging Face API and GitHub API were unreachable from the shell this run, so all fetches went through WebFetch; the GitHub API returned 403 for image-blaster, and its first-trending date (May 15) comes from Trendshift and VP Land. Runway's news page returned navigation only, Google's blog listing showed no dated posts for the window, and the ComfyUI releases page returned unreliable dates, so those were skipped. OpenAI, Adobe, ElevenLabs, Suno, Stability, BFL, Luma, fal and Replicate had nothing creative inside the window. Midjourney, Krea, Ideogram, Pika, HeyGen, Udio and Recraft surfaced nothing dated via WebSearch. Adversarial fact-check ran (Sonnet subagent) and caught: an unverifiable press attribution for image-blaster (replaced with World Labs' own page), an incomplete list of the services image-blaster calls (ElevenLabs added), a rounded Irodori size (now about 4.5 GB), a missing "Sound: none" in the pixelart8x prompt, and two interpretive lines on the outfit add-on (reworded to match the card). GitHub's HTML page showed 4.8k stars for image-blaster against shields.io's 6.3k; the briefing uses shields.io per house rules. Stability's 4Director paper (camera and object control for video) was skipped as a lead because its repo says code is "Coming soon." Trendshift deltas are daily figures from its homepage; star totals are shields.io.