Creative AI Briefing: Monday, October 5, 2026
Open a layered PSD in a free editor, let an AI agent apply a curves adjustment layer and a smart sharpen, and export the finished PNG without touching Photoshop. That became possible this weekend with PhotoCraft, a from-scratch rebuild of Photoshop's core editing in Rust that ships an agent socket on day one. It arrived alongside Ideogram 4 squeezed onto a 16 GB card, a lyrics-to-song model running in a Chrome tab, and voice cloning on an iPhone's own chip. We found nothing new from the major labs in the window. The community spent that time moving the creative stack onto machines you already own, and the license under each file now decides what you are allowed to sell.
New models
PhotoCraft puts a Photoshop-style editor on your machine, with a door an AI agent can walk through. The storytold/photocraft repo was created on September 30, and v0.1.0 builds for macOS, Windows and Linux landed on October 2 (00:09 UTC), with a v0.1.1-rc.4 pre-release later the same day. It is not a model. It is a layered image editor with masks, 16 adjustment-layer types, type and vector tools, layer styles, editable smart filters, and 8, 16 and 32-bit color in RGB, CMYK and Lab. The README claims 134 of 135 real-world PSD test files round-trip byte for byte, and says anything it does not model yet is kept verbatim rather than dropped. The creative-AI angle is the control layer: every menu item runs one of 500+ commands from a single registry, and the same commands are exposed through a command line and an MCP server (photocraft-cli mcp), so Claude or another agent can edit a PSD exactly as a person clicking menus would. The code is dual-licensed MIT or Apache 2.0, so client work is fine; the name and logo sit under a separate brand license. The honest catch: the README calls it "early alpha," it has no generative fill or AI image tools at all, and its parity page reports "626 of 626 menu items live," which, as we read it, measures that each menu command exists in the registry, not that each tool matches Photoshop's output. Stars: 1.3k per shields.io (GitHub's own API shows far fewer), +341 on Trendshift today.
Image
Ideogram 4 now fits a 16 GB card, and runs in Chrome, as a community build. MarcinEU/Ideogram-4-INT4-AWQ-GPTQ-G32-embed8-16GB-VRAM-ONNX (created October 4, 21:58 UTC) compresses the 9.3B-parameter model to 4-bit weights: a 14.74 GiB download against 25.66 GiB for Ideogram's own fp8 release, and about 14,000 MiB of peak graphics memory at 1024 × 1024 on the author's card. On the author's RTX 5070 Ti, the card reports 3.31 seconds per step through DirectML, about 10.8 through WebGPU in Chrome, and 553.65 on CPU, or about 110 seconds per image on the fastest route. The --chunked-prefill flag drops that peak to 11,519 MiB for smaller cards, at the same speed per step. Two honest catches the author publishes himself: a grey blocking panel is "a memorized mode of the base model," not a real filter, and shows up far less on detailed prompts (13.8% against 36.2%); and the built-in caption writer invents signage on 12 of 27 prompts that never ask for text. The license is the Ideogram Non-Commercial Model Agreement, and this is an unofficial conversion "not produced, endorsed or validated by Ideogram." The browser version runs from a local server (python usage/web/serve.py), not a hosted page.
Outfit-level try-on with a layering order you write out. ArtmeScienceLab/Garments2Look-LoRA (created October 3, 21:52 UTC) is a pair of add-ons for Qwen-Image-Edit-2509 from the team behind the CVPR 2026 Garments2Look paper and its 98,012-outfit dataset. Where Sunday's outfit swap handled one garment, this one dresses a person in a whole look from a collage: top, sweater, pants, shoes, bag and belt, with a numbered prompt that says what goes over what ("Layering Order: (1) -> (6) -> (2)"). One file repaints a grey-masked region, the other replaces an existing outfit. The catch: the repo states no license for the add-ons, the card warns that accessory detail and pose preservation "may vary," and there is no ComfyUI workflow yet.
Video
A world model you can steer, cut down to fit a big Mac. ogtsvc/lingbot-world-fast-diffusers-int8 (created October 4) shrinks LingBot-World Fast's video transformer from 74.2 GB to 18.6 GB, and the full folder from 86.1 GB to 30.5 GB. Apache 2.0, so commercial use is open. The catch is the hardware: on an M3 Ultra at 640 × 352 it holds about 31 GB once loaded and 44 to 46 GB while generating, so plan on a 64 GB machine. Second-day work, dated by its own creation, not a new model.
Research only, no creator tool yet: VGGT-Diff (October 4) builds new camera angles from six posed photos on top of Wan 2.1, under a non-commercial research license.
Audio and music
A full lyrics-to-song model now runs in a browser tab. mrfakename/YuE2-3B-ONNX (created October 5, 02:24 UTC) converts m-a-p's YuE2 3B (created September 9) into a roughly 2.1 GB package for onnxruntime-web with WebGPU, so you paste lyrics and style notes and the song is made on your own graphics card, nothing uploaded. The author links a demo Space, mrfakename/yue2-webgpu, which refused our requests this run, so treat it as unconfirmed. Catches from the card: guidance is "only approximated," 4-bit rounding changes the output slightly, the same seed will not reproduce a PyTorch render, and the license is CC BY-NC 4.0, so no paid releases.
Voice cloning that runs on an iPhone's Neural Engine. tinytrashlabs/Qwen3-TTS-0.6B-Base-ANE (created October 5, 03:13 UTC) is a Core ML build of Qwen3-TTS 0.6B that clones a voice from a few seconds of audio plus its transcript and, per the card, speaks at about real time on an iPhone 15 Pro at roughly 5% CPU. About 1.9 GB, iOS 18 or macOS 15, Apache 2.0. On the Mac side, sucrette/chatterbox-turbo-t3-coreai (October 4, MIT) ports Resemble's Chatterbox Turbo to Apple's Core AI framework on macOS 27, capped at 80 tokens of text per call and needing its separate voice decoder. Both are developer bundles, not apps, and the first card's own rule applies to both: "Only clone your own voice, or a voice you have permission to use."
Open and local
The local story today is ownership. A PSD editor that costs nothing and takes commands, a top-tier typography model on a mid-range card, a song model and a voice model that never send your audio anywhere. The tradeoff moved too: these run on your hardware, but the Ideogram, YuE2 and VGGT-Diff files all carry non-commercial terms.
- storytold/photocraft: a free, local, Photoshop-compatible editor an agent can drive through MCP. New repo (September 30), first builds October 2, +341 on Trendshift, 1.3k stars per shields.io (repo).
- storytold/artcraft: the same team's desktop "IDE for interactive AI image and video creation," staging 3D scenes and posed mannequins before sending them to cloud models such as FLUX, Kling, Veo and Seedance. Not new, but +812 on Trendshift, 2.3k stars per shields.io; generation runs through cloud providers, some needing their own keys (repo).
- OpenCut-app/OpenCut: the open-source CapCut alternative, where the clips from all of the above get cut. +1,874 on Trendshift, 92k stars per shields.io (repo).
- neilsonnn/image-blaster: one photo in, an explorable 3D scene out, run through Claude Code. Back on Trendshift's list, 9k stars per shields.io (repo).
- chenkangjie1123/VGGT-Diff: code for the six-photo new-angle research above. 54 stars per shields.io, non-commercial (repo).
Creative workflows
1. Batch-finish a folder of PSDs from the command line with PhotoCraft (photocraft-cli).
The steps. Install a v0.1.x build from the releases page. Run the README's own example on one file first: photocraft-cli run wave.psd --cmd filter.sharpen.smartSharpen --params '{"amount":80}' --cmd layer.newAdjustmentLayer.curves --params '{"points":[[0,0],[64,48],[192,212],[255,255]]}' --out wave-final.png. Once the look is right, loop the same line over a folder in your shell. To hand it to an agent instead, start photocraft-cli mcp and connect it as an MCP server.
How it works. The menus, the command line and the MCP server all call the same 500+ commands, so a recipe you test by hand is the recipe the script or agent runs.
Why it is good. A curves layer added this way is still an adjustment layer, so the client file stays editable instead of baked flat.
Where it breaks. Early alpha. One of 135 test PSDs failed to round-trip, so open a few outputs in Photoshop before trusting a batch. No generative fill. And "menu item live" is not the same as "matches Photoshop's math," so check your sharpening by eye.
2. Render Ideogram 4 locally on a 16 GB card and write your own caption, with the INT4 ONNX build.
The steps. Download the repo, run npm install, then node runtime/cli.js --package . --prompt "your idea" -o out.png to see the default. For anything with type, write a structured caption and pass it with --caption-file caption.json, which skips the caption writer (about 90 seconds saved). Use --preset turbo (12 steps) for drafts and quality (48) for finals, --seed to lock a look, and --chunked-prefill if your card is under 16 GB.
How it works. Ideogram 4 reads a structured caption, not a sentence. The bundled writer turns your idea into one, and that is where the invented signs come from. Writing the caption yourself removes the middleman.
Why it is good. Poster layouts with exact text, on a card you own, with no per-image fee.
Where it breaks. Non-commercial license, the grey "memorized mode" panel on short prompts (the Node runtime re-renders automatically), Firefox fails at 1024 × 1024, and the author claims no accuracy figure against the full model.
3. Dress a model in a full look with Garments2Look.
The steps. Make a collage of every piece (top, layer, bottoms, shoes, bag, belt). For repainting, grey out the clothing on the person photo in flat 128 grey. Number each item in the prompt with styling notes ("partially unbuttoned, tucked-in"), then end with the layering order. Run the repo's scripts/inference/inference.py --task inpainting (or editing) with the matching add-on, at the card's 40 steps, guidance 4.0, seed 123.
How it works. The add-ons learned from 20,000 outfit pairs where the numbered list matched the garments, so the numbers tell the model which collage piece goes where and which sits on top.
Why it is good. A lookbook comp in one pass instead of six separate swaps.
Where it breaks. No stated add-on license, fine accessory detail drifts, no ComfyUI node, and the mask is on you.
Worth testing
- PhotoCraft v0.1.1-rc.4: open your messiest PSD and save it back. Tradeoff: alpha software, keep a backup.
- Ideogram 4 INT4 in Chrome: run the local web page on a 16 GB card. Tradeoff: about three times slower than DirectML, and non-commercial.
- YuE2 WebGPU demo: a sung demo with no upload. Tradeoff: we could not confirm the Space is live, and output is non-commercial.
- ArtCraft: block a shot with mannequins before spending video credits. Tradeoff: generation is cloud-billed, and some models need your own API keys.
What actually matters from today's signal
The Photoshop rebuild is the story to watch, and not because it beats Photoshop today. It does not. It matters because it ships with an agent interface from version 0.1, while the big editors bolt AI on as a chat panel inside a subscription. If an agent can open a real PSD, add adjustment layers and save it back with layers intact, the repetitive half of retouching and versioning becomes a script you own. That is a bigger change to a working week than another image model.
Everything else today points the same way: Ideogram 4 on a mid-range card, a song model in a browser, voice cloning on a phone. The cost of running good tools locally fell again, and the cost now sits in memory you already bought rather than credits you rent.
The counter-signal is the paperwork. The Ideogram build, the song model and the new-angle research are all non-commercial, Garments2Look states no license, and the voice-cloning files depend on you having consent. Local does not mean free to sell. Test everything, but check the license before anything leaves your machine for a client.
Source access notes: Adversarial fact-check ran (Sonnet subagent) and caught: a wrong "same afternoon" for PhotoCraft's second build (now "later the same day", flagged as pre-release), an overreaching "use it for anything" (the brand has its own license), the parity reading stated as fact (now marked as our interpretation), Ideogram figures (WebGPU 10.8 s/step, peak memory in MiB, chunked mode at the same speed rather than "different results"), the LingBot 64 GB line stated as the card's (now our estimate), an unsourced "new" for Core AI, a miscounted non-commercial tally, and an absence claim in the opening (softened). Trendshift deltas could not be independently confirmed. Every Hugging Face date is createdAt from the API (/api/models/<id>, cache-busted) or from API listings sorted by creation date; the cloud shell could not reach Hugging Face directly, so all fetches went through WebFetch. The Hugging Face Spaces API and the YuE2 demo page returned 401, so the demo is listed as unconfirmed. GitHub's API worked for PhotoCraft's creation date but returned 403 for ArtCraft; GitHub's HTML and API star counts for PhotoCraft (16 and 77) lag far behind shields.io (1.3k) and Trendshift's +341 delta, and the briefing uses shields.io per house rules. Trendshift showed an implausible +69,240 delta for image-blaster, so no delta is given for it. OpenAI, Google, Adobe, ElevenLabs, fal, Replicate, BFL, Hugging Face's blog and Suno had nothing creative inside the window (Suno's Speech beta, October 1, ran in earlier briefings). Runway's and Stability's news pages returned no dated posts, Luma's showed only 2025 items, and Midjourney, Krea, Ideogram, Pika, HeyGen, Udio and Recraft surfaced nothing dated via WebSearch. ComfyUI's releases page returned stale content. akhaliq/MiniMax-H3-Claymation-Style-LoRA (October 5) was skipped because it has no model card. sky-meilin/JoyAI-Image-Edit-Plus-Diffusers (October 5) was skipped as a re-upload of JD's June 22 release.