Creative AI Briefing: Monday, September 28, 2026
Take a portrait you shot at f/5.0, click the eye you meant to focus on, and get it back rendered as if you had shot it at f/2.8. That landed Sunday afternoon with a paper behind it and code you can run. It is one of four things released in the last two days doing the same kind of work: reaching back into a picture that already exists and handing you a camera decision you thought was locked at capture. Focus and aperture. Camera angle. Whether an edit moves the frame under your feet. Depth and normals as clean maps. None of these are new models, and that is the point. The generators stopped being the bottleneck a while ago. What you could not do was aim them.
Image
AnyBokeh changes the focus point and the f-number of a photo after the fact, using optics rather than a blur filter. itsmag11/AnyBokeh went up September 27 at 16:50 UTC with code and an arXiv paper from S-Lab at Nanyang Technological University, accepted at NeurIPS 2026. Two adapters stacked on FLUX.1-Fill-dev: the first reads how far every pixel sits from the sharp plane, the second re-renders at whatever focus point and aperture you name. A local tool (focus_picker.py) serves the image in a browser so you click the spot instead of guessing coordinates. Catches: about 27 GB of graphics memory end to end, S-Lab plus FLUX.1 [dev] non-commercial licensing so client work is out, and both the training code and the UnrealBokeh dataset still marked to-do. Ten stars this morning, which tells you how early this is.
Someone finally measured how far Qwen's edits drift, and shipped the fix. ausboss/Qwen-Image-2.1-Consistency-LoRA, created September 27 at 21:27 UTC, is a small add-on that stops Qwen-Image-2.1 redrawing an edit a few percent taller or nudged sideways. His held-out numbers: across 36 restyles, the worst corner landed 24.3 pixels off without the add-on and 1.6 with it, and restyles arriving within 3 pixels went from 3 percent to 75. On local edits, the picture repainted outside the thing you asked to change fell from 12.8 percent to 7.5. He publishes the floor too: decoding a picture through Qwen's own image writer and back changes about 2 percent, so 7.5 is progress, not perfection. Two checkpoints ship; the later one lines up tighter and washes colour out of painted restyles, and he says so. Qwen Research License, so it inherits the base model's non-commercial terms, which cover research and evaluation and nothing you get paid for.
And you can move the camera around a still image by building a rough 3D stand-in first. lilylilith/QI_2.1_AnyAngle, created this morning at 05:08 UTC under Apache 2.0, trained on real Blender renders plus orbit shots generated with MiniMax H3. Workflow below. Its failure mode is documented with a picture of it failing, which is rarer than it should be.
One model, five kinds of control map. FunAILab/Open-Vision-Banana (September 26, 23:51 UTC) swaps a fine-tuned transformer into the FLUX.2 [klein] base 9B pipeline and does semantic, instance and referring segmentation plus depth and surface normals, all selected by the prompt, with your own classes and colours. If you keep a pile of separate preprocessor models for control work, that is one download instead of five. FLUX non-commercial.
Audio and music
Thin. becausereasons/yue2-trbdr-folk-troubadour, created September 27 at 22:18 UTC, is a style add-on for YuE2 trained with ai-toolkit and aimed at American folk, singer-songwriter and protest material, harmonica included. CC BY-NC 4.0, so a sketchpad rather than a delivery tool. It matters mainly as evidence that people now train music styles the same casual way they train picture styles.
Open and local
The other thread this weekend runs underneath the first: the runtime keeps leaving the workstation. Two web exports landed September 27 within seven minutes of each other, both from Saimon8420. restoreformer-pp-web is blind face restoration in a browser tab on WebGPU, Apache 2.0, storing its big weights at half precision and converting them back inside the graph, which halves the download to roughly 147 MB and lets it run on graphics chips with no half-precision shader support. It powers the free Doetra Photo Restorer. Its companion realesr-general-x4v3-web is a 4.87 MB general upscaler under BSD-3, small enough to ship on any page.
On phones, Gestura published its whole conversion stack September 25: text to video with synchronised audio, on the handset. The honesty is the useful part. On a Samsung Galaxy S26, a five-second clip at 256 by 256 takes about 431 seconds, and two reference images push past ten minutes. Read the licence first: the MiniMax H3 Community terms exclude the European Union, the United Kingdom, South Korea and the United States from the granted territory.
- itsmag11/AnyBokeh: refocus and restop a photograph you already took, with a picker so you click the focus point instead of typing coordinates. Created September 27; code repo at 10 stars, so expect rough edges.
- ausboss/Qwen-Image-2.1-Consistency-LoRA: makes Qwen edits land on the original's frame. Created September 27, trained with ostris/ai-toolkit (12k stars).
- lilylilith/QI_2.1_AnyAngle: point the camera anywhere around a still image without the style drifting. Created this morning, Apache 2.0, 12 likes in its first hours.
- Saimon8420/restoreformer-pp-web: face restoration in a browser tab, no install, nothing uploaded. Created September 27, Apache 2.0, from RestoreFormerPlusPlus (291 stars).
- mlx-community/Ming-Image-0.1-Design-Layer-8bit: flat artwork split back into editable layers on Apple Silicon, in 4-bit, 8-bit and full-precision builds. Created September 26 and 27, MIT.
- mlx-community/VOSR2-fp16: video super-resolution for Macs. Created September 26, Apache 2.0.
- alibabagroup/SparkWan2.1-T2V-14B-720P-0.97Sparsity: Alibaba's own sparse-attention rebuilds of Wan, including a three-step variant, all Apache 2.0. Created September 24.
Creative workflows
1. Refocus a photograph you already shot. Files: itsmag11/AnyBokeh, stages in stage1/ and stage2/, driven by inference_full.py from the GitHub repo.
The steps. Accept the FLUX.1-Fill-dev licence on its model page and sign in; the base model is gated. Run python focus_picker.py --image_path yourphoto.jpg, open the local page it prints, click the point you want sharp, copy the two numbers. Then run inference_full.py with --focus_x, --focus_y, the f-number it was shot at and the one you want. Do not know the original f-number? Drop both and use --bokeh_scale 2.0 to double the blur or 0.5 to halve it. The two stages also run separately, which matters: save the stage-one maps once and re-render that frame at a dozen focus points without paying for the analysis again.
How it works. The first pass estimates how far every pixel sits from the sharp plane and roughly what is near and far. The second repaints the image from those maps plus your target, so blur follows the scene's geometry rather than a cut-out around the subject. No calibration pass, because the relation it fits between blur size and distance is a straight line whose slope it reads off your picture.
Why it is good. It fixes the one mistake Lightroom cannot. A Gaussian blur behind a cut-out subject has never fooled anyone who looks at photographs, and working from an optical model means edges and out-of-focus highlights behave the way a lens behaves.
Where it breaks. About 27 GB of graphics memory for the full run puts it on a rented card. S-Lab plus FLUX non-commercial keeps it out of paid work. Training code and dataset are unreleased, so nobody outside the lab can retrain or verify it. And the output is a re-render, not a correction: everything goes through the model, so expect small changes you did not ask for.
2. Measure whether your edit model is moving the picture, then stop it. Files: qwen-image-2.1-consistency.safetensors (start with the step-1500 file, not the 2000).
The steps. Run a baseline test first: one picture, ask Qwen-Image-2.1 for a watercolour, lay the result over the original in a difference blend, and see where a hard edge lands. Then add a LoraLoaderModelOnly node straight after the model loader at strength 1.0. Feed your picture into Text Encode Qwen Image 2.1 as image_1 with resolution set to 0 and the edit instruction as the prompt. Sample on that node's latent output: 25 steps, CFG 1, euler/simple, denoise 1. Decode, then pass through Split Image with Alpha, because this decoder returns four channels. Re-run the test and compare.
How it works. Trained on pairs where the drift had already been measured and removed, so every target sits exactly on its source's frame. The set includes reverse pairs (turn this watercolour back into a photo), which is why it holds in both directions.
Why it is good. Drift is why your before-and-after sliders never line up and your edit passes will not composite. It also stops the model repainting the storefront sign when you asked for a different cardigan. Every number is measured on held-out pictures with seeds and settings stated, and the author publishes the floor his own method cannot beat.
Where it breaks. Sample on a latent of any size other than the encoder's output and the model zooms by the size ratio; nothing undoes that. Strengths below 1.0 let drift back. Comic and anime restyles redraw every outline, so shapes still move a few pixels; worst case 12.9 at a corner. Untested above roughly one megapixel, with fast few-step add-ons, or at guidance above 1. And it keeps a weak edit weak.
3. Put the camera somewhere else without losing the style. Files: lilylilith/QI_2.1_AnyAngle.
The steps. Generate a rough 3D version of your scene with a splat or mesh generator, import it into Blender, place a camera where you want the new shot, render one frame. That render will look coarse, which is fine. Feed both pictures into Qwen-Image-2.1 with this add-on at strength 1.0 and the prompt Change the camera angle from <image2> to <image1>. Guidance 3.0 and 20 steps or more for finals.
How it works. The rough render is the instruction, not the output. It carries the geometry of the new angle and the model repaints your original's style onto it, so the style holds where it drifts when you just ask for an angle in words.
Why it is good. Apache 2.0, which is unusual here, and it gives arbitrary angles rather than a menu of presets.
Where it breaks. Everything rests on the 3D stand-in: a coarse or confused splat puts objects in the wrong place, and the card shows this happening to a table and a ponytail. Small or soft faces come back deformed. Minutes per shot, not seconds.
Worth testing
- Doetra Photo Restorer, free, in the browser, nothing uploaded. Tradeoff: faces must be found and aligned first, so group shots and profiles are hit or miss, and a restoration is an invention, not a recovery.
FunAILab/Open-Vision-Bananafor control maps. Tradeoff: FLUX non-commercial, and a 9B picture model doing segmentation costs far more than a dedicated one.mlx-community/Ming-Image-0.1-Design-Layer-4bitfor pulling flat artwork into layers on a Mac. Tradeoff: 4-bit is the build that fits comfortably and the one most likely to smear fine type.alibabagroup/SparkWan2.1-T2V-14B-720P-0.95Sparsity-3Step, three steps at 720p, Apache 2.0. Tradeoff: no published creator-facing timings, so budget a test afternoon.
What actually matters from today's signal
Three of the four image releases this weekend are corrections, not capabilities. AnyBokeh corrects focus. The consistency add-on corrects frame drift. AnyAngle corrects camera position. That is what a maturing tool space looks like, and notice where the correction work comes from: a university lab, an independent trainer, and someone who mostly posts on Civitai. No large lab shipped anything this weekend that touched these problems.
The consistency add-on deserves a second look for a reason that has nothing to do with Qwen. Its author did what almost nobody doing this work does: defined a number, measured his baseline, measured his fix, and published the floor his method cannot go below. Twenty-four pixels of drift was always there. Everyone using that model has been fighting it without knowing what they were fighting. That lesson carries to whatever you use. Overlay your output on your input and look at where the edges land, because the thing you cannot name, you cannot fix.
The counter-signal is licensing, and it is getting worse rather than better. AnyBokeh is non-commercial twice over. Open-Vision-Banana is non-commercial. The consistency add-on inherits a research-only base. The Android video stack excludes four of the largest creative markets on earth from its granted territory. You can learn from all of it and ship almost none of it, and the gap between what is downloadable and what is usable is now the first thing to check before you build a week around a new toy.
Source access notes: Hugging Face listing endpoints served stale data on the first pass (the text-to-image feed returned results dated July 27); all listings were re-fetched with cache busting and every date here comes from the API's createdAt field. blog.comfy.org is a Substack JavaScript wall and returned no posts. The GitHub releases API returned an empty body for ComfyUI, so no ComfyUI release is claimed. Star counts are shields.io totals fetched this morning with cache busting. Midjourney, Civitai and Kling were unreachable and are not cited. Adobe's blog was read; its most recent creative items (Topaz Labs acquisition completed September 23, Premiere on Android September 22) fall outside this window and are noted rather than covered. A fact-check pass ran against this draft. During research it caught that a browser music bundle surfacing September 27 is a mirror of a May 22 export rather than a new release, so that item was cut rather than dated forward, and the demo repository its README points at (github.com/lsb/stable-audio-3-small-music-onnx) does not resolve. Against the draft it verified every date, licence, file size, star count and measured figure, and found one error: the Qwen Research License was called evaluation-only when it covers research too. Corrected. Sections omitted for lack of material: Video, Writing. Sections omitted for lack of material: Video, Writing.