FervorCreative AI
Live Latest 28.09.26 · morning 86 tools tracked 272 workflows indexed 226 topics Hot: Qwen-Image-2.1, ComfyUI, MiniMax H3

Instruction-based image editors redraw the whole frame a few percent off its original position, which is why edit passes never composite cleanly, and the fix starts with a five-minute measurement anyone can run.

Qwen-Image-2.1 Consistency LoRAQwen-Image-2.1ComfyUI-AusBossai-editingcomfyuiimage-gencreative-workflowslora-finetuning

Your AI Image Editor Is Moving the Picture. Here Is How to Measure It

Ask a modern editing model for a watercolour version and it hands back the whole scene a few percent taller. A new set of measurements shows how far, and a small add-on puts it back.

Ask Qwen-Image-2.1 to turn a photograph into a watercolour and it does. It also hands the picture back about five percent taller, nudged sideways, and differently every seed. A waistband that sat on one line in the original no longer sits on it. A horizon rises 25 pixels. The tip of a traffic cone drops 21. Nothing in the interface tells you this happened. Nothing in the model card mentions it. You find out when you drop the two images into a before-and-after slider and the slider looks wrong.

Most people using these models have felt this and filed it under "AI is weird." It is not weird. It is measurable, it is consistent in kind if not in amount, and somebody finally sat down and measured it.

The numbers come from ausboss, who published both the measurement and a fix on September 27. Across 36 restyles on pictures the model had never seen in training, the worst corner of the frame landed a median of 24.3 pixels away from where it belonged. Worst case: 61.1 pixels. Only 3 percent of those edits came back within 3 pixels of where they started.

That is the number I want you to sit with. Three percent. If you have ever tried to stack two edit passes, or comp a restyle back over a plate, or hand a client a slider, ninety-seven times out of a hundred the geometry was against you and you did not know it.

Why this happens, in plain terms

When you feed a picture into an instruction-based editor, the model does not paint on top of your picture. It reads the picture, reads your instruction, and draws a new picture from the beginning. The new one is meant to look like yours with one change. In practice it comes out framed slightly differently, the way a second photographer standing a step back would frame it.

There is a second, related problem. Ask for a new cardigan and the model also repaints the hair, the storefront sign and the traffic light behind your subject. It was never told to keep those. It was told to make a picture of someone in a different cardigan.

Both problems have the same root and they get measured differently, which matters. Drift is the whole frame moving. Repainting is pixels changing in places you did not ask about. If you line the edit up with the original first and then measure what changed, you separate them. That alignment step is what makes the second number honest, and it is the step most comparisons skip.

Measure yours first

Do this before you install anything. It takes five minutes and it works for whatever editor you use, not just this one.

  1. Pick one picture with a hard, unambiguous horizontal edge in it. A windowsill, a shoreline, a table edge. Something you can point at.
  2. Run one restyle. "Turn this into a watercolour painting" is a good stress test because restyles drift far more than local edits do.
  3. Open both images as layers in whatever you use, put the edit on top, and set the top layer to a difference blend.
  4. Look at your chosen edge. If it were in the same place, it would vanish. It will not vanish. Measure how far it moved, in pixels.
  5. Run it twice more with different seeds and measure again.

You now have a number and, more usefully, a spread. If your three runs disagree with each other by 20 pixels, no amount of careful prompting will fix that, because it is not a prompt problem.

I would do this even if you never intend to install a fix. Knowing the size of the error changes how you plan a job. Under a few pixels, you can ignore it. At 25 pixels, any workflow that assumes the edit lines up with the plate is already broken and you should design around it.

Fix one: train the drift out

The add-on file is qwen-image-2.1-consistency.safetensors. Two versions ship. Start with the step-1500 one.

Wiring it into a ComfyUI edit graph:

  1. Drop a LoraLoaderModelOnly node straight after your model loader. Strength 1.0. Anything lower lets drift back in, which the author tested and says plainly.
  2. In Text Encode Qwen Image 2.1, your picture goes in as image_1, resolution is set to 0, and your edit instruction is the prompt.
  3. Sample on that node's own canvas output, not a fresh blank one. 25 steps, guidance 1, euler/simple, denoise 1.
  4. Decode, then run the result through Split Image with Alpha, because this model's decoder returns four channels.

Step 3 has a trap in it. If you sample on a blank canvas of any other size instead of the encoder's own output, the model zooms by the ratio between the two sizes, and no add-on on earth can undo that. This catches people constantly. The size has to come from the encoder.

With the step-1500 file loaded, that 24.3-pixel median drops to 1.6, and the share of restyles landing within 3 pixels goes from 3 percent to 75. Repainting outside the edited object falls from 12.8 percent of pixels to 7.5. The step-2000 file lines up tighter still, 0.9 pixels median and 97 percent within 3, and it pays for that by washing colour out of painted restyles: whiter paper, less ink. The author measured that too rather than hiding it, and he recommends 1500.

One number in his card deserves special credit. He reports that simply sending a picture through the model's own encoder and decoder, with no edit at all, changes about 2 percent of its pixels. That is the floor. Nothing built on this model can do better. Publishing the floor you cannot beat, right next to the number you achieved, is the single most trustworthy thing in the whole release.

Fix two: put it back afterwards

The same author's ComfyUI pack, ComfyUI-AusBoss, ships a node called Realign to Source that goes at the problem from the other end. It takes the edit and the picture it came from, measures the zoom and shift between them, and warps the edit back onto the source's frame at the source's size. It also reports what it measured, which makes it a measuring instrument as much as a fix.

Two things to know. It corrects the whole-frame move, not shapes that a restyle genuinely redrew somewhere new, so the result lines up closely rather than pixel for pixel. And an edit made at the picture's own size will have pushed a thin strip out of view, so warping it back leaves an empty margin. The node marks that strip with a mask so you can crop or fill it. The tidier route is to pad the picture with mirrored edges before the edit and hand the node the padding information, so the edges come back as real picture.

Which fix you want depends on what you are doing. The add-on prevents the problem and costs you a slot in your model stack. The node repairs the problem and costs you a resampling pass. For a batch of restyles headed for compositing, I would use both: the add-on so the drift is 1.6 pixels instead of 24, and the node to take the last of it out.

Where all of this breaks

The add-on was tested at around one megapixel with those exact settings. Above that, with fast few-step add-ons, or at guidance above 1, nobody has checked. The author says so rather than implying broad coverage.

Comic and anime restyles redraw every outline from scratch, so individual shapes still move a few pixels under their own power. The worst held-out restyle was 12.9 pixels off at a corner even with the fix loaded. If your work is heavily stylised line art, you get less out of this than a photographer does.

It does not make weak edits stronger. If the base model will not do the thing you asked, this changes nothing about that. It only controls where the result lands.

The instructions it learned were English. The licence is the base model's Qwen Research License, which means research and evaluation, not paid work. That last point is worth reading twice before you build a client pipeline on it, and it applies to the base model too, not only the add-on.

And the two sources disagree slightly on how bad the underlying drift gets. The model card's examples run to about 5 percent. The realign node's documentation says up to about 12 percent. Both are that author's own measurements on his own material, so treat the range as the range and measure your own.

The part that outlives this model

Qwen-Image-2.1 will be replaced. The add-on will stop mattering. The measurement will not.

Every instruction-based editor redraws the whole frame, which means every one of them can drift, and almost none of them report it. Overlay your output on your input, set a difference blend, and look at where the edges land. Five minutes, no install, works on anything.

The thing you cannot name, you cannot fix. Twenty-four pixels of drift has been sitting in everybody's work for a year, and the only reason it is fixable this week is that somebody bothered to put a number on it. Go put a number on yours.


Medium metadata

Title: Your AI Image Editor Is Moving the Picture. Here Is How to Measure It

Subtitle: Ask a modern editing model for a watercolour version and it hands back the whole scene a few percent taller. A new set of measurements shows how far, and a small add-on puts it back.

Tags: ComfyUI, AI Image Editing, Generative AI, Creative Workflow, Open Source AI

Estimated read time: 8 minutes