FervorCreative AI
Live Latest 05.10.26 · morning 110 tools tracked 374 workflows indexed 284 topics Hot: MiniMax H3, fal, Humaneness Voice Small

AI try-on fails on full outfits because nothing tells the model what sits on top of what, and Garments2Look fixes that with a numbered item list and a written layering order, which makes a whole-look comp a single edit, at the cost of rented hardware, a hand-made mask and an add-on license nobody has stated yet.

Garments2Look LoRAQwen-Image-Edit-2509Garments2Look datasetimage-genai-editinglora-finetuningprompt-craftcreative-workflows

Garments2Look: Dress a Whole Outfit in One AI Edit by Writing the Layering Order

A new pair of add-ons for Qwen-Image-Edit-2509 puts a full look on a model from a collage of product shots. The trick is a prompt that reads like a stylist's notes. Here is the workflow, the cost and where it still fails.

Ask any AI image editor to put a shirt on a person and it will usually manage. Ask it to put on a shirt, a sweater over it, trousers, loafers, a belt around the waist and a bag on the shoulder, and something goes wrong. The belt ends up under the sweater. The shirt that should be tucked in hangs loose. The bag merges with the sleeve.

The researchers behind Garments2Look say exactly this about current methods: in their paper's words, existing approaches "struggle with complete outfit visualization, proper layering inference, and styling accuracy, resulting in misalignment and visual artifacts." Their answer is not a bigger model. It is a better way of saying what you want.

The interesting part of this release is a prompt format, not the files. You number every item, add a styling note to each one, and end with a line that spells out the order they stack in, like a dresser's checklist. My position: that format is the reusable idea here, and it is worth learning even if these particular add-ons get replaced next month.

What shipped

Garments2Look-LoRA went up on Hugging Face on October 3. It contains two small add-on files for Qwen-Image-Edit-2509, Alibaba's open image-editing model. Each add-on teaches the base model one job.

The inpainting add-on fills in a region you have greyed out on a photo of a person. You keep everything else and repaint only the clothes.

The editing add-on takes a person who is already dressed and replaces the outfit outright, no mask needed.

Both take two pictures: the person, and a collage of every item in the look. The team behind it published the work as a CVPR 2026 paper, alongside a dataset the README puts at 98,012 outfit records (the paper's abstract cites 80,000 outfit pairs). The add-ons were trained on 20,000 of those examples.

The licensing is mixed, so get it straight before you start. The base model is Apache 2.0. The Garments2Look code and dataset are Apache 2.0, which permits commercial use. The add-on files themselves have no license yet: the team's own README says no explicit adapter license has been specified yet. Until that changes, treat this as something to test, not something to put in a client deliverable.

Why a numbered list works

Think about how a stylist briefs a dresser on set. Not "make her look smart casual." Something like: white shirt, half unbuttoned, tucked in. Grey cardigan over it, open. Belt over the shirt, under the cardigan. Loafers. Bag on the left shoulder.

That brief carries two kinds of information a normal prompt drops. It ties each instruction to one specific garment. And it says what goes over what.

Garments2Look's prompt carries both. Here is the example from the model card, word for word:

Keep the woman's identity, pose, background in Figure 1 unchanged, wearing the outfit in Figure 2, include (1) a top (partially unbuttoned, tucked-in), (2) a sweater (unbuttoned), (3) pants, (4) loafers, (5) a bag, (6) a belt (worn around waist). Layering Order: (1) -> (6) -> (2).

Read the last line. Top first, then the belt, then the sweater. So the belt sits over the shirt and under the open sweater, which is exactly what a stylist would do.

The numbers work because the collage is assembled in the same order as the list. The repo notes that missing collages are generated "from reference item images in prompt order." So item (1) in your sentence and the first piece in your collage are the same garment, and the model learned during training to match them up. You are not hoping the model guesses which grey knit is the sweater. You are pointing at it.

The first line does the other job. "Keep the woman's identity, pose, background in Figure 1 unchanged" tells the model what is off limits. Every good editing prompt has that sentence.

Put this into practice

The model card publishes one tested setup: 40 steps (how many passes the model takes to refine the picture), guidance 4.0 (how strictly it follows your words) and seed 123 (a number that makes a run repeatable). Start there and change one thing at a time.

1. Rent the right hardware. The team validated this on an NVIDIA H200, a data-center card you rent by the hour from a cloud provider, and they publish no figure for smaller cards. Unless you own something very large, plan on renting one for a session. Set up the environment the README describes (Python 3.10, PyTorch 2.7.1), download Qwen-Image-Edit-2509, and download the two add-on files.

2. Build the collage in prompt order. Lay out a clean product shot of each item, in the order you will list them. Top, sweater, pants, shoes, bag, belt. Plain backgrounds help the model see where each item ends. Write the order down so the collage and the sentence never drift apart.

3. Pick your mode. For a photo where the clothes are wrong but everything else is right, use inpainting. Paint the clothing area flat grey, the exact middle grey with a value of 128 on a 0 to 255 scale, and leave the face, hands and background alone. The project's own masks are slightly enlarged past the garment edge, so give yourself a margin. For a person who is already dressed and you want to start over, use editing and skip the mask.

4. Write the brief. Open with the keep sentence. Number every item, matching the collage. Add a styling note in brackets to anything that is not worn the default way: tucked in, unbuttoned, sleeves rolled, worn around the waist. End with "Layering Order:" and the chain of numbers for the items that overlap. You do not need to list the shoes and bag in the chain if they do not cover anything.

5. Run it. The repo's script takes the task, the base model folder, the add-on file, the two images and the prompt file:

python scripts/inference/inference.py --task inpainting --model-dir models/Qwen-Image-Edit-2509 --lora models/Garments2Look-LoRA/Qwen-Image-Edit-2509-LoRA-2-refer-20k-inpainting-epoch-1.safetensors --origin input.png --ootd ootd.png --prompt-file prompt.txt --output output.png --seed 123 --steps 40 --cfg-scale 4.0

For editing, change --task to editing and swap in the editing add-on file.

6. Change one thing, then compare. Keep the seed fixed and move only the layering line. Run it as (1) -> (6) -> (2), then (1) -> (2) -> (6). If the belt moves from under the sweater to over it, you have proof the model is reading your order, and you know how far you can push it.

7. Check the details at full size. Look at the buttons, the buckle, the bag strap and the hems. Those are where this approach drifts.

Where it breaks

No add-on license. This is the big one. The dataset and code are commercially usable, the add-on files are not yet licensed at all. Do not use the output in paid work until the team publishes terms.

Fine details drift. The card says plainly that "fine accessory details, garment fidelity, styling, and pose preservation may vary." A buckle may change shape. A logo on a bag may turn into a smear. A product page needs the exact product, and this is not there yet.

The hardware bar is high. Validated on an H200, nothing published below that. That makes every test a rented session, which changes how you iterate: batch your ideas before you start the clock.

The mask is your job. For inpainting, a sloppy mask gives a sloppy edge. The project's own masks are slightly dilated; copy that habit.

No ComfyUI node or hosted demo that I could find. You are running a Python script. If you live in a node graph, wait for a port or build one.

Poses do not adapt. The keep sentence locks the pose, which is what you want for consistency and the wrong thing if the outfit calls for a different stance. A long coat on a person mid-stride may hang strangely.

Two references only. The method takes one person and one collage. If you want a second angle of the same look, you are running it again and hoping the details match.

The habit worth keeping

Strip away the files and what is left is a way of briefing a machine the way you would brief a person: name each item, say how it is worn, say what goes on top. That works because it removes guesswork from the part of the job the model has always found hardest.

Try it on one look you have already shot, where you know what the right answer is. Write the numbered brief, run it, then flip the layering order and run it again. If the belt moves when you tell it to, you have a new way to comp a lookbook before the samples arrive. If it does not, you have learned exactly where the model stops listening, and that is the line to watch when the next version lands.


Medium metadata

Title: Garments2Look: Dress a Whole Outfit in One AI Edit by Writing the Layering Order

Subtitle: A new pair of add-ons for Qwen-Image-Edit-2509 puts a full look on a model from a collage. The numbered prompt that controls what sits on top, the full workflow, and the license gap you need to know about.

Tags: Fashion, AI Image Editing, Virtual Try-On, Generative AI, Photography

Estimated read time: 8 minutes