FervorCreative AI
Live Latest 11.10.26 · morning 133 tools tracked 463 workflows indexed 344 topics Hot: ComfyUI, MiniMax H3, Qwen-Image-2.1

Akatz Labs' free H3 Person Remover workflow removes a person from a short clip by never letting the video model see them, so the result is decided by two human jobs, painting a clean first frame and checking the mask, while short rerollable chunks keep a bad patch cheap, and a regional license excludes most professional readers by default.

H3 Person Remover LoRAMiniMax H3SAM 3.1H3 RelayAkatz Labsvideo-genai-editingcomfyuicreative-workflowslocal-creative-ailicensing-provenance

Remove a Person From Video With AI: Why the Frame You Paint by Hand Decides the Whole Shot

Clean plates are one of the oldest chores in post. A stranger crosses the back of a perfect take, and somebody spends an afternoon painting them out frame by frame, or the shot gets cut.

A free workflow released on October 8 handles that chore with an open video model, and the clever part is what the model never sees. It never sees the person. It sees a green hole where they were, and a single frame you cleaned up yourself, and it fills the hole from there.

That design choice moves the craft. The quality of the result depends less on the AI than on two jobs that stay in human hands: painting out the first frame, and checking the mask. Get those right and the rest is waiting.

What it does

The workflow is the H3 Person Remover from Akatz Labs. It runs in ComfyUI, the free node-based app many people use for local image and video generation, on top of MiniMax H3, a downloadable video model that makes picture and sound together.

You give it a short clip and a clean version of the clip's first frame, with the person already removed. You type a description of who to remove ("man in gray shirt"). It returns the clip with that person gone and the background rebuilt behind them, moving with the camera.

Your soundtrack survives too, if you wire it back in. More on that below.

How it works, in plain terms

Three separate pieces do three separate jobs, and seeing them apart is what makes the workflow easy to debug.

Finding the person. A tracking model called SAM 3.1 reads your description, finds the person in the first frame, and follows them through the clip. Everywhere they appear gets painted solid green.

Knowing what belongs there. Your clean first frame tells the video model what the scene looks like without the person. It is the only ground truth the model gets.

Filling the hole over time. The video model regenerates the green area in short chunks, and each chunk inherits the last 18 frames of picture and sound from the one before it, so the fill stays continuous from chunk to chunk.

The add-on itself does surprisingly little of the hard thinking. As the workflow notes put it, you remove the person from the first frame "externally with an image editor or image-editing model before starting this workflow." Photoshop, an AI editor, a clone stamp: whatever gets you a believable empty frame.

The first frame is the shot

Here is the position this whole article rests on: the clean first frame is the most important input you will make, and it deserves the care you would give a matte painting.

The reason is how the chunks chain. Every later chunk inherits from the one before it, and the first chunk inherits from your painted frame. The author says it directly: "A poor clean first frame can propagate through later windows." A smudge on a wall in frame one becomes a smudge that follows the camera for five seconds.

The workflow notes add one more requirement: the clean frame "must match the source frame's size and composition." No crop, no reframe, no color grade that the rest of the clip does not have.

So spend your time there. Rebuild the background where the person stood with real texture. Match the grain. Check edges at full size. Ten extra minutes in an image editor saves an hour of rerolling video.

The mask is the second human job

The other place quality is decided is the mask, and the workflow lets you check it before spending any render time.

The notes describe a preview trick: set the final save node to Never, run the workflow, and you get only the green mask. Then look, specifically, at the things the notes name: "Inspect hands, hair, shadows, and other people."

Those are the four places masks fail. A hand that swings out of the tracked shape stays in the shot as a floating hand. Hair edges leave a ghost outline. A shadow on the floor is not the person, so the tracker skips it, and you get a shadow walking across the room by itself. And if two people match your description, the tracker may pick the wrong one, so make the description more specific ("man in gray shirt on the left").

When the mask looks right, set the save node back to Always and render.

Short chunks keep a bad patch cheap

Video models have a familiar way of wasting your evening. You render a whole clip, one second in the middle goes wrong, and you render the whole clip again and hope.

This workflow avoids that. It renders in chunks the notes call windows, and the Person Remover node shows each finished window as a preview card with its own Reroll button. In the workflow's words, "Reroll keeps the earlier prefix and rebuilds the selected window and its continuation." The good part stays good. Only the broken part and what follows it get redone.

Window lengths come in fixed steps (22, 39, 56, 73 frames). The author recommends starting at 22, because longer windows use more memory without reliably looking better. The companion node pack, H3 Relay, manages the windows and saves your accepted ones, so rerunning the same inputs returns the take you already approved.

One warning from the notes: those saved takes "live in the local ComfyUI runtime." Move the workflow to another machine and the approved windows do not come with it.

Put this into practice

Pick one short shot you already own, ideally a locked or gently moving camera, about five seconds, one person to remove.

  1. Conform the clip. 24 frames per second, with width and height both divisible by 32. Trim to a single continuous shot.
  2. Paint the first frame. Export frame one, remove the person in your image editor, match grain and edges, and export at the exact same size.
  3. Install the pieces. Update ComfyUI until it has native MiniMax H3 and SAM 3.1 support. Download H3-Person-Remover-V1.safetensors into ComfyUI/models/loras/, and load the workflow file examples/Person-Remover-Window-Reroll-V1.json.
  4. Load your two inputs. The clip goes in Load Video, your clean frame in Load Image.
  5. Describe the person and preview the mask. Check hands, hair, shadows and anyone else nearby.
  6. Run the first test at the recommended settings. The notes suggest a 22-frame window, 12 sampling passes and the add-on at full strength (1.0) for a first try.
  7. Review window by window. Reroll only the windows that fail.
  8. Put your sound back. The output is silent by default. Connect Get Video Components (audio) to Create Video (audio) on the clean-output branch, and your original soundtrack rides along.
  9. Difference-check the result. Stack the output over the original in your editor in Difference blend mode. Anything that lights up outside where the person was is a change you did not ask for.

That last step matters for a reason the next section explains.

Where it breaks

The whole frame is regenerated. The author is plain about this: "H3 regenerates the full scene." Pixels outside the green area are not locked to your original. They usually come out close, and close is the trap, because you will not notice a softened sign or a shifted window frame until a client does. Check with a Difference blend every time.

It trusts your mask completely. Whatever the mask misses stays: a hand, a shadow, a reflection in a shop window. A larger mask catches more, but asks the model to invent more background.

Camera moves and cuts break it. Strong camera motion or hard cuts can break continuity between windows. This is a tool for short continuous shots.

You need a serious card. The author validated the workflow on an RTX 4090 and calls that "a tested configuration, not a minimum hardware specification." The model card publishes no timings and no hosted demo, so plan on your own high-end NVIDIA card or a rented one.

The license may exclude you outright. This is the one to read before you start. The add-on is distributed under the MiniMax H3 Community License, and per the model card, "The standard territorial grant excludes the US, EU, UK, and Republic of Korea." If you work in those places, you need separate authorization from MiniMax before using it. That is not fine print. For most professional readers of this article, it is the headline.

Rerolls change what follows. Rerolling a window rebuilds every window after it too, so a fix in the middle means checking the end again.

The paint-out never really went away

AI was supposed to make clean plates disappear as a job. This workflow suggests something more honest: it makes the job smaller and moves it to the front. One carefully painted frame and one carefully checked mask now carry a whole shot, where they used to be the first of a hundred frames.

Before you install anything, try the human half. Pick a shot, paint frame one as if it were the only frame anyone would ever see, and ask whether it would hold up. If it would, the model has something good to work from. If it would not, no amount of rerolling will fix it.


Medium metadata

Title: Remove a Person From Video With AI: Why the Frame You Paint by Hand Decides the Whole Shot

Subtitle: A free ComfyUI workflow built on MiniMax H3 rebuilds the background behind a removed person in short, rerollable chunks. The results depend on two jobs that stay human, and a license whose standard grant excludes the US, EU, UK and South Korea.

Tags: Video Editing, Visual Effects, Generative AI, ComfyUI, Post Production

Estimated read time: 8 minutes