FervorCreative AI
Live Latest 10.10.26 · morning 128 tools tracked 448 workflows indexed 333 topics Hot: ComfyUI, MiniMax H3, ArtCraft

The useful part of TEIN's free Genjutsu workflow is less the character swap than its order of operations, which hides the performer from the video model, makes you approve a mask and then a draft before any full-quality render, and puts the original soundtrack back at the end, a habit worth copying for any paid AI video tool.

Genjutsu Open Source WorkflowTEINSeedance 2.5EnhancorSAM 3Demucsvideo-genai-editingcreative-workflowsavatar-gen

AI Video Actor Replacement That Keeps Your Shot and Your Soundtrack, and Makes You Approve Twice Before You Pay

The hard part of replacing a person in a video with AI is not the new face. Every major video model can invent a convincing character now. The hard part is everything around the face: the camera move you planned, the room you lit, the timing of the performance and the music under it. Ask a video model to "replace the dancer with a knight" and you often get a lovely knight in a slightly different room, moving to a slightly different beat, with a soundtrack that is no longer yours.

A small free project published this week takes an unusual route around that. The Genjutsu Open Source Workflow, released October 6 by Sirio Berati of TEIN under the MIT license, never shows the video model your performer at all. It shows the model a moving silhouette of them, placed in the real background, and asks it to paint a new character into that shape. Then it hands you the original audio back, untouched.

That trick is clever. But the part I would steal for every AI video job is simpler. The workflow will not send anything to the paid video model until you approve the mask (the cutout outline of your performer), and it will not render the full-quality version until you approve a draft.

What it makes

You give it a clip and pictures of the character you want. It gives you back the same shot, same camera, same background and same soundtrack, with your performer replaced by that character, at 1920x1080.

It runs as a small app on your own computer that you open in a browser. It needs no graphics card, because the heavy lifting happens on paid online services: Replicate for the preparation steps and Enhancor's hosted version of ByteDance's Seedance video model for the generation. You bring your own accounts for both.

The repo's validation notes describe one full live run: the HD export came out at 1920x1080, and "all 189 original audio packets match byte-for-byte," meaning the soundtrack you hear at the end is literally the file you started with.

How it works, in plain terms

Think of it as a compositing pipeline where AI does each step a junior artist might do.

First, it separates the sound. An open-source audio tool called Demucs, run here as a hosted service, splits the clip's audio into vocals and music. The vocals are shifted up three semitones, with timing unchanged. The project does not say why. One reasonable guess is that it keeps the model from treating the original singer's voice as something to copy, but that is my guess, not the author's statement.

Second, it turns your shot into a depth picture. A depth map shows how far each point is from the camera as color. Your performer becomes a colored shape with real volume and real motion, but no face, no skin and no costume.

Third, it cuts the performer out. Meta's SAM 3, a tool that finds and outlines objects in video, draws a mask around the person in every frame.

Fourth, it builds a hybrid frame. The colored depth version of your performer is laid over the original, untouched background. So the video model sees your real room, your real camera move and a faceless moving figure where the actor stood.

Fifth, it asks the model for a draft. The hybrid clip and your character pictures go to Seedance as a draft request. The model keeps the motion and the room and invents the character.

Last, it puts your audio back. Whatever sound the model returns is thrown away and replaced with the original track.

Why the depth trick works: the model has nothing to copy from your performer except how they move. That is exactly what you wanted to keep.

The two gates are the real lesson

Here is the order of approval, straight from the project's agent rules:

  1. "Mandatory visual mask QA before EVERY generation." You look at the mask on the full clip. The rules add that "automated checks and sampled contact sheets alone do not establish full visual approval," and "never approve a known-bad mask."
  2. A draft renders. You watch it.
  3. "Never automatically purchase an HD upgrade: the user approves that step." The 1080p version is a separate, separately billed request that reuses the approved draft's settings.

Anyone who has paid for AI video knows why this matters. The expensive failure is not a bad generation. It is a bad generation you could have predicted from a bad mask, rendered at full price, three times, because you were hoping.

If the mask leaks onto a doorframe, the knight will grow out of the doorframe. You can see that in the mask in ten seconds. You cannot unspend the money after the HD render.

That habit transfers to any tool. Check the input the model will actually see. Judge a draft. Pay for full quality last.

Put this into practice

The repo's README lists the exact steps. In plain terms:

1. Set up once. Download the project, run setup.sh, and add two keys to a plain settings file called .env: an Enhancor key for the video model and a Replicate key for the preparation steps. Both accounts need money on them. You also need Python 3.12 or newer and a small free audio tool called Rubber Band. The author tested a clean install on an Apple Silicon Mac; other systems are untested, and Windows users are told to use WSL2, Windows' built-in Linux layer.

2. Launch it. Run start.command and open http://127.0.0.1:8770 in your browser.

3. Pick the right shot. One clear subject, reasonably centered, with a background you like as it is. Busy crowds and people crossing in front of each other make masks fail.

4. Review the mask, all of it. Scrub the whole clip, not a few frames. If the outline grabs a chair or loses a hand, the coding agent that runs the workflow can retry the cutout on the original color footage. Approve only when it is clean.

5. Read the draft as a motion test. Is the timing right? Does the character stay on the figure? Is the camera move intact? Judge motion and placement here, not fine texture.

6. Approve HD only once. If the draft is wrong, fix the mask or the reference pictures and draft again. Do not "see if 1080p fixes it."

7. Check the final against your edit. The original audio comes back, but sync it against your timeline anyway.

There is also a faster mode for a single person talking to camera. It tracks the face with Google's MediaPipe and overlays a face mesh instead of running depth and masking. The project calls it experimental and built for "a single centered speaker."

What it costs

The project does not publish prices. Its notes say "hosted availability, price and access are account-dependent," and the main README's claim that Enhancor offers "the lowest price in the market" comes with no figures.

For a sense of scale only: Segmind, a different provider, published its own Seedance 2.5 rates in August at about $0.11 a second at 480p and about $0.24 a second at 720p, so a 10-second 720p clip billed $2.38 in its test. Segmind's article gives no 1080p rate, and Enhancor's prices may differ. Check your own dashboard before the first run, and add Replicate's charges for the preparation steps on top.

Where it breaks

Your footage leaves your machine, publicly. To hand the clip to the video service, the app uploads it to a public file host: tmpfiles, which deletes files after about six hours, or Catbox as a fallback. The README says both hosts "expose media through public URLs," and Catbox keeps files indefinitely. Do not run unreleased client footage, or anything under NDA, through it as built.

It is a single-user tool. The author says it needs proper sign-in and a job queue before anyone hosts it for others.

Paid calls are not retried for you. If a request times out, check whether it went through before resubmitting, or you may pay twice.

It is young and lightly tested. Created October 6. Other operating systems and most failure cases are, in the notes' words, "not exhaustively validated," and "no claim of bug-free operation is made."

It does not settle consent. The MIT license covers the code, and the project says including its example videos "does not grant rights to third-party footage, likenesses, music, logos or branding." Replacing a real performer, or using a real person as your character reference, still needs their permission. The tool will not ask for you.

It cannot fix a bad shot. Overlapping people, fast whip pans and partial bodies at the frame edge all make masks fail, and a bad mask makes a bad swap.

The habit outlasts the tool

Projects like this one tend to change fast, and the provider underneath may change too. What will last is the order of operations: hide what you do not want copied, check the exact input the model will see, judge a draft, and only then pay for full quality, with your original sound put back at the end.

Try it on one shot you already like, with a performer who has said yes. Spend your attention on the mask review. That is where this workflow, and most AI video jobs, are won or lost.


Medium metadata

Title: AI Video Actor Replacement That Keeps Your Shot and Your Soundtrack, and Makes You Approve Twice Before You Pay

Subtitle: TEIN's free Genjutsu workflow swaps the person in a video while keeping the camera move, the room and the original audio. The smartest part is the order it makes you work in.

Tags: Video Editing, Generative AI, Filmmaking, Visual Effects, AI Tools

Estimated read time: 8 minutes