Image & Video

Face swapping a 4K close-up: why the identity only half transferred, and the two-model fix

A video face-swap model blended the new face with the original one no matter what we asked. Measuring identity with a face-recognition model showed why, and putting a classic face swapper in front of it took similarity from 0.52 to 0.83.

Pixedi AI Lab
  • 7 min read

SummaryUnder a minute

The short version

A video face-swap model blended the new face with the original one on a 4K close-up. An ArcFace score put the result halfway between the two people, and more sampling steps made it worse. Running a classic face swapper first, so the video model never sees the original face, raised identity to 0.83 on the full clip in 365 s.

Key takeaways

  1. Cropping only the face instead of sending the full 1080p frame cut render time about 3x (42.1 s vs 125.2 s for 56 frames).
  2. On a 4K close-up the crop window hit a silent 1024 px cap and the paste-back left a hard line across the hair.
  3. The video model's output scored 0.52 against the reference subject and still 0.35 against the original person: a blend of the two.
  4. A classic face swapper (HyperSwap) as a pre-pass fixed it: 0.82 to 0.84 with the video model on top, 192 frames of 4K in 365 s.
Two frames side by side: the original man in a dark coat in a library, and the same shot with a different face, curly hair and a beard.

What we wanted to find out

The goal was to replace one person's face in an existing video with a different person, keeping the original performance: head movement, expression, gaze, lip movement, and the scene's light. The test footage is an 8-second shot generated with Veo 3.1 in Google Flow: a man in a dark coat speaking straight into the camera in a library at night. It was generated at 1080p, cleaned up with SeedVR2 and scaled to 4K, so the face fills a large part of a 3840 x 2160 frame. That is the hard case for face swapping: a big face, low-key light, and a direct gaze that shows every error.

The swap itself is done by MiniMax H3, a video model with a face-swap LoRA. It takes a reference picture of the new person plus the original clip as a motion source, and generates the clip again with the new face. Everything ran on one RTX 4090.

The reference set

The new identity came from a consenting adult subject photographed against a plain white background from fifteen angles: straight on, three-quarter views, both profiles, looking up, looking down, and the back of the head. Five of these views (front, both three-quarters, one profile, one turned down) were used as identity sources in the tests below.

Fifteen photos of the same man from different angles against a white background.
The reference set: one subject, fifteen angles, flat studio light.

What we tried

1. Crop the face, not the frame

Sending a full frame to the video model is slow. Instead, a face detector (SCRFD) finds every face, a recognition model (ArcFace) groups them by person, and only the chosen person's head is cut out of each frame in a window that follows it. After the swap, the window is pasted back at the exact same pixel position. On this clip the tracking step found one person in all 192 frames in 6 s. These timings were measured during development on a 56-frame 1080p clip:

Input (56 frames)Time
Full 1080p frame125.2 s
Standard 0.5 MP resize40.1 s
768 px face window, no resize42.1 s

The face window costs the same as the standard resize but keeps the face at full resolution. Window sizes are tiered at 512, 768 and 1024 px. Changing tier in the middle of a shot caused a visible jump. Overlapping the parts caused ghosting. Handing the last frame of one part to the next as a second reference helped only a little. Using one window size per shot removed the seam completely, so seams now only fall on scene cuts. A simple ellipse mask let the original hair show around the new face, so it was replaced with a head-parsing mask (BiSeNet) that covers face, hair, ears and neck of both the old and the new head.

2. First run on 4K footage: the silent cap

The first full run on the 4K clip, with the original window rule.

The face box on this close-up is about 730 x 1,050 px. The window rule was face size times 1.5, so it needed roughly 1,575 px, but the tier list stopped at 1024 and the code silently returned 1024. The model received a crop from the forehead to just below the chin, with the top of the head cut off.

Left: a tight crop of the original face. Right: the model output with a beard and different hair.
Left: what the model received. Right: what it returned. It added a beard and new hair, but only inside the window.

The model rebuilt the hair inside the window, but the original hair above the window stayed. On paste-back, the window edge became a hard horizontal line across the head: new hair below, original hair above. The run took 246 s for 192 frames.

A close-up where a straight line crosses the hair: different hair below it, the original hair above.
The window edge: the swapped part ends in a straight line, the original hair continues above it.

3. Fix the window, then guard it

The window now takes the whole head (face size times 2.0). When that is larger than 1024 px, the crop is taken at full size (2,160 px on this close-up), shrunk to 1024 for the model, and enlarged again on paste-back. Two guards were added. An alignment check runs face detection on the model's input and output. If the face moved more than 25% of its size or changed size outside 0.7 to 1.4, the part is retried with another seed and, if it fails again, left as the original. On this clip it passed on the first try (shift 0.4%, size ratio 1.03). The mask now fades to zero before the window edge, so a line or a square can no longer appear.

The line was gone and the head pose followed the original. But the face looked like a mix of the two people.

4. More sampling steps

We tried the 8-step distillation LoRA with 8 steps instead of 4. It took 416 s instead of 254 s and was worse: the original face came back through.

Two rows of three close-ups: original, 4-step result, 8-step result.
Left to right: original, 4 steps, 8 steps. More steps pulled the result back toward the original person.

How we measured it

Looking at frames was not enough, because a new beard makes a face look different without changing who it is. We used ArcFace (w600k R50), the same recognition model that groups faces in the first step. Each output face is turned into an embedding and compared with the average embedding of the five reference photos using cosine similarity. Every number below is the mean of four frames spread across the clip.

For scale: the five reference photos score 0.84 to 0.92 against their own average. The original person scores -0.01. The limit is that this measures identity only. It says nothing about gaze, grain or lighting, so every result was also checked by eye in side-by-side video.

ResultSimilarity to subjectSimilarity to original
Step 2, silent cap0.460.50
Step 3, fixed window, 4 steps0.520.35
Step 4, 8 steps0.370.55

The swap looked done, with the subject's beard and hair. To a recognition model it was halfway between the two people.

What happened

A scene-matched reference from an image editor (failed)

The idea: use Qwen-Image-Edit 2509 to place the reference subject into a still from the shot, then give that still to the video model as a second reference that already carries the scene's light and angle. Each image took 134 to 149 s.

SetupSimilarity to subjectResult
Scene as image 1, subject photos as images 2-30.00Added a mustache and a red cast, still the original person (0.88)
Subject photo as image 1, scene as image 20.00Copied the scene almost exactly (0.93 to the original)
Face in the scene erased with gray0.29A new face came through, but as a bald oval pasted on the head
Three pairs of images: the input scene and the edited result for each setup.
Left of each pair: the input scene. Right: the edit. Only the erased face let a new identity in.

A face swapper as a pre-pass (worked)

Classic face swappers work differently: they inject the identity embedding directly, frame by frame. They are fast and strong on identity, weak on resolution and hair. So the order was reversed. A face swapper runs first on the cropped clip, and the video model gets that result as its motion source. The original face is no longer visible to it.

Face swapper (pixel boost 512)SimilarityTime (192 frames)
HyperSwap 1a 2560.7928 s
HyperSwap 1b 2560.8025 s
HyperSwap 1c 2560.7925 s
HyperSwap 1a, eyes left from original0.7718 s

All four kept the original hair and gave only a thin beard. Because the face swapper reuses the original pixels around the face, light and grain stayed close to the scene. The person here looks straight into the lens, so gaze was kept by all four.

Five close-ups: the original and four face swapper variants.
Left to right: original, HyperSwap 1a, 1b, 1c, 1a with original eyes.
Left to right: original, video model alone (0.54), face swapper alone (0.79), face swapper then video model (0.82).

The prompt decides what the video model adds

With the pre-pass in place, the video model received two references (front and three-quarter photos) and one of four prompts. On this clip all four brought in the subject's curly hair and full beard, even the bare trigger word Faceswap. The descriptive prompts scored slightly higher on identity.

PromptSimilarityTime
"Faceswap" only0.82242 s
Short official template0.84243 s
Detailed, scene-matching0.84242 s
General, identity-from-reference0.83248 s
Left to right: original, face swapper only, then the video model with "Faceswap", the official template and the detailed prompt.

The last prompt is written to work for any subject and any scene: identity comes from the reference pictures, light and image quality from the footage, performance from the video. It is 0.01 below the best score here and is the one we use by default.

Faceswap. <Subject 1> is the person shown in the reference pictures <Picture 1> and <Picture 2>. Replace the person in <Video 1> with <Subject 1>. Copy the identity of <Subject 1> exactly from the reference pictures: face shape, eyes and eye color, eyebrows, nose, mouth and lips, teeth, ears, skin tone and skin details, hairstyle, hair color and hairline, beard, mustache and any facial hair, glasses or accessories on the face. Do not copy the lighting, pose or background of the reference pictures. Blend <Subject 1> naturally into the scene of <Video 1>: use the same light direction, light color, shadows and highlights that fall on the original person, the same color tone and color grading, the same contrast, focus, blur, noise, grain and compression of the footage, so the face looks filmed with the same camera in the same shot. Keep the performance of <Video 1>: head pose and movement, facial expression, eye movements and blinks, lip and mouth movement, timing. Do not change the body, clothes or background.
Left to right: original, detailed prompt, general prompt.

The full clip

The whole 8-second, 192-frame 4K clip went through the final setup: head-window crop (3 s), HyperSwap 1c pre-pass (28 s), the video model at 4 steps with two references and the general prompt (251 s), then head-mask paste-back (86 s). Total: 365 s. The result scores 0.83 against the reference subject and 0.11 against the original person.

Left: original (AI-generated). Right: final result (AI-altered). Audio is from the original.

What we learned

  • Measure identity, don't eyeball it: a new beard and new hair fooled the eye; a recognition score showed the face was still half the original person. Without the number we would have kept tuning steps and prompts.
  • A video face-swap model follows the face it can see: hiding the original face, with a gray patch or a fast face swapper, is what lets the new identity through, while more sampling steps made it worse.
  • Combine the two model families: the classic swapper brings identity and keeps the scene's pixels, and the video model brings hair, beard, resolution and temporal consistency that the swapper can't.
  • Silent caps are bugs: a size function that quietly returned 1024 for a 1,575 px need caused the visible failure. Every limit now either scales properly or stops the run.

Still open: the video model also reads the clip's audio. On shots where the original person is silent it can make the new face talk; this clip, where he speaks throughout, does not test that. The next step is feeding the model silence for those parts.

The test clip was generated with AI and shows no real person. The reference subject consented to the use of their photos. All face-swapped frames and clips in this note are AI-altered and shown for technical comparison only.

Ask AI about this AI Lab note

Opens the assistant in a new tab with this page as the source.

Keep reading

Want this handled for your business?

Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.