Image & Video
Face swapping a 4K close-up: why the identity only half transferred, and the two-model fix
A video face-swap model blended the new face with the original one no matter what we asked. Measuring identity with a face-recognition model showed why, and putting a classic face swapper in front of it took similarity from 0.52 to 0.83.
SummaryUnder a minute
The short version
A video face-swap model blended the new face with the original one on a 4K close-up. An ArcFace score put the result halfway between the two people, and more sampling steps made it worse. Running a classic face swapper first, so the video model never sees the original face, raised identity to 0.83 on the full clip in 365 s.
Key takeaways
- Cropping only the face instead of sending the full 1080p frame cut render time about 3x (42.1 s vs 125.2 s for 56 frames).
- On a 4K close-up the crop window hit a silent 1024 px cap and the paste-back left a hard line across the hair.
- The video model's output scored 0.52 against the reference subject and still 0.35 against the original person: a blend of the two.
- A classic face swapper (HyperSwap) as a pre-pass fixed it: 0.82 to 0.84 with the video model on top, 192 frames of 4K in 365 s.

What we wanted to find out
The goal was to replace one person's face in an existing video with a different person, keeping the original performance: head movement, expression, gaze, lip movement, and the scene's light. The test footage is an 8-second shot generated with Veo 3.1 in Google Flow: a man in a dark coat speaking straight into the camera in a library at night. It was generated at 1080p, cleaned up with SeedVR2 and scaled to 4K, so the face fills a large part of a 3840 x 2160 frame. That is the hard case for face swapping: a big face, low-key light, and a direct gaze that shows every error.
The swap itself is done by MiniMax H3, a video model with a face-swap LoRA. It takes a reference picture of the new person plus the original clip as a motion source, and generates the clip again with the new face. Everything ran on one RTX 4090.
The reference set
The new identity came from a consenting adult subject photographed against a plain white background from fifteen angles: straight on, three-quarter views, both profiles, looking up, looking down, and the back of the head. Five of these views (front, both three-quarters, one profile, one turned down) were used as identity sources in the tests below.

What we tried
1. Crop the face, not the frame
Sending a full frame to the video model is slow. Instead, a face detector (SCRFD) finds every face, a recognition model (ArcFace) groups them by person, and only the chosen person's head is cut out of each frame in a window that follows it. After the swap, the window is pasted back at the exact same pixel position. On this clip the tracking step found one person in all 192 frames in 6 s. These timings were measured during development on a 56-frame 1080p clip:
| Input (56 frames) | Time |
|---|---|
| Full 1080p frame | 125.2 s |
| Standard 0.5 MP resize | 40.1 s |
| 768 px face window, no resize | 42.1 s |
The face window costs the same as the standard resize but keeps the face at full resolution. Window sizes are tiered at 512, 768 and 1024 px. Changing tier in the middle of a shot caused a visible jump. Overlapping the parts caused ghosting. Handing the last frame of one part to the next as a second reference helped only a little. Using one window size per shot removed the seam completely, so seams now only fall on scene cuts. A simple ellipse mask let the original hair show around the new face, so it was replaced with a head-parsing mask (BiSeNet) that covers face, hair, ears and neck of both the old and the new head.
2. First run on 4K footage: the silent cap
The face box on this close-up is about 730 x 1,050 px. The window rule was face size times 1.5, so it needed roughly 1,575 px, but the tier list stopped at 1024 and the code silently returned 1024. The model received a crop from the forehead to just below the chin, with the top of the head cut off.

The model rebuilt the hair inside the window, but the original hair above the window stayed. On paste-back, the window edge became a hard horizontal line across the head: new hair below, original hair above. The run took 246 s for 192 frames.

3. Fix the window, then guard it
The window now takes the whole head (face size times 2.0). When that is larger than 1024 px, the crop is taken at full size (2,160 px on this close-up), shrunk to 1024 for the model, and enlarged again on paste-back. Two guards were added. An alignment check runs face detection on the model's input and output. If the face moved more than 25% of its size or changed size outside 0.7 to 1.4, the part is retried with another seed and, if it fails again, left as the original. On this clip it passed on the first try (shift 0.4%, size ratio 1.03). The mask now fades to zero before the window edge, so a line or a square can no longer appear.
The line was gone and the head pose followed the original. But the face looked like a mix of the two people.
4. More sampling steps
We tried the 8-step distillation LoRA with 8 steps instead of 4. It took 416 s instead of 254 s and was worse: the original face came back through.

How we measured it
Looking at frames was not enough, because a new beard makes a face look different without changing who it is. We used ArcFace (w600k R50), the same recognition model that groups faces in the first step. Each output face is turned into an embedding and compared with the average embedding of the five reference photos using cosine similarity. Every number below is the mean of four frames spread across the clip.
For scale: the five reference photos score 0.84 to 0.92 against their own average. The original person scores -0.01. The limit is that this measures identity only. It says nothing about gaze, grain or lighting, so every result was also checked by eye in side-by-side video.
| Result | Similarity to subject | Similarity to original |
|---|---|---|
| Step 2, silent cap | 0.46 | 0.50 |
| Step 3, fixed window, 4 steps | 0.52 | 0.35 |
| Step 4, 8 steps | 0.37 | 0.55 |
The swap looked done, with the subject's beard and hair. To a recognition model it was halfway between the two people.
What happened
A scene-matched reference from an image editor (failed)
The idea: use Qwen-Image-Edit 2509 to place the reference subject into a still from the shot, then give that still to the video model as a second reference that already carries the scene's light and angle. Each image took 134 to 149 s.
| Setup | Similarity to subject | Result |
|---|---|---|
| Scene as image 1, subject photos as images 2-3 | 0.00 | Added a mustache and a red cast, still the original person (0.88) |
| Subject photo as image 1, scene as image 2 | 0.00 | Copied the scene almost exactly (0.93 to the original) |
| Face in the scene erased with gray | 0.29 | A new face came through, but as a bald oval pasted on the head |

A face swapper as a pre-pass (worked)
Classic face swappers work differently: they inject the identity embedding directly, frame by frame. They are fast and strong on identity, weak on resolution and hair. So the order was reversed. A face swapper runs first on the cropped clip, and the video model gets that result as its motion source. The original face is no longer visible to it.
| Face swapper (pixel boost 512) | Similarity | Time (192 frames) |
|---|---|---|
| HyperSwap 1a 256 | 0.79 | 28 s |
| HyperSwap 1b 256 | 0.80 | 25 s |
| HyperSwap 1c 256 | 0.79 | 25 s |
| HyperSwap 1a, eyes left from original | 0.77 | 18 s |
All four kept the original hair and gave only a thin beard. Because the face swapper reuses the original pixels around the face, light and grain stayed close to the scene. The person here looks straight into the lens, so gaze was kept by all four.

The prompt decides what the video model adds
With the pre-pass in place, the video model received two references (front and three-quarter photos) and one of four prompts. On this clip all four brought in the subject's curly hair and full beard, even the bare trigger word Faceswap. The descriptive prompts scored slightly higher on identity.
| Prompt | Similarity | Time |
|---|---|---|
| "Faceswap" only | 0.82 | 242 s |
| Short official template | 0.84 | 243 s |
| Detailed, scene-matching | 0.84 | 242 s |
| General, identity-from-reference | 0.83 | 248 s |
The last prompt is written to work for any subject and any scene: identity comes from the reference pictures, light and image quality from the footage, performance from the video. It is 0.01 below the best score here and is the one we use by default.
Faceswap. <Subject 1> is the person shown in the reference pictures <Picture 1> and <Picture 2>. Replace the person in <Video 1> with <Subject 1>. Copy the identity of <Subject 1> exactly from the reference pictures: face shape, eyes and eye color, eyebrows, nose, mouth and lips, teeth, ears, skin tone and skin details, hairstyle, hair color and hairline, beard, mustache and any facial hair, glasses or accessories on the face. Do not copy the lighting, pose or background of the reference pictures. Blend <Subject 1> naturally into the scene of <Video 1>: use the same light direction, light color, shadows and highlights that fall on the original person, the same color tone and color grading, the same contrast, focus, blur, noise, grain and compression of the footage, so the face looks filmed with the same camera in the same shot. Keep the performance of <Video 1>: head pose and movement, facial expression, eye movements and blinks, lip and mouth movement, timing. Do not change the body, clothes or background.
The full clip
The whole 8-second, 192-frame 4K clip went through the final setup: head-window crop (3 s), HyperSwap 1c pre-pass (28 s), the video model at 4 steps with two references and the general prompt (251 s), then head-mask paste-back (86 s). Total: 365 s. The result scores 0.83 against the reference subject and 0.11 against the original person.
What we learned
- Measure identity, don't eyeball it: a new beard and new hair fooled the eye; a recognition score showed the face was still half the original person. Without the number we would have kept tuning steps and prompts.
- A video face-swap model follows the face it can see: hiding the original face, with a gray patch or a fast face swapper, is what lets the new identity through, while more sampling steps made it worse.
- Combine the two model families: the classic swapper brings identity and keeps the scene's pixels, and the video model brings hair, beard, resolution and temporal consistency that the swapper can't.
- Silent caps are bugs: a size function that quietly returned 1024 for a 1,575 px need caused the visible failure. Every limit now either scales properly or stops the run.
Still open: the video model also reads the clip's audio. On shots where the original person is silent it can make the new face talk; this clip, where he speaks throughout, does not test that. The next step is feeding the model silence for those parts.
The test clip was generated with AI and shows no real person. The reference subject consented to the use of their photos. All face-swapped frames and clips in this note are AI-altered and shown for technical comparison only.
Ask AI about this AI Lab note
Opens the assistant in a new tab with this page as the source.
Keep reading
AI Lab · Image & VideoChecking AI images against measured numbers for a dark, warm lookOctober 9, 2026
AI Lab · Image & VideoGrading GPT and Gemini color casts, and why color belongs in the promptOctober 8, 2026
AI Lab · Image & VideoSeedVR2 vs ESRGAN on one RTX 4090: what video upscaling costsOctober 6, 2026
Want this handled for your business?
Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.