Image & Video
MiniMax H3 at 4K on one RTX 4090: 4.9x faster local video and SyncTile
We tested every reported speed-up for MiniMax H3 on one RTX 4090, then split the 4K frame into tiles like a 3D renderer. Two changes cut a clip from 562 s to 115 s, and SyncTile, tiles that move through the sampling steps together, reached 4K with less shimmer than MMH3 Split Upscale or SeedVR2.
SummaryUnder a minute
The short version
We measured every reported speed-up for MiniMax H3 on one RTX 4090. Step caching cost picture, while an 8-step LoRA and int8 attention gave 4.9x with little change. For 4K, SyncTile blends nine tiles after every sampling step: no seams, no eye drift, a 5 s clip in about 11 minutes, and less shimmer than MMH3 Split Upscale or SeedVR2.
Key takeaways
- On MiniMax H3, an 8-step distillation LoRA plus int8 attention took a 5.9 s clip from 562 s to about 115 s (4.9x) while staying close to the original picture.
- Step caching (EasyCache, LazyCache, TeaCache) saved up to 2.9x but drifted the composition; SageAttention 2.2 ran 2x slower on this model.
- Four separate runs stitched together made four different scenes; tiles refined from one shared base clip lined up on stills but disagreed in motion.
- SyncTile, which blends the tiles after every sampling step, removed seams and eye drift and reached 3840 x 2160: about 11 minutes for a 5 s MiniMax H3 clip.
- Against MMH3 Split Upscale and SeedVR2 7B, SyncTile was the fastest and had the least shimmer in motion; SeedVR2 looked sharpest on a still frame but its texture boiled in playback.

What we wanted to find out
We generate short video clips locally with MiniMax H3, an open-weights video model that also produces sound. On one RTX 4090, a 5.9-second MiniMax H3 clip at 1344 x 768 took 562 s with the default 25 sampling steps. That is fine for one shot and painful for a commercial that needs ten variations. The GPU also stops at roughly that resolution, so anything sharper had to come from an upscaler.
Two questions came out of that:
- How much faster can MiniMax H3 run on the same GPU without visibly changing the picture?
- Can several separate MiniMax H3 runs be stitched into one 4K canvas, the way 3D renderers split a frame into tiles?
What we tried
Every speed test used the same two prompts and the same seed: a crowded medieval bazaar where an earthquake sends people running (fast motion), and a close-up of an old shepherd speaking Turkish by a fire (a face, lip movement, sound). Each clip was 141 frames at 1344 x 768. Later tests added a fast drone shot through a red-rock canyon with a kayaker, and a fisheye close-up of a café owner reacting to his laptop, both 5 seconds.
The candidates were the speed-ups people report for MiniMax H3 and similar video models:
- Step caching (EasyCache, LazyCache, TeaCache): skip a model call when the result barely changes from the previous step.
- Faster attention kernels: SageAttention, and an int8 attention backend already built into the software we run MiniMax H3 in.
- Fewer steps: an 8-step distillation LoRA instead of the default 25.
- Starting small: run the first steps at half resolution, then enlarge with the MiniMax H3 latent upscaler.
How we measured it
Time is wall-clock for the whole run, model already loaded. For speed tests we compared each clip with the clip from the same seed without the speed-up, using SSIM (1.00 means identical). For 4K we added a shimmer score: on a full-resolution center crop, we isolate the finest detail and measure how much it changes from frame to frame. A plain resize scores about 1.5. Higher means texture that boils in motion. Every result was also watched in motion, full screen, because single frames turned out to be misleading.
What happened
Speed on a single 1344 x 768 MiniMax H3 clip
| Setup | Time | Speed-up | SSIM (motion / face) |
|---|---|---|---|
| 25 steps, no speed-up | 562 s | 1.00x | reference |
| EasyCache, threshold 0.20 | 325 to 369 s | 1.52x to 1.73x | 0.85 / 0.96 |
| LazyCache, threshold 0.20 | 346 to 391 s | 1.44x to 1.63x | 0.85 / 0.96 |
| TeaCache, threshold 0.25 | 195 s | 2.87x | 0.58 / 0.75 |
| int8 attention | 304 to 307 s | 1.85x | 0.70 / 0.91 |
| 8 steps (distillation LoRA) | 199 s | 2.8x | own reference |
| 8 steps + int8 attention | 114 to 117 s | 4.9x | 0.69 / 0.92 against 8 steps |

Caching was the disappointment. EasyCache helped on the calm face shot and much less on the busy one. TeaCache was the fastest cache, but it rebuilt the composition: same prompt, same seed, different framing, softer detail. On the 8-step setup it was worse, adding lights that were not in the prompt. We also found that this TeaCache node never reset its step counter when the same settings ran twice in a row, so the second clip silently ran with no caching at all. We patched it locally before measuring again.

SageAttention 2.2 ran twice as slow on MiniMax H3 (1,127 s) and produced a bit-identical clip. The int8 attention backend skips nothing, it just computes each step faster, and the scene stays the same. Stacked on the 8-step LoRA it brings a MiniMax H3 clip from 562 s to about 115 s.
Starting small was the fastest of all (60 s, 9.4x), but the small first pass decides the composition, so the clip no longer matches the full-size run. The 32B text encoder was not a cost here either: a new prompt took 114 s and a repeated one 117 s.
Four separate runs, one canvas
The literal tiling idea does not work. We generated each quarter of a 2688 x 1536 frame on its own, with a prompt describing that quarter. Each MiniMax H3 run built a complete scene of its own: four different men in four different cafés.

Tiles only line up if they share a map. So we made one small clip of the whole scene first, enlarged it, cut it into four overlapping 1472 x 896 tiles, and ran each tile through MiniMax H3 again from a partly noised state (the last 3 of 8 steps). A still frame showed no seam. In motion, the seams ran between the eyes and each tile made its own small decision about the pupils: a slight squint, and a head that moved differently on each side of the seam.

We tested four ways to hold the tiles together:
| Method (2688 x 1536, 5 s) | Time | Result |
|---|---|---|
| Four tiles, each finished on its own | 430 to 438 s | Moving objects shift, eyes can disagree |
| Plus a fifth tile centered on the seams | 541 to 550 s | Fixes the center, slowest |
| Plus "frequency lock" (coarse shape from the base clip) | 430 to 438 s | Double edges on moving objects |
| Synchronized tiles | 243 to 272 s | No seam, one kayaker, both eyes agree |


We call the winning method SyncTile. It applies the MultiDiffusion idea to MiniMax H3's joint audio-video latent. Tiles are still processed one at a time, so the GPU never holds more than one, but they move through the steps together: after every step all tiles are blended in latent space, and the next step starts from that shared result. Every tile starts from one shared noise image, and the finished canvas is decoded in one piece, so there is no pixel seam to hide.

SyncTile at 4K
For 3840 x 2160 the base clip is not enlarged as pixels, which had produced blocky texture. It is encoded and enlarged straight to 4K by the MiniMax H3 latent upscaler, then refined in nine SyncTile tiles (3 x 3) and decoded once.
| 5 s at 3840 x 2160 | Last 3 of 8 steps | Last 4 of 8 steps |
|---|---|---|
| Café close-up | 542 s | 650 s |
| Drone canyon | 549 s | 776 s |
With the 8-step base clip (about 100 s), a 5-second MiniMax H3 clip reaches 4K in roughly 11 minutes on one RTX 4090. Four steps added time and small shifts, so we use three.

SyncTile against other 4K routes
Before calling it useful we ran the same two base clips through two existing routes: MMH3 Split Upscale, a community node for MiniMax H3 that finishes tiles one after another, and SeedVR2 7B, a dedicated video upscaler.
| 4K, 5 s (café / drone) | Time | Fidelity to base (SSIM) | Shimmer (resize ≈ 1.5) |
|---|---|---|---|
| SyncTile | 542 / 549 s | 0.87 / 0.81 | 1.62 / 1.60 |
| MMH3 Split Upscale | 605 / 601 s | 0.84 / 0.79 | 1.68 / 1.66 |
| SeedVR2 7B | 2,097 / 2,085 s | 0.96 / 0.95 | 3.00 / 2.45 |
Split Upscale produced more fine texture and moved the face slightly against the base clip. SeedVR2 ran out of memory at 4K on default settings and needed block swapping, which took it to about 35 minutes. On a single frame SeedVR2 looked the sharpest and a standard sharpness score agreed. Played back, its fine texture boiled from frame to frame and looked visibly AI-made. The shimmer score caught what the sharpness score had rewarded as detail.

SyncTile stays closest to a plain resize in motion while adding real detail. The last step is a light unsharp mask (5 x 5, amount 1.0), which we matched to a manual sharpening pass in an editor. It lifts fine detail by about a quarter and keeps shimmer at 2.1, still well below SeedVR2.
What we learned
- Speed-ups that skip work are not free. On MiniMax H3, step caching saved time by letting the scene drift. The safe wins did the same work faster: a distilled LoRA and int8 attention, 4.9x together.
- Test reported gains on your own setup. SageAttention and text-encoder caching both have good public numbers. Here one was slower and the other saved nothing.
- Tiles need a shared map and a shared clock. One base clip gives the tiles one composition, and blending them after every step keeps two eyes in two tiles looking the same way.
- Judge video in motion. The four-tile result and SeedVR2 both looked best as stills and failed in playback. A sharpness score can reward noise; a frame-to-frame shimmer score did not.
Where this is used now
Two MiniMax H3 settings are now our defaults. Drafts and most clips use 8 steps with int8 attention. When a project needs a large, sharp picture, the clip goes through SyncTile to 4K with the last 3 of 8 steps and a light unsharp mask. The code, the exact settings and the 4K demo clips are open source on GitHub: alaintural/synctile-h3 (MIT). So far SyncTile has been tested on 5-second clips. Longer clips make every tile heavier, so they are the next thing to measure.
All clips in this note were generated with AI and show no real people.
Sources
Ask AI about this AI Lab note
Opens the assistant in a new tab with this page as the source.
Keep reading
AI Lab · Image & VideoFace swapping a 4K close-up: why the identity only half transferred, and the two-model fixOctober 5, 2026 AI Lab · Image & VideoChecking AI images against measured numbers for a dark, warm lookOctober 9, 2026
AI Lab · Image & VideoGrading GPT and Gemini color casts, and why color belongs in the promptOctober 8, 2026
Want this handled for your business?
Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.