Image & Video

MiniMax H3 at 4K on one RTX 4090: 4.9x faster local video and SyncTile

We tested every reported speed-up for MiniMax H3 on one RTX 4090, then split the 4K frame into tiles like a 3D renderer. Two changes cut a clip from 562 s to 115 s, and SyncTile, tiles that move through the sampling steps together, reached 4K with less shimmer than MMH3 Split Upscale or SeedVR2.

Pixedi AI Lab
  • 7 min read

SummaryUnder a minute

The short version

We measured every reported speed-up for MiniMax H3 on one RTX 4090. Step caching cost picture, while an 8-step LoRA and int8 attention gave 4.9x with little change. For 4K, SyncTile blends nine tiles after every sampling step: no seams, no eye drift, a 5 s clip in about 11 minutes, and less shimmer than MMH3 Split Upscale or SeedVR2.

Key takeaways

  1. On MiniMax H3, an 8-step distillation LoRA plus int8 attention took a 5.9 s clip from 562 s to about 115 s (4.9x) while staying close to the original picture.
  2. Step caching (EasyCache, LazyCache, TeaCache) saved up to 2.9x but drifted the composition; SageAttention 2.2 ran 2x slower on this model.
  3. Four separate runs stitched together made four different scenes; tiles refined from one shared base clip lined up on stills but disagreed in motion.
  4. SyncTile, which blends the tiles after every sampling step, removed seams and eye drift and reached 3840 x 2160: about 11 minutes for a 5 s MiniMax H3 clip.
  5. Against MMH3 Split Upscale and SeedVR2 7B, SyncTile was the fastest and had the least shimmer in motion; SeedVR2 looked sharpest on a still frame but its texture boiled in playback.
A café owner with curly hair and a mustard apron leaning toward the camera in a colorful café, generated with MiniMax H3 and taken to 4K with SyncTile.

What we wanted to find out

We generate short video clips locally with MiniMax H3, an open-weights video model that also produces sound. On one RTX 4090, a 5.9-second MiniMax H3 clip at 1344 x 768 took 562 s with the default 25 sampling steps. That is fine for one shot and painful for a commercial that needs ten variations. The GPU also stops at roughly that resolution, so anything sharper had to come from an upscaler.

Two questions came out of that:

  • How much faster can MiniMax H3 run on the same GPU without visibly changing the picture?
  • Can several separate MiniMax H3 runs be stitched into one 4K canvas, the way 3D renderers split a frame into tiles?

What we tried

Every speed test used the same two prompts and the same seed: a crowded medieval bazaar where an earthquake sends people running (fast motion), and a close-up of an old shepherd speaking Turkish by a fire (a face, lip movement, sound). Each clip was 141 frames at 1344 x 768. Later tests added a fast drone shot through a red-rock canyon with a kayaker, and a fisheye close-up of a café owner reacting to his laptop, both 5 seconds.

The candidates were the speed-ups people report for MiniMax H3 and similar video models:

  • Step caching (EasyCache, LazyCache, TeaCache): skip a model call when the result barely changes from the previous step.
  • Faster attention kernels: SageAttention, and an int8 attention backend already built into the software we run MiniMax H3 in.
  • Fewer steps: an 8-step distillation LoRA instead of the default 25.
  • Starting small: run the first steps at half resolution, then enlarge with the MiniMax H3 latent upscaler.

How we measured it

Time is wall-clock for the whole run, model already loaded. For speed tests we compared each clip with the clip from the same seed without the speed-up, using SSIM (1.00 means identical). For 4K we added a shimmer score: on a full-resolution center crop, we isolate the finest detail and measure how much it changes from frame to frame. A plain resize scores about 1.5. Higher means texture that boils in motion. Every result was also watched in motion, full screen, because single frames turned out to be misleading.

What happened

Speed on a single 1344 x 768 MiniMax H3 clip

SetupTimeSpeed-upSSIM (motion / face)
25 steps, no speed-up562 s1.00xreference
EasyCache, threshold 0.20325 to 369 s1.52x to 1.73x0.85 / 0.96
LazyCache, threshold 0.20346 to 391 s1.44x to 1.63x0.85 / 0.96
TeaCache, threshold 0.25195 s2.87x0.58 / 0.75
int8 attention304 to 307 s1.85x0.70 / 0.91
8 steps (distillation LoRA)199 s2.8xown reference
8 steps + int8 attention114 to 117 s4.9x0.69 / 0.92 against 8 steps
Three versions of the same frame of an old man by a fire: 8 steps, 8 steps with int8 attention, and 8 steps with TeaCache, which adds stray lights.
Same frame, three setups: 8 steps (199 s), 8 steps + int8 attention (114 s), 8 steps + TeaCache (112 s).

Caching was the disappointment. EasyCache helped on the calm face shot and much less on the busy one. TeaCache was the fastest cache, but it rebuilt the composition: same prompt, same seed, different framing, softer detail. On the 8-step setup it was worse, adding lights that were not in the prompt. We also found that this TeaCache node never reset its step counter when the same settings ran twice in a row, so the second clip silently ran with no caching at all. We patched it locally before measuring again.

Two frames of a crowded medieval bazaar side by side, with different camera framing.
Left: 25 steps (562 s). Right: TeaCache 0.15 (219 s). Same prompt and seed, different composition.

SageAttention 2.2 ran twice as slow on MiniMax H3 (1,127 s) and produced a bit-identical clip. The int8 attention backend skips nothing, it just computes each step faster, and the scene stays the same. Stacked on the 8-step LoRA it brings a MiniMax H3 clip from 562 s to about 115 s.

Left: 25 steps, 562 s. Right: 8 steps + int8 attention, 114 s. Same prompt and seed, with sound.

Starting small was the fastest of all (60 s, 9.4x), but the small first pass decides the composition, so the clip no longer matches the full-size run. The 32B text encoder was not a cost here either: a new prompt took 114 s and a repeated one 117 s.

Four separate runs, one canvas

The literal tiling idea does not work. We generated each quarter of a 2688 x 1536 frame on its own, with a prompt describing that quarter. Each MiniMax H3 run built a complete scene of its own: four different men in four different cafés.

A grid of four separately generated close-ups of four different laughing men in four different cafés.
Four quarters generated on their own and placed side by side: four scenes, not one.
Four quarters generated separately: each run builds its own scene.

Tiles only line up if they share a map. So we made one small clip of the whole scene first, enlarged it, cut it into four overlapping 1472 x 896 tiles, and ran each tile through MiniMax H3 again from a partly noised state (the last 3 of 8 steps). A still frame showed no seam. In motion, the seams ran between the eyes and each tile made its own small decision about the pupils: a slight squint, and a head that moved differently on each side of the seam.

Close-up of a man whose left eye looks to the side while his right eye looks forward.
Four tiles finished on their own: the seam runs between the eyes and they disagree.

We tested four ways to hold the tiles together:

Method (2688 x 1536, 5 s)TimeResult
Four tiles, each finished on its own430 to 438 sMoving objects shift, eyes can disagree
Plus a fifth tile centered on the seams541 to 550 sFixes the center, slowest
Plus "frequency lock" (coarse shape from the base clip)430 to 438 sDouble edges on moving objects
Synchronized tiles243 to 272 sNo seam, one kayaker, both eyes agree
Four crops of the same eyes: the base clip, four tiles, synchronized tiles, and four tiles with a center tile.
Same frame. Top row: base clip, four tiles. Bottom row: synchronized tiles, four tiles plus center tile.
Three crops of a kayaker on a canyon river; in the third the paddle appears twice.
Base, four tiles, frequency lock. The lock doubles the moving paddle.

We call the winning method SyncTile. It applies the MultiDiffusion idea to MiniMax H3's joint audio-video latent. Tiles are still processed one at a time, so the GPU never holds more than one, but they move through the steps together: after every step all tiles are blended in latent space, and the next step starts from that shared result. Every tile starts from one shared noise image, and the finished canvas is decoded in one piece, so there is no pixel seam to hide.

Three crops of a kayaker on a canyon river from the base clip, four tiles and synchronized tiles.
Base, four tiles, synchronized tiles, at the point where the seams cross.

SyncTile at 4K

For 3840 x 2160 the base clip is not enlarged as pixels, which had produced blocky texture. It is encoded and enlarged straight to 4K by the MiniMax H3 latent upscaler, then refined in nine SyncTile tiles (3 x 3) and decoded once.

5 s at 3840 x 2160Last 3 of 8 stepsLast 4 of 8 steps
Café close-up542 s650 s
Drone canyon549 s776 s

With the 8-step base clip (about 100 s), a 5-second MiniMax H3 clip reaches 4K in roughly 11 minutes on one RTX 4090. Four steps added time and small shifts, so we use three.

Two close-ups of the same eyes; the lower one shows sharper lashes and skin texture.
Top: 2x SyncTile result scaled to 4K. Bottom: SyncTile at native 4K.
MiniMax H3 drone shot taken to 3840 x 2160 with SyncTile, shown scaled to 1080p. 549 s for 5 s.

SyncTile against other 4K routes

Before calling it useful we ran the same two base clips through two existing routes: MMH3 Split Upscale, a community node for MiniMax H3 that finishes tiles one after another, and SeedVR2 7B, a dedicated video upscaler.

4K, 5 s (café / drone)TimeFidelity to base (SSIM)Shimmer (resize ≈ 1.5)
SyncTile542 / 549 s0.87 / 0.811.62 / 1.60
MMH3 Split Upscale605 / 601 s0.84 / 0.791.68 / 1.66
SeedVR2 7B2,097 / 2,085 s0.96 / 0.953.00 / 2.45

Split Upscale produced more fine texture and moved the face slightly against the base clip. SeedVR2 ran out of memory at 4K on default settings and needed block swapping, which took it to about 35 minutes. On a single frame SeedVR2 looked the sharpest and a standard sharpness score agreed. Played back, its fine texture boiled from frame to frame and looked visibly AI-made. The shimmer score caught what the sharpness score had rewarded as detail.

Two 4K crops of the same face: the lower one sits slightly higher in the frame with different proportions.
Same frame at 4K. Top: SyncTile. Bottom: MMH3 Split Upscale, more texture but the face has shifted.
Same 960 x 1080 region of the 4K café clip, 1:1. Left: SeedVR2 7B. Right: SyncTile. Watch the skin and background texture in motion.

SyncTile stays closest to a plain resize in motion while adding real detail. The last step is a light unsharp mask (5 x 5, amount 1.0), which we matched to a manual sharpening pass in an editor. It lifts fine detail by about a quarter and keeps shimmer at 2.1, still well below SeedVR2.

A 1920 x 1080 window cut 1:1 from the middle of the 4K café clip, after SyncTile and the light unsharp mask.

What we learned

  • Speed-ups that skip work are not free. On MiniMax H3, step caching saved time by letting the scene drift. The safe wins did the same work faster: a distilled LoRA and int8 attention, 4.9x together.
  • Test reported gains on your own setup. SageAttention and text-encoder caching both have good public numbers. Here one was slower and the other saved nothing.
  • Tiles need a shared map and a shared clock. One base clip gives the tiles one composition, and blending them after every step keeps two eyes in two tiles looking the same way.
  • Judge video in motion. The four-tile result and SeedVR2 both looked best as stills and failed in playback. A sharpness score can reward noise; a frame-to-frame shimmer score did not.

Where this is used now

Two MiniMax H3 settings are now our defaults. Drafts and most clips use 8 steps with int8 attention. When a project needs a large, sharp picture, the clip goes through SyncTile to 4K with the last 3 of 8 steps and a light unsharp mask. The code, the exact settings and the 4K demo clips are open source on GitHub: alaintural/synctile-h3 (MIT). So far SyncTile has been tested on 5-second clips. Longer clips make every tile heavier, so they are the next thing to measure.

All clips in this note were generated with AI and show no real people.

Sources

  1. SyncTile for MiniMax H3 (GitHub)

Ask AI about this AI Lab note

Opens the assistant in a new tab with this page as the source.

Keep reading

Want this handled for your business?

Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.