Image & Video

SeedVR2 vs ESRGAN on one RTX 4090: what video upscaling costs

A 640x480 clip taken to 1080p and then 4K on a single local GPU, with every step timed. Where the minutes and memory go, and why the 4K step isn't worth running.

Pixedi AI Lab
  • 5 min read

SummaryUnder a minute

The short version

Two upscalers were compared on a single RTX 4090. The restoration model took a 640x480 clip to 1080p quickly, while the lighter sharpening model took much longer to reach 4K and the output looked almost identical to the 1080p version. Keeping memory tiling on keeps the job fast, and most of the time goes into compressing and expanding frames.

Key takeaways

  1. SeedVR2 took the 640x480 clip with its 451 frames to 1440x1080 in 9.0 min at 0.83 fps, with VRAM peaking at 15.2 GB.
  2. RealESRGAN x4 then needed 31.5 min to take the 1080p output to 4K, and the result looked almost identical to the 1080p input.
  3. ESRGAN only sharpens existing pixels and adds no new detail, so the full 40.5 min two-step route was dropped.
  4. Turning VAE tiling off spilled memory into system RAM and made decode 17x slower, so tiling stays on.
  5. In SeedVR2, VAE decode and encode took 54% and 22% of the time, while the diffusion model itself took only 21%.
A desktop workstation with a large graphics card visible through its side panel in a dim home office.

The question

Old 640x480 footage looks soft and blocky on modern screens. The question here is what it actually costs to pull a clip like that up to 1080p, and then to 4K, on local hardware, and whether the final jump to 4K changes what a viewer sees.

An out-of-focus monitor next to a small older laptop on a wooden desk in daylight.

Everything runs on one RTX 4090 with 24 GB of graphics memory. Since nothing leaves the local machine, the only costs tracked are minutes and GPU memory. The test asks how long each step takes, how much memory headroom it needs, and whether the 4K step adds anything visible or only adds pixels.

The setup

The test ran on September 20, 2026 on a single workstation. Two upscalers were compared. SeedVR2 uses diffusion to find damage and rebuild detail, and it looks at several frames at once so the picture doesn't jump between frames. The 7B version in fp8 was used because it fits memory better, with the 3B version kept for quick tests. RealESRGAN x4 is an older, lighter network that enlarges each frame and sharpens what is already there.

The clip was 640x480 with 451 frames. The plan was to run it through SeedVR2 to reach 1440x1080, then hand that output to RealESRGAN x4 for the jump to 4K, so the heavier model does the restoration and the lighter one only enlarges.

VAE tiling matters a lot here. Before the main model starts, a component called a VAE compresses each frame into a compact latent form and later expands it back, and tiling splits the frame into smaller chunks so that step doesn't take all the memory. A tile size of 1536 works when there is headroom, and 768 keeps things running once memory gets tight.

SeedVR2 processes frames in groups, which it calls batches, and the group size has to be one more than a multiple of four: 5, 9, 13, 17, 21, 33 and so on. Sizes in between don't work. For short shots the batch is set close to the shot length, because that is where frame-to-frame consistency comes from. A three-frame overlap between groups plus color correction keeps skin tones and skies from drifting between groups.

The first failure had nothing to do with video. The software refused to start because the upscaler add-on prints emoji on startup, and the Windows console was set to cp1254, a legacy code page that can't encode them. The encoding errors chained until the app crashed, and the trace pointed at the logging code instead of the add-on. Switching the console to UTF-8 at the top of the launch script fixed it, and that line should stay there.

Measurement

Each step recorded wall-clock time, frames per second and peak VRAM (the GPU's own memory). For SeedVR2 the time was also split by phase, to show which part is actually slow. For the 4K step the output was compared visually with its 1080p input.

These numbers are specific to one clip on one machine, so a different length or GPU will move them. The quality check for the 4K step was visual only, with no score behind it, and not every setting was tested, so a faster configuration may exist.

Results

The restoration step

SeedVR2 took the 640x480 clip to 1440x1080 in 9.0 minutes for all 451 frames at 0.83 fps, and VRAM peaked at 15.2 GB, which leaves a comfortable margin on a 24 GB GPU.

The 4K step

RealESRGAN x4 took the 1080p output to 4K in 31.5 minutes with VRAM peaking at 8.0 GB. It was lighter on the hardware than the restoration step but slower on this clip, and side by side, the 4K version looked almost identical to the 1080p source.

That follows from what the model does. It sharpens the pixels it receives and generates no new detail, so a clean 1080p input leaves little for it to fix, and the result is a bigger file that looks about the same. The full two-step route took 40.5 min, and since the second half added nothing visible, it was dropped.

The tiling mistake

A second clip, 1344x768 with 141 frames and output at 1890x1080, was run with VAE tiling switched off on the theory that splitting frames was unnecessary overhead. Memory use exceeded the GPU's capacity and spilled into system RAM, which is far slower for this workload. The log states that VRAM overflowed and was paged to system RAM, and that line is a reliable sign the settings are wrong. Without tiling, decode took 13 to 14 minutes per batch, and with tiling it took 4 minutes per batch. With tiling back on, the whole clip finished in 4.0 minutes at 0.59 fps, with memory peaking at 18.1 GB.

An open desktop computer showing memory modules beside a large graphics card on a workbench.

Where the time goes

The phase breakdown says more than the total. In SeedVR2, VAE decode took 54% of the time and encode another 22%, which left 21% for the DiT to do the actual restoration. The model doing the restoration gets the smallest share, and compressing and expanding frames take most of the rest.

Where SeedVR2 spends its time%
  • VAE decode54%
  • VAE encode22%
  • DiT model21%

Source: Measurement on one RTX 4090

Takeaways

Choose the tool from what the source looks like. Small, damaged footage needs a restoration model that rebuilds detail, while footage that is already clean only needs more pixels, and a light sharpener adds those without making the picture better.

Look at the output before committing the whole job. The 4K step looked reasonable on paper and took 31.5 minutes, and it added nothing visible. Testing a few seconds first shows whether a step is worth its time.

Tiling looks like overhead, but on a 24 GB GPU it is what keeps the job out of system RAM, so a memory spill warning should be fixed before any other tuning.

Profile before tuning. Speeding up SeedVR2 by working on the model itself would target only 21% of the time. The phase split suggests that encode and decode are where a speedup would have to come from, though that hasn't been tested yet.

A person at a desk comparing two printed video frames side by side in window light.

A better process would test every stage on a few seconds of footage and compare by eye before running a full clip, and repeat the measurements on more clips so the numbers rest on more than two examples.

For low-resolution footage of this kind, go straight to 1080p with the restoration model, keep tiling on, and stop there. The 4K sharpening step isn't worth running by default because it adds nothing visible. When 4K is required, treat it as a separate decision and check whether the larger file actually shows anything new.

Ask AI about this AI Lab note

Opens the assistant in a new tab with this page as the source.

Keep reading

Want this handled for your business?

Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.