Local AI
What fits on one RTX 4090, measured on real jobs
Peak video memory for common local AI jobs on a single 24 GB GPU: which jobs leave room on paper, which take the whole GPU, and which settings should stay fixed.
SummaryUnder a minute
The short version
Peak video memory was measured for common local AI jobs on one RTX 4090 with 24 GB. A language model, video restoration and video generation all claim most or all of the GPU, while a classic upscaler and an image-to-3D model sit at the light end. Speed tricks and memory-saving settings both needed their own memory checks before they could be trusted.
Key takeaways
- A 27B language model took 17 GB and stayed 100% on the GPU at 16k context, and past 16384 tokens it slipped onto the CPU.
- The 7B fp8 restoration model hit 15.2 to 18.1 GB, while the classic upscaler needed only 8.0 GB and the image-to-3D model came in at 7.3 GB.
- Local video generation at 1344x768 used 24.6-25.5 GB, which is the ceiling of the GPU.
- The latent upscaler cut generation time from 54.1 to 29.5 seconds, a 1.84x speedup, but it also raised VRAM usage, so it saves time and costs memory.
- With tiling switched off for the decode step, each chunk of frames took 13-14 minutes, against 4 minutes with tiling back on.

The question
The test machine is a Windows 11 Pro workstation with an i9-14900KF, 96 GB of system RAM and a single RTX 4090 with 24 GB of video memory. On a setup like that, video memory is the budget every other decision answers to. The size of a model file says surprisingly little about what a job needs once it starts, because settings like context length and output resolution add memory on top of the model itself. The goal was a plain list to plan around, showing which jobs leave room on paper and which take the whole GPU.
Running out of memory rarely means a crash. Usually the job slows down or quietly starts using the CPU instead, and that drag is heavy. One step that ran in 4 minutes stretched to 13-14 minutes for each chunk of frames when memory got tight, so the time cost of overflowing video memory is the thing that hurts most.
The jobs
Each measurement comes from a typical local AI job run with realistic inputs.

A local language model
The 27B language model runs with a 16k token context, which is the limit for how much text it can hold in view at once.
Video restoration and upscaling
For footage that needs repair, the job is a 7B video restoration model stored in fp8, a compact 8-bit format that keeps the file small enough to load without trouble. When footage is already clean and just needs more pixels, a classic upscaler does the job instead, raising the resolution without inventing new detail.
The first hour on September 20, 2026 went to a setup issue that had nothing to do with memory. The image tool refused to start because the restoration add-on prints emoji when it loads, and the Windows console was set to a legacy code page instead of UTF-8. The error message pointed at the logging code rather than the add-on, which made the real cause hard to spot. Switching the console to UTF-8 at startup fixed it, and those lines now sit at the top of the startup script with a comment saying not to delete them.
Image to 3D and a set of detector models
An image-to-3D model was measured too. Separately, a set of AI text detector models runs on the same machine, and the interesting part there is how they load when memory is tight.

Local video generation
The local video model produces short clips at 1344x768. It was also run with a latent upscaler, a smaller model that raises the resolution while the video is still in the compressed internal form the generator works in, before it becomes actual frames.
Measurement
The tracked number is peak video memory during each job, since that is what decides whether it fits. For the language model the extra check was whether it stays on the GPU at its full context length, because that is usually where layers start sliding over to the CPU. For the latent upscaler, generation was timed with and without it.
The limits are plain. Everything comes from one machine, so other drivers and settings will move the numbers. Peak memory depends on the input and on settings like resolution and context length, which is why some results are ranges. Output quality was not scored here, only one comparison was timed where speed was the point, and two jobs sharing the GPU were not tested.
Results
The 27B language model took 17 GB and stayed 100% on the GPU at 16k context. On paper there is memory left beside it, but its working memory grows with the conversation, so it runs alone. The edge showed up directly: past 16384 tokens of context, the cache it keeps for the conversation grew enough to push the model onto the CPU, and it slowed down. The chat front end now clips the context at 16384 so nobody has to remember.
Video restoration with the 7B fp8 model peaked at 15.2-18.1 GB. The two ends came from two real clips. An old 640x480 clip of 451 frames, enlarged to 1440x1080, peaked at 15.2 GB. A 1344x768 clip of 141 frames, enlarged to 1890x1080, peaked at 18.1 GB. The second clip was shorter, but both its input and its output were larger, so two clips can't say which of those moved the peak.
Turning off the memory-saving tiling for the decode step, which turns the model's compressed output back into frames, turned out to be a bad move, even though that step took the largest share of the time in the first measurement. Tiling looked slow, so it was switched off, but at 1890x1080 output the memory overflowed into system RAM and decoding each chunk of frames stretched to 13-14 minutes. With tiling back on, the same work took 4 minutes. The tool had logged a warning that memory was being paged to system RAM, and that warning is now treated as a sign that a setting is wrong.
The classic upscaler and the image-to-3D model need only 8.0 GB and 7.3 GB respectively, so on paper they leave room on a 24 GB GPU for another task, though two jobs sharing the GPU were never actually measured.
The local video generation job at 1344x768 is the odd one, at 24.6 to 25.5 GB, which is technically more than the 24 GB listed for the GPU. There's no good explanation yet for a reading above the spec, but the practical fact is clear: while this job runs, the GPU is full and nothing else fits.
- 27B language model at 16k context17 GB
- Video restoration, smaller clip15.2 GB
- Video restoration, larger clip18.1 GB
- Classic upscaler8 GB
- Image-to-3D model7.3 GB
- Video generation at 1344x768, top reading25.5 GB
Source: Workstation measurements on one RTX 4090
The latent upscaler cut generation time 1.84x, from 54.1 to 29.5 s, which adds up across many clips. It also raised video memory use, so it buys speed with memory.
- Without latent upscaler54.1 s
- With latent upscaler29.5 s
Source: Workstation measurements on one RTX 4090
The detector models load in bf16, a 16-bit format that keeps more precision, only when at least 17.5 GB is free. With less free memory they load in 8-bit, which takes less space, and one after another instead of together, so the work still completes in sequence.
Takeaways
Lined up by measured peak memory, the jobs split into two groups. The heavy ones, everything above the classic upscaler and the image-to-3D model, take the whole GPU. The lighter pair leaves space on paper, but pairing one with another job is still unmeasured, so that check comes before running them together.

Any setting that promises speed needs a memory check, and that includes turning off features meant to save memory. The latent upscaler raises video memory use, so assuming it saves space would have broken the plan. The tiling issue taught the same lesson from the other side, which is why memory gets measured again after every change that claims to make things faster.
A setup that loads several models needs a fallback, and the detector models are the example: they switch to 8-bit and load one after another when free memory drops below 17.5 GB. The work finishes even when memory is tight, at the cost of speed.
Starting over, the better habit is to record time and memory for every job from the start, with the exact settings written next to each peak. Some of the ranges here are honest but hard to explain line by line, and the video generation reading above 24 GB still needs pinning down.
Settings in use
The chat front end caps the context at 16384, and tiling stays on for restoration decoding.
When the detector models load, they check free memory and drop to 8-bit if it's tight, so the job runs slower but still runs.
Ask AI about this AI Lab note
Opens the assistant in a new tab with this page as the source.
Keep reading
Want this handled for your business?
Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.


