Image & Video
Checking AI images against measured numbers for a dark, warm look
68 reference images measured, their tone turned into pass or fail limits, and every AI image checked against them. What the numbers caught, what they missed, and the color problem that only a correction step fixed.
SummaryUnder a minute
The short version
A set of 12 AI-generated lifestyle images had to share one dark, warm editorial look. A measured reference set was turned into tone limits, and every image was checked automatically. The check caught drift in brightness and color, a correction step fixed a repeated cast, and people still reviewed what numbers cannot see.
Key takeaways
- 68 reference images were measured pixel by pixel and their dark, warm look was written down as numbers.
- Every generated image was checked on five tone readings and made again with a plain-English correction note if it failed, up to 3 tries.
- One model added an amber cast to almost every frame, and making the image again did not fix it, while a correction step after generation did.
- The check only sees tone, so hands, faces, composition and details leaked from the reference image still need a person.

The question
A set of 12 AI-generated lifestyle images is meant to share one look: the calm, dark, warm editorial photography seen on a well-known consumer electronics site. The images sit next to each other on the same page, so a single frame that comes out too bright or too orange breaks the whole set.
The usual way to keep a set consistent is a person looking at each image and deciding whether it "feels right". That works for a handful of images, but the judgment shifts with the screen in use and with how many images have already been reviewed that day. So the question is simple: can the look be written down as numbers, and can every generated image be checked against them before anyone spends time reviewing it?
The setup
The starting point was 68 images from that site, measured pixel by pixel on a 0 to 255 scale. Anything under 45 was tagged as shadow, 100 to 160 as midtone, and over 200 as highlight, which shows exactly how the light behaves in the target frames.
The images fell into four families: dark warm interiors, daylight scenes, products on white, and scenes built around a device screen. The dark warm interiors were chosen as the target. In those, the shadows lean toward brown and orange with almost no blue. One office scene had shadows at 29, 15 and 7 on the red, green and blue channels, and another came in at 31, 16 and 8. That reads as dark chocolate, and there is no pure black anywhere in those frames. The midtones are warm and beige (148, 125 and 99 in the office scene), while the brightest parts stay close to a neutral cream around 232, 223 and 206, so only the darker half of the image is warmed.
- Red29 count
- Green15 count
- Blue7 count
Source: Pixel measurement of the reference set
Mean brightness in that family sat between 0.18 and 0.27, and the white point (the brightest value in the image) was clipped on purpose. In the office scene the brightest value was 165, and in a classroom scene it was 149, so even the window light never blows out. Roughly a quarter of each frame sits in darkness, with the light falling on a face and the hands holding a device. In the same office scene the labels on the shelf behind the person are almost readable but stay soft, which is what a 35-50mm lens at around f/2-2.8 tends to give.

Saturation was a small surprise. The reference lifestyle images measured at 0.41 to 0.57 saturation, which is fairly colorful, but the colors were earthy (mustard, rust, olive, burgundy) with no neon or pure blue anywhere. Often the only cool color in the frame was a single piece of clothing, like a blue shirt, used as an accent.

Images were generated with two general-purpose image models, one image at a time. Each image then went through a correction pass that pulls the light toward neutral cream, pulls saturation toward earth tones, warms the shadows and adds a thin layer of grain.
Measurement
The check is a measurement gate: a small script that opens the finished image, takes five readings and answers pass or fail. The readings are mean brightness, the 99th-percentile highlight (the level only the very brightest pixels exceed), the share of the frame that is dark, shadow warmth measured as red minus blue in the shadows, and mean saturation. Its job is to stop an image that drifts too bright or too saturated before review time is spent on it.
Each style has its own limits. For the dark warm style, a mean luminance above 0.42 fails, a 99th-percentile highlight above 242 fails as blown highlights, and saturation above 0.60 fails.
When an image fails, the check writes a short correction note in plain English, such as a line saying the highlights are too bright or the shadows are too cold. That note is added to the written instruction for the image model (usually called the prompt), and the image is made again, up to 3 tries.
The yardstick has limits. It measures tone and nothing else, so an image can pass every number and still have stiff hands or an odd crop. It is tuned to one style family, so the same thresholds would wrongly fail a good bright daylight image. It is also built from one site's photography, which is a taste decision, so it says nothing about whether that taste fits a given use.
Results
The dark warm interior style, paired with the office scene from the measured set as the reference image, gave the best result.
The more useful finding was about color. One of the two models pushed an amber, golden-hour cast into almost every frame. Making the image again with a correction note seemed like the obvious fix, and it failed, because the shift came back each time. The correction pass after generation fixed it by pulling the light back to neutral cream. The other model had the opposite habit and gave cold, bluish shadows, and the same correction pass warmed them.
A few frames came out far too dark. For those, the midtones were brightened and the black point (the darkest value in the image) was pulled back down, so the image opened up while keeping the deep shadows that define the style.
Details also leaked from the reference image into the result. If the reference person wore a beanie or had a beard, the generated person sometimes did too, even when the instruction didn't ask for it. The check can't see that, so a person had to.
The tooling was not always reliable either. On September 29, 2026, the second model's full-size downloads kept arriving as empty files of 0 bytes, and clicks failed with an error saying the button was not "enabled and stable". The image was still visible in the chat window, so the 1024 preview was captured by drawing it onto a canvas in the browser, the chat address was saved, and the full-size file was fetched later. In the end only 2 of the 12 images came from that model.
Takeaways
Measure the reference before writing a single instruction. Numbers for shadows, midtones and highlights give the check something concrete to compare against, and they also supply the plain words that go into a correction note, such as "shadows too cold".
Split problems into the ones a model repeats and the ones it gets wrong at random. Random misses respond well to another try with a correction note. Repeated habits, like a warm cast in every frame, belong in a correction step after generation.
Keep a person on everything the numbers can't see. A tone check only covers tone, and hands, faces, composition and leaked details still need human eyes before an image goes on a page.
A stronger version would measure framing as well as color from the start, for example where the light falls in the frame and how much empty space sits around the subject.

Cropping for layout
Some of the images were cropped tall to 556x778 for a narrow layout slot, with the crop centered on the subject. A tall crop from a wide frame cuts away most of the dark surround, so it is worth running the gate again on the cropped file, since the dark-area share and mean brightness both change with the crop.
Ask AI about this AI Lab note
Opens the assistant in a new tab with this page as the source.
Keep reading
AI Lab · Image & VideoGrading GPT and Gemini color casts, and why color belongs in the promptOctober 8, 2026
AI Lab · Image & VideoSeedVR2 vs ESRGAN on one RTX 4090: what video upscaling costsOctober 6, 2026
AI Lab · 3D & WebTaking a static one-page site from 60 to 81 on mobile LighthouseNovember 3, 2026
Want this handled for your business?
Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.