Image & Video

Grading GPT and Gemini color casts, and why color belongs in the prompt

The same image brief went to a GPT model and a Gemini model, came back with opposite color casts, and one shared grade fixed both. What that grade did, where the process broke, and why color now lives in the instruction instead.

Pixedi AI Lab
  • 5 min read

SummaryUnder a minute

The short version

A GPT model and a Gemini model got the same brief for a batch of web images and returned opposite color casts: amber from GPT, cold blue shadows from Gemini. A shared grade fixed both for that batch, while Gemini's failed downloads limited it to 2 of 12 images. Local color work was later dropped, and color is now written into the instruction.

Key takeaways

  1. GPT added an amber, golden-hour shift to every frame, and regenerating the image did not fix it.
  2. Gemini went the other way with cold, blue shadows, and the same grade warmed them.
  3. One grade that pulled light toward neutral cream, muted saturation to earth tones and added fine grain brought both models onto one palette.
  4. Only 2 of 12 images in the batch came from Gemini on September 29, 2026, because its full-size downloads came back as 0 bytes.
  5. Adding vividness in local post-processing was later dropped, so color is now set in the instruction and the model's raw output is final.
Two prints of the same interior on a wooden table, one with an amber tint and one with cool blue shadows.

The question

A set of 12 images for one web page has to look as if it came from a single photographer, with warm interiors and a palette that holds steady from one section to the next. The test gave the same written brief to two image models, one from the GPT family and one from the Gemini family, to see how far apart they land on color.

A page gets read as a whole, so if one image glows orange and the next feels like a cold morning, the page starts to look stitched together. Underneath that sits a second question, which turned out to matter more over time: should a color fix live in how the image is described before it exists, or in a correction step on the local machine afterward?

All images in this note are AI-generated test frames made to illustrate the color behavior, and none shows a real space.

The setup

Both models got the same brief, and the images were worked through one model at a time in a fixed order, each image with its own scene, style and shape written down in advance. Of the visual styles tried, a dark, warm interior paired with a reference image held up best, though small details from a reference could leak into the new frame, so a beanie or a beard would sometimes turn up on a person who wasn't meant to have one.

Both models came back with a color cast, meaning an unwanted tint that sits across the whole frame. GPT's cast was amber: every frame looked as if it had been shot in the last light before sunset, which can be pleasant in a single picture but turns white walls orange across a full page. It first showed on a plain white wall behind a workbench, which came out the color of weak tea however the wall was described. Regenerating the same image didn't help, because the tint came back each time.

A room with white walls that look orange under strong amber light.

Gemini went the other way, with cold, blue shadows that made the darker corners of a warm wooden room read as cool and a little clinical even though the materials in the scene were warm.

A wooden room whose shadows carry a cool blue tint despite the warm materials.

To bring both onto one palette, a small grade script applied the same color correction to every image. It pulls the light toward a neutral cream, mutes saturation toward earth tones and warms the shadows, then adds a fine grain for a photographic texture. On GPT images the grade took the amber out, and on Gemini images it warmed the blue shadows. When a frame came out too dark for its spot, a gamma adjustment lifted it and the black point was pulled back down so the shadows didn't wash out to gray. Each file was backed up before editing, and the finished versions were saved in a compact web format (AVIF). Some images were cropped to a tall 556x778 frame, with a focus point set for each one so the important part of the scene stayed inside the crop.

Measurement

After the grade, each image went through an automated check that compared it against reference measurements for the target look. A failed image was made again with a short correction note added to the written instruction, up to 3 times, then graded and checked once more.

That check has limits. It sat after the grade, so it showed whether the corrected image hit the target, which says less about how far off each model's raw output was. It also only knows the targets it is given, so every image still had to be viewed in place on the page, next to its neighbors, to judge whether the highlights read as cream and the shadows felt warm enough for the room. No per-image scores from that session are available, so the results below are described in words.

Results

The GPT cast was the most predictable part of the exercise, since it appeared on every frame and survived regeneration, and the grade removed it each time. A fault that consistent is easier to handle than a random one, because one correction can be applied everywhere.

The Gemini cast also responded to the same grade, which warmed the cold shadows until those frames sat comfortably next to the GPT images.

On September 29, 2026, Gemini's full-size downloads, which on a normal day arrive at 1536x2752 in about 12 seconds, came back as 0 bytes, and clicking the download control sometimes threw an error saying it wasn't "enabled and stable". The finished image was still visible in the chat window, so the 1024 preview could be saved by drawing it onto a canvas in the browser, and the chat address was noted so the full-size file could be fetched later from the same conversation. The failed downloads left empty files behind, which were deleted. There is no confirmed cause for any of this, and it stays an open question. In practice, only 2 of 12 images in the batch came from Gemini.

A later attempt pushed further in local post-processing, adding vividness with an S-curve (a contrast adjustment that deepens shadows and lifts highlights) and a vibrance boost. That direction was dropped, and the method changed: color is now set in the written instruction given to the image model, and the model's raw output is final. The grade worked for that batch, but it is no longer part of the process.

Takeaways

A steady color cast belongs to the model

When the same tint shows up on every frame and regeneration doesn't clear it, treat it as a trait of that model. Asking again usually wastes time, so the choice is between correcting it the same way on every image or describing the light differently from the start.

One shared target matched two models in that batch

The grade worked on both casts because it pushed everything toward one target, a neutral cream light with earthy saturation. Correcting every image toward the same target is what made a mixed batch look like one set. The part worth keeping is the target itself: decide on it before the first image, and write it into the instruction each model receives.

Four interior prints pinned in a row that share the same cream highlights and earthy colors.

Keep a second model ready for the same brief

Two models working from the same brief gave a fallback when one model's downloads came back empty. For any batch of site images, keep two models briefed the same way, and save the chat address whenever a download fails.

Decide early where color lives

The cleaner approach is to describe the light and palette in the instruction to the model from day one and plan the batch without a grading stage. Color is then set in the written instruction, and the model's raw output is the file that goes on the page.

Ask AI about this AI Lab note

Opens the assistant in a new tab with this page as the source.

Keep reading

Want this handled for your business?

Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.