Automation

A number check that blocks unsourced figures in generated reports

An automatic check that fails any report containing a figure with no measured source behind it: how it works, what it still misses, and the wording problem it did not catch in the first edition.

Pixedi AI Lab
  • 5 min read

SummaryUnder a minute

The short version

Recurring reports drafted by a language model need a guarantee that no figure reaches a reader without a measured source. Every figure is stored with the query that produced it, and any report containing an unmatched number fails automatically. The check worked on numbers, but it missed internal jargon in 11 of 12 first-edition reports, so wording now has its own check.

Key takeaways

  1. Every percentage, dollar amount and number of 10 or more in a report must match a stored figure, or the report fails.
  2. Each stored figure keeps its sentence, value, period, source ids and the exact SQL query that produced it.
  3. When the check fails, the model receives the list of unmatched numbers and gets exactly one retry.
  4. The number check missed internal jargon, which appeared in 11 of 12 first-edition reports, so a separate vocabulary check was added.
  5. Figures based on samples under 30 are marked preliminary automatically.
A printed report on a desk with pencil check marks beside its figures, next to a coffee mug and a laptop.

The problem

In recurring written reports where a language model does most of the drafting, the model writes fluent prose, and it writes fluent numbers just as easily, including numbers nobody ever measured. When a measured figure and an invented one sit in the same confident sentence, the reader has no way to tell them apart. Someone takes a figure at face value and passes it on, and an invented percentage ends up behind a real decision.

The design question here is narrow and practical: can a report be made mechanically unable to reach a reader with a figure that has no measured source behind it? And once that works, what does the check still miss?

The design

The work splits into two halves. The first half has nothing to do with writing. Before any drafting happens, the figures are computed from the data and each one is stored as a small record. A record keeps the sentence the figure belongs in, the value itself, the period it covers, the ids of the sources it came from, and the exact SQL query that produced it. If anyone later asks where a number came from, the answer is the query, which can be read and repeated.

The second half is the drafting. The model receives the stored figures and writes the report around them. Before the report goes anywhere, an automatic number check reads the finished text and pulls out every number that matters. The rule is simple to state: every percentage, every dollar amount and every number of 10 or more has to match a stored figure. Years are left out, since a report mentioning 2026 makes no claim with that number. Small counting numbers below the threshold are skipped as well, because in a phrase like "three steps" they are ordinary words and carry no finding.

The check is a short script. It takes two inputs, the finished report and the sector the report covers, and compares the text against that sector's stored figures, which live together in one file per sector. Before comparing, it strips dollar signs, percent signs and thousands commas, and then rounds both sides to one decimal place, so a figure stored with more precision still matches the way a person would naturally write it. Anything shaped like a four-digit year from the last two centuries passes untouched. A number in the text counts as matched when it equals a stored value or a number inside a stored figure's sentence.

First-edition reports and leaked internal jargoncount
  • First-edition reports12 count
  • Reports with leaked internal jargon11 count

Source: Report generation log, September 30, 2026

A failing report produces a single line: a failure message followed by every number that has no matching stored figure, written the way it appeared in the text (dollar sign and percent sign included), sorted and separated by commas. The script also exits with an error, so the report is not treated as finished. It does not try to fix anything itself. That line goes back to the model as the error, and the model gets exactly one retry to rewrite those sentences.

An analyst compares a printed page with handwritten notes at a standing desk.

Two smaller rules sit next to the check. Any figure based on a sample under 30 is marked preliminary automatically, so the reader sees that label beside a number that might otherwise look more solid than it is. And when a report type has no edition from the month before, the model is told it is writing a first edition, so it doesn't write "unchanged" or refer to an earlier edition that never existed.

Measurement

The yardstick for the number check is binary: every qualifying number in the final text either traces back to a stored figure with its query attached, or the report fails. Anyone with access to the stored figures can confirm a pass or a fail independently.

Its limits are worth stating plainly. The check confirms that a number exists among the stored figures, and it has no way to confirm that the number sits next to the right claim, so a real value attached to the wrong sentence would pass. It ignores small numbers by design, which means an invented "seven out of nine" style claim would slip through. It also reads only numbers, and anything wrong with the words around them is outside its view. That third limit showed up in the very first edition.

A red pencil resting on printed pages with one line circled.

Results

The first edition of the reports was produced on September 30, 2026. The number check behaved as designed and looked at numbers only, so it missed a problem it had no way of seeing: internal shop talk had leaked into the text. The leaked words were the system's own working vocabulary. The internal nickname for a stored figure showed up in the reports, and so did the internal name for the whole reporting process. It happened in 11 of 12 first-edition reports.

None of those reports would have failed the number check, because the figures in them were sourced. The trouble was in the language around the figures. A reader meeting unexplained system vocabulary would reasonably wonder what they were looking at, and correct numbers do little to help with that.

The fix was a separate vocabulary check. It holds a list of internal words that must never reach a reader, along with the plain words to use instead: a stored result is a "figure", a monthly report is an "edition", and the old internal confidence label is now "preliminary". References to the system's own processing steps are on the list too. The same guidance goes to the model before it writes. If any listed word appears in a finished report anyway, the report stays as a draft and does not go out.

A quiet conference room table with a single printed report left on it.

Takeaways

Store the query with the figure

Keeping the exact SQL query next to each figure turns a question about a number into a lookup, and nobody has to remember how a figure was built weeks earlier.

Make the error message specific

The model receives the exact list of numbers with no matching stored figure, which tells it which sentences need to change. It gets one retry and no more. That limit was set deliberately, and how it compares with allowing several attempts hasn't been measured.

Give each kind of failure its own check

The number check was built for figures, and it does that job. Wording is a separate kind of failure, and leaked internal vocabulary needs its own list and its own check. A better order is to read a few drafts aloud to someone outside the project and write the vocabulary list before the first edition goes out.

Label small samples automatically

Anything based on a sample under 30 is marked preliminary without anyone having to decide, and the reader gets a fair signal about how much weight a figure can carry.

Current state

Every recurring report now passes through both checks. The number check fails any report with a percentage, dollar amount or number of 10 or more that has no matching stored figure, and the vocabulary check holds back any report that still contains internal words. The first-edition rule switches itself off from October, once each report type has an earlier month to compare against.

Ask AI about this AI Lab note

Opens the assistant in a new tab with this page as the source.

Keep reading

Want this handled for your business?

Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.