Automation

Monthly industry reports from 329 cataloged public data sources

A monthly job that turns free public data into industry reports and checks every number in the text against a stored figure before human review. The setup, how it was measured and where it fell short.

Pixedi AI Lab
  • 5 min read

SummaryUnder a minute

The short version

A monthly job collects free public data, turns it into stored figures, has a language model write industry reports from those figures only, and checks every number before a person sees the draft. All 12 September deep reports passed the number and voice checks, but internal system labels leaked into most first drafts and needed a separate fix.

Key takeaways

  1. The catalog holds 329 free public data sources across 12 industries plus a shared set, and 246 were working at launch.
  2. Each figure is stored with its sentence, value, period, source ids and the exact query that produced it, so every number in a draft can be checked against one that already exists.
  3. All 12 September deep reports passed both the number check and the voice check, and repeating the rewritten analysis reproduced 301 of 301 existing figures exactly.
  4. Internal system words leaked into 11 of 12 first drafts, which led to a separate vocabulary check.
  5. Nothing is published automatically, and every draft waits for human approval.
Printed statistical tables, a closed laptop and a coffee mug on a desk in morning light.

The question

Can a monthly industry report be built from free public data without anyone assembling it by hand, and without a language model inventing a single figure along the way? The second half is the hard part. A report lives or dies on its numbers, and one wrong figure in a paragraph is usually enough to make a reader doubt the rest.

So the test here is narrower than "can a model write a report." The model gets only verified figures, writes the prose, and afterwards every number in the text has to be traceable to a real query against real public data.

The setup

The data layer is a catalog of 329 free public data sources across 12 industries, plus a shared set that several industries draw from, such as government employment and wage tables. Each source is pulled by a small connector that matches the way the publisher shares it, so open data portals, spreadsheets, CSV files and federal statistics services each get their own handler.

Archive boxes and binders on metal shelves in a bright records room.

On top of the catalog sits a single script that runs the monthly job in order. It collects the sources that are due, analyzes them into figures, hands those figures to a language model to write the draft, and checks the draft. A draft that passes goes onto a list and waits for human approval. Nothing is published automatically.

Hands sorting printed report pages into two piles, one marked with colored tabs.

Most of the design work goes into the stored figure. Each one keeps the sentence it belongs to, the value, the period it covers, the ids of the sources it came from and the exact database query that produced it. The model receives those figures and a short list of what changed since the month before, with no web access and no other data to draw from. Any number it writes is supposed to already exist as a stored figure, and the checks below confirm that it does.

Sources that don't cooperate are handled explicitly. A server that is down for maintenance gets retried 3 times, and if it still fails, the job leaves it for next month and doesn't record a fresh pull. Sources that need an access key that isn't available yet, or a file that has to be downloaded by hand, are skipped and written to a list, so the gap is always visible.

Measurement

Every draft goes through two automatic checks first. The number check reads every number in the text and matches it against the stored figures, so a number with no match fails the draft. The voice check looks for phrasing that reads as machine-written, including long dashes. A third, a vocabulary check, was added later for a reason covered below. A failed draft goes back to the model once with the errors attached, and a draft that still fails stays marked as a draft.

The analysis side was tested on its own. After the analysis code was rewritten, it was repeated against the same data and every figure was compared to the ones already stored. A simulated next month then checked whether the whole sequence holds up with a new period.

These yardsticks have limits. The number check proves that a number traces back to a stored figure. It can't tell whether the publishing agency released a bad table, or whether the sentence around the number draws a fair conclusion from it. The voice check is a rule set that catches common patterns, so a draft can pass it and still read flat. A simulated month is a rehearsal, and the first real one says more.

Results

Of the 329 sources in the catalog, 246 were working at launch. The rest are not live yet for tracked reasons, such as a missing access key or a file that has to be downloaded by hand. They stay in the catalog with their status, so the gap shows up every month.

Public data sources at launchcount
  • Sources in the catalog329 count
  • Working at launch246 count

Source: Public data source catalog

All 12 September deep reports passed both the number check and the voice check. A deep report is the longer written analysis for an industry, as distinct from the shorter monthly updates the same job also produces. Repeating the rewritten analysis reproduced 301 of 301 existing figures exactly, and the simulated next month finished without errors.

The surprise was language. Internal labels from the system itself, the shorthand used for a stored figure and for the prior month's results, leaked into 11 of 12 first drafts. The number check didn't catch it because nothing was wrong with the numbers, and the voice check didn't catch it because the sentences read fine. The model was echoing the labels it had been handed along with the figures. A separate vocabulary check now flags those words, and it sits next to the other two before a draft can reach review.

A person reviewing a printed draft at a standing desk and marking a word with a pen.

There was also a quieter bug in the data. One employment dataset was spread across 50 URLs that ended in the same file name, so until September 29, 2026 the job was writing those 50 downloads into 5 files. The table ended up repeated 10 times over a single period, and nothing crashed. Naming each file from its full path fixed it.

Takeaways

The most reusable habit is to keep the query with the number. If only the value is stored, it can't be checked later without redoing the work. With the sentence, the period, the sources and the query kept together, checking a draft comes down to looking each number up, and a rewrite of the analysis can be compared with the old one figure by figure.

The leak in 11 of 12 drafts came from the input side. The labels used to hand over the figures were internal working shorthand, and the model repeated them in the prose. A cleaner design gives the model reader-facing labels from day one and has the vocabulary check in place before the first draft.

A clean finish is not proof that the data is right. The repeated employment table never raised anything, so a job that ended without errors said nothing about that table, and the fix turned out to be a small change in how files are named.

The checks have a ceiling too. They can show that a number is real and that a sentence avoids certain patterns, but they can't judge whether a conclusion is fair to the reader, and no rule found so far does that.

Current state

The monthly job works end to end when started by hand. A scheduled task that starts it each month is prepared but not switched on, and the drafts wait for human approval either way. The open work is on the data side: the missing access keys and the manual downloads, so more of the 329 sources go live.

Ask AI about this AI Lab note

Opens the assistant in a new tab with this page as the source.

Keep reading

Want this handled for your business?

Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.