Local AI

Moving AI text closer to a human voice, measured

A local rewrite loop that scores English drafts on 3 open AI-text detectors and a reader-side check, while holding numbers, names and meaning in place. What moved the scores, what didn't, and where the measuring stick falls short.

Pixedi AI Lab
  • 5 min read

SummaryUnder a minute

The short version

The test: can a local rewrite loop move AI-written English toward a human voice without changing the facts? On one text, a first-person rewrite met the under 35% target on 3 detectors, while a company-voice version, cut short, failed on one. Detectors turned out to be a partial yardstick, so a reader-side check was added and a person still reads everything.

Key takeaways

  1. Each round, a local 27B model writes 8 rewrites from the original text, and only candidates that keep every number and name and stay at or above 0.80 meaning similarity are scored.
  2. On one marketing text, the first-person version scored 0, 19 and 32 on 3 detectors, under the 35% target, while a company-voice version that was cut short failed at 94 on one detector.
  3. The strictest detector also flagged 5 of 7 human-written marketing pages from 2016, so a high score can point to brochure language as much as to a model.
  4. Drafts that pass every detector can still read as machine-written, so a reader-side check for contrast framing, mirrored sentences and strings of short sentences now sits next to the detectors.
  5. Real numbers, places, names and sentences in the author's own words remain the strongest lever, and a person still reads every draft before it's published.
A desktop computer tower beside a desk with marked-up printed pages and a coffee cup in afternoon light.

The question

A first draft from a language model often slides into brochure language, and a reader notices it and stops reading. The question here is how far a local setup can move an AI-written English draft toward something that reads like a person wrote it, with a number for "how far" so the result isn't judged by feel.

The second half of the question matters just as much. A rewrite that sounds warmer but quietly changes a price or bends the meaning of a sentence is worse than the stiff draft it started from. So any method worth keeping has to hold the facts still while it changes the voice.

The setup

Everything runs on one workstation with a single high-end consumer GPU, and nothing leaves the machine. The same GPU also handles image generation, so before a rewrite session the image queue has to be empty, and when the work finishes the graphics memory is released again. The rewriter is a local 27B model, large enough to write decent English and small enough to run on consumer hardware. Scoring one page and drawing a sentence-by-sentence heat map of where the AI signal sits takes about 50 seconds.

The work happens in rounds, where a round is one attempt at a better version. In each round the model takes the original text and writes 8 different rewrites of it at temperature 1.05, a setting that makes it a bit more willing to pick unusual words. Every round starts again from the original, so a bad choice in one candidate doesn't carry into the next. Texts under 3000 characters are rewritten in one piece, and longer ones paragraph by paragraph.

The inside of a desktop computer tower showing a large GPU, sitting on a wooden floor beside a desk.

The first approach fixed the draft one sentence at a time, going after whichever sentence scored worst, and on the test text the strictest detector stayed at 100%. Rewriting the whole text, or one paragraph at a time for longer pieces, and then choosing the best of the 8 is what started to work.

Each candidate has to pass a few filters before it can win. The first is a fabrication check, which compares the numbers and proper nouns in the rewrite with the ones in the original. A candidate is thrown out if it adds a number or a name that wasn't there, and also if it drops one that was. Then meaning is compared: a small model scores how close the rewrite is to the original, and anything under 0.80 similarity is dropped. Simpler filters remove dashes, lists of three, exclamation marks and question marks.

The rewriter also takes a voice prompt. The default is a company voice prompt that describes the business from the outside, and the variant tested here is a first-person voice prompt, written the way someone explains their own job to a neighbor over the fence. Choosing between them is an editorial decision, because a first-person rewrite speaks on someone's behalf.

Measurement

The yardstick is 3 open AI-text detectors hosted on the same machine. Two of them look at how predictable each next word is to a language model, and the third is a classifier trained to label text as human or AI. Among the candidates that survive the filters, the one with the lowest scores wins, and the target is under 35% on all 3.

Before trusting them, all 3 were run on a public set of 50 human-written and 50 machine-written texts, and they separated the two groups well. The limits are real, though. How these scores line up with the commercial detectors people actually use hasn't been measured, so there is no claim that a text will pass every detector out there. The classifier also called 11 of the 50 human texts AI, and on a 2016 batch of human-written marketing copy from home service companies it flagged 5 of the 7 pages, one of them at 100%. The page it scored lowest, at 6%, was plainly written and anchored in the company's own history, with a founding date and a named person right in the copy. So a high score can mean "sounds like a brochure" as easily as "a model wrote this."

The bigger limit showed up later. Drafts that cleared all 3 detectors still read as machine-written when a person sat down with them. Detectors mostly measure how predictable the word choices are, and people notice patterns: the setup that flips into a reversal, two sentences that mirror each other, a dramatic string of short lines, and lines that sound wise but say nothing.

A person at a kitchen table reading a printed page with a pen in hand.

That led to a reader-side check, a second set of tests that looks for those patterns directly. It flags contrast framing, mirrored sentences, strings of short sentences and filler lines that carry no observation, and a candidate that fails it is dropped even if the detectors liked it.

Results

On one AI-written marketing text, the first-person version came in at 0, 19 and 32 on the 3 detectors, which put it under the 35% target on each. The company-voice version of the same text came in at 18 and 20 on the first two and failed at 94 on the third.

Detector scores for one text, first-person and company voice%
  • First-person voice, detector A0%
  • First-person voice, detector B19%
  • First-person voice, detector C32%
  • Company voice, detector A18%
  • Company voice, detector B20%
  • Company voice, detector C94%

Source: Local detector test

That result needs care. It's one text, and the company-voice output was cut short, which can move a score on its own, so the gap can't be put down to the voice alone. What it does show is that on this text the first-person version met the target and the company-voice version didn't, and the first-person prompt gave the rewriter room to be specific and a little informal.

The fabrication check is built to block new numbers and names and to keep the ones the original had. What it can't catch is a soft shift in meaning, like a rewrite claiming a service "is always available" when the original said nothing of the kind. Nothing in that sentence counts as a new number or name, so it would sail through, which is why every output still gets read by a person before it's published.

Takeaways

Raw material does most of the work

The strongest lever is still raw material: real numbers, places, names and sentences in the author's own words. The human page that scored lowest in the test, at 6%, was the one carrying its founding date and a named person, and a rewriter can't produce a detail like that if it was never given one. It is deliberately blocked from trying.

An open notebook with handwritten notes on a workbench next to a tape measure and a pencil.

Score the text two ways

Passing a detector and reading like a person turned out to be two separate tests, because the detectors look at word predictability and readers react to sentence patterns. Both reports now sit side by side for every draft.

Rewrite whole passages

On the test text, repairing one sentence at a time left the strictest detector at 100%. Generating several versions from the original and choosing among them is the approach that stuck.

What to do differently

Build the reader-side check on day one, before trusting any detector score, and calibrate against commercial detectors much earlier. That comparison is still open.

Current use

Every English draft gets a detector report and a reader-side report next to each other, and then a person reads it before anything is published. A draft that scores high goes through the rewrite loop, and the first-person voice prompt is used only where a first-person text is appropriate.

Ask AI about this AI Lab note

Opens the assistant in a new tab with this page as the source.

Keep reading

Want this handled for your business?

Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.