Local AI
Three AI text detectors, measured on human and machine copy
Three open AI text detectors run locally, checked against known human and machine writing, then pushed down with rewrites. What the numbers tracked, where they misled, and why a human reader still has the last word.
SummaryUnder a minute
The short version
Three open detectors looked near perfect on 50 human and 50 GPT-4 texts, but one of them flagged real human writers, including most of a set of 2016 marketing pieces. Rewrites lowered the scores, yet a human reader still called the passing drafts too AI, so detector numbers alone can't decide whether copy reads as human.
Key takeaways
- On a calibration set of 50 human texts and 50 GPT-4 texts, the three detectors scored 0.983, 0.980 and 0.992 AUROC.
- The detector that looked most accurate still flagged 11 of 50 human texts, and it called 5 of 7 human-written marketing pieces from 2016 AI.
- It was reacting to corporate brochure language: an electrician's text with a start date, a place and the name of the person who started the business scored only 6%.
- Rewriting sentence by sentence left scores at 100%, while best-of-8 rewrites at whole-text or paragraph level worked.
- Drafts that passed all 3 detectors were still judged "too AI" by a human reader, so a reader-side check now sits next to the detector scores.

The question
The goal was a way to measure how machine-like a piece of English marketing copy sounds, and then to lower that number with a fully local setup, so every check is free and no draft leaves the machine. An AI text detector is the obvious tool, but it was unclear whether a low score means a reader would actually take the copy as human.
That matters in both directions. If these tools flag honest human writing, a low score becomes a target that means nothing, and a high score can scare people away from copy that was fine.
The setup
The setup went in on September 25, 2026, with three open detectors running on a single workstation GPU. Binoculars and Fast-DetectGPT measure how predictable each word is to a pair of open language models, a base model and its instruction-tuned sibling: Falcon 7B for one and Llama 3 8B for the other, both loaded in bf16 half precision. The third tool, desklib, is a trained classifier that reads the text and returns a probability that a machine wrote it. The model pairs load one after another, scoring a single text takes about 50 seconds, and nothing stays in graphics memory once a job finishes. The same GPU also handles image generation, so in practice a text check waits until queued image jobs are done, and checks fit best between image batches. Beyond the scores, each check produces a map of the text with every sentence colored by how machine-like it looks, which turned out to be the most useful view.
A rewriting loop on a larger local model sits on top of the scores. Each pass goes back to the source text and produces 8 new versions at a sampling temperature of 1.05, handling anything under 3000 characters in one go and splitting longer pieces by paragraph. Any candidate that breaks the writing rules or introduces numbers and names that weren't in the original is discarded. A small embedding model then checks that the meaning stayed close, with a floor of 0.80 similarity. Survivors are scored by all three detectors and the lowest score wins, with the target of keeping all three under 35%.

Measurement
The first step was to see whether the detectors can tell people from machines at all. A calibration set of 50 human texts and 50 GPT-4 texts came from a public research dataset, and performance was measured with AUROC, a number between 0 and 1 that shows how well the detector ranks machine text above human text across every possible cutoff, where a perfect detector scores 1.
The second step used marketing copy written by people in 2016, before anyone drafted with chatbots. Only desklib was run on this part, so the numbers that follow apply to that one tool. Machine drafts were also rewritten in several styles to see which changes brought the scores down.
These checks have limits. A set of 50 plus 50 texts is small, and its topics and styles look little like small business copy. The open detectors haven't been compared with commercial checkers, so nothing here says how those tools would judge the same text.
Results
The calibration looked excellent
Binoculars scored 0.983, Fast-DetectGPT 0.980 and desklib 0.992, which made desklib look like a solid tool to rely on.
- Binoculars1.0 AUROC
- Fast-DetectGPT1.0 AUROC
- desklib1.0 AUROC
Source: Local detector measurement
desklib flagged human writers
The texts told a different story. In the same set of 50 human pieces, desklib called 11 of them AI while its overall accuracy score stayed high. That is possible because the metric tracks how the detector ranks texts against each other, and it ignores whether a specific cutoff lands on a real person. Those errors weren't tracked for the other two tools, so the data here covers desklib only.
The marketing test made the pattern clearer. Out of 7 human-written marketing texts from 2016, desklib called 5 of them AI. A house cleaning company's brochure, written by a person years before chatbot drafting existed, scored a flat 100%. Read side by side, the detector seemed to be reacting to corporate brochure language. The one human text that stayed low was an electrician's brochure that said when the business started, where it worked and who started it, and it scored only 6%. That is a single text, so it is an observation worth testing further.

Sentence-by-sentence rewriting went nowhere
The first attempt at lowering scores on machine drafts rewrote one sentence at a time, and desklib stayed at 100%. The likely reason is that the detector judges the texture of a whole passage, so polishing single sentences leaves the pattern in place. Rewriting the whole text at once, or a paragraph at a time, and keeping the best of 8 candidates did bring scores down.
The voice moved the numbers
One machine-written marketing piece was rewritten in two voices. In a corporate voice, Binoculars gave it 18 and Fast-DetectGPT 20, while desklib scored it 94, which missed the target, and that version also came out truncated. Written as a first-person account by the person running the business, the scores dropped to 0, 19 and 32, all under 35%. A first-person voice only works honestly when the person it speaks for actually stands behind the words, so it can't be applied by default.
A human reader still called the passing drafts too AI
The most important result came next: drafts that passed all 3 detectors were still judged "too AI" by a human reader. The detectors measure how predictable the words are, while a person reading the copy notices rhetorical habits, such as a setup built on contrast, two sentences that mirror each other, a string of short punchy lines or filler that says nothing, and very little of that shows up in word predictability.
Takeaways
The 0.992 score looked perfect and hid the fact that the same detector called 11 of 50 human samples AI. Before running these checks on real work, test them on older human writing from the same trade, since that is where the false flags show up.
At least one of these detectors behaves less like an AI checker and more like a cliche meter. That makes it a poor tool for judging a writer, but a useful one for editing, since generic brochure language needs trimming whoever wrote it. The strongest lever is raw material: a specific date, a place, a name, or the subject's own words.
Rewriting whole paragraphs or the entire text, with several candidates to choose from, works far better than polishing one sentence at a time.
A human reader belongs in the loop from the start, before trusting three green scores, and a comparison with commercial checkers is still open. The rewriting loop has a gap too: its checks catch invented numbers and names but miss soft new claims, like "the phone rings all day", when the meaning drifts slightly, so a person still has to read every draft.
Recommended setup
Run a draft through the three detectors, then through a reader-side check for the patterns a person notices, like mirrored sentences or a string of short lines, and drop any rewrite candidate that fails it. Neither step can tell whether the text contains a real observation or an unexpected detail, so that judgment stays with a human editor.

Ask AI about this AI Lab note
Opens the assistant in a new tab with this page as the source.
Keep reading
Want this handled for your business?
Start with the free site audit: speed, search, mobile, security, local presence and email, in plain English.


