On the same July week that Substack handed its 50 million-plus readers a button to scan writing for signs of a machine, a small research lab published the fine print that button can't show: the scanners miss a meaningful slice of AI text, and they miss the most when the machine is told to write like a real person.
Substack rolled out its AI-detection feature on July 22, partnering with the startup Pangram to let anyone check a post, Note, reply, or comment longer than 100 words. The tool returns not a yes-or-no ruling but an estimate — a probability of how much of the text appears to be AI-generated. The company was careful to frame it that way. Detection tools, Substack acknowledged in its own announcement, "do not guarantee perfect accuracy," and the platform is not auto-enforcing anything based on a scan. Readers can request a check; creators can run Pangram on their own drafts before publishing, add an optional "How I make this" statement disclosing their process, and even flag or remove scan results they believe are wrong. The feature launched on web and iOS, with Android promised later.
The timing was almost too neat. One week earlier, on July 15, the nonprofit research group Epoch AI published a study that put Pangram and two competitors — GPTZero and Originality.ai — through exactly the kind of stress test the Substack rollout invites. The results are a caution label for the whole enterprise.
The one-in-five problem
Epoch's researcher Jaeho Lee assembled 495 human-written passages of roughly 500 words each, drawn from 99 authors spanning blogging, fiction, and scientific writing, all published before ChatGPT existed. Then, for each author, the team fed a frontier model — Claude Opus 4.8, GPT-5.5, or Gemini 3.1 Pro — five real samples of that author's work and asked it to write something new in the same voice.
Under ordinary conditions, the detectors were nearly flawless. On AI text generated from a plain few-word prompt, false-negative rates were at most 0.7 percent. On genuine human writing, Pangram and GPTZero flagged nothing at all; only Originality.ai stumbled, misclassifying 19 of 495 human passages — 3.8 percent — as machine-made, every one of them writing that predated ChatGPT.
But when the models imitated a specific author's style, the floor fell out. Across the three tools, false-negative rates ran from 10 percent (Pangram) to 11 percent (GPTZero) to 18 percent (Originality.ai) — an average of about 13 percent, or roughly one style-imitated passage in eight slipping through. Epoch summarized it plainly: the detectors "missed roughly one in five passages imitating a specific author's style."
The vulnerability concentrated in the genre where these tools are arguably deployed most: scientific writing. There, the detectors failed on about 26 percent of imitated passages on average. The single worst case was Gemini-generated scientific text, which Pangram missed 48 percent of the time and GPTZero 36 percent. Fiction, by contrast, was easy to catch — Pangram missed just one imitated fiction passage out of 99.
"When we gave models five samples of a specific author's work and asked them to mimic it, an average of 38 of 297 of the resulting passages went undetected," Lee wrote, noting the detectors "performed particularly poorly on mimicked scientific writing." Epoch pinned the versions it tested — Pangram 3.3.2, GPTZero's May 2026 base model, Originality's Turbo 3.0.2 — precisely because, as the report notes, detectors "can be updated by their makers without notice."
Why it matters
The asymmetry is the whole story. Prompting a model to imitate an author takes one line of instruction and a few reference paragraphs. Detecting that imitation is a hard statistical problem that gets harder every time the models improve. This is an arms race in which the attacker's cost is trivial and the defender's cost is enormous — and the defender is always a step behind, testing last month's model against this month's evasion.
That structural imbalance is exactly why a detection score should never be treated as evidence for consequences. A number that misses one in five disguised passages — and nearly half of some scientific text — cannot bear the weight of a failing grade, a rejected pitch, a fired writer, or a public accusation. Substack, to its credit, built its feature around this reality: it calls the output an estimate, lets creators contest results, and declines to enforce automatically. That posture should be the template, not the exception.
Institutions leaning on detection as a disciplinary tripwire are building on sand. Universities that run student essays through these classifiers, journals screening submissions, editors vetting freelancers — Epoch's data shows their tools are weakest precisely against a motivated user who samples a target style, and weakest of all on the academic prose they most need to police. The better adaptations are structural: process-based verification like drafts, version history, and oral defense; disclosure norms of the "How I make this" variety; and a hard institutional rule that a probability is a prompt for a conversation, never a substitute for one. Detection can flag; it cannot convict.
What to watch
Watch whether Substack's estimate-not-verdict framing survives contact with its users, or whether readers start treating a Pangram percentage as a verdict the company explicitly says it isn't — and whether creators begin gaming the "How I make this" disclosure. Watch the version numbers: Epoch tested June 2026 detectors against mid-2026 models, and both sides will move. If the next study runs the same test against the models released this fall and the miss rate climbs, the case for using detection as evidence collapses further. And watch the institutions — schools, journals, employers — that adopt these scanners anyway. The question is no longer whether detectors work in the lab. It is whether the people using them will accept that a tool built to estimate can be trusted to accuse. The evidence says it can't.
“When we gave models five samples of a specific author's work and asked them to mimic it, an average of 38 of 297 of the resulting passages went undetected.”— Jaeho Lee, Researcher, Epoch AI