AntiGPT

How AI detectors actually work, and why they misfire

By Vitalik Hakim · August 5, 2026 · 6 min read

Every mainstream AI detector is answering one question: does this text look like something a language model would produce? They approach it two ways, and most combine them.

Perplexity: how surprised is the model?

Perplexity measures how well a language model predicts the next word in your text. If the model finds each word highly probable (exactly what it would have written), perplexity is low, and the detector reads that as machine-like. Human writing tends to surprise the model more: an odd word choice, an unusual turn, a metaphor from somewhere the model wouldn't reach. Low perplexity, high suspicion.

The flaw is immediate: plenty of human writing is highly predictable. Formal instructions, legal boilerplate, a five-paragraph essay written to a rubric: all low-perplexity, all human. So perplexity alone flags careful, conventional writers.

Burstiness: how much does the rhythm vary?

Burstiness measures variation: in sentence length, in complexity, in perplexity across the text. Humans are bursty: a long sentence, then three words. Models trained to be helpful and clear tend toward uniformity, every sentence a similar shape. Detectors read low burstiness as machine-like.

This is why the single most effective thing you can do to make writing read as human (to a detector and to a person) is vary your rhythm. It's also why humanizing works on the signal at all: it deliberately reintroduces variation.

Classifiers: trained to spot the style

Newer detectors add a machine-learning classifier trained on large sets of human and AI text. These can be more accurate than perplexity alone, but they inherit their training data's blind spots and go stale as models change. A classifier trained on last year's model output is weaker against this year's.

Why they misfire

Put together, the failure modes are predictable: short texts (not enough signal), formulaic genres (legitimately low perplexity), heavily edited text (mixed signal), and writing by non-native English speakers (a documented pattern across vendors), whose vocabulary is often more predictable. Responsible vendors say plainly that a score is not proof and shouldn't be used alone to accuse anyone.

That's the honest frame for any detector, ours included: it's an estimate of a statistical property, useful as feedback, never a verdict. Which is also why no tool can promise a text will pass one.