AI detectors feel like magic — paste text, get a verdict. Underneath, they're statistics engines measuring how predictable your writing is. Once you understand the three measurements, both the false positives and the evasions make sense.
Perplexity: the predictability score
A language model can calculate how "surprised" it is by each next word. "The cat sat on the ___" → "mat" is unsurprising (low perplexity). AI text is generated by picking probable words, so it scores consistently low perplexity. Human writing is messier — odd word choices, idioms, small tangents — and that unpredictability reads as human.
Burstiness: the rhythm test
Humans vary wildly: a 40-word sentence, then a 5-word one. AI settles into a monotone of 15–25 word sentences, every paragraph the same shape. Detectors compute the variance of sentence lengths — low variance is the single most reliable AI tell.
Vocabulary fingerprints
Models overuse a known word set — "delve", "tapestry", "furthermore", "seamlessly", "in today's fast-paced world". Detectors keep lists. Three "moreover"s on one page is practically a confession.
Why detectors get it wrong
Non-native speakers and technical writers naturally produce low-perplexity, low-burstiness text — which is why false accusations happen constantly. Detector verdicts are probabilities, not proof.
What this means for your content
To pass, text needs the statistical texture of human writing: varied sentence lengths, diverse vocabulary, zero AI-tell phrases. You can edit that in by hand — or use the free PlugNest AI Humanizer, which restructures all three signals in one click and includes its own detector so you can verify the result. For the practical workflow, see our step-by-step humanizing guide.