How we test our output against detectors: our method, and what we'll publish
By Vitalik Hakim · August 12, 2026 · 5 min read
A humanizer that won't show its homework isn't worth trusting. Here is exactly how we evaluate our own output, and the rule we hold ourselves to when we report it.
The method
We keep a fixed evaluation set: samples across genres (essays, articles, emails, product copy) and sources (several AI models, plus human controls). On a schedule we run each sample through multiple detectors, humanize it, and run it again, recording the before/after estimated-AI score for each detector and each mode.
Two things matter about that design. First, human controls: if our own 'human' samples get flagged, the detector is noisy and any 'improvement' is suspect. Second, dates and versions: detectors change without notice, so a result is only meaningful attached to the day and the detector version it was measured on.
What we will and won't say
We will publish measured pass rates (the share of the set that moved below a detector's 'likely AI' threshold) with the date, the detectors, and the method. We will update them as detectors change. We will not print a number we haven't measured, and we will never say 'undetectable' or 'guaranteed', because neither is a property any text can have.
When a detector update reverses a result, we'll say so. The point of measuring is to tell the truth about a moving target, not to freeze a flattering snapshot.
Why this is on the roadmap, not the homepage yet
Standing up a rigorous, repeatable eval, with real human controls and versioned detectors, is worth doing properly rather than quickly. Until it's producing numbers we'd stake our name on, the site doesn't show a pass-rate figure. When it does, it'll be dated, sourced, and here on the blog first.