There is a version of the AI detection debate that is about accuracy, and it is the less interesting version. Accuracy questions have engineering answers: more data, better models, higher thresholds. Wait long enough and someone improves the number.
The version that matters is about distribution. When a detector is wrong, who does it tend to be wrong about?
The distribution is not random
In 2023 a Stanford group ran essays by native and non-native English speakers through seven commercial detectors. The native writing came back mostly clean. The non-native writing was flagged at rates so high that the authors described the tools as unsuitable for use on that population. The mechanism is not mysterious. Detectors lean on measures of lexical variety and sentence-level unpredictability, and a competent writer working in a second language tends to produce exactly what those measures read as machine-like: a narrower vocabulary deployed carefully, and sentences that do not take risks.
Since then a fairly consistent picture has emerged from institutions and independent testers. Plain, structured, careful prose gets flagged. So does writing by students who were taught to a template, which describes a great deal of secondary education. So does writing by some autistic students, whose prose can be precise and low in the kind of stylistic noise the tools treat as human.
Every one of those groups is a group that already has to work harder to be believed. Handing them an extra burden of proof is not a rounding error.
Why raising the threshold does not fix it
The obvious response is to demand more confidence before flagging. This helps, and it trades one failure for another. Push the threshold up and the detector stops catching the students who are actually outsourcing their work, which was the entire point. Push it down and you are back to accusing your international students.
This is not a bug that gets engineered away, because the underlying signal is style, and style is not authorship. A confident, careful, slightly formal paragraph can come from a language model or from a nineteen year old who learned English from textbooks and edits carefully. There is no amount of training data that separates those two, because they are not different in the text. They are different in what happened before the text.
The thing a detector cannot see
Here is the question an instructor actually has: does this student understand the work they handed in? A detector cannot answer that. It was never designed to. It answers a proxy question about the statistical shape of a document, and then the institution treats the proxy as though it were the real one.
The proxy fails in both directions and the failures are asymmetric in consequence. A missed case is a student who got away with something. A false positive is a student in a disciplinary hearing defending work they wrote, often without the vocabulary or the confidence to push back, sometimes on a visa.
What we would do instead
We are not neutral here, so take this as a stated position rather than a finding: the way out is to stop asking documents to prove themselves and start asking students to.
- Ask about the work, not about the writing. A five minute conversation about a paper produces evidence that survives a hearing in a way a percentage never has.
- Make the check universal rather than suspicion-triggered. A check everyone does is not an accusation. A check triggered by a score is.
- Keep the human at the end of it. Software should surface evidence and a person should make the judgment, particularly when a judgment has consequences.
None of this is novel. It is how doctoral defenses, clinical exams and bar qualifications have worked for a very long time. What changed is that the conversation used to be affordable only where the stakes were highest.
Related: AI detection was never going to save the essay and how teachers actually know you used ChatGPT.