When schools first felt the ground move under written assessment, the instinctive response was a tool: software that reads a text and estimates whether a machine wrote it. It is easy to see the appeal. Detection promises to restore the old world with one click, no redesign required. We spent our last post on why that old world is gone. This one is about why the click does not work.
Start with reliability. Detector accuracy is inconsistent across text types, and the scores can be defeated by paraphrasing tools or a quick "humanizing" pass. A safeguard that fails after one trivial extra step is not a safeguard for anything high-stakes. The clearest verdict came from the least suspect source: OpenAI shipped its own AI-text classifier and then retired it, citing low accuracy. The company with the most information about how its models write concluded it could not reliably recognize that writing.
Then there is the question of who gets flagged. Reviews of this area keep raising the same concern: prose by non-native English speakers is flagged disproportionately, because some of the statistical features detectors read as machine-like also characterize careful second-language writing. Think about what that means in a classroom. The students most likely to be accused are the ones writing in their second language, and a probability score gives them nothing to appeal. There is no evidence behind the number, only the number.
There is no evidence behind the number, only the number.
But suppose the technical problems were solved. Imagine a detector that is never wrong and never biased. It would still be answering the wrong question. A detector asks whether a machine touched the text. A teacher needs to know whether the student understands the work. Those are different questions, and the gap between them is where real learning lives. A student can use AI heavily and understand the result deeply. Another can submit text no machine touched and understand none of it. Detection cannot tell these students apart, and a policy built on it quietly turns the classroom into a contest of evasion, teacher against student, when the whole point was supposed to be learning.
There is a growing consensus in the assessment literature that the productive move is to stop policing the artifact after the fact and start designing tasks where demonstrating understanding is the task. That sounds abstract until you notice that education already has a format that does exactly this. It is older than the essay, older than the university, and it fell out of use for a reason that has nothing to do with how well it works.
That format, and the reason it disappeared from most classrooms, is the subject of the next post.
Next in this series: The oldest exam in the world is coming back →