Every semester this question turns up in student forums, usually at two in the morning: can they tell? The answers you get are mostly wrong in both directions. One camp insists detectors are magic. The other insists nothing can be proven. Neither is how it actually goes.
We build software in this space, so we have read a lot of misconduct paperwork. Here is what is really happening on the other side of the desk.
The detector score is rarely the reason
Instructors do run submissions through detectors, and detectors do return numbers. But most experienced instructors have learned not to trust those numbers on their own, and many institutions have quietly told staff not to open a case on a score alone. The reason is simple: the scores are wrong often enough to ruin someone's year. Research out of Stanford in 2023 found that detectors flagged writing by non-native English speakers as AI-generated at strikingly higher rates than writing by native speakers, and that finding has been replicated in spirit many times since. A tool that punishes you for having learned English second is not a tool you can build a hearing on.
So when a case does move forward, the score is almost never the evidence. It is the thing that made someone look more closely.
What actually gives it away
Voice discontinuity. Instructors read a lot of student writing, and they build a model of how each student sounds without meaning to. A paragraph that is fluent, balanced, and slightly airless sitting between two paragraphs that are messier and more specific reads like a splice. You do not need software to hear it. You need to have read the last three assignments.
Citations that do not exist, or exist and say something else. This one is brutal and common. A model produces a plausible reference, the reference is checked, and it turns out the paper is real but argues the opposite, or is not real at all. Nothing else in academic work fails this cleanly.
Specificity that stops at the door. The work describes the method beautifully and cannot say why that method. It names the dataset and does not know how many rows it has. Generic competence with no local detail is a strong signal because your own work is full of local detail whether you notice it or not.
The tell is not that the writing is too good. It is that the writing is unattached to a person who made choices.
Process evidence. Version history, drafts, a document created eleven minutes before the deadline with 2,400 words pasted in at once. Instructors increasingly ask for this, and it is far more persuasive in a hearing than any percentage.
And the oldest one: the conversation. A teacher asks you about your own paper, and either you can talk about it or you cannot. This is why oral checks are coming back, and it is the part of the answer students most consistently underestimate.
What this means if you are the student
The honest advice is not what you probably expect, because it is not really about tools.
- If you used a model to help you think and you understand what you handed in, you are almost certainly fine, and you should be able to say so plainly when asked.
- If you cannot explain a paragraph in your own submission without rereading it, that is the risk, and no detector needs to be involved for it to surface.
- Check every citation yourself, every time. It is the single highest-yield thing on this list.
- Keep your drafts. Version history has exonerated more students than it has convicted.
And if you are the instructor
The uncomfortable part of all this is that the reliable signals are labour-intensive and the cheap signal is unreliable. Reading version history for 120 students is not a plan. Neither is trusting a number that misfires on your international students.
Which is why the interesting question stopped being how to detect AI text and became how to verify understanding at a scale a real course can carry. Those are different questions, and only the second one has a good answer.
If you want the longer argument for that, we wrote it here: AI detection was never going to save the essay.