The enquiries we get from schools follow a pattern. Turnitin's AI indicator produced a number, someone acted on the number, a student contested it, and the conversation that followed was uncomfortable enough that the institution started looking for something better.
The instinct is to shop for a more accurate detector. Before doing that, it is worth being precise about what went wrong, because a more accurate detector usually does not fix it.
What the detector actually gave you
A detection score is a statement about text: this document looks statistically more like machine output than human output. It is not a statement about a person, and it does not come with the thing an academic misconduct process needs, which is evidence a panel can examine and a student can answer.
That gap is why the hearings go badly. The student says they wrote it. The instructor has a percentage. Neither of those is checkable by the other party, and the process has nowhere to go except who seems more credible, which is exactly the situation academic due process exists to avoid.
Swapping detectors changes the number. It does not change the fact that a number is not evidence.
The realistic options
Another detector. Cheapest to adopt, no workflow change. Same failure mode, and the published false-positive behaviour on non-native English writing is a live liability. If your student body is international, this is the option with the legal and reputational exposure.
Process evidence. Version history, draft capture, keystroke logs. Considerably more persuasive than a score, genuinely useful in a hearing, and it has two costs people underestimate: it is surveillance, which students feel, and someone has to read it.
Supervised writing. In-class essays, blue books, invigilated exams. Reliable and, for many courses, a real regression in what you can assess. You get to see them write, on a narrower task, under time pressure that measures something other than understanding.
Verification by conversation. Ask the student about their own submission and evaluate whether they can account for it. This produces an artefact a panel can review, does not require trusting a probability, and applies equally to everyone rather than only to whoever tripped a threshold.
We build the fourth thing, so weigh that accordingly. The argument for it does not depend on our product though: an oral check is what medical, doctoral and licensing assessment already use when being wrong is expensive.
Questions worth asking any vendor
- What exactly does your output claim, and would you be comfortable defending that claim at a disciplinary hearing?
- What is your false-positive behaviour on writing by non-native English speakers, and can I see the evidence rather than a marketing sentence?
- Does the tool make a decision, or does it hand a person the material to decide with?
- What does the student see and what can they contest?
- If the underlying model changes next quarter, what happens to results produced this term?
That last one catches more vendors than you would expect, and it matters: assessment records outlive the software that produced them.
If you want our answers to those five, they are on the Viva standard, written down in public so they can be held against us.