This series has made three claims so far. The essay no longer reliably evidences the thinking behind it. Detection cannot repair that, and answers the wrong question anyway. And the oral exam, the format that evidences thinking directly, left mass education over cost rather than validity. Which brings us to the live question: conversational AI can now hold a structured spoken conversation for a fraction of what an examiner's hour costs. Does that reopen the door?
The honest answer, and the one we hold ourselves to, is: plausibly, and it has to be tested. It is worth being precise about why it is plausible. The benefits of oral assessment do not come from the examiner's presence as such. They come from two mechanisms. Retrieval: the student has to pull reasoning out of their own head, in real time, with no window to paste into. Dialogue: follow-up questions find the edges of understanding and force tacit reasoning into the open. An AI interviewer preserves the first mechanism almost by construction. The second it preserves partially: genuinely contingent follow-ups, generated from what the student just said, are not a script. Whether that partial preservation captures enough of what a human examiner does is exactly the open question.
We want to be equally precise about what we are not claiming. An AI interlocutor is not an expert human examiner. It does not bring a discipline's judgment or a mentor's rapport, and nothing we build treats it as if it did. The realistic comparison is not AI versus the expert viva that mass education has never been able to afford. It is AI-assisted oral evidence versus the status quo: an unsupervised written artifact that, in the GenAI era, may evidence very little at all.
That is also why the right architecture is a triage layer, not a judge. The interview generates evidence: a transcript, the student's own spoken explanation, the places where their account of the work diverged from the work itself. A teacher reads that evidence and makes every call that matters. The AI never adjudicates; it surfaces. Done this way, an AI error becomes a flag a human reviews rather than a grade a student is stuck with, and scarce teacher attention lands where it is actually needed.
If the hypothesis fails, we would rather find out first.
The risks deserve naming, because they are real. Spoken formats raise anxiety for some students, and accent robustness and accommodations are design requirements, not afterthoughts; a system that swapped detection's bias for a new one would be no progress at all. Some learning is not well evidenced orally, and a short conversation should sit alongside written work, not replace it. And any assessment format invites gaming, so identity and administration conditions have to be engineered deliberately. None of these strikes us as disqualifying. Each is a design constraint we build against, and an empirical question someone should hold us to.
Which is why the last thing to say is about evidence. Our team has written up the full argument in a working paper, and a first classroom pilot is underway now, comparing AI-mediated interview outcomes against teachers' independent judgments of the same students' understanding, with fairness across language backgrounds examined explicitly. Results will be reported in a separate empirical paper, whichever direction they point. If the hypothesis fails, we would rather find out first. If it holds, then the oldest exam in the world stops being a luxury good, and every student gets asked the one question that has always mattered: can you walk me through your thinking?