Research · August 4, 2026

Can AI give every student a viva?

Viva Team · 7 min read

연구 · 2026년 8월 4일

AI는 모든 학생에게 구술시험을 줄 수 있을까요?

Viva 팀 · 7분 분량

This series has made three claims so far. The essay no longer reliably evidences the thinking behind it. Detection cannot repair that, and answers the wrong question anyway. And the oral exam, the format that evidences thinking directly, left mass education over cost rather than validity. Which brings us to the live question: conversational AI can now hold a structured spoken conversation for a fraction of what an examiner's hour costs. Does that reopen the door?

The honest answer, and the one we hold ourselves to, is: plausibly, and it has to be tested. It is worth being precise about why it is plausible. The benefits of oral assessment do not come from the examiner's presence as such. They come from two mechanisms. Retrieval: the student has to pull reasoning out of their own head, in real time, with no window to paste into. Dialogue: follow-up questions find the edges of understanding and force tacit reasoning into the open. An AI interviewer preserves the first mechanism almost by construction. The second it preserves partially: genuinely contingent follow-ups, generated from what the student just said, are not a script. Whether that partial preservation captures enough of what a human examiner does is exactly the open question.

We want to be equally precise about what we are not claiming. An AI interlocutor is not an expert human examiner. It does not bring a discipline's judgment or a mentor's rapport, and nothing we build treats it as if it did. The realistic comparison is not AI versus the expert viva that mass education has never been able to afford. It is AI-assisted oral evidence versus the status quo: an unsupervised written artifact that, in the GenAI era, may evidence very little at all.

That is also why the right architecture is a triage layer, not a judge. The interview generates evidence: a transcript, the student's own spoken explanation, the places where their account of the work diverged from the work itself. A teacher reads that evidence and makes every call that matters. The AI never adjudicates; it surfaces. Done this way, an AI error becomes a flag a human reviews rather than a grade a student is stuck with, and scarce teacher attention lands where it is actually needed.

If the hypothesis fails, we would rather find out first.

The risks deserve naming, because they are real. Spoken formats raise anxiety for some students, and accent robustness and accommodations are design requirements, not afterthoughts; a system that swapped detection's bias for a new one would be no progress at all. Some learning is not well evidenced orally, and a short conversation should sit alongside written work, not replace it. And any assessment format invites gaming, so identity and administration conditions have to be engineered deliberately. None of these strikes us as disqualifying. Each is a design constraint we build against, and an empirical question someone should hold us to.

Which is why the last thing to say is about evidence. Our team has written up the full argument in a working paper, and a first classroom pilot is underway now, comparing AI-mediated interview outcomes against teachers' independent judgments of the same students' understanding, with fairness across language backgrounds examined explicitly. Results will be reported in a separate empirical paper, whichever direction they point. If the hypothesis fails, we would rather find out first. If it holds, then the oldest exam in the world stops being a luxury good, and every student gets asked the one question that has always mattered: can you walk me through your thinking?

Read about our research program →

이 시리즈는 지금까지 세 가지 주장을 했어요. 에세이는 더 이상 그 뒤의 사고를 신뢰성 있게 증명하지 못해요. 탐지는 그걸 고치지 못하고, 애초에 잘못된 질문에 답하고 있어요. 그리고 사고를 직접 증명하는 형식인 구술시험은, 타당성이 아니라 비용 때문에 대중 교육을 떠났어요. 그래서 지금 살아 있는 질문에 도착해요. 대화형 AI는 이제 시험관 한 시간 비용의 극히 일부로 구조화된 음성 대화를 진행할 수 있어요. 그 문이 다시 열리는 걸까요?

정직한 답, 그리고 저희가 스스로에게 요구하는 답은 이거예요. 그럴듯하다, 그리고 반드시 검증되어야 한다. 왜 그럴듯한지는 정확히 말할 가치가 있어요. 구술 평가의 효과는 시험관이 그 자리에 있다는 사실 자체에서 나오는 게 아니에요. 두 가지 메커니즘에서 나와요. 인출: 학생은 자기 머릿속에서 실시간으로 추론을 꺼내야 하고, 붙여넣을 창이 없어요. 대화: 후속 질문이 이해의 경계를 찾아내고, 암묵적 추론을 밖으로 끌어내요. AI 인터뷰어는 첫 번째 메커니즘을 구조적으로 거의 그대로 보존해요. 두 번째는 부분적으로 보존해요. 학생이 방금 한 말에서 만들어지는 진짜 맥락형 후속 질문은 대본이 아니니까요. 그 부분적 보존이 인간 시험관이 하는 일의 충분한 몫을 담아내는지, 그게 바로 열려 있는 질문이에요.

저희가 주장하지 않는 것도 똑같이 정확히 말할게요. AI 대화 상대는 전문가 인간 시험관이 아니에요. 학문 분야의 판단력도, 멘토의 교감도 갖고 있지 않고, 저희가 만드는 어떤 것도 그런 척하지 않아요. 현실적인 비교 대상은 대중 교육이 한 번도 감당해 본 적 없는 전문가 구술시험 대 AI가 아니에요. AI가 도운 구술 증거 대 현재 상태, 그러니까 생성형 AI 시대에 거의 아무것도 증명하지 못할 수 있는 무감독 과제물이에요.

그래서 올바른 구조는 판정자가 아니라 분류 레이어예요. 인터뷰는 증거를 만들어요. 대화 기록, 학생이 직접 말한 설명, 그리고 학생의 설명이 제출물과 어긋난 지점들이요. 교사가 그 증거를 읽고 중요한 판단을 전부 내려요. AI는 결코 판정하지 않아요. 드러낼 뿐이에요. 이렇게 하면 AI의 오류는 학생에게 박히는 성적이 아니라 사람이 검토하는 플래그가 되고, 부족한 교사의 시간은 정말 필요한 곳에 쓰여요.

가설이 틀렸다면, 저희가 먼저 알아내는 편이 나아요.

위험도 이름을 붙여 둘 가치가 있어요. 실재하니까요. 말로 하는 형식은 일부 학생에게 불안을 키우고, 억양에 대한 견고함과 배려는 나중에 붙이는 게 아니라 설계 요구사항이에요. 탐지의 편향을 새로운 편향으로 바꾸는 시스템이라면 전혀 진보가 아니에요. 어떤 학습은 말로 잘 증명되지 않으니, 짧은 대화는 글로 된 결과물을 대체하는 게 아니라 그 옆에 놓여야 해요. 그리고 모든 평가 형식은 우회를 부르니, 신원 확인과 시험 환경은 의도적으로 설계되어야 해요. 이 중 무엇도 결격 사유는 아니라고 봐요. 각각은 저희가 맞춰 설계해야 할 제약이고, 누군가 저희에게 책임을 물어야 할 실증적 질문이에요.

그래서 마지막으로 할 이야기는 증거에 대한 거예요. 저희 팀은 이 논증 전체를 워킹 페이퍼로 정리했고, 첫 교실 파일럿이 지금 진행 중이에요. AI가 진행한 인터뷰 결과를, 같은 학생의 이해에 대한 교사의 독립적 판단과 비교하고, 언어 배경에 따른 공정성도 명시적으로 살펴보고 있어요. 결과는 어느 방향을 가리키든 별도의 실증 논문으로 보고할 거예요. 가설이 틀렸다면, 저희가 먼저 알아내는 편이 나아요. 가설이 맞다면, 세상에서 가장 오래된 시험은 더 이상 사치재가 아니게 되고, 모든 학생이 언제나 가장 중요했던 그 한 가지 질문을 받게 돼요. 네 생각을 나에게 설명해 줄래?

저희 연구 프로그램 살펴보기 →

Pilot this with your class.

We are onboarding pilot partners for the fall term. Setup takes a class period, not a semester.

이번 학기, 우리 반에서 파일럿해요.

가을 학기 파일럿 파트너를 모집하고 있어요. 셋업은 한 학기가 아니라 수업 한 시간이면 충분해요.