What generative AI broke in assessment, why detection cannot fix it, and how the oldest exam in the world holds up in a class of two hundred. Written by the team as we work it out, and grounded in the research behind our paper.
생성형 AI가 평가에서 무엇을 부쉈는지, 탐지가 왜 그걸 고칠 수 없는지, 그리고 세상에서 가장 오래된 시험이 200명 수업에서 어떻게 버티는지. 저희가 답을 찾아가며 쓰고 있고, 논문의 연구를 바탕으로 하고 있어요.
Not the detector score. The things that actually give it away are older and more human than that, and most of them have nothing to do with software.
Oral assessment left mass education over cost, not validity. Conversational AI moves that cost. What we claim, what we do not, and how we intend to find out.
The viva never went away where the stakes are highest. It left the classroom for one reason, and that reason was never about how well it measures learning.
Most schools shopping for a different detector are trying to solve an evidence problem with a probability tool. Here is how the options actually compare.
Detectors are unreliable, they flag the wrong students, and even a perfect one would answer the wrong question. The instructor's question is about understanding.
For two centuries a good essay was evidence of good thinking. Generative AI quietly broke that link, and most of education has not caught up with what it means.
False positives are not spread evenly. They land hardest on non-native English writers, on autistic students, and on anyone whose prose is plain. That is a fairness problem, not a tuning problem.
Panels do not weigh evidence the way software vendors assume. What persuades, what collapses, and why the burden sits where it does.
The IB already requires teachers to authenticate student work through discussion. The framework anticipated this problem better than most.
Lockdown browsers and webcam proctoring solve a narrow problem at a high trust cost. There is a cheaper trade available.
Most syllabus AI policies fail for the same three reasons. Here is what to fix, plus wording you can paste and adapt.
Anxiety, accents, stammers, processing differences. The objection is legitimate, and most of it is answerable by design rather than by exemption.
The AI conversation has been dominated by essays. In STEM the problem set went first, and the fix is not a harder problem set.
Peer evaluation forms measure social dynamics as much as contribution. There is a more direct question, and it takes about four minutes per student.
A practical plan for oral assessment at scale: what to check, who to check, how long, and the scheduling mistakes that sink it.
How you introduce it decides how it goes. A practical script for week one, and the three sentences that cause most of the trouble.
Two decades of retrieval practice research says that pulling knowledge out produces learning, not just measures it. Oral checks sit exactly on that mechanism.
Notes from reading a lot of interview transcripts: what separates a question that reveals understanding from one that reveals vocabulary.
We argue the essay stopped certifying learning. That is not the same as arguing it stopped being worth doing, and the difference matters.
The surveys keep saying the same thing about why students hand in generated work, and it is mostly about time, fear and the shape of the assignment.
The marking load was already unsustainable before generative AI. What changed is that the work being marked stopped carrying the same information.
A short, practical definition of the viva voce: where it came from, where it is still required, and what makes one good rather than merely stressful.
탐지기 점수가 아니에요. 실제로 드러나는 신호는 그보다 오래되고 사람다운 것들이고, 대부분 소프트웨어와 무관해요.
구술 평가가 대중 교육을 떠난 건 타당성이 아니라 비용 때문이었어요. 대화형 AI가 그 비용을 움직여요. 저희가 주장하는 것과 주장하지 않는 것, 그리고 검증 계획까지.
판돈이 가장 큰 곳에서 구술시험은 사라진 적이 없어요. 교실을 떠난 이유는 단 하나였고, 그 이유는 측정을 잘하고 못하고와 무관했어요.
다른 탐지기를 찾는 학교 대부분은 근거의 문제를 확률 도구로 풀려고 하고 있어요. 선택지들이 실제로 어떻게 비교되는지 정리했어요.
탐지기는 신뢰하기 어렵고, 엉뚱한 학생을 표시하고, 완벽해진다 해도 잘못된 질문에 답해요. 선생님의 질문은 이해에 대한 거예요.
두 세기 동안 좋은 에세이는 좋은 사고의 증거였어요. 생성형 AI가 그 연결을 조용히 끊었고, 교육 대부분은 아직 그 의미를 따라잡지 못했어요.
오탐은 고르게 분포하지 않아요. 영어가 모국어가 아닌 학생, 자폐 스펙트럼 학생, 문장이 담백한 학생에게 집중돼요. 이건 조정의 문제가 아니라 공정성의 문제예요.
위원회는 소프트웨어 업체가 가정하는 방식으로 근거를 저울질하지 않아요. 무엇이 설득하고 무엇이 무너지는지, 그리고 입증 책임이 왜 그 자리에 있는지.
IB는 이미 교사가 대화를 통해 학생 작업의 진정성을 확인하도록 요구해요. 이 프레임워크는 다른 어떤 것보다 이 문제를 잘 예상했어요.
락다운 브라우저와 웹캠 감독은 좁은 문제를 높은 신뢰 비용으로 풀어요. 더 싼 교환이 있어요.
강의계획서의 AI 정책 대부분이 같은 세 가지 이유로 실패해요. 무엇을 고쳐야 하는지와 그대로 붙여 쓸 수 있는 문구를 정리했어요.
불안, 억양, 말더듬, 정보 처리의 차이. 이 반론은 정당하고, 대부분은 면제가 아니라 설계로 답할 수 있어요.
AI 논의는 에세이가 지배했어요. 이공계에서는 문제 풀이가 먼저 무너졌고, 해법은 더 어려운 문제 풀이가 아니에요.
동료 평가지는 기여만큼이나 인간관계를 재요. 더 직접적인 질문이 있고, 학생당 4분쯤 걸려요.
규모가 큰 수업에서 구술 평가를 운영하는 실용적인 계획이에요. 무엇을 확인하고 누구를 확인하며 얼마나 걸리는지, 그리고 이걸 무너뜨리는 일정 실수까지.
어떻게 소개하느냐가 결과를 결정해요. 1주 차를 위한 실용적인 대본과, 문제 대부분을 일으키는 세 문장이에요.
20년의 인출 연습 연구는 지식을 꺼내는 행위가 학습을 측정할 뿐 아니라 만들어 낸다고 말해요. 구술 확인은 정확히 그 지점에 있어요.
인터뷰 전사를 많이 읽으며 얻은 메모예요. 이해를 드러내는 질문과 어휘를 드러내는 질문은 무엇이 다를까요.
저희는 에세이가 학습을 보증하지 못하게 되었다고 말해요. 그건 에세이가 할 가치를 잃었다는 말과 다르고, 그 차이가 중요해요.
학생들이 왜 생성된 결과물을 제출하는지에 대해 설문은 계속 같은 이야기를 해요. 대부분 시간, 두려움, 그리고 과제의 형태에 대한 이야기예요.
채점 부담은 생성형 AI 이전에도 이미 지속 가능하지 않았어요. 달라진 건, 채점 대상이 예전과 같은 정보를 담지 않게 되었다는 점이에요.
구술시험에 대한 짧고 실용적인 정의예요. 어디에서 왔고, 지금도 어디에서 필수이며, 무엇이 좋은 구술시험을 만드는지 정리했어요.
A real Viva interview and the instructor report it produces. Three minutes, no sign-up.