A four-part series on what generative AI broke in assessment, why detection cannot fix it, and the very old idea that might. Written by the team, grounded in the research behind our working paper.
생성형 AI가 평가에서 무엇을 부쉈는지, 탐지가 왜 그걸 고칠 수 없는지, 그리고 그 답이 될지도 모르는 아주 오래된 아이디어에 대한 4부작 시리즈예요. 저희 워킹 페이퍼의 연구를 바탕으로 팀이 직접 썼어요.
Oral assessment left mass education over cost, not validity. Conversational AI moves that cost. What we are claiming, what we are not, and how we intend to find out.
The viva never went away where the stakes are highest. It left the classroom for one reason, and that reason was never about how well it measures learning.
Detectors are unreliable, they flag the wrong students, and even a perfect one would answer the wrong question. The instructor's question is about understanding, not authorship of text.
For two centuries a good essay was evidence of good thinking. Generative AI quietly broke that link, and most of education has not caught up with what that means.
Best read in order: start with Part 1 and the argument builds from there.
구술 평가가 대중 교육을 떠난 건 타당성이 아니라 비용 때문이었어요. 대화형 AI가 그 비용을 움직여요. 저희가 주장하는 것, 주장하지 않는 것, 그리고 검증 계획까지.
판돈이 가장 큰 곳에서 구술시험은 사라진 적이 없어요. 교실을 떠난 이유는 단 하나였고, 그 이유는 측정을 잘하고 못하고와는 무관했어요.
탐지기는 신뢰하기 어렵고, 엉뚱한 학생을 플래그하고, 완벽해진다 해도 잘못된 질문에 답할 뿐이에요. 교사의 질문은 글의 저자가 아니라 이해에 대한 거예요.
두 세기 동안 좋은 에세이는 좋은 사고의 증거였어요. 생성형 AI가 그 연결고리를 조용히 끊었고, 교육 대부분은 아직 그 의미를 따라잡지 못했어요.
순서대로 읽는 걸 추천해요. 1부부터 시작하면 논증이 차곡차곡 쌓여요.
A real Viva interview and the instructor report it produces. Three minutes, no sign-up.