Academic Integrity · April 8, 2026

The students AI detectors get wrong are the ones you least want to be wrong about

Viva Team · 7 min read

A sieve that lets the marked shapes through and holds the plain one
A sieve that lets the marked shapes through and holds the plain one
학문적 진정성 · 2026년 4월 8일

AI 탐지기가 틀리는 학생은 하필 가장 틀리면 안 되는 학생이에요

Viva 팀 · 7분 분량

표시된 것들은 통과시키고 담백한 것 하나를 붙드는 체
표시된 것들은 통과시키고 담백한 것 하나를 붙드는 체

There is a version of the AI detection debate that is about accuracy, and it is the less interesting version. Accuracy questions have engineering answers: more data, better models, higher thresholds. Wait long enough and someone improves the number.

The version that matters is about distribution. When a detector is wrong, who does it tend to be wrong about?

The distribution is not random

In 2023 a Stanford group ran essays by native and non-native English speakers through seven commercial detectors. The native writing came back mostly clean. The non-native writing was flagged at rates so high that the authors described the tools as unsuitable for use on that population. The mechanism is not mysterious. Detectors lean on measures of lexical variety and sentence-level unpredictability, and a competent writer working in a second language tends to produce exactly what those measures read as machine-like: a narrower vocabulary deployed carefully, and sentences that do not take risks.

Since then a fairly consistent picture has emerged from institutions and independent testers. Plain, structured, careful prose gets flagged. So does writing by students who were taught to a template, which describes a great deal of secondary education. So does writing by some autistic students, whose prose can be precise and low in the kind of stylistic noise the tools treat as human.

Every one of those groups is a group that already has to work harder to be believed. Handing them an extra burden of proof is not a rounding error.

Why raising the threshold does not fix it

The obvious response is to demand more confidence before flagging. This helps, and it trades one failure for another. Push the threshold up and the detector stops catching the students who are actually outsourcing their work, which was the entire point. Push it down and you are back to accusing your international students.

This is not a bug that gets engineered away, because the underlying signal is style, and style is not authorship. A confident, careful, slightly formal paragraph can come from a language model or from a nineteen year old who learned English from textbooks and edits carefully. There is no amount of training data that separates those two, because they are not different in the text. They are different in what happened before the text.

The thing a detector cannot see

Here is the question an instructor actually has: does this student understand the work they handed in? A detector cannot answer that. It was never designed to. It answers a proxy question about the statistical shape of a document, and then the institution treats the proxy as though it were the real one.

The proxy fails in both directions and the failures are asymmetric in consequence. A missed case is a student who got away with something. A false positive is a student in a disciplinary hearing defending work they wrote, often without the vocabulary or the confidence to push back, sometimes on a visa.

What we would do instead

We are not neutral here, so take this as a stated position rather than a finding: the way out is to stop asking documents to prove themselves and start asking students to.

None of this is novel. It is how doctoral defenses, clinical exams and bar qualifications have worked for a very long time. What changed is that the conversation used to be affordable only where the stakes were highest.

Related: AI detection was never going to save the essay and how teachers actually know you used ChatGPT.

AI 탐지 논쟁에는 정확도에 대한 버전이 있어요. 그리고 그건 덜 흥미로운 버전이에요. 정확도 문제에는 공학적인 답이 있으니까요. 데이터를 늘리고, 모델을 개선하고, 기준을 높이면 돼요. 시간이 지나면 누군가는 숫자를 개선해요.

중요한 건 분포에 대한 버전이에요. 탐지기가 틀릴 때, 주로 누구에 대해 틀리나요?

분포는 무작위가 아니에요

2023년 스탠퍼드 연구진은 영어 원어민과 비원어민의 에세이를 상용 탐지기 일곱 개에 넣어 봤어요. 원어민 글은 대체로 깨끗하게 나왔어요. 비원어민 글은 표시되는 비율이 너무 높아서, 저자들은 이 도구들이 그 집단에 쓰기에 적합하지 않다고 썼어요. 원리는 이상하지 않아요. 탐지기는 어휘의 다양성과 문장의 예측 불가능성 같은 지표에 기대는데, 제2언어로 조심스럽게 잘 쓰는 사람은 정확히 그 지표들이 기계 같다고 읽는 글을 써요. 좁은 어휘를 정확하게 쓰고, 모험하지 않는 문장을 쓰죠.

그 뒤로 기관과 독립 테스터들에게서 꽤 일관된 그림이 나왔어요. 담백하고 구조가 뚜렷하고 조심스러운 문장이 표시돼요. 템플릿에 맞춰 배운 학생의 글도 그렇고요. 중등교육의 상당 부분이 그렇죠. 정확하고 문체적 잡음이 적은 일부 자폐 스펙트럼 학생의 글도 마찬가지예요.

여기 나온 집단은 모두 이미 믿어지기 위해 더 애써야 하는 집단이에요. 그들에게 입증 책임을 더 얹는 건 반올림 오차가 아니에요.

기준을 높인다고 해결되지 않는 이유

당연한 대응은 표시하기 전에 더 높은 확신을 요구하는 거예요. 도움이 되긴 하지만 하나의 실패를 다른 실패와 바꾸는 거예요. 기준을 올리면 실제로 남에게 맡긴 학생을 놓치게 되는데, 그게 애초의 목적이었죠. 내리면 다시 유학생을 의심하게 되고요.

이건 공학으로 없앨 수 있는 결함이 아니에요. 밑에 깔린 신호가 문체이고, 문체는 저자가 아니니까요. 자신 있고 조심스럽고 살짝 격식 있는 문단은 언어 모델에서 나올 수도 있고, 교과서로 영어를 배우고 꼼꼼히 고쳐 쓰는 열아홉 살에게서 나올 수도 있어요. 아무리 학습 데이터를 늘려도 둘은 갈라지지 않아요. 텍스트 안에서는 다르지 않거든요. 다른 건 텍스트 이전에 일어난 일이에요.

탐지기가 볼 수 없는 것

선생님이 실제로 가진 질문은 이거예요. 이 학생이 제출한 작업을 이해하고 있나요? 탐지기는 그 질문에 답하지 못해요. 애초에 그러라고 만들어지지 않았어요. 문서의 통계적 형태에 대한 대리 질문에 답할 뿐인데, 기관은 그 대리 질문을 진짜 질문처럼 다뤄요.

대리 질문은 양쪽으로 실패하고, 두 실패의 결과는 대칭이 아니에요. 놓친 경우는 넘어간 학생 한 명이에요. 오탐은 자기가 쓴 글을 변호하러 징계 자리에 앉은 학생이에요. 반박할 어휘도 자신감도 없는 경우가 많고, 비자가 걸려 있기도 해요.

저희라면 이렇게 할 거예요

저희는 중립이 아니니 발견이 아니라 입장으로 읽어 주세요. 출구는 문서에게 스스로를 증명하라고 하는 걸 멈추고, 학생에게 묻기 시작하는 거예요.

새로운 이야기는 아니에요. 박사 디펜스, 임상 시험, 변호사 자격 시험이 아주 오래 그렇게 해 왔어요. 달라진 건, 예전에는 그 대화를 판돈이 가장 큰 곳에서만 감당할 수 있었다는 점이에요.

함께 읽기: AI 탐지는 에세이를 구하지 못해요, 선생님은 ChatGPT를 쓴 걸 어떻게 알아챌까요.

See what an oral check actually looks like.

A real interview and the report an instructor receives. Three minutes, no sign-up.

구술 확인이 실제로 어떤 모습인지 보세요.

실제 인터뷰와 선생님이 받는 리포트예요. 3분이면 되고 가입도 필요 없어요.