Academic Integrity · August 18, 2026

How teachers actually know you used ChatGPT

Viva Team · 6 min read

A lens searching a page while the answer arrives as a spoken question
A lens searching a page while the answer arrives as a spoken question
학문적 진정성 · 2026년 8월 18일

선생님은 ChatGPT를 쓴 걸 어떻게 알아챌까요

Viva 팀 · 6분 분량

페이지를 들여다보는 렌즈, 그러나 답은 말로 건네지는 질문으로 도착해요
페이지를 들여다보는 렌즈, 그러나 답은 말로 건네지는 질문으로 도착해요

Every semester this question turns up in student forums, usually at two in the morning: can they tell? The answers you get are mostly wrong in both directions. One camp insists detectors are magic. The other insists nothing can be proven. Neither is how it actually goes.

We build software in this space, so we have read a lot of misconduct paperwork. Here is what is really happening on the other side of the desk.

The detector score is rarely the reason

Instructors do run submissions through detectors, and detectors do return numbers. But most experienced instructors have learned not to trust those numbers on their own, and many institutions have quietly told staff not to open a case on a score alone. The reason is simple: the scores are wrong often enough to ruin someone's year. Research out of Stanford in 2023 found that detectors flagged writing by non-native English speakers as AI-generated at strikingly higher rates than writing by native speakers, and that finding has been replicated in spirit many times since. A tool that punishes you for having learned English second is not a tool you can build a hearing on.

So when a case does move forward, the score is almost never the evidence. It is the thing that made someone look more closely.

What actually gives it away

Voice discontinuity. Instructors read a lot of student writing, and they build a model of how each student sounds without meaning to. A paragraph that is fluent, balanced, and slightly airless sitting between two paragraphs that are messier and more specific reads like a splice. You do not need software to hear it. You need to have read the last three assignments.

Citations that do not exist, or exist and say something else. This one is brutal and common. A model produces a plausible reference, the reference is checked, and it turns out the paper is real but argues the opposite, or is not real at all. Nothing else in academic work fails this cleanly.

Specificity that stops at the door. The work describes the method beautifully and cannot say why that method. It names the dataset and does not know how many rows it has. Generic competence with no local detail is a strong signal because your own work is full of local detail whether you notice it or not.

The tell is not that the writing is too good. It is that the writing is unattached to a person who made choices.

Process evidence. Version history, drafts, a document created eleven minutes before the deadline with 2,400 words pasted in at once. Instructors increasingly ask for this, and it is far more persuasive in a hearing than any percentage.

And the oldest one: the conversation. A teacher asks you about your own paper, and either you can talk about it or you cannot. This is why oral checks are coming back, and it is the part of the answer students most consistently underestimate.

What this means if you are the student

The honest advice is not what you probably expect, because it is not really about tools.

And if you are the instructor

The uncomfortable part of all this is that the reliable signals are labour-intensive and the cheap signal is unreliable. Reading version history for 120 students is not a plan. Neither is trusting a number that misfires on your international students.

Which is why the interesting question stopped being how to detect AI text and became how to verify understanding at a scale a real course can carry. Those are different questions, and only the second one has a good answer.

If you want the longer argument for that, we wrote it here: AI detection was never going to save the essay.

학기마다 학생 커뮤니티에 이 질문이 올라와요. 보통 새벽 두 시에요. 들키나요? 그런데 돌아오는 답은 양쪽 다 틀린 경우가 많아요. 한쪽은 탐지기가 마법이라고 하고, 다른 쪽은 절대 증명 못 한다고 해요. 실제로는 둘 다 아니에요.

저희는 이 분야에서 소프트웨어를 만들고 있어서 학업 부정 관련 문서를 꽤 많이 읽었어요. 책상 반대편에서 실제로 무슨 일이 일어나는지 정리해 볼게요.

탐지기 점수는 이유가 되는 경우가 드물어요

선생님들도 탐지기를 돌리고, 탐지기는 숫자를 내놓아요. 하지만 경험 있는 선생님 대부분은 그 숫자만 믿지 않는 법을 배웠고, 많은 기관이 점수 하나로 사건을 열지 말라고 조용히 안내하고 있어요. 이유는 간단해요. 그 점수가 한 사람의 한 해를 망칠 만큼 자주 틀리거든요. 2023년 스탠퍼드 연구는 영어가 모국어가 아닌 사람의 글을 탐지기가 AI 생성물로 표시하는 비율이 눈에 띄게 높다는 걸 보여줬어요. 영어를 나중에 배웠다는 이유로 벌을 주는 도구 위에 징계 절차를 세울 수는 없어요.

그래서 실제로 사건이 진행될 때, 점수가 증거가 되는 일은 거의 없어요. 점수는 누군가 더 자세히 들여다보게 만든 계기일 뿐이에요.

실제로 드러나는 것들

목소리의 끊김이에요. 선생님은 학생 글을 아주 많이 읽고, 의도하지 않아도 학생마다 어떤 문체인지 감을 쌓아요. 거칠고 구체적인 두 문단 사이에 매끄럽고 균형 잡혔지만 어딘가 공기가 빠진 문단이 끼어 있으면 이어 붙인 자국처럼 읽혀요. 소프트웨어는 필요 없어요. 지난 세 번의 과제를 읽었으면 충분해요.

존재하지 않거나, 존재하지만 다른 말을 하는 인용이에요. 이건 잔인할 만큼 자주 나와요. 모델이 그럴듯한 참고문헌을 만들어 내고, 확인해 보면 논문이 아예 없거나 정반대 주장을 하고 있어요. 학문 작업에서 이렇게 깔끔하게 무너지는 신호는 별로 없어요.

문 앞에서 멈추는 구체성이에요. 방법론은 훌륭하게 설명하는데 왜 그 방법인지는 말하지 못해요. 데이터셋 이름은 대는데 행이 몇 개인지는 몰라요. 로컬한 디테일 없는 일반적인 유능함은 강한 신호예요. 본인이 직접 한 작업에는 의식하지 못한 디테일이 잔뜩 묻어 있으니까요.

신호는 글이 너무 잘 쓰였다는 게 아니에요. 그 글이 선택을 한 사람과 붙어 있지 않다는 거예요.

과정의 흔적이에요. 버전 기록, 초고, 마감 11분 전에 만들어진 문서에 2,400단어가 한 번에 붙여 넣어진 기록 같은 것들이요. 요즘 선생님들이 점점 더 이걸 요구하고, 청문 자리에서는 어떤 퍼센트보다 훨씬 설득력이 있어요.

그리고 가장 오래된 방법이 있어요. 대화예요. 선생님이 학생에게 본인 글에 대해 묻고, 학생은 이야기할 수 있거나 없거나예요. 구술 확인이 돌아오고 있는 이유이고, 학생들이 가장 자주 과소평가하는 부분이기도 해요.

학생이라면

솔직한 조언은 아마 예상과 다를 거예요. 사실 도구에 대한 이야기가 아니거든요.

선생님이라면

불편한 지점은, 믿을 만한 신호는 손이 많이 가고 손쉬운 신호는 믿을 수 없다는 거예요. 120명의 버전 기록을 읽는 건 계획이 아니에요. 유학생에게 자주 오작동하는 숫자를 믿는 것도 마찬가지고요.

그래서 흥미로운 질문은 AI 텍스트를 어떻게 탐지하느냐가 아니라, 실제 수업이 감당할 수 있는 규모로 이해를 어떻게 확인하느냐로 옮겨갔어요. 둘은 다른 질문이고, 좋은 답이 있는 쪽은 두 번째예요.

더 긴 논증은 여기에 써 두었어요. AI 탐지는 에세이를 구하지 못해요.

See what an oral check actually looks like.

A real interview and the report an instructor receives. Three minutes, no sign-up.

구술 확인이 실제로 어떤 모습인지 보세요.

실제 인터뷰와 선생님이 받는 리포트예요. 3분이면 되고 가입도 필요 없어요.