Practice · October 14, 2025

Problem sets were the first casualty and nobody wrote about it

Viva Team · 4 min read

A correct curve with nothing behind it
A correct curve with nothing behind it
실무 · 2025년 10월 14일

가장 먼저 무너진 건 문제 풀이였는데 아무도 그 얘길 안 했어요

Viva 팀 · 4분 분량

뒤에 아무것도 없는 정확한 곡선
뒤에 아무것도 없는 정확한 곡선

Almost all of the public argument about AI and assessment has been about writing, which is odd, because the first assessment format to fall over was the weekly problem set. A model that can produce a competent literature review can certainly produce a correct integral with working shown, or a functioning implementation of a data structure with comments explaining the choices.

In computer science this happened first and hardest, and the response in many departments was to make the problems harder. That did not work and could not: difficulty is exactly the axis on which models have improved fastest.

Why the usual fixes underperform

Harder problems. Raises the bar for the honest student more than for the dishonest one. The student using a model is not solving the problem, so its difficulty is not their constraint.

Novel problems. Helps for a term. Then the problems are online, or the model is better, or both.

More weight on the exam. Concentrates everything into one high-anxiety session and tells students the coursework is theatre. It also removes the formative value that made problem sets worth setting.

The problem set was never valuable because the answers were valuable. It was valuable because producing them changed the student.

What survives

The formats that hold are the ones where the artefact was never the whole point.

That last one is worth sitting with. A great deal of the integrity problem in STEM is a grading-weight problem: we attached marks to an activity whose value was formative, then acted surprised when students optimised for the marks.

The oral version of this is short. Five minutes on a submitted problem set, questions drawn from the student's own working, and the answer to who understands their solution stops being a guess.

AI와 평가에 대한 공개 논쟁은 거의 전부 글쓰기에 대한 것이었어요. 이상한 일이에요. 가장 먼저 무너진 평가 형식은 주간 문제 풀이였거든요. 그럴듯한 선행연구 정리를 만들어 내는 모델이라면, 풀이 과정을 보여주는 정적분이나 선택을 설명하는 주석이 달린 자료구조 구현도 당연히 만들어요.

컴퓨터 과학에서 가장 먼저, 가장 세게 일어났고, 많은 학과의 대응은 문제를 더 어렵게 내는 것이었어요. 그건 통하지 않았고 통할 수도 없었어요. 난이도야말로 모델이 가장 빠르게 좋아진 축이니까요.

흔한 해법이 힘을 못 쓰는 이유

더 어려운 문제. 정직하지 않은 학생보다 정직한 학생의 문턱을 더 높여요. 모델을 쓰는 학생은 문제를 풀고 있는 게 아니라서, 난이도가 그 학생의 제약이 아니에요.

새로운 문제. 한 학기는 도움이 돼요. 그다음엔 문제가 온라인에 올라가 있거나, 모델이 더 좋아졌거나, 둘 다예요.

시험 비중 늘리기. 모든 걸 불안도 높은 한 번의 자리로 몰아넣고, 학생에게 과제는 연극이라고 알려줘요. 문제 풀이를 낼 가치가 있게 만들었던 형성적 가치도 함께 사라지고요.

문제 풀이가 가치 있었던 건 답이 가치 있어서가 아니에요. 답을 만들어 내는 과정이 학생을 바꿔 놓았기 때문이에요.

살아남는 것

버티는 형식은 결과물이 애초에 전부가 아니었던 형식이에요.

마지막 항목은 곱씹어 볼 만해요. 이공계 진정성 문제의 상당 부분은 채점 비중 문제예요. 형성적 가치를 지닌 활동에 점수를 붙여 놓고, 학생이 점수에 최적화하자 놀란 척한 거죠.

이걸 구술로 하면 짧아요. 제출된 문제 풀이에 대해 5분, 질문은 학생 본인의 풀이 과정에서 뽑고요. 그러면 누가 자기 풀이를 이해하는지가 더 이상 추측이 아니게 돼요.

See what an oral check actually looks like.

A real interview and the report an instructor receives. Three minutes, no sign-up.

구술 확인이 실제로 어떤 모습인지 보세요.

실제 인터뷰와 선생님이 받는 리포트예요. 3분이면 되고 가입도 필요 없어요.