Every instructor who has assigned group work knows the shape of the complaint. Four names on a document, two people wrote it, and the peer evaluation forms come back saying everyone contributed equally because nobody wants to be the person who sank a classmate.
That was true before generative models. What changed is the cover story. "I did the literature review section" used to require doing a literature review. It is now a sentence anyone can say about a section they did not write, produced by a group member who also did not write it.
Why peer evaluation underperforms
The instrument has a well-known set of problems: friendship inflation, retaliation fear, and a strong tendency for everyone to award the middle score. Weighted variants and confidential forms help at the margin. None of them fix the underlying issue, which is that you are asking students to make a social judgment and then treating the result as a measurement.
You are not trying to find out who students think worked hardest. You are trying to find out who can account for the work.
The more direct question
Ask each member to explain the part they say they did. Not to justify their teammates, not to allocate credit, just to talk through their own contribution: what they chose, what they tried that did not work, why the section is structured the way it is.
This is unusually informative, and it has a property peer evaluation lacks: the student is being asked about themselves, so a bad outcome is not a betrayal of anyone. Free riders reveal themselves without a classmate having to name them, which is precisely the dynamic that breaks peer forms.
It also surfaces something more useful than a free rider list. When two students both claim the same component and one of them can describe the decisions inside it and the other cannot, you have learned something specific. When a quiet student turns out to be the one who can explain the analysis, you have learned something better.
Keeping it proportionate
A few practical notes from instructors doing this.
- Four to five minutes per student is enough for a component-level check. This is not a viva on the whole project.
- Ask everyone, including the students you are confident about. A targeted version is an accusation.
- Have them name their part before the deadline, not after. A claim made in advance is worth more than one reconstructed under questioning.
- Grade the individual understanding separately from the group product. Both numbers are real and they answer different questions.
The unglamorous truth is that group work was always assessed on trust, and trust worked reasonably well when producing the artefact required doing the work. That assumption is what broke. The response does not need to be more forms.
If you want the mechanics of how we do this, it is written up at group projects.