Research

Evidence behind the platform

Everything we build is grounded in research on how oral assessment reveals understanding — and how it doesn't. We share our work openly, mark clearly what is still a hypothesis, and hold ourselves accountable to the evidence.

연구

플랫폼을 뒷받침하는 근거

저희가 만드는 모든 것은 구술 평가가 어떻게 이해를 드러내는지, 그리고 어떤 한계가 있는지에 대한 연구에 기반해요. 저희의 연구를 투명하게 공유하고, 아직 가설인 부분은 분명히 표시하며, 그 근거에 책임을 지고 있어요.

Questions we are studying

What we ask

Does oral assessment surface understanding the essay alone hides?

How we test

Are AI scores within the same range as a panel of experienced instructors?

What we watch

Do conversation scores differ across language background, disability, or socioeconomic indicators?

What we are learning

Does knowing an oral defense is coming change how students engage with their work before they submit?

저희가 살펴보는 질문들

저희가 묻는 것

구술 평가가 글만으로는 보이지 않는 이해를 드러낼까?

저희가 시험하는 것

AI 점수가 경험 많은 교사 패널과 같은 범위 안에 들어올까?

저희가 살피는 것

언어 배경, 장애 여부, 사회경제적 지표에 따라 점수가 달라지지는 않을까?

저희가 배우는 것

구술 평가가 예정되어 있다는 사실이, 학생이 제출 전에 자기 글을 다루는 방식을 바꿀까요?

Working papers

What we are studying

Beyond Detection: The Case for Scalable Oral Assessment as a Response to Generative AI in Education

Viva Research · Working paper, 2026 (under revision)

A critical review of why oral assessment — long sidelined in mass education because it does not scale — deserves reconsideration as a response to generative AI. We synthesize the evidence on oral versus written assessment, the cognitive science of retrieval and dialogue, and early work on AI-mediated oral assessment, and argue that conversational AI may relax the cost constraint that once confined oral exams to elite settings. We are explicit that this central claim is a hypothesis to be tested — not an established result — and we set out the research agenda that would test it. Not yet peer-reviewed.

AI-Mediated Oral Assessment in Practice: A Multi-School Study

Viva Research · In preparation · Pilot data collection underway

The empirical companion to the review above. With partner schools, we are studying how AI-mediated oral assessment compares to experienced instructors' judgment (validity), whether outcomes differ across student groups (equity), and how the format shapes the way students prepare and learn. Data collection is underway; we will report findings — and their limitations — only when the evidence supports them.

How we work

Our research principles

📐

Transparent methods

We document our research methodology — study design, measures, and analysis — and share it with partner schools on request.

🎯

Instructor calibration

AI scores are calibrated against a panel of experienced instructors, with ongoing review of agreement rates.

⚖️

Adverse impact monitoring

We actively test for score disparities across language background, disability status, and socioeconomic indicators. If we find them, we investigate them.

🧑‍🏫

No closed-loop grading

Viva scores are advisory signals. Final grade decisions remain with the instructor. We are explicit about this in every interface.

저희의 방식

연구 원칙

📐

투명한 방법론

연구 방법론(연구 설계, 측정 방법, 분석)을 문서화하고, 파트너 학교에는 요청 시 공유해요.

🎯

교사 캘리브레이션

AI 점수는 경험 많은 교사 패널과의 비교로 보정되며, 일치율을 지속적으로 검토해요.

⚖️

불공평한 영향 모니터링

언어적 배경, 장애 여부, 사회경제적 지표에 따른 점수 격차가 있는지 적극적으로 검증해요. 격차가 발견되면 원인을 추적해요.

🧑‍🏫

닫힌 채점 루프 없음

Viva 점수는 참고 신호일 뿐이에요. 최종 채점은 항상 교사가 결정해요. 모든 화면에서 이 점을 명확히 하고 있어요.

Partner with our research team.

We collaborate with universities, school districts, and independent researchers working on assessment, integrity, and AI in education.

저희 리서치 팀과 함께해요.

평가, 진정성, 그리고 교육 속 AI를 연구하는 대학, 교육청, 독립 연구자분들과 협력하고 있어요.