Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes
Researchers have developed an AI system that can grade handwritten physics assessments with high accuracy. The system used a variant of the GPT-5.5 model to evaluate over 10,000 submissions from students across three different exams. The AI was able to match official scores and outcomes in most cases, including selecting the same team for the final Olympiad selection as human graders. However, the system struggled with exact partial-credit grading, particularly in experimenta
Researchers have developed an AI system that can grade handwritten physics assessments with high accuracy. The system used a variant of the GPT-5.5 model to evaluate over 10,000 submissions from students across three different exams. The AI was able to match official scores and outcomes in most cases, including selecting the same team for the final Olympiad selection as human graders. However, the system struggled with exact partial-credit grading, particularly in experimental work. To improve accuracy, the researchers used detailed rubrics and had the AI grade submissions twice, once with and once without revised instructions. The study suggests that reliable AI grading can be achieved but requires careful development of rubrics and should be used as a second reader or audit tool under human control.
---
Why it matters: This research matters to engineers and researchers in AI because it demonstrates the potential for large-scale AI grading of handwritten assessments, which could have significant implications for education and assessment. The study's findings on the importance of detailed rubrics and careful development of AI grading systems will be relevant to those working on similar projects.
Source: https://arxiv.org/abs/2608.20521
This article was originally published at: https://arxiv.org/abs/2608.20521