AI

Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme

A new benchmark called THPT-Ladder has been introduced to evaluate language models on Vietnam's National High School Graduation Examination. The benchmark takes into account the exam's non-additive grading scheme, where partial knowledge is not rewarded proportionally. In contrast, standard accuracy metrics assume that each correct response earns equal credit. The authors found that using the official rubric instead of proportional credit can result in a model's apparent comp
A new benchmark called THPT-Ladder has been introduced to evaluate language models on Vietnam's National High School Graduation Examination. The benchmark takes into account the exam's non-additive grading scheme, where partial knowledge is not rewarded proportionally. In contrast, standard accuracy metrics assume that each correct response earns equal credit. The authors found that using the official rubric instead of proportional credit can result in a model's apparent competence being overestimated by 0.020 to 0.159 points per question. This discrepancy can significantly impact a model's ranking among candidates. The results suggest that standard benchmarks may not accurately reflect a model's performance on exams with non-additive grading schemes. --- Why it matters: This matters because language models are often evaluated using benchmarks that assume proportional credit, which may not accurately reflect their performance in real-world scenarios where non-additive grading schemes are used. Source: https://arxiv.org/abs/2608.18336

This article was originally published at: https://arxiv.org/abs/2608.18336