AI

What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

Researchers from Xiaonan Xu and Wenjing Wu analyzed the migration of commercial large language model APIs by looking at item-level behavior rather than aggregate scores. They found that while some items improve with newer models, others regress or remain unchanged. This suggests that relying solely on aggregate benchmark scores can lead to incomplete information about the actual performance of a model. The study used three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol produ
Researchers from Xiaonan Xu and Wenjing Wu analyzed the migration of commercial large language model APIs by looking at item-level behavior rather than aggregate scores. They found that while some items improve with newer models, others regress or remain unchanged. This suggests that relying solely on aggregate benchmark scores can lead to incomplete information about the actual performance of a model. The study used three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence and classified 900 public benchmark items as reliably improved, regressed, equivalent, or inconclusive. --- Why it matters: This matters because it highlights the limitations of relying on aggregate scores when evaluating model performance. Engineers working with commercial LLM APIs need to consider item-level behavior to make informed decisions about migrations. Source: https://arxiv.org/abs/2608.17719

This article was originally published at: https://arxiv.org/abs/2608.17719