AI

Execution-grounded evaluation reveals hidden failures in language-model calculations for environmental science

Researchers have developed a new benchmark called AtmosCoder-Bench to evaluate the performance of large language models in environmental science. Unlike previous evaluations that only score final answers, this benchmark makes the calculation process visible and transparent. The study found that existing evaluation methods can inflate accuracy by up to 12 percentage points due to multiple-choice formats, and that many language models fail to apply known formulas consistently t
Researchers have developed a new benchmark called AtmosCoder-Bench to evaluate the performance of large language models in environmental science. Unlike previous evaluations that only score final answers, this benchmark makes the calculation process visible and transparent. The study found that existing evaluation methods can inflate accuracy by up to 12 percentage points due to multiple-choice formats, and that many language models fail to apply known formulas consistently throughout multi-step computations. This suggests that even state-of-the-art models may not be reliable in certain situations, highlighting the need for expert oversight. --- Why it matters: This study matters because it reveals hidden failures in language-model calculations that can have significant consequences for environmental science applications. By making the calculation process visible and transparent, researchers can identify areas where models are failing and improve their performance accordingly. Source: https://arxiv.org/abs/2608.18726

This article was originally published at: https://arxiv.org/abs/2608.18726