AI

DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning

Researchers propose a new method called DirEAG to improve confidence estimation in large language models for mathematical reasoning. Current methods struggle to calibrate verbalized confidence, which is essential for reliable model performance. The proposed method aggregates evidence from multiple prompts and models to produce calibrated soft evidence over candidate answers and a null state. Experiments show that DirEAG achieves better calibration than existing methods while
Researchers propose a new method called DirEAG to improve confidence estimation in large language models for mathematical reasoning. Current methods struggle to calibrate verbalized confidence, which is essential for reliable model performance. The proposed method aggregates evidence from multiple prompts and models to produce calibrated soft evidence over candidate answers and a null state. Experiments show that DirEAG achieves better calibration than existing methods while maintaining competitive answer selection. --- Why it matters: This matters because accurate confidence estimation is crucial in AI applications, especially in high-stakes domains like mathematical reasoning. Engineers working on large language models can benefit from this research by improving the reliability and trustworthiness of their models. Source: https://arxiv.org/abs/2608.20717

This article was originally published at: https://arxiv.org/abs/2608.20717