When Text and Numbers Disagree: Evidence Arbitration in Large Language Models
Large language models (LLMs) often encounter conflicting evidence from text and numbers when making decisions. Researchers have created a synthetic benchmark to study how LLMs arbitrate between these sources. The results show that current LLMs rely on heuristic strategies, such as preferring text or temporal recency over other factors, which can lead to errors in decision-making.
Large language models (LLMs) often encounter conflicting evidence from text and numbers when making decisions. Researchers have created a synthetic benchmark to study how LLMs arbitrate between these sources. The results show that current LLMs rely on heuristic strategies, such as preferring text or temporal recency over other factors, which can lead to errors in decision-making.
---
Why it matters: This matters because it highlights a potential flaw in the way LLMs integrate multiple types of evidence, which could impact their reliability and trustworthiness in real-world applications.
Source: https://arxiv.org/abs/2608.20116
This article was originally published at: https://arxiv.org/abs/2608.20116