AI

Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study

Researchers have studied the performance of majority voting in Large Language Models (LLMs) to improve answer accuracy. They found that this method can backfire on hard questions and its effectiveness varies erratically. The study decomposes the agreement index into a mechanical component, which is related to the per-case answer preference, and a residual component. The results show that the per-case answer preference explains most of the held-out test-run agreement index in
Researchers have studied the performance of majority voting in Large Language Models (LLMs) to improve answer accuracy. They found that this method can backfire on hard questions and its effectiveness varies erratically. The study decomposes the agreement index into a mechanical component, which is related to the per-case answer preference, and a residual component. The results show that the per-case answer preference explains most of the held-out test-run agreement index in some cases, but not all. This work highlights the importance of understanding how LLMs arrive at their answers. --- Why it matters: This study matters because it sheds light on the limitations of majority voting in LLMs and provides insights into how these models make decisions. Understanding this can help improve the accuracy and reliability of language models, which are increasingly used in applications such as question-answering systems. Source: https://arxiv.org/abs/2608.18795

This article was originally published at: https://arxiv.org/abs/2608.18795