When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
Researchers have created a benchmark to test how accurately Vision-Language Models (VLMs) and humans follow their own decision-making rules. The Graded Color Attribution dataset consists of simple drawings with varying color coverage. Participants, including VLMs, are asked to state a threshold rule for assigning colors to objects. The results show that VLMs systematically violate their own stated rules, while human participants remain faithful to theirs. This discrepancy sug
Researchers have created a benchmark to test how accurately Vision-Language Models (VLMs) and humans follow their own decision-making rules. The Graded Color Attribution dataset consists of simple drawings with varying color coverage. Participants, including VLMs, are asked to state a threshold rule for assigning colors to objects. The results show that VLMs systematically violate their own stated rules, while human participants remain faithful to theirs. This discrepancy suggests that VLMs' introspective self-knowledge is flawed and may lead to unreliable deployment in high-stakes situations.
---
Why it matters: This study matters because it challenges the assumption that VLM reasoning failures are due to difficulty with complex tasks. Instead, it highlights a fundamental issue with how VLMs understand their own decision-making processes, which has significant implications for AI development and deployment.
Source: https://arxiv.org/abs/2604.06422
This article was originally published at: https://arxiv.org/abs/2604.06422