AI

Which Source Wins? Task-Dependent Reliance in Vision-Language Models

Researchers studied how vision-language models (VLMs) handle conflicting information between images and text. They created a controlled setup where one source was degraded while the other remained clean, and tracked how the model's reliance shifted. The results show that VLMs' reliance on modality varies across tasks, evidence structures, models, and evaluation settings. For example, in arithmetic problems, five out of six models relied more strongly on text than images when
Researchers studied how vision-language models (VLMs) handle conflicting information between images and text. They created a controlled setup where one source was degraded while the other remained clean, and tracked how the model's reliance shifted. The results show that VLMs' reliance on modality varies across tasks, evidence structures, models, and evaluation settings. For example, in arithmetic problems, five out of six models relied more strongly on text than images when it was degraded. However, in chart-report conflicts, all six models relied more strongly on visual sources. This reversal persisted even after adjusting for accuracy loss or replacing charts with plain table images. --- Why it matters: These findings are important because they reveal that VLMs' behavior is not fixed, but rather depends on the specific task and context. Understanding how these models adapt to conflicting information can help improve their performance in real-world applications. Source: https://arxiv.org/abs/2608.17205

This article was originally published at: https://arxiv.org/abs/2608.17205