Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages
Researchers have created a benchmark called VSysBench to evaluate how well multimodal large language models (MLLMs) follow system messages. These messages govern model behavior in production deployments and are crucial for ensuring the model's actions align with user expectations. The benchmark assesses MLLMs' ability to comply with system messages, particularly in scenarios where there is conflict between user input and model output. The study found that imposing system mess
Researchers have created a benchmark called VSysBench to evaluate how well multimodal large language models (MLLMs) follow system messages. These messages govern model behavior in production deployments and are crucial for ensuring the model's actions align with user expectations. The benchmark assesses MLLMs' ability to comply with system messages, particularly in scenarios where there is conflict between user input and model output. The study found that imposing system messages can significantly reduce base task accuracy, with some models struggling more than others to follow vision-grounded constraints.
---
Why it matters: This research matters because it highlights the challenges of integrating system messages into MLLMs, which are increasingly used in production environments. Understanding how these models respond to system messages is crucial for their deployment and use in applications where user expectations must be met.
Source: https://arxiv.org/abs/2608.19207
This article was originally published at: https://arxiv.org/abs/2608.19207