When Contextual Inference Fails: Cancelability in Interactive Instruction Following
Researchers have created an interactive benchmark called Build What I Mean (BWIM) to test how well language models can follow instructions in a collaborative setting. In BWIM, a speaker gives underspecified instructions and the model must decide whether to make a contextual inference or ask for clarification. The study found that while language models are good at detecting when a speaker is unreliable, they often fail to use this awareness to adjust their behavior. Instead of
Researchers have created an interactive benchmark called Build What I Mean (BWIM) to test how well language models can follow instructions in a collaborative setting. In BWIM, a speaker gives underspecified instructions and the model must decide whether to make a contextual inference or ask for clarification. The study found that while language models are good at detecting when a speaker is unreliable, they often fail to use this awareness to adjust their behavior. Instead of asking for clarification efficiently, models tend to over-ask or avoid questions altogether.
---
Why it matters: This matters because it shows that current language models struggle with contextual reasoning in interactive settings, which is an important aspect of human communication. Understanding how to improve these abilities could lead to more effective and efficient collaboration between humans and machines.
Source: https://arxiv.org/abs/2603.19997
This article was originally published at: https://arxiv.org/abs/2603.19997