Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
Researchers evaluated how well large language models (LLMs) can ask questions and gather information to answer complex queries. They created a controlled test suite with over 9,000 tasks across various domains, including math, biology, and medicine. The results show that LLMs struggle to identify the amount of missing information and often stop acquiring necessary data before reaching a final answer. This study highlights the limitations of current LLM evaluations, which focu
Researchers evaluated how well large language models (LLMs) can ask questions and gather information to answer complex queries. They created a controlled test suite with over 9,000 tasks across various domains, including math, biology, and medicine. The results show that LLMs struggle to identify the amount of missing information and often stop acquiring necessary data before reaching a final answer. This study highlights the limitations of current LLM evaluations, which focus on generating accurate answers rather than seeking information effectively.
---
Why it matters: This research matters because it reveals the shortcomings of large language models in handling complex queries, where they fail to ask the right questions and gather sufficient information. As AI systems are increasingly used for tasks that require multi-turn interactions, understanding their limitations is crucial for improving their performance and developing more effective evaluation metrics.
Source: https://arxiv.org/abs/2608.14808
This article was originally published at: https://arxiv.org/abs/2608.14808