AI

ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation

Researchers have proposed a new benchmark called ATP-Bench for evaluating the ability of Multimodal Large Language Models (MLLMs) to generate interleaved text and images. The benchmark includes 7,702 question-and-answer pairs across eight categories and 25 visual-critical intents, with human-verified queries and ground truths. To evaluate agentic planning in MLLMs, a Multi-Agent system called MAM is proposed, which assesses tool-call precision and identifies missed opportunit
Researchers have proposed a new benchmark called ATP-Bench for evaluating the ability of Multimodal Large Language Models (MLLMs) to generate interleaved text and images. The benchmark includes 7,702 question-and-answer pairs across eight categories and 25 visual-critical intents, with human-verified queries and ground truths. To evaluate agentic planning in MLLMs, a Multi-Agent system called MAM is proposed, which assesses tool-call precision and identifies missed opportunities for tool use without requiring ground-truth references. Experiments on 10 state-of-the-art MLLMs show that they struggle with coherent interleaved planning and exhibit significant variations in tool-use behavior. --- Why it matters: This matters to researchers working on AI models that can generate text and images, as it provides a new benchmark for evaluating their ability to do so coherently. The results of the experiments also highlight areas where these models need improvement. Source: https://arxiv.org/abs/2603.29902

This article was originally published at: https://arxiv.org/abs/2603.29902