AI

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

Researchers have developed a new benchmark called LongRCA Bench to help diagnose problems in long-horizon AI agents. These agents can perform tasks over hundreds of steps, but when they fail, it's often unclear where the mistake occurred. The LongRCA Bench includes 1,140 failed trajectories across five domains and provides human-labeled data for identifying responsible roles and root causes. A new method called Root-Cause Trajectory Attribution (RCTA) has been developed to he
Researchers have developed a new benchmark called LongRCA Bench to help diagnose problems in long-horizon AI agents. These agents can perform tasks over hundreds of steps, but when they fail, it's often unclear where the mistake occurred. The LongRCA Bench includes 1,140 failed trajectories across five domains and provides human-labeled data for identifying responsible roles and root causes. A new method called Root-Cause Trajectory Attribution (RCTA) has been developed to help identify these issues without requiring training. RCTA achieves better results than existing methods in identifying responsible roles and exact root steps. --- Why it matters: This matters because long-horizon AI agents are increasingly used in applications such as autonomous vehicles, robotics, and supply chain management. Accurate diagnosis of failures is crucial to improve the reliability and safety of these systems. Source: https://arxiv.org/abs/2608.15242

This article was originally published at: https://arxiv.org/abs/2608.15242