AI

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

Researchers have developed ATBench, a new benchmark for evaluating the safety of language models. The benchmark includes 1,000 trajectories with varying levels of risk and failure modes, making it more realistic than existing benchmarks. It's designed to capture long-horizon risks that can emerge over multiple interactions. The data is filtered using rule-based and LLM-based methods, and then audited by humans to ensure quality.
Researchers have developed ATBench, a new benchmark for evaluating the safety of language models. The benchmark includes 1,000 trajectories with varying levels of risk and failure modes, making it more realistic than existing benchmarks. It's designed to capture long-horizon risks that can emerge over multiple interactions. The data is filtered using rule-based and LLM-based methods, and then audited by humans to ensure quality. --- Why it matters: This matters because language models are increasingly being used in real-world applications where safety is crucial. ATBench provides a more realistic evaluation of agent safety, which can help developers identify potential risks and improve the overall safety of their models. Source: https://arxiv.org/abs/2604.02022

This article was originally published at: https://arxiv.org/abs/2604.02022