MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection
Researchers have created a comprehensive benchmark for detecting malicious behavior in AI agents. The MaliciousSkillBench consolidates data from 13 public sources and reduces over 8,000 raw records to 9,740 unique skills, with 7,505 identified as malicious and 2,235 as benign. The benchmark evaluates the performance of three learned text detectors and three off-the-shelf skill scanners, finding that reliable detection requires both broad cross-source coverage and evaluation m
Researchers have created a comprehensive benchmark for detecting malicious behavior in AI agents. The MaliciousSkillBench consolidates data from 13 public sources and reduces over 8,000 raw records to 9,740 unique skills, with 7,505 identified as malicious and 2,235 as benign. The benchmark evaluates the performance of three learned text detectors and three off-the-shelf skill scanners, finding that reliable detection requires both broad cross-source coverage and evaluation methods that balance attack detection and benign flagging.
---
Why it matters: This matters to AI researchers because it provides a standardized tool for evaluating the effectiveness of malicious agent skill detection techniques. Accurate detection is crucial for preventing AI agents from being exploited by malicious actors, which can have serious consequences in areas like finance, healthcare, and national security.
Source: https://arxiv.org/abs/2608.19901
This article was originally published at: https://arxiv.org/abs/2608.19901