AI

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

Researchers have developed a framework called ACES (Agentic Continuous Evaluation of Skills) to evaluate the effectiveness of skills and product capability packages in enterprise agent programs. ACES runs paired live trials with and without a target skill, normalizes trajectories into a standard format, and grades six default runtime metrics. The framework has been tested on 145 real skills from internal enterprise repositories and public catalogs, showing that traditional sc
Researchers have developed a framework called ACES (Agentic Continuous Evaluation of Skills) to evaluate the effectiveness of skills and product capability packages in enterprise agent programs. ACES runs paired live trials with and without a target skill, normalizes trajectories into a standard format, and grades six default runtime metrics. The framework has been tested on 145 real skills from internal enterprise repositories and public catalogs, showing that traditional scan-only gates are insufficient to measure the value of skills. Instead, ACES provides a more comprehensive evaluation of skills, including their added value for specific tasks and workflows. --- Why it matters: This matters because it addresses a critical challenge in AI development: evaluating the effectiveness of skills and capabilities in real-world applications. By providing a standardized framework for continuous evaluation, ACES enables developers to make data-driven decisions about skill deployment and improvement. Source: https://arxiv.org/abs/2608.20614

This article was originally published at: https://arxiv.org/abs/2608.20614