AI

TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

A new system called TRACE aims to improve the reliability of large language models (LLMs) in real-world applications. Current LLMs often struggle with consistency and limit-awareness, meaning they may behave differently across repeated trials or fail to recognize when a request is too complex. TRACE addresses this issue by creating a 'Skill Bank' that organizes modular skills for tool use and behavioral guidelines. The system iteratively improves these skills through self-evo
A new system called TRACE aims to improve the reliability of large language models (LLMs) in real-world applications. Current LLMs often struggle with consistency and limit-awareness, meaning they may behave differently across repeated trials or fail to recognize when a request is too complex. TRACE addresses this issue by creating a 'Skill Bank' that organizes modular skills for tool use and behavioral guidelines. The system iteratively improves these skills through self-evolution, leading to consistent performance gains. In experiments, TRACE improved the consistency of GPT-5.5 models by 34.6 percentage points and achieved first place on a hidden test set using GPT-5.6-Sol. --- Why it matters: This matters because current LLMs often fail to translate their potential into real-world performance, leading to inconsistent results and safety concerns. TRACE's ability to improve consistency and limit-awareness could make large language models more reliable for use in applications like in-car assistants. Source: https://arxiv.org/abs/2608.22793

This article was originally published at: https://arxiv.org/abs/2608.22793