StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
Researchers have developed a new agent-native runtime called StateM that improves the execution system around an existing AI model without changing its weights. StateM organizes execution around durable states and phase-local context, which leads to significant performance gains on several benchmarks. On Terminal-Bench 2.1, StateM raises the accuracy of GPT-5.6 Sol Ultra from 91.9% to 95.3%. The runtime also improves the performance of other models, including DeepSeek-V4 Flas
Researchers have developed a new agent-native runtime called StateM that improves the execution system around an existing AI model without changing its weights. StateM organizes execution around durable states and phase-local context, which leads to significant performance gains on several benchmarks. On Terminal-Bench 2.1, StateM raises the accuracy of GPT-5.6 Sol Ultra from 91.9% to 95.3%. The runtime also improves the performance of other models, including DeepSeek-V4 Flash and GPT-5.6 Luna, with minimal adaptation costs. The authors claim that StateM can be used to make learned controls explicit and enforceable through stateful controls.
---
Why it matters: This matters because it shows how a new execution system can significantly improve the performance of existing AI models without requiring changes to their weights or architecture. This could have important implications for the development of more efficient and effective AI systems, particularly in applications where accuracy is critical such as natural language processing and computer vision.
Source: https://arxiv.org/abs/2608.15089
This article was originally published at: https://arxiv.org/abs/2608.15089