Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
Researchers have developed a new method for on-policy self-distillation in agentic reinforcement learning. The approach, called Influence Calibration for Self-Distillation (ICSD), aims to address the issue of trust-utility mismatch by measuring the impact of token emphasis on policy objectives. ICSD calculates allocation weights based on the response of RL surrogate contributions to teacher-directed output perturbations. This method improves performance over existing methods
Researchers have developed a new method for on-policy self-distillation in agentic reinforcement learning. The approach, called Influence Calibration for Self-Distillation (ICSD), aims to address the issue of trust-utility mismatch by measuring the impact of token emphasis on policy objectives. ICSD calculates allocation weights based on the response of RL surrogate contributions to teacher-directed output perturbations. This method improves performance over existing methods in various tasks, including ALFWorld and WebShop, with significant reductions in objective-opposed tokens and increases in cosine compatibility with RL gradients.
---
Why it matters: This matters because it addresses a key limitation of current on-policy self-distillation methods, which rely solely on teacher trust. By calibrating the influence of token emphasis, ICSD can improve performance and reduce errors in tasks like language understanding and decision-making.
Source: https://arxiv.org/abs/2608.14945
This article was originally published at: https://arxiv.org/abs/2608.14945