MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents
Researchers have proposed a new method for assigning credits in long-horizon reinforcement learning. The approach, called MileGPO, identifies meaningful intermediate milestones by analyzing on-policy rollouts and weights these candidates based on outcome-based confidence. This allows the model to refine trajectory-level signals into step-level credits more effectively. Experiments show that MileGPO achieves state-of-the-art performance on two benchmark tasks.
Researchers have proposed a new method for assigning credits in long-horizon reinforcement learning. The approach, called MileGPO, identifies meaningful intermediate milestones by analyzing on-policy rollouts and weights these candidates based on outcome-based confidence. This allows the model to refine trajectory-level signals into step-level credits more effectively. Experiments show that MileGPO achieves state-of-the-art performance on two benchmark tasks.
---
Why it matters: This matters because assigning credits in long-horizon reinforcement learning is a challenging problem, and improving credit assignment can lead to better performance and more efficient training of agents.
Source: https://arxiv.org/abs/2608.19803
This article was originally published at: https://arxiv.org/abs/2608.19803