SPyCE: Skill-Policy Co-evolution for Multimodal Agents
Researchers propose SPyCE (Skill-Policy Co-evolution), a framework that helps multimodal agents learn skills and policies simultaneously. This approach distills trajectories into reusable skills, which are updated throughout training. The policy model conditions on retrieved skills to guide its rollouts, while the skill library evolves using valuable rollouts generated by the policy. Experiments show SPyCE outperforms both RL-based and memory-based baselines across eight benc
Researchers propose SPyCE (Skill-Policy Co-evolution), a framework that helps multimodal agents learn skills and policies simultaneously. This approach distills trajectories into reusable skills, which are updated throughout training. The policy model conditions on retrieved skills to guide its rollouts, while the skill library evolves using valuable rollouts generated by the policy. Experiments show SPyCE outperforms both RL-based and memory-based baselines across eight benchmarks.
---
Why it matters: This matters because it provides a new paradigm for building capable multimodal agents that can learn from experience and improve over time. It has implications for applications such as robotics, autonomous vehicles, and human-computer interaction.
Source: https://arxiv.org/abs/2607.13854
This article was originally published at: https://arxiv.org/abs/2607.13854