EXIMO: VLM Guided Exploration of VLA Policies
Researchers from the University of Freiburg and the Max Planck Institute for Intelligent Systems have proposed an algorithm called EXIMO to efficiently fine-tune robot policies for new tasks. The current approach to robotic manipulation relies on large vision-language-action models, but collecting data for these models is expensive and time-consuming. EXIMO operates in three stages: exploration, imitation, and optimization. During the exploration phase, a vision language mode
Researchers from the University of Freiburg and the Max Planck Institute for Intelligent Systems have proposed an algorithm called EXIMO to efficiently fine-tune robot policies for new tasks. The current approach to robotic manipulation relies on large vision-language-action models, but collecting data for these models is expensive and time-consuming. EXIMO operates in three stages: exploration, imitation, and optimization. During the exploration phase, a vision language model acts as a planner to break down complex problems into shorter ones. The algorithm then collects an orchestrated dataset and fine-tunes the policy using residual off-policy reinforcement learning. In experiments, EXIMO outperformed existing approaches in terms of sample efficiency and final performance.
---
Why it matters: This matters because it addresses the challenge of finetuning large vision-language-action models for new tasks, which is a significant problem in robotic manipulation. Engineers can use EXIMO to improve the efficiency and effectiveness of their robot policies without relying on expensive human labor or sample-inefficient reinforcement learning methods.
Source: https://arxiv.org/abs/2608.19891
This article was originally published at: https://arxiv.org/abs/2608.19891