AI

ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

Researchers have developed a new AI model called ForeTime-VLA that can predict and manipulate moving objects on a conveyor belt. The model uses a combination of vision, language, and action to anticipate contact events and improve dynamic manipulation. It distills a future-aware representation from a frozen teacher model while remaining causal at inference. The model was tested on a dataset with 768 matched windows per split and showed a significant improvement in test MAE (m
Researchers have developed a new AI model called ForeTime-VLA that can predict and manipulate moving objects on a conveyor belt. The model uses a combination of vision, language, and action to anticipate contact events and improve dynamic manipulation. It distills a future-aware representation from a frozen teacher model while remaining causal at inference. The model was tested on a dataset with 768 matched windows per split and showed a significant improvement in test MAE (mean absolute error) and L2 loss. In real-robot evaluation, ForeTime-VLA achieved high grasp success rates compared to other models. --- Why it matters: This matters because it can improve the efficiency of dynamic manipulation tasks, such as picking up objects on a conveyor belt, by predicting future events and reducing the need for explicit imagination or video-scale teachers. Source: https://arxiv.org/abs/2608.20735

This article was originally published at: https://arxiv.org/abs/2608.20735