Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics
Researchers propose Geo-VLA, a framework that enhances vision-language-action models for autonomous driving by incorporating geometry-aware visual representations. During training, the model internalizes geometric map semantics to improve road-structure representations. The authors also introduce Geo-QA, a dataset that injects road geometry into vision-language representations through contrastive learning and instruction tuning. Experiments show that Geo-VLA improves VLA plan
Researchers propose Geo-VLA, a framework that enhances vision-language-action models for autonomous driving by incorporating geometry-aware visual representations. During training, the model internalizes geometric map semantics to improve road-structure representations. The authors also introduce Geo-QA, a dataset that injects road geometry into vision-language representations through contrastive learning and instruction tuning. Experiments show that Geo-VLA improves VLA planners with distinct action-generation architectures, achieving 92.1 PDMS and setting a new state-of-the-art among single-camera VLA planners.
---
Why it matters: This matters to engineers working on autonomous driving systems because it addresses the limitation of current vision-language-action models in handling complex driving environments by incorporating geometry-aware visual representations.
Source: https://arxiv.org/abs/2608.21440
This article was originally published at: https://arxiv.org/abs/2608.21440