GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting
Researchers have developed a new framework called GS-VLA to improve the robustness of Vision-Language-Action (VLA) policies to changes in camera viewpoints. Current VLA performance relies on identical training and deployment camera configurations, but even small displacements can significantly reduce performance. The proposed framework uses Gaussian-based novel-view synthesis to adapt to viewpoint shifts without retraining the policy. This approach improves performance across
Researchers have developed a new framework called GS-VLA to improve the robustness of Vision-Language-Action (VLA) policies to changes in camera viewpoints. Current VLA performance relies on identical training and deployment camera configurations, but even small displacements can significantly reduce performance. The proposed framework uses Gaussian-based novel-view synthesis to adapt to viewpoint shifts without retraining the policy. This approach improves performance across various axes, including policy architectures, unseen task suites, and perturbation scales, with a lightweight visual module that recovers lost performance.
---
Why it matters: This matters because VLA policies are widely used in robotics and other applications where camera viewpoints can change unpredictably. The ability to adapt to these changes without retraining the policy is crucial for reliable deployment in real-world scenarios.
Source: https://arxiv.org/abs/2608.19066
This article was originally published at: https://arxiv.org/abs/2608.19066