EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints
Researchers have proposed a framework called EcoVLA for energy-efficient inference of Vision-Language-Action models in real-time. These models are used in Embodied AI and require significant computational resources. The authors suggest that device-edge co-inference is a promising solution, but existing approaches have limitations. EcoVLA introduces a unified abstraction over different VLA paradigms and formulates a joint latency and energy prediction model to enable rapid eva
Researchers have proposed a framework called EcoVLA for energy-efficient inference of Vision-Language-Action models in real-time. These models are used in Embodied AI and require significant computational resources. The authors suggest that device-edge co-inference is a promising solution, but existing approaches have limitations. EcoVLA introduces a unified abstraction over different VLA paradigms and formulates a joint latency and energy prediction model to enable rapid evaluation of candidate schemes. It continuously selects the energy-optimal scheme satisfying real-time constraints with millisecond-level overhead. Experimental results show that EcoVLA improves system energy efficiency by up to 236% compared to existing co-inference approaches.
---
Why it matters: This research matters because it addresses a significant challenge in deploying Embodied AI systems, which require both real-time control and energy efficiency. The proposed framework has the potential to enable more widespread adoption of these models in robotic systems.
Source: https://arxiv.org/abs/2608.15502
This article was originally published at: https://arxiv.org/abs/2608.15502