HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
Researchers have proposed a framework called HiViS to improve the efficiency of Vision-Language Models (VLMs). These models are used for tasks like image captioning and visual question answering. The issue with VLMs is that they contain redundant visual tokens, which can slow down processing. HiViS hides these unnecessary tokens from the model's drafting process, allowing it to focus on text-based information. This approach improves performance in terms of speed and accuracy.
Researchers have proposed a framework called HiViS to improve the efficiency of Vision-Language Models (VLMs). These models are used for tasks like image captioning and visual question answering. The issue with VLMs is that they contain redundant visual tokens, which can slow down processing. HiViS hides these unnecessary tokens from the model's drafting process, allowing it to focus on text-based information. This approach improves performance in terms of speed and accuracy. According to the authors, their method achieves significant improvements over existing methods.
---
Why it matters: This matters because VLMs are increasingly used for real-world applications, such as image captioning and visual question answering. Improving their efficiency can lead to faster and more accurate processing times, which is crucial for tasks that require quick decision-making or high-volume processing.
Source: https://arxiv.org/abs/2509.23928
This article was originally published at: https://arxiv.org/abs/2509.23928