TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling
Researchers have developed a method called TempJail to break into large vision-language models (LVLMs) by manipulating subtitles in videos. They found that the effectiveness of previous text- and image-based jailbreak methods depends not only on the meaning of the information, but also on how it is presented over time. The team used this insight to create a black-box video-based jailbreak framework that constructs query-aligned dialogue-style subtitle sequences and optimizes
Researchers have developed a method called TempJail to break into large vision-language models (LVLMs) by manipulating subtitles in videos. They found that the effectiveness of previous text- and image-based jailbreak methods depends not only on the meaning of the information, but also on how it is presented over time. The team used this insight to create a black-box video-based jailbreak framework that constructs query-aligned dialogue-style subtitle sequences and optimizes their temporal scheduling. They tested TempJail on four LVLMs and two datasets, achieving an attack success rate 53-18 percentage points higher than the strongest baseline.
---
Why it matters: This matters because it shows how vulnerable large vision-language models are to attacks that manipulate video content. Engineers working with these models need to be aware of this vulnerability and consider implementing additional security measures.
Source: https://arxiv.org/abs/2608.19737
This article was originally published at: https://arxiv.org/abs/2608.19737