A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models
Researchers have compared five different methods to defend against 'vision-language model jailbreaks', which occur when a model is tricked into producing incorrect results. The study found that no single method works best across all models and architectures. However, one approach, called the 'image conditioning shift', performed well on certain models and was also non-transferable between different architectures. This suggests that defences should be tailored to specific lang
Researchers have compared five different methods to defend against 'vision-language model jailbreaks', which occur when a model is tricked into producing incorrect results. The study found that no single method works best across all models and architectures. However, one approach, called the 'image conditioning shift', performed well on certain models and was also non-transferable between different architectures. This suggests that defences should be tailored to specific language decoder families.
---
Why it matters: This matters because it highlights the need for more targeted and architecture-specific defence strategies against vision-language model attacks. Engineers working in this field will need to consider these findings when developing new defence methods.
Source: https://arxiv.org/abs/2607.27910
This article was originally published at: https://arxiv.org/abs/2607.27910