Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs
Researchers have developed a method to use benign natural images as 'jailbreak anchors' to attack Vision-Language Models (VLMs). They created MemJack, a framework that uses these images to launch attacks on VLMs. The team found that with this approach, they can achieve an 71.48% success rate against one model, and up to 90% with more resources. This method is concerning because it shows that current safety-aligned VLMs have vulnerabilities.
Researchers have developed a method to use benign natural images as 'jailbreak anchors' to attack Vision-Language Models (VLMs). They created MemJack, a framework that uses these images to launch attacks on VLMs. The team found that with this approach, they can achieve an 71.48% success rate against one model, and up to 90% with more resources. This method is concerning because it shows that current safety-aligned VLMs have vulnerabilities.
---
Why it matters: This matters to researchers in AI because it highlights the potential for attacks on Vision-Language Models using everyday images, which could compromise their safety and security. The development of MemJack also raises questions about the robustness of these models and the need for more effective defensive measures.
Source: https://arxiv.org/abs/2604.12616
This article was originally published at: https://arxiv.org/abs/2604.12616