COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models
A new safety gate for multimodal large language models (MLLMs) called COMIC has been proposed to address a structural weakness in current defenses. Current multimodal jailbreaks can be safe when viewed individually but become unsafe when combined with visual targets. COMIC identifies the requested operation and reference type, constructs candidate targets, grounds plausible referents, and evaluates safety over explicit operation-target pairs. The results show that COMIC impro
A new safety gate for multimodal large language models (MLLMs) called COMIC has been proposed to address a structural weakness in current defenses. Current multimodal jailbreaks can be safe when viewed individually but become unsafe when combined with visual targets. COMIC identifies the requested operation and reference type, constructs candidate targets, grounds plausible referents, and evaluates safety over explicit operation-target pairs. The results show that COMIC improves robustness while preserving benign utility and practical efficiency.
---
Why it matters: This matters to researchers in AI because it highlights a critical flaw in current multimodal defenses and proposes a new approach to address the issue of reference-dependent failure modes.
Source: https://arxiv.org/abs/2608.17234
This article was originally published at: https://arxiv.org/abs/2608.17234