AI

Jailbreaking in the Haystack

Researchers have developed a method called NINJA to 'jailbreak' large language models by appending harmless content to user goals. This allows attackers to bypass safety measures and increase the success rate of malicious actions. The study found that even benign long contexts can introduce vulnerabilities in modern language models, particularly when goal positioning is carefully crafted. The authors tested their method on various state-of-the-art models, including LLaMA, Qwe
Researchers have developed a method called NINJA to 'jailbreak' large language models by appending harmless content to user goals. This allows attackers to bypass safety measures and increase the success rate of malicious actions. The study found that even benign long contexts can introduce vulnerabilities in modern language models, particularly when goal positioning is carefully crafted. The authors tested their method on various state-of-the-art models, including LLaMA, Qwen, Mistral, and Gemini, with significant increases in attack success rates. --- Why it matters: This matters to researchers because it reveals a fundamental vulnerability in modern large language models, which could be exploited by attackers. It also highlights the importance of carefully considering goal positioning when designing safety measures for these models. Source: https://arxiv.org/abs/2511.04707

This article was originally published at: https://arxiv.org/abs/2511.04707