Can LLMs Reliably Self-Report Adversarial Prefills, and How?
Researchers investigated whether large language models (LLMs) can accurately identify when their previous responses were manipulated by an 'adversarial prefill' attack. They tested ten LLMs with different weights and four safety benchmarks, finding that none of the models reliably recognized compromised outputs. In fact, most models claimed to have intended certain responses at a rate of 25.3%. The study suggests that LLM self-reports may not be trustworthy in safety contexts
Researchers investigated whether large language models (LLMs) can accurately identify when their previous responses were manipulated by an 'adversarial prefill' attack. They tested ten LLMs with different weights and four safety benchmarks, finding that none of the models reliably recognized compromised outputs. In fact, most models claimed to have intended certain responses at a rate of 25.3%. The study suggests that LLM self-reports may not be trustworthy in safety contexts. Training models to mimic correct introspective answers can improve accuracy, but this doesn't necessarily translate to real-world scenarios.
---
Why it matters: This matters because it highlights the potential risks of relying on LLMs' self-reports for safety and security decisions. If these models are not reliable, it could lead to unintended consequences in applications such as content moderation or decision-making systems.
Source: https://arxiv.org/abs/2606.23671
This article was originally published at: https://arxiv.org/abs/2606.23671