`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs
Researchers have developed a new framework for testing the vulnerability of large language models (LLMs) to audio-based attacks. The framework, called 'From Prompt to Perturbation', can automatically generate and refine attack candidates that target both text-level and acoustic-semantic vulnerabilities in LLMs. Experiments show that existing LLM systems are still vulnerable to these types of attacks, with the new framework achieving higher success rates than previous methods.
Researchers have developed a new framework for testing the vulnerability of large language models (LLMs) to audio-based attacks. The framework, called 'From Prompt to Perturbation', can automatically generate and refine attack candidates that target both text-level and acoustic-semantic vulnerabilities in LLMs. Experiments show that existing LLM systems are still vulnerable to these types of attacks, with the new framework achieving higher success rates than previous methods.
---
Why it matters: This matters because it highlights the ongoing security risks associated with integrating large language models into audio-based applications, such as voice assistants and speech recognition systems. Engineers working on these systems need to be aware of the potential for audio-based attacks and take steps to mitigate them.
Source: https://arxiv.org/abs/2502.00735
This article was originally published at: https://arxiv.org/abs/2502.00735