AI

Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks

Researchers have created a benchmark to evaluate the safety of code generated by large language models in real-world tasks. The benchmark, called SUSVIBES, consists of 186 feature requests from open-source projects that contain vulnerable implementations. Twelve widely used coding agents were tested on this benchmark and found to perform poorly in terms of software security. While some agents produced functionally correct solutions, only a small percentage were secure. The st
Researchers have created a benchmark to evaluate the safety of code generated by large language models in real-world tasks. The benchmark, called SUSVIBES, consists of 186 feature requests from open-source projects that contain vulnerable implementations. Twelve widely used coding agents were tested on this benchmark and found to perform poorly in terms of software security. While some agents produced functionally correct solutions, only a small percentage were secure. The study's findings raise concerns about the adoption of 'vibe coding' in security-sensitive applications. --- Why it matters: This matters because it highlights potential risks associated with using AI-generated code in critical systems, which could have serious consequences if vulnerabilities are not addressed. Source: https://arxiv.org/abs/2512.03262

This article was originally published at: https://arxiv.org/abs/2512.03262