AI

Structured but Fragile: On the Limits of LLMs in Cybersecurity Decision-Making

Researchers studied whether large language models (LLMs) can perform structured security reasoning in cybersecurity workflows. They used real-world threat scenarios to test LLMs' ability to select security controls and minimize attacker success. The results show that LLMs are competent when given explicit attack-graph structure, but their capabilities are fragile and sensitive to framing. Small changes in prompts or relabeling strategies can significantly alter rankings. This
Researchers studied whether large language models (LLMs) can perform structured security reasoning in cybersecurity workflows. They used real-world threat scenarios to test LLMs' ability to select security controls and minimize attacker success. The results show that LLMs are competent when given explicit attack-graph structure, but their capabilities are fragile and sensitive to framing. Small changes in prompts or relabeling strategies can significantly alter rankings. This has implications for the design of AI-assisted security decision-support systems. --- Why it matters: These findings matter because they highlight the limitations of LLMs in cybersecurity decision-making. Engineers and researchers need to understand these limitations to develop more robust AI-assisted security systems that can effectively support human decision-makers. Source: https://arxiv.org/abs/2608.20966

This article was originally published at: https://arxiv.org/abs/2608.20966