AI

Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda

Large Language Models (LLMs) are expanding their capabilities beyond code completion to participate in security-sensitive workflows. However, the evidence for these systems is divided between software engineering evaluations and software security evaluations. A new survey synthesizes existing work across various tasks, adaptation mechanisms, and evaluation design. The review highlights the limitations of current approaches and identifies recurring validity threats. It conclud
Large Language Models (LLMs) are expanding their capabilities beyond code completion to participate in security-sensitive workflows. However, the evidence for these systems is divided between software engineering evaluations and software security evaluations. A new survey synthesizes existing work across various tasks, adaptation mechanisms, and evaluation design. The review highlights the limitations of current approaches and identifies recurring validity threats. It concludes that LLM capability should be judged as an assurance case supported by task-appropriate evidence, rather than a single benchmark score. --- Why it matters: This matters to AI researchers because it highlights the need for more comprehensive evaluation frameworks that consider both functional correctness and security. The survey's findings have implications for the development of more robust and trustworthy LLMs. Source: https://arxiv.org/abs/2608.21107

This article was originally published at: https://arxiv.org/abs/2608.21107