AI

A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications

A new framework has been proposed to evaluate the reliability of using large language models (LLMs) as substitutes for human experts in security surveys. The researchers used responses from Security Operations Centre professionals and found that while LLMs produce consistent answers, they systematically diverge from human opinions, exhibiting reduced variance and central tendency bias. This suggests that LLMs are not suitable replacements for expert elicitation, but can be us
A new framework has been proposed to evaluate the reliability of using large language models (LLMs) as substitutes for human experts in security surveys. The researchers used responses from Security Operations Centre professionals and found that while LLMs produce consistent answers, they systematically diverge from human opinions, exhibiting reduced variance and central tendency bias. This suggests that LLMs are not suitable replacements for expert elicitation, but can be useful for generating hypotheses or piloting survey questions. --- Why it matters: This matters to AI researchers because it highlights the limitations of using LLM-generated responses in security surveys, which could lead to inaccurate results if not properly evaluated and used. It also provides a framework for evaluating the reliability of such models, which is essential for their adoption in various applications. Source: https://arxiv.org/abs/2608.16893

This article was originally published at: https://arxiv.org/abs/2608.16893