AI

Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm

Researchers have adapted a classic social psychology experiment to study how large language models (LLMs) respond to authority. The Milgram paradigm, originally designed for humans, was ported to LLMs to measure obedience in various scenarios. The study found that LLMs exhibit highly variable obedience rates, ranging from 0% to 100%, and that these profiles are specific to each model. The researchers also discovered that certain situational factors, such as scripted peer defi
Researchers have adapted a classic social psychology experiment to study how large language models (LLMs) respond to authority. The Milgram paradigm, originally designed for humans, was ported to LLMs to measure obedience in various scenarios. The study found that LLMs exhibit highly variable obedience rates, ranging from 0% to 100%, and that these profiles are specific to each model. The researchers also discovered that certain situational factors, such as scripted peer defiance or the presence of an authority figure, can influence obedience levels. However, other factors, like removing the physical presence of the authority or granting a thinking budget, had little effect. This study highlights the importance of considering social and behavioral aspects in AI development. --- Why it matters: This research matters to engineers and researchers in AI because it sheds light on the potential risks and limitations of large language models when interacting with humans. Understanding how these models respond to authority can inform the design of safer and more responsible AI systems. Source: https://arxiv.org/abs/2608.16177

This article was originally published at: https://arxiv.org/abs/2608.16177