Incoherent by Design? On the Moral Self-Consistency of LLMs
Researchers have found that large language models (LLMs) often fail to consistently apply moral principles across different scenarios. In a study of three major philosophical schools of thought - deontology, utilitarianism, and virtue ethics - the authors constructed sets of equivalent situations with varying frames and evaluated responses from multiple LLMs. The results showed substantial inconsistency, with up to 78% of contradictions across scenarios. This instability has
Researchers have found that large language models (LLMs) often fail to consistently apply moral principles across different scenarios. In a study of three major philosophical schools of thought - deontology, utilitarianism, and virtue ethics - the authors constructed sets of equivalent situations with varying frames and evaluated responses from multiple LLMs. The results showed substantial inconsistency, with up to 78% of contradictions across scenarios. This instability has real-world implications, as it can shape human reasoning and decision-making. The study suggests that demonstrating internal incoherence is a necessary step towards aligning AI systems with human values.
---
Why it matters: This research matters because it highlights the importance of ensuring that AI systems can consistently apply moral principles. If LLMs cannot reliably maintain coherence, they may inadvertently shape human reasoning and decision-making in unpredictable ways, undermining efforts to align AI with human values.
Source: https://arxiv.org/abs/2608.15354
This article was originally published at: https://arxiv.org/abs/2608.15354