Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
Researchers have explored how theory of mind and prosocial beliefs can be used to align the behaviors of large language models (LLMs) with human norms in negotiation tasks. They initialized LLM agents with different prosocial beliefs and reasoning methods, including chain of thought and varying levels of theory-of-mind reasoning. The study involved 2,700 simulations using various LLMs, including o3-mini and DeepSeek-R1 Distilled Qwen 32B. The results showed that theory-of-min
Researchers have explored how theory of mind and prosocial beliefs can be used to align the behaviors of large language models (LLMs) with human norms in negotiation tasks. They initialized LLM agents with different prosocial beliefs and reasoning methods, including chain of thought and varying levels of theory-of-mind reasoning. The study involved 2,700 simulations using various LLMs, including o3-mini and DeepSeek-R1 Distilled Qwen 32B. The results showed that theory-of-mind reasoning enhances behavioral alignment with human norms, decision-making consistency, and negotiation outcomes. Fair proposers and responders accepting offers were the most consistent with their strategic reasonings, while all agents showed strong consistencies with human beliefs when rejecting offers. The study's findings advance understanding of theory of mind's role in human-AI interaction and cooperative decision-making.
---
Why it matters: This research matters because it can help improve the negotiation capabilities of AI systems, making them more effective collaborators in complex social interactions.
Source: https://arxiv.org/abs/2505.24255
This article was originally published at: https://arxiv.org/abs/2505.24255