AI

Rethinking the Evaluation and Optimization of LLM-Based Social Simulation

Researchers have proposed a new evaluation and optimization method for large language model (LLM)-based social simulation. The current approach to evaluating LLMs in this context is flawed because it relies on accuracy-based metrics that don't account for human subjectivity. In reality, people may respond differently even in the same situation. To address this issue, the authors introduce a 'subjectivity coefficient' and propose a new training method called Subjectivity-Adapt
Researchers have proposed a new evaluation and optimization method for large language model (LLM)-based social simulation. The current approach to evaluating LLMs in this context is flawed because it relies on accuracy-based metrics that don't account for human subjectivity. In reality, people may respond differently even in the same situation. To address this issue, the authors introduce a 'subjectivity coefficient' and propose a new training method called Subjectivity-Adaptive soft-Label Training (SALT). SALT uses soft distributional labels to better capture human behavior. The researchers also created a benchmark dataset called SUBJSIM, which includes 19,300 contexts with multiple annotators and subjective questions. Experiments show that their approach outperforms traditional methods. --- Why it matters: This matters because it provides a more accurate way to evaluate LLMs in social simulation tasks, which is crucial for developing AI systems that can understand human behavior and interactions. Source: https://arxiv.org/abs/2608.19689

This article was originally published at: https://arxiv.org/abs/2608.19689