VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models
Researchers have proposed a new method for controlling emotions in language models called VA-DPO. This approach allows for more precise emotional expression by specifying desired affect as a continuous point in the Valence-Arousal plane, rather than using discrete labels like 'happy' or 'angry'. The method is an extension of Direct Preference Optimization and involves training a model to hit a target affect with minimal distance. Initial results show significant improvements
Researchers have proposed a new method for controlling emotions in language models called VA-DPO. This approach allows for more precise emotional expression by specifying desired affect as a continuous point in the Valence-Arousal plane, rather than using discrete labels like 'happy' or 'angry'. The method is an extension of Direct Preference Optimization and involves training a model to hit a target affect with minimal distance. Initial results show significant improvements in mean VA distance to the target, with valence/arousal correlation lifted to 0.93 and 0.75 respectively.
---
Why it matters: This matters because it enables more nuanced emotional expression in language models, which could have applications in areas like human-computer interaction, chatbots, or even creative writing assistants.
Source: https://arxiv.org/abs/2608.20374
This article was originally published at: https://arxiv.org/abs/2608.20374