Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding
Researchers have found that speech models (SpeechLMs) can be contaminated by their own previous beliefs when interpreting audio streams. This 'previous-belief contamination' (PBC) causes the model to distort its current perception of emotions, leading to reduced accuracy and flipped predictions. The team proposes a training-free framework called EmoUpdate to address PBC, which separates current-audio perception from historical state revision through three components: an acous
Researchers have found that speech models (SpeechLMs) can be contaminated by their own previous beliefs when interpreting audio streams. This 'previous-belief contamination' (PBC) causes the model to distort its current perception of emotions, leading to reduced accuracy and flipped predictions. The team proposes a training-free framework called EmoUpdate to address PBC, which separates current-audio perception from historical state revision through three components: an acoustic firewall, a causal belief filter, and a decontamination operator.
---
Why it matters: This matters because it highlights the potential for speech models to be biased by their own previous predictions, leading to inaccurate emotion understanding. Engineers working on streaming emotion understanding need to consider this issue when designing and training their models.
Source: https://arxiv.org/abs/2608.20769
This article was originally published at: https://arxiv.org/abs/2608.20769