AI

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

Researchers have developed a new benchmark called Hear2Act to evaluate how prosody affects task-oriented dialogue. The benchmark includes 480 scenarios where users convey concerns through either words or prosody. Two audio-capable language models were tested using Hear2Act, and the results show that when lexical evidence is insufficient, prosody can significantly impact downstream decisions. However, this effect largely disappears when the concern is explicitly mentioned in t
Researchers have developed a new benchmark called Hear2Act to evaluate how prosody affects task-oriented dialogue. The benchmark includes 480 scenarios where users convey concerns through either words or prosody. Two audio-capable language models were tested using Hear2Act, and the results show that when lexical evidence is insufficient, prosody can significantly impact downstream decisions. However, this effect largely disappears when the concern is explicitly mentioned in the utterance. --- Why it matters: This matters to engineers working on AI assistants because it highlights the importance of considering prosody in dialogue systems, particularly when there's ambiguity or uncertainty in user input. The results suggest that audio-capable models can recover information from speech but need an explicit intermediate representation to reliably carry it into action. Source: https://arxiv.org/abs/2608.19515

This article was originally published at: https://arxiv.org/abs/2608.19515