When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
Researchers have found that current audio-language models are not effectively using additional clinical context to improve speech recognition for people with dysarthric speech. They created a benchmark to test how well these models perform when given various types of clinical information, such as diagnosis labels and clinician-derived ratings. The results show that while some models can slightly improve performance with this context, they often degrade or have no significant
Researchers have found that current audio-language models are not effectively using additional clinical context to improve speech recognition for people with dysarthric speech. They created a benchmark to test how well these models perform when given various types of clinical information, such as diagnosis labels and clinician-derived ratings. The results show that while some models can slightly improve performance with this context, they often degrade or have no significant impact. However, the study also found that fine-tuning certain models using a mixture of clinical prompt formats can lead to significant improvements in speech recognition accuracy.
---
Why it matters: This research matters because it highlights the limitations of current audio-language models in handling atypical speech and emphasizes the need for more inclusive ASR systems. Engineers working on these models will want to consider how to better leverage multimodal context to improve performance, particularly for individuals with dysarthric speech.
Source: https://arxiv.org/abs/2605.02782
This article was originally published at: https://arxiv.org/abs/2605.02782