AI

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades

Researchers analyzed how errors from automatic speech recognition (ASR) systems propagate through ASR-LLM cascades in Korean spoken question answering. They found that the impact of these errors is consistent across different language models and can be attributed to single-character transcription mistakes, which can significantly affect downstream performance. The study also suggests that using a large audio language model may mitigate this issue by directly processing audio
Researchers analyzed how errors from automatic speech recognition (ASR) systems propagate through ASR-LLM cascades in Korean spoken question answering. They found that the impact of these errors is consistent across different language models and can be attributed to single-character transcription mistakes, which can significantly affect downstream performance. The study also suggests that using a large audio language model may mitigate this issue by directly processing audio input. --- Why it matters: This research matters because it highlights the importance of accurate ASR in spoken question answering systems, particularly for languages like Korean where small transcription errors can have significant consequences. Understanding how to address these issues is crucial for improving the performance and reliability of such systems. Source: https://arxiv.org/abs/2605.17443

This article was originally published at: https://arxiv.org/abs/2605.17443