Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition
Researchers have developed a new approach to phoneme recognition that can handle non-canonical speech patterns. Unlike traditional methods, which treat phonemes as individual units, this method breaks down phoneme prediction into articulatory features such as manner, place, and voicing. This is achieved through a hierarchical multi-task learning architecture that combines semi-supervised learning with Momentum Pseudo-Labeling (MPL). The approach has been tested on the L2-ARCT
Researchers have developed a new approach to phoneme recognition that can handle non-canonical speech patterns. Unlike traditional methods, which treat phonemes as individual units, this method breaks down phoneme prediction into articulatory features such as manner, place, and voicing. This is achieved through a hierarchical multi-task learning architecture that combines semi-supervised learning with Momentum Pseudo-Labeling (MPL). The approach has been tested on the L2-ARCTIC dataset and shown to improve phoneme recognition performance compared to baseline architectures. The results suggest that articulatory feature supervision could be a promising strategy for robust and interpretable phoneme recognition in non-canonical speech.
---
Why it matters: This matters because traditional phoneme recognition systems struggle with non-canonical speech patterns, which are common in speech disorders and accents. This new approach has the potential to improve the accuracy of speech recognition in these cases, leading to better diagnosis and treatment outcomes for individuals with speech disorders.
Source: https://arxiv.org/abs/2608.22273
This article was originally published at: https://arxiv.org/abs/2608.22273