AI

The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP

Researchers have developed a method to detect bias in mental health natural language processing (NLP) models. They propose the Divergence Hypothesis, which suggests that NLP models can be misled by 'lexical interference' - where human annotators and automated labeling systems prioritize different linguistic features. The team introduces TSS, a diagnostic framework that breaks down text into three channels: lexical character n-grams, morpho-syntactic features, and psycholingui
Researchers have developed a method to detect bias in mental health natural language processing (NLP) models. They propose the Divergence Hypothesis, which suggests that NLP models can be misled by 'lexical interference' - where human annotators and automated labeling systems prioritize different linguistic features. The team introduces TSS, a diagnostic framework that breaks down text into three channels: lexical character n-grams, morpho-syntactic features, and psycholinguistic style. They found that adding lexical features to the style channel can reduce accuracy on human-labeled data but not on auto-labeled data. To address this issue, they developed a statistic called Degree of Divergence (DoD), which measures the difference in performance between models trained on different labeling sources. The study suggests that NLP models may be prone to 'shortcut learning' and recommends using TSS as a diagnostic tool to audit label-source bias before making generalization claims. --- Why it matters: This matters because mental health NLP models can have serious consequences if they are biased or inaccurate. Detecting and mitigating these issues is crucial for ensuring the reliability of such models, particularly in high-stakes applications like clinical decision-making. Source: https://arxiv.org/abs/2608.20353

This article was originally published at: https://arxiv.org/abs/2608.20353