Trust Stack for Mental Health AI: A Survey of Calibration across Human, Interaction, and AI Layers
Researchers propose a three-layer framework for building trustworthy AI systems that support mental health. The framework separates human-oriented trust from interaction- and AI-level trustworthiness. A survey of 61 papers identifies a calibration gap between what users perceive as trustworthy and actual safety, highlighting the need to shift focus from maximizing perceived trust to calibrating it with demonstrated trustworthiness. This work aims to bridge the gap between dif
Researchers propose a three-layer framework for building trustworthy AI systems that support mental health. The framework separates human-oriented trust from interaction- and AI-level trustworthiness. A survey of 61 papers identifies a calibration gap between what users perceive as trustworthy and actual safety, highlighting the need to shift focus from maximizing perceived trust to calibrating it with demonstrated trustworthiness. This work aims to bridge the gap between different communities working on mental health AI, including NLP, HCI, and regulatory fields.
---
Why it matters: This research matters because it addresses a critical issue in mental health AI: ensuring that users trust the systems they interact with. By developing a framework for trustworthy AI, engineers can create more effective and safe support systems for individuals struggling with mental health issues.
Source: https://arxiv.org/abs/2604.20166
This article was originally published at: https://arxiv.org/abs/2604.20166