AI

Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs

Researchers have found that current safety alignment training for Large Language Models (LLMs) is heavily biased towards English. This can lead to stereotype-reinforcing outputs and propagation of harmful biases to non-English speaking communities. A new benchmark, called INCLUDE, has been introduced to address this cross-lingual gap. The benchmark evaluates LLMs against socio-cultural biases in six languages, including Hindi, Bengali, and Marathi. The results show that Benga
Researchers have found that current safety alignment training for Large Language Models (LLMs) is heavily biased towards English. This can lead to stereotype-reinforcing outputs and propagation of harmful biases to non-English speaking communities. A new benchmark, called INCLUDE, has been introduced to address this cross-lingual gap. The benchmark evaluates LLMs against socio-cultural biases in six languages, including Hindi, Bengali, and Marathi. The results show that Bengali models have the highest bias scores, while English models have lower bias scores when open-sourced but higher bias scores when closed-sourced. --- Why it matters: This matters to AI researchers because it highlights a critical failure mode in LLMs deployed in multilingual environments. It also underscores the need for more inclusive and culturally sensitive training data to prevent the propagation of biases. Source: https://arxiv.org/abs/2608.18131

This article was originally published at: https://arxiv.org/abs/2608.18131