Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms
Researchers have created a benchmark for evaluating the safety of language models in Bengali, a language spoken by over 220 million people worldwide. The benchmark, called BanglaSafe, consists of 879 prompts that cover various culturally grounded harm categories and are designed to test the limits of current language models. Evaluations showed that many popular LLMs struggle with Bengali content, with over half of responses being unsafe or partially unsafe. Interestingly, the
Researchers have created a benchmark for evaluating the safety of language models in Bengali, a language spoken by over 220 million people worldwide. The benchmark, called BanglaSafe, consists of 879 prompts that cover various culturally grounded harm categories and are designed to test the limits of current language models. Evaluations showed that many popular LLMs struggle with Bengali content, with over half of responses being unsafe or partially unsafe. Interestingly, the researchers found that the writing style used in the prompt had a significant impact on the model's response, with formal requests being more likely to elicit harmful content than informal ones.
---
Why it matters: This matters because it highlights the limitations of current language models and their inability to generalize across languages and cultures. It also underscores the need for more culturally sensitive and linguistically diverse benchmarks to ensure that AI systems are safe and responsible in real-world applications.
Source: https://arxiv.org/abs/2608.22335
This article was originally published at: https://arxiv.org/abs/2608.22335