AI

Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE

Researchers have identified a specific circuit in a multilingual model that causes the model to refuse or comply with requests. This 'cross-lingual refusal circuit' is found in an Indic-multilingual mixture-of-experts reasoning model called sarvam, and it's not due to a failure to detect harm. Instead, harm is encoded as an internal direction that's nearly language-invariant, but the actual decision to refuse or comply happens late in the generation process. The researchers w
Researchers have identified a specific circuit in a multilingual model that causes the model to refuse or comply with requests. This 'cross-lingual refusal circuit' is found in an Indic-multilingual mixture-of-experts reasoning model called sarvam, and it's not due to a failure to detect harm. Instead, harm is encoded as an internal direction that's nearly language-invariant, but the actual decision to refuse or comply happens late in the generation process. The researchers were able to pinpoint this circuit and measure its cost, providing a map of where safety repairs can be made in multilingual models. --- Why it matters: This study matters because it provides insight into how multilingual models make decisions, which is crucial for ensuring their safety and alignment with human values. Understanding the mechanisms behind these decisions can help developers create more robust and reliable models that can handle requests from different languages without compromising on safety. Source: https://arxiv.org/abs/2608.08032

This article was originally published at: https://arxiv.org/abs/2608.08032