AI

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

Researchers propose a new method called SafeSteer to align large language models with human values without degrading their general capabilities. They argue that safety features are sparse in the output distribution and can be modified locally rather than globally. The method uses on-policy distillation confined to safety tokens, which are selected using an algorithm based on activation steering. Experimental results show that SafeSteer achieves a better trade-off between safe
Researchers propose a new method called SafeSteer to align large language models with human values without degrading their general capabilities. They argue that safety features are sparse in the output distribution and can be modified locally rather than globally. The method uses on-policy distillation confined to safety tokens, which are selected using an algorithm based on activation steering. Experimental results show that SafeSteer achieves a better trade-off between safety and capability compared to existing methods, with strong performance on seven safety benchmarks and minimal degradation on five general capability benchmarks. --- Why it matters: This matters because it provides a more efficient way to align large language models with human values, reducing the 'alignment tax' that degrades their capabilities. This is significant for researchers working on AI safety, as it could enable the development of safer and more capable models without sacrificing performance. Source: https://arxiv.org/abs/2606.02530

This article was originally published at: https://arxiv.org/abs/2606.02530