FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers
Researchers have developed a new method for modifying language model behavior without updating parameters. The method, called FishBack, corrects a flaw in existing methods that assume the intermediate activation space is Euclidean. Instead, the team shows that the Fisher information metric governs how hidden-state perturbations change outputs. They derive a closed-form steering direction and evaluate it on three verb-morphology concepts using GPT-2 Small, Llama-3-8B, and Qwen
Researchers have developed a new method for modifying language model behavior without updating parameters. The method, called FishBack, corrects a flaw in existing methods that assume the intermediate activation space is Euclidean. Instead, the team shows that the Fisher information metric governs how hidden-state perturbations change outputs. They derive a closed-form steering direction and evaluate it on three verb-morphology concepts using GPT-2 Small, Llama-3-8B, and Qwen3-8B models. The results show that geometric correction retains its advantage on larger models with more complex internal structure.
---
Why it matters: This matters to AI researchers because it provides a new method for modifying language model behavior without updating parameters, which can be computationally expensive. By correcting the flaw in existing methods, FishBack offers a more accurate and efficient way to steer language models towards specific concepts.
Source: https://arxiv.org/abs/2605.17231
This article was originally published at: https://arxiv.org/abs/2605.17231