SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization
Researchers have developed SAEVerbalizer, a framework that generates explanations for features extracted by sparse autoencoders (SAEs) from large language model (LLM) representations. This is achieved through a process of representation verbalization, where the SAE decoder directions are injected into the LLM's representations and fine-tuned to produce natural-language explanations. The approach addresses limitations in current methods, which rely on external observation and
Researchers have developed SAEVerbalizer, a framework that generates explanations for features extracted by sparse autoencoders (SAEs) from large language model (LLM) representations. This is achieved through a process of representation verbalization, where the SAE decoder directions are injected into the LLM's representations and fine-tuned to produce natural-language explanations. The approach addresses limitations in current methods, which rely on external observation and can be computationally inefficient. Experiments show that the learned verbalization capability generalizes to unseen features and transfers across separately trained SAE dictionaries.
---
Why it matters: This matters because it enables more efficient and effective explanation of complex AI models' behavior, which is crucial for understanding their decision-making processes and building trust in their outputs.
Source: https://arxiv.org/abs/2608.13538
This article was originally published at: https://arxiv.org/abs/2608.13538