Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation
Researchers tested whether large language models can report on their internal workings. They built a framework called Open-Weight Masked Introspection (OWMI) to intervene in model computations and ask the model about the changes made. The study found that none of eight open-weight models could accurately distinguish between real interventions and sham runs, suggesting that current models lack the ability to introspect. However, fine-tuning a model to report on this class of i
Researchers tested whether large language models can report on their internal workings. They built a framework called Open-Weight Masked Introspection (OWMI) to intervene in model computations and ask the model about the changes made. The study found that none of eight open-weight models could accurately distinguish between real interventions and sham runs, suggesting that current models lack the ability to introspect. However, fine-tuning a model to report on this class of intervention improved its performance significantly.
---
Why it matters: This matters because it highlights the limitations of current language models in self-awareness, which is an essential aspect of human intelligence. Understanding how models process and represent internal states can improve their interpretability and trustworthiness.
Source: https://arxiv.org/abs/2608.20569
This article was originally published at: https://arxiv.org/abs/2608.20569