AI

Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation

Researchers tested whether large language models can report on their internal workings. They built a framework called Open-Weight Masked Introspection (OWMI) to intervene in model computations and ask the model about the changes made. The study found that none of eight open-weight models could accurately distinguish between real interventions and sham runs, suggesting that current models lack the ability to introspect. However, fine-tuning a model to report on this class of i
Researchers tested whether large language models can report on their internal workings. They built a framework called Open-Weight Masked Introspection (OWMI) to intervene in model computations and ask the model about the changes made. The study found that none of eight open-weight models could accurately distinguish between real interventions and sham runs, suggesting that current models lack the ability to introspect. However, fine-tuning a model to report on this class of intervention improved its performance significantly. --- Why it matters: This matters because it highlights the limitations of current language models in self-awareness, which is an essential aspect of human intelligence. Understanding how models process and represent internal states can improve their interpretability and trustworthiness. Source: https://arxiv.org/abs/2608.20569

This article was originally published at: https://arxiv.org/abs/2608.20569