Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints
Researchers have been trying to understand what large language models (LLMs) really know and how they use that knowledge. To do this, they looked at six frozen LLMs and tested four properties: can the model decode local geometric relations? Can it generate correct answers for those relations? Does the model's behavior change when given new information? And can the model be controlled to produce specific outputs? The study found that while pretraining improves the model's abil
Researchers have been trying to understand what large language models (LLMs) really know and how they use that knowledge. To do this, they looked at six frozen LLMs and tested four properties: can the model decode local geometric relations? Can it generate correct answers for those relations? Does the model's behavior change when given new information? And can the model be controlled to produce specific outputs? The study found that while pretraining improves the model's ability to decode local relations, much of its knowledge about sketch-level constraints is already present in the initial representation. However, the model often fails to express this knowledge or control its output. This means that just because a model can 'decode' information doesn't mean it will use that information correctly.
---
Why it matters: This study matters to researchers and engineers working on large language models because it highlights the limitations of current understanding of how these models work. By distinguishing between what models encode and what they can actually do, this research provides valuable insights for improving model performance and behavior.
Source: https://arxiv.org/abs/2608.17843
This article was originally published at: https://arxiv.org/abs/2608.17843