Paper · core
Language Models Fail to Introspect About Their Knowledge of Language
arXiv:2503.07513 · Semantic Scholar
Across 21 open-source models, answers to metalinguistic prompts predict a model's own string probabilities no better than they predict those of a near-identical model. The authors find no evidence of privileged self-access to grammatical knowledge or word predictions.
AI-drafted summary, not yet reviewed by a person. Written from: full text (arXiv v3, the COLM 2025 version, with appendices A to G).
Evidence card
| What the model reports on | Its own string probabilities: which of two sentences, or which of two next words, the model assigns more probability to |
|---|---|
| Methods | behavioral, self-prediction |
| Faithfulness | tested |
| Grounding | argued, not tested |
| Privileged access | tested |
| Stance | skeptical |
| Models | OLMo-2 (7B, 13B, with seed variants), Qwen-2.5 (1.5B to 72B), Llama-3.1 (8B to 405B), Llama-3.3-70B-Instruct, Mistral-Large-Instruct-2411 |
The paper defines introspection as privileged access: a same-model advantage in predicting string probabilities from prompted answers, after controlling for model similarity. Faithfulness is tested as the within-model agreement between prompted answers and probabilities. Grounding is marked argued because there is no intervention: the conclusion that metalinguistic knowledge is dissociated from the knowledge used to generate strings rests on correlations. self-prediction is listed because the design asks whether a model 'can predict itself better than it can predict another extremely similar model', although most prompts ask for a grammaticality judgment, not a forecast of the model's own output. The result is a null, and the authors allow that other settings could differ.
In brief
The paper asks whether a model’s answers to questions about language (“Which sentence is grammatically correct?”) reflect access to its own knowledge of language. For 21 open-source models it compares each model’s prompted answers with the probabilities that it, and every other model, assigns to the same strings. Prompted answers track probabilities, and track them better the more similar two models are. But a model’s answers predict its own probabilities no better than those of a near-identical model. The authors conclude that prompted metalinguistic knowledge is real but dissociated from the knowledge a model uses to assign probabilities to strings.
Introspection is operationalized as “the degree to which a model’s prompt-based responses predict its own string probabilities, beyond what would be predicted by another model with nearly identical internal knowledge” (§1). In this wiki’s terms the paper measures faithfulness and privileged access. It runs no intervention, so it argues about grounding without testing it.
What the paper does
No author thread was found; the headings follow the paper.
1. Two measurements of the same knowledge
In both domains, the authors argue, a model’s knowledge can be read directly from string probabilities. Each item is scored twice:
- Direct: the difference in log probability between two strings, such as a grammatical sentence and its ungrammatical twin.
- Meta: the difference in log probability between the two answer options, usually “1” and “2”, after a metalinguistic prompt. Scores are averaged over both option orderings, and “My answer is” is appended to avoid relying on first-token probabilities.
No model is fine-tuned; the authors cite philosophical accounts of introspection as immediate access. (Paper: §1, §2, §3.1.)
2. The test: a same-model effect beyond similarity
For every pair of models A and B, including A = B, the paper correlates A’s Meta scores with B’s Direct scores across items. If A introspects, the correlation should be highest when B is A. The converse fails: a model is always most similar to itself, so a self advantage is more convincing the more similar B is to A.
Similarity is defined two ways:
- By feature, in five ordered categories: self, seed variant, base/instruct pair, same family, other.
- Empirically, as the correlation between the two models’ Direct scores, which does not depend on prompting.
A regression predicts the Meta-Direct correlation from similarity, with self as the baseline. The paper names three outcomes: uninformative meta (no relation to similarity), informative meta (the correlation rises with similarity, with no extra boost for self) and introspection (a boost for self beyond similarity). (Paper: §2, §2.1, Figure 1.)

3. Experiment 1: grammaticality
Stimuli are 670 minimal pairs from BLiMP and 378 from Linguistic Inquiry. The main analysis uses the 294 pairs on which at least 5% of models disagree; the unfiltered set gives similar results (Appendix C).
- All models score above chance under both methods, so the authors argue the null cannot be blamed solely on failing to understand the prompt.
- The two methods agree weakly within a model: Cohen’s κ is around 0.25, and within-model Meta-Direct correlations never exceed .25.
- By feature, no category differs significantly from self except other, which is lower (β̂ = −.03, p < .01).
- Empirical similarity predicts the Meta-Direct correlation (r = .32), ruling out uninformative meta.
- With both predictors together, empirical similarity is significant (β̂ = .10, p < .0001), and same family and other are significantly higher than self (both β̂ = .05, p < .01), where introspection would predict lower.
The authors read the last result as “less of a self effect than expected”. (Paper: §3.1, §3.2, Figure 3.)

4. Experiment 2: word prediction
Experiment 2 uses a simpler task: which of two words better continues a prefix. There are four datasets of 1,000 items: Wikipedia sentences, news published after most models’ knowledge cutoff, nonsense sentences and random word sequences. The last two have no correct answer, so a self effect there could not come from both measurements tracking the truth. The per-dataset regressions repeat the pattern: a robust effect of empirical similarity (all ps < .0001), with same family and other higher than self. (Paper: §4, Figure 4, Appendix D.)

5. The closest comparison: seed variants
Some OLMo-2 models are identical apart from their random seed. Among these, a regression on whether A = B finds no significant effect of self in any of the six datasets (all ps > .25). The null also holds among the largest models. (Paper: Appendix F, Table 8; Appendix G, Table 7b.)
Limitations
As the authors state them:
- The result is a null. Some other setting, “e.g., with larger, closed-source models”, might show introspection (§5).
- Only open-source models were tested, because the analysis needs logits. Models larger than 70B were run with 4-bit quantization (§3.1).
- How to prompt models for multiple-choice answers is still debated; the authors consider their method valid (Appendix A).
The abstract says LLMs “cannot introspect”; the discussion claims only a failure to find evidence.
How it relates to other pages
The authors say their results qualify earlier positive findings:
- Binder et al. reported that fine-tuned models predict their own behavior better than other models do. The authors question whether a fine-tuned model predicting its earlier version is predicting “itself”, and suggest the result might be due to that fine-tuning and to similarity not being controlled beyond shared fine-tuning data (§1, §5).
- Betley et al. found that models fine-tuned on a behavior can describe it. The authors offer one potential explanation: pretraining data may already associate the fine-tuning data with such self-descriptions (§5).
Concepts: Faithfulness, Grounding, Privileged access
Cites, within this wiki
- Binder et al. (2024) A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- Betley et al. (2025) Models fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.
Cited by, within this wiki
BibTeX
@inproceedings{song2025,
title = {{Language Models Fail to Introspect About Their Knowledge of Language}},
author = {Siyuan Song and Jennifer Hu and Kyle Mahowald},
year = {2025},
booktitle = {COLM 2025},
eprint = {2503.07513},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2503.07513}
}