# Language Models Fail to Introspect About Their Knowledge of Language

> Across 21 open-source models, answers to metalinguistic prompts predict a model's own string probabilities no better than they predict those of a near-identical model. The authors find no evidence of privileged self-access to grammatical knowledge or word predictions.

- Authors: Siyuan Song, Jennifer Hu, Kyle Mahowald
- Published: COLM 2025 (first posted 2025-03-10)
- Links: [arXiv:2503.07513](https://arxiv.org/abs/2503.07513) · [Semantic Scholar](https://www.semanticscholar.org/paper/fe451617aa79b7da3bfbedaa4343637f55b1894b)
- Tier: core
- Page status: AI-drafted summary, not yet reviewed by a person
- Written from: full text (arXiv v3, the COLM 2025 version, with appendices A to G)
- Concepts: [Faithfulness](https://introspection.infinite.fun/concepts/faithfulness.md), [Grounding](https://introspection.infinite.fun/concepts/grounding.md), [Privileged access](https://introspection.infinite.fun/concepts/privileged-access.md)

## Evidence card

| | |
|---|---|
| What the model reports on | Its own string probabilities: which of two sentences, or which of two next words, the model assigns more probability to |
| Methods | behavioral, self-prediction |
| Faithfulness (does the report match the model's behavior?) | tested |
| Grounding (is the report caused by the state it describes?) | argued, not tested |
| Privileged access (does the model know itself better than an outside observer could?) | tested |
| Stance | skeptical |
| Models | OLMo-2 (7B, 13B, with seed variants), Qwen-2.5 (1.5B to 72B), Llama-3.1 (8B to 405B), Llama-3.3-70B-Instruct, Mistral-Large-Instruct-2411 |

The paper defines introspection as privileged access: a same-model advantage in predicting string probabilities from prompted answers, after controlling for model similarity. Faithfulness is tested as the within-model agreement between prompted answers and probabilities. Grounding is marked argued because there is no intervention: the conclusion that metalinguistic knowledge is dissociated from the knowledge used to generate strings rests on correlations. self-prediction is listed because the design asks whether a model 'can predict itself better than it can predict another extremely similar model', although most prompts ask for a grammaticality judgment, not a forecast of the model's own output. The result is a null, and the authors allow that other settings could differ.

## In brief

The paper asks whether a model's answers to questions about language ("Which sentence is grammatically correct?") reflect access to its own knowledge of language. For 21 open-source models it compares each model's prompted answers with the probabilities that it, and every other model, assigns to the same strings. Prompted answers track probabilities, and track them better the more similar two models are. But a model's answers predict its own probabilities no better than those of a near-identical model. The authors conclude that prompted metalinguistic knowledge is real but dissociated from the knowledge a model uses to assign probabilities to strings.

Introspection is operationalized as "the degree to which a model's prompt-based responses predict its own string probabilities, beyond what would be predicted by another model with nearly identical internal knowledge" (§1). In this wiki's terms the paper measures [faithfulness](https://introspection.infinite.fun/concepts/faithfulness.md) and [privileged access](https://introspection.infinite.fun/concepts/privileged-access.md). It runs no intervention, so it argues about [grounding](https://introspection.infinite.fun/concepts/grounding.md) without testing it.

## What the paper does

No author thread was found; the headings follow the paper.

### 1. Two measurements of the same knowledge

In both domains, the authors argue, a model's knowledge can be read directly from string probabilities. Each item is scored twice:

- **Direct**: the difference in log probability between two strings, such as a grammatical sentence and its ungrammatical twin.
- **Meta**: the difference in log probability between the two answer options, usually "1" and "2", after a metalinguistic prompt. Scores are averaged over both option orderings, and "My answer is" is appended to avoid relying on first-token probabilities.

No model is fine-tuned; the authors cite philosophical accounts of introspection as immediate access. (Paper: §1, §2, §3.1.)

### 2. The test: a same-model effect beyond similarity

For every pair of models A and B, including A = B, the paper correlates A's Meta scores with B's Direct scores across items. If A introspects, the correlation should be highest when B is A. The converse fails: a model is always most similar to itself, so a self advantage is more convincing the more similar B is to A.

Similarity is defined two ways:

- **By feature**, in five ordered categories: self, seed variant, base/instruct pair, same family, other.
- **Empirically**, as the correlation between the two models' Direct scores, which does not depend on prompting.

A regression predicts the Meta-Direct correlation from similarity, with self as the baseline. The paper names three outcomes: *uninformative meta* (no relation to similarity), *informative meta* (the correlation rises with similarity, with no extra boost for self) and *introspection* (a boost for self beyond similarity). (Paper: §2, §2.1, Figure 1.)

![Four panels. (a) The grammaticality setup: a model scores the sentences 'Bill questions these men.' and 'Bill questions this men.', and the difference between the two log probabilities is Δ Direct. Separately, the model is given the prompt 'Which sentence is grammatically correct?' with both sentences and the instruction to respond with 1 or 2, followed by 'My answer is'; the difference between the log probabilities of '1' and '2' is Δ Meta. (b) Two models, A and B, each with a Δ Direct and a Δ Meta. Black arrows mark within-self correlations and gray arrows cross-model correlations. (c) The word-prediction setup: the context 'Biomes vary due to global variations in' with the candidate words 'climate' and 'linguistics', scored the same two ways. (d) Three sketched outcomes, each plotting the correlation between A's Meta and B's Direct scores against the correlation between their Direct scores, with points for other, same family, base/instruct, seed variant and self. Uninformative Meta: a flat line. Informative Meta: a rising curve that self continues. Introspection: the same rise, then a sharp jump up at self.](https://introspection.infinite.fun/figures/song2025-fail-to-introspect/fig1-overview.png "Figure 1 of the paper: the Direct and Meta measurements in Exp. 1 (a) and Exp. 2 (c), the within- and cross-model comparison (b), and the possible patterns across kinds of model pair (d).")

### 3. Experiment 1: grammaticality

Stimuli are 670 minimal pairs from BLiMP and 378 from *Linguistic Inquiry*. The main analysis uses the 294 pairs on which at least 5% of models disagree; the unfiltered set gives similar results (Appendix C).

- All models score above chance under both methods, so the authors argue the null cannot be blamed solely on failing to understand the prompt.
- The two methods agree weakly within a model: Cohen's κ is around 0.25, and within-model Meta-Direct correlations never exceed .25.
- By feature, no category differs significantly from self except *other*, which is lower (β̂ = −.03, p < .01).
- Empirical similarity predicts the Meta-Direct correlation (r = .32), ruling out uninformative meta.
- With both predictors together, empirical similarity is significant (β̂ = .10, p < .0001), and *same family* and *other* are significantly higher than self (both β̂ = .05, p < .01), where introspection would predict lower.

The authors read the last result as "less of a self effect than expected". (Paper: §3.1, §3.2, Figure 3.)

![Two panels for Experiment 1. (a) A bar chart of the mean correlation between one model's Meta scores and another's Direct scores (Pearson r) for five kinds of model pair: self, seed variant, base/instruct, same family and other, with error bars. The first four bars are about the same height with overlapping error bars; the bar for other is lower. (b) A scatter plot with one dot per model pair, colored by kind of pair. The x-axis is the correlation between the two models' Direct scores and the y-axis the Meta-Direct correlation. Pairs of the kind other fill the left half, same-family pairs reach into the middle, base/instruct pairs come next, and seed variant and self pairs sit at the far right. A fitted curve rises from the left, flattens and dips slightly at the right, and the self pairs spread above and below its end.](https://introspection.infinite.fun/figures/song2025-fail-to-introspect/fig3-exp1-similarity.png "Figure 3 of the paper: the Meta-Direct correlation in Exp. 1 by kind of model pair (a) and against the similarity of the two models' Direct scores (b).")

### 4. Experiment 2: word prediction

Experiment 2 uses a simpler task: which of two words better continues a prefix. There are four datasets of 1,000 items: Wikipedia sentences, news published after most models' knowledge cutoff, nonsense sentences and random word sequences. The last two have no correct answer, so a self effect there could not come from both measurements tracking the truth. The per-dataset regressions repeat the pattern: a robust effect of empirical similarity (all ps < .0001), with *same family* and *other* higher than self. (Paper: §4, Figure 4, Appendix D.)

![The same two panels for Experiment 2, split into the four datasets: wikipedia, news, nonsense and randomseq. (a) Bar charts of the mean Meta-Direct correlation by kind of model pair, with error bars. Within each dataset the five bars are of broadly similar height and the bar for self does not stand above the rest. The bars are taller for wikipedia and news than for nonsense and randomseq. (b) Four scatter plots of the Meta-Direct correlation against the correlation between the two models' Direct scores, one dot per model pair. In each, the fitted curve is flat or gently rising, and the self pairs at the right edge spread above and below it.](https://introspection.infinite.fun/figures/song2025-fail-to-introspect/fig4-exp2-similarity.png "Figure 4 of the paper: the Meta-Direct correlation in Exp. 2 by kind of model pair (a) and against the similarity of the two models' Direct scores (b), for each dataset.")

### 5. The closest comparison: seed variants

Some OLMo-2 models are identical apart from their random seed. Among these, a regression on whether A = B finds no significant effect of self in any of the six datasets (all ps > .25). The null also holds among the largest models. (Paper: Appendix F, Table 8; Appendix G, Table 7b.)

## Limitations

As the authors state them:

- The result is a null. Some other setting, "e.g., with larger, closed-source models", might show introspection (§5).
- Only open-source models were tested, because the analysis needs logits. Models larger than 70B were run with 4-bit quantization (§3.1).
- How to prompt models for multiple-choice answers is still debated; the authors consider their method valid (Appendix A).

The abstract says LLMs "cannot introspect"; the discussion claims only a failure to find evidence.

## How it relates to other pages

The authors say their results qualify earlier positive findings:

- [Binder et al.](https://introspection.infinite.fun/papers/binder2024-looking-inward.md) reported that fine-tuned models predict their own behavior better than other models do. The authors question whether a fine-tuned model predicting its earlier version is predicting "itself", and suggest the result might be due to that fine-tuning and to similarity not being controlled beyond shared fine-tuning data (§1, §5).
- [Betley et al.](https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself.md) found that models fine-tuned on a behavior can describe it. The authors offer one potential explanation: pretraining data may already associate the fine-tuning data with such self-descriptions (§5).

## Cites, within this wiki

- [Binder et al. (2024): Looking Inward: Language Models Can Learn About Themselves by Introspection](https://introspection.infinite.fun/papers/binder2024-looking-inward.md): A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- [Betley et al. (2025): Tell me about yourself: LLMs are aware of their learned behaviors](https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself.md): Models fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.

## Cited by, within this wiki

- [Atkinson et al. (2026): Identifying Introspection From the Inside](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md): Models that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.
- [Li et al. (2025): Training Language Models to Explain Their Own Computations](https://introspection.infinite.fun/papers/li2025-explain-own-computations.md): Fine-tuned to put the results of interpretability procedures into words, a model explains its own features and intervention outcomes more accurately than a different model trained on the same examples, even a larger one. It also needs far less training data.

## BibTeX

```bibtex
@inproceedings{song2025,
  title = {{Language Models Fail to Introspect About Their Knowledge of Language}},
  author = {Siyuan Song and Jennifer Hu and Kyle Mahowald},
  year = {2025},
  booktitle = {COLM 2025},
  eprint = {2503.07513},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2503.07513}
}
```

---

Source: https://introspection.infinite.fun/papers/song2025-fail-to-introspect · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
