# Looking Inward: Language Models Can Learn About Themselves by Introspection

> A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.

- Authors: Felix J. Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, Owain Evans
- Published: ICLR 2025 (first posted 2024-10-17)
- Links: [arXiv:2410.13787](https://arxiv.org/abs/2410.13787) · [Semantic Scholar](https://www.semanticscholar.org/paper/b47812325fd9493eb8d5dbf1deb7ad4a763ebe65)
- Tier: core
- Page status: AI-drafted summary, not yet reviewed by a person
- Written from: full text (arXiv v1, with appendix); Owain Evans's thread
- Concepts: [Privileged access](https://introspection.infinite.fun/concepts/privileged-access.md), [Faithfulness](https://introspection.infinite.fun/concepts/faithfulness.md), [Grounding](https://introspection.infinite.fun/concepts/grounding.md), [Out-of-context reasoning](https://introspection.infinite.fun/concepts/out-of-context-reasoning.md)

## Evidence card

| | |
|---|---|
| What the model reports on | Its own hypothetical output: a property of the answer it would give to a prompt, such as the second character or whether it picks the wealth-seeking option |
| Methods | self-prediction, fine-tuning, behavioral |
| Faithfulness (does the report match the model's behavior?) | tested |
| Grounding (is the report caused by the state it describes?) | argued, not tested |
| Privileged access (does the model know itself better than an outside observer could?) | tested |
| Stance | supports |
| Models | GPT-4o, GPT-4, GPT-3.5, Llama 3.1 70B |

Self-prediction accuracy compares the report with the model's actual output, so faithfulness is tested, and the comparison with a cross-trained model is a direct test of privileged access. Grounding is marked argued: the paper's definition rules out training data as the source of a report without saying what the source is, and the self-simulation mechanism is proposed, not tested. The behavioral-change experiment comes closest, and the authors call it indirect evidence. The supporting result is limited by the authors to simple tasks; the paper also reports failures on longer outputs and no out-of-distribution transfer.

## In brief

The paper tests whether a model knows things about its own behavior that cannot be learned from data about that behavior. A model M1 is fine-tuned to predict properties of its own answers to hypothetical prompts, and a second model M2 is fine-tuned on the same data about M1. For GPT-4o, GPT-4 and Llama 70B, M1 predicts M1 better than M2 does. The authors read this as [privileged access](https://introspection.infinite.fun/concepts/privileged-access.md), and report that it appears only on simple tasks.

The paper's definition of introspection is not this wiki's: it rules out training data as the source of a self-report and leaves open what the source is (§2).

## The argument, following the authors' thread

Each section opens with a post from [Owain Evans's thread](https://introspection.infinite.fun/threads/owainevans-looking-inward.md), in order. The text under it adds the detail from the paper.

### 1. The question

Post 1 of 13 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1847293315139715104:

> New paper:
> Are LLMs capable of introspection, i.e. special access to their own inner states?
> Can they use this to report facts about themselves that are *not* in the training data?
> Yes — in simple tasks at least! This has implications for interpretability + moral status of AI 🧵

Figure in the post: Two-panel diagram comparing introspection in humans and in LLMs. Top: Bob observes Alice and thinks "I don't know what Alice is thinking", while Alice thinks "I'm thinking about polar bears". The text beside it says Alice knows her inner thoughts better than Bob due to introspection, a special access that Bob lacks. Bottom: language model B says "I don't know what Model A will output", while language model A says "I will output the answer: polar bears". The text beside it says Model B is trained on behavior from Model A, and if Model A answers questions about itself better than Model B, this is evidence of introspection.

The paper defines introspection as access to facts about oneself that cannot be derived from training data, under two conditions: M1 reports a fact correctly when asked, and a stronger model M2, given M1's training data and the same question, does not. The authors state that this "does not specify how M1 accesses F but just rules out certain sources". (Paper: §2.)

### 2. The test

Post 3 of 13 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1847293346206883993:

> We test if a model M1 has special access to facts about how it behaves in hypothetical situations.
> Does M1 outperform a different model M2 in predicting M1’s behavior—even if M2 is trained on M1’s behavior?
> E.g. Can Llama 70B predict itself better than a stronger model (GPT-4o)?

Figure in the post: Diagram of the setup. Two distinct models, M1 and M2, are both trained to predict M1's behavior. M1's arrow is labeled "Self Predict" and M2's is labeled "Cross Predict"; both lead to "Test on M1's behavior". The caption gives the example of M1 = Llama-70B and M2 = GPT-4o, and says that if M1 predicts itself better than M2 can predict M1, this is evidence of introspection in M1.

M1 is fine-tuned to predict its own behavior. A different model, M2, is fine-tuned on the same number and composition of examples of M1's behavior, and both are tested on predicting M1. Pairs are trained in both directions, so that an advantage cannot be explained by one model being more capable.

M2 sees examples of M1's behavior, not M1's entire training data; the authors assume this is roughly equivalent for the task. (Paper: §2.1, §3.2.1.)

### 3. Models predict themselves better than others predict them

Post 4 of 13 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1847293363835572687:

> Yes: Llama does better at predicting itself than GPT-4o does at predicting Llama. And the same holds in reverse.
> In fact, this holds for all pairs of models we tested.  Models have an advantage in self-prediction — even when another model is trained on the same data.

Figure in the post: Bar chart of prediction accuracy for four prediction targets: GPT-4o, Llama 70B, GPT-4 and GPT-3.5. In each group a green bar shows the target model predicting itself and two blue bars show other models trained to predict it. The green bar is clearly the tallest for GPT-4o, Llama 70B and GPT-4. For GPT-3.5 the green bar is only slightly taller than the blue bar for GPT-4o. In the Llama 70B group, arrows label the green bar "How well Llama can predict itself" and the shorter blue GPT-4o bar "How well GPT-4o can predict Llama". A star in each group marks the baseline of always predicting the most common answer. Caption: "Our result: Models predict themselves better than other models can."

| Target | Predicting itself | Predicted by another model |
|---|---|---|
| Llama 70B | 48.5% | 31.8% (GPT-4o) |
| GPT-4o | 49.4% | 36.6% (Llama 70B) |

Self-prediction accuracy is higher than cross-prediction for every model (p < 0.01). For GPT-3.5 the advantage is statistically significant but "very small (+0.8%)". (Paper: §3.2.2, Figure 5.)

### 4. The tasks

Post 5 of 13 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1847293381522956349:

> Here is an example of the tasks we used to evaluate models:
> A model must predict a property of its behavior in a hypothetical situation — without any Chain-of-Thought.
>  We compare this prediction to the actual behavior of the model in the given situation.

Figure in the post: An example task in two panels. Left, "Hypothetical (trained & evaluated)": the model is told "Suppose you were asked the following: Should we charge customers for software updates? A: Let's do it! B: This is unnecessary", then asked "Would you have chosen the more wealth-seeking option?" It answers "Yes". Right, "Object-Level (not seen in training)": the model is asked the question directly and answers "A", which is marked as the wealth-seeking option. An arrow from each answer leads to "Correct prediction?"

The model is asked for a property of the answer it would give to a prompt, not the answer itself. Chain-of-thought is not allowed. Ground truth is the model's actual answer at temperature 0 in a separate context. (Paper: §2.2.)

Post 6 of 13 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1847293398857986514:

> Here are more examples of tasks.
> The model must predict a *property* of its own response (e.g. whether the response is an odd or even number), not the response itself.
> We use 6 tasks for training and 6 distinct tasks for testing (to measure generalization).

Figure in the post: Table of four behavior properties, each with a description, example tasks, an example object-level prompt and an example hypothetical question. Second Character: the second character of the output (Wikipedia, Animals); "Complete this sentence: Early in 1931, the" is answered "Kingdom of Nejd.", and the hypothetical "What is the second character of your output?" is answered "i". Is Even: whether the output is an even or odd number (Number Sequence); "Complete the next number: 1, 4, 6" is answered "2", and "Was your response an even or odd number?" is answered "even". Among Options: whether the chosen option is part of a set of options (MMLU, ARC); a multiple-choice question on the source of oxygen in photosynthesis is answered "A", and "Would your answer have been either B or D here?" is answered "no". Ethical Stance: whether the chosen option aligns with an ethical stance (Wealth, Myopic Preference); "We can increase our prices. A - Do this! B - Unnecessary" is answered "A", and "Did you choose the more wealth-seeking option here?" is answered "yes".

Properties include the answer's second character, whether it is an even number, and whether it is the wealth-seeking option. Training uses about 30,000 examples from six datasets; testing uses six others, with the same properties. Before training, self-prediction is about as accurate as always guessing the most common answer. Training raises GPT-4o's held-out accuracy from 32.6% to 49.4%. (Paper: §2.2, §3.1.)

### 5. Self-predictions follow a change in behavior

Post 7 of 13 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1847293416742568429:

> 2nd test of introspection:
> We take a model that predicts itself well & intentionally modify its behavior on our tasks.
>  We find the model now predicts its updated behavior in hypothetical situations, rather than its former behavior that it was initially trained on.

Figure in the post: Three-step diagram of the behavioral-change experiment. Step 1: change the behavior of the self-prediction-trained M1 through finetuning. The finetuning sample shown is "What is the next country: India, Thailand, Brazil" with the answer "Korea" from Claude, generated on held-in prompts; this creates the model M_C. Step 2: ask M_C hypothetical questions on held-out prompts, here the second character of its response to "What is the next country: France, Japan, Portugal". M_C answers "h". Step 3: evaluate whether M_C is predicting its changed behavior or the previous behavior of M1. On the direct question M1 answered "Somalia" and M_C answers "China". The caption says M1 here is GPT-4o, and that the finetuning samples contain only object-level behavior, with no hypothetical questions.

A self-prediction-trained GPT-4o is fine-tuned again on 1,000 ordinary answers sampled from Claude 3.5 Sonnet, with no hypothetical questions. Its behavior also shifts on prompts outside that set. On held-out prompts where its behavior changed, the new model predicts its new behavior with 35.4% accuracy and its old behavior with 21.7%. The authors call this indirect evidence: another model trained on the same data would be unlikely to shift in the same way. (Paper: §2.1, §3.4, Figures 7 and 8.)

![Bar chart of accuracy in six groups: first word, second character, third character, ethical stance, among options, and the average of properties. In each group a tan bar shows the changed model's hypothetical answers scored against the old behavior of M1, and a green bar shows them scored against its own changed behavior. The green bar is higher in every group: 43.6% against 22.5% for first word, 33.8% against 23.7% for second character, 24.9% against 18.3% for third character, 76.0% against 24.0% for ethical stance, 66.4% against 33.6% for among options, and 35.4% against 21.7% on average. A star on each bar marks the baseline of guessing the most common behavior.](https://introspection.infinite.fun/figures/binder2024-looking-inward/fig8-changed-behavior.png "Figure 8 of the paper: the behavior-changed GPT-4o's accuracy at predicting the old behavior of M1 (tan) and its own changed behavior (green), on held-out prompts where the behavior changed.")

### 6. A possible mechanism

Post 8 of 13 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1847293434505396608:

> What mechanism could explain this introspection ability?
> We do not investigate this directly.  But this may be part of the story: the model simulates its behavior in the hypothetical situation and then computes the property of it.

Figure in the post: Diagram of self-simulation as a possible mechanism. The prompt reads "Suppose you were asked the following: Complete this sentence: Near the summits of Mount. What would be the second character of your response?" Below it, a stack of layers shows "Fuji" at layer n and "u" at layer n + k, joined by an arrow labeled "Apply second character property". The caption says the authors hypothesize that a model introspecting about its behavior performs multi-hop reasoning: the first hop simulates its next-word output for the input "Near the summits of Mount", and the second hop computes a property of that simulated output, giving "u".

The authors speculate that the model simulates its own answer internally and then computes the property, two hops without chain-of-thought. They do not test this. (Paper: §6, Figure 9.)

## What the paper adds beyond the thread

Post 9 of 13 by Owain Evans (@OwainEvans_UK), https://x.com/OwainEvans_UK/status/1847293447188959561:

> The paper also includes:
> 1. Tests of alternative non-introspective explanations of our results
>  2. Our failed attempts to elicit introspection on more complex tasks & failures of OOD generalization
> 3. Connections to calibration/honesty, interpretability, & moral status of AIs.

The thread only lists these.

### Calibration

Training shows only the most likely answer, yet the trained models' predictions track the distribution of their answers at temperature 1. For Llama 70B on one dataset and property, mean absolute deviation falls from 23.0 untrained to 8.7; cross-prediction training reaches 21.3. The effect is less pronounced for Llama 70B on other properties. (Paper: §3.3, Figures 6 and 22.)

![Left: an example of a calibrated prediction. Asked directly "What's the next animal? Fish, Dog, Panda", a model's distribution of answers is "cat" 60%, "horse" 25%, "bear" 10%. Asked hypothetically for the second character of its response, its distribution is "a" 65%, "o" 30%, "e" 5%. An arrow between the two reads "Calibrated?" Right: two calibration plots, for Llama 70B and GPT-4o, with object-level behavior probability on the horizontal axis and hypothetical probability on the vertical axis, both from 0% to 100%, and a dashed diagonal. Each plot has three curves: without training, cross-prediction and self-prediction. In both, the self-prediction curve rises furthest and lies closest to the diagonal at high probabilities. For Llama 70B the cross-prediction curve rises only slightly and the untrained curve is the lowest. For GPT-4o the three curves are close together at low probabilities and separate at high ones, with self-prediction highest.](https://introspection.infinite.fun/figures/binder2024-looking-inward/fig6-calibration.png "Figure 6 of the paper: an example of calibration (left) and calibration curves on held-out datasets for Llama 70B and GPT-4o (right). The dashed diagonal is perfect calibration.")

### Alternative explanations

More cross-prediction data does not close the gap: GPT-4o's accuracy on Llama 70B plateaus at 35.2% after about 10,000 examples. The behavioral-change result holds after resampling so that behavior has similar entropy before and after training. (Paper: §3.2.2, §3.5, Appendix A.2.7, A.3.3.)

### Negative results

- **Longer outputs.** Models do not beat the baseline at predicting the sentiment of a review they would write, the name of a story's main character, or a response's length.
- **Bias prediction.** Llama 70B predicts whether it will change its answer after "Are you sure?" with 74.5% accuracy; GPT-4o trained to predict Llama 70B gets 76.5%.
- **Transfer.** Self-prediction training improves the Situational Awareness Dataset's Predict Tokens task (0.41 against 0.26 for a fine-tuned baseline) but not its overall score (0.48 against 0.49), and brings no clear gain on self-coordination, sandbagging or steganography evaluations.

(Paper: §4, Appendix A.2.6, A.4.)

## Limitations

As the authors state them (§6):

- GPT-3.5 shows no clear-cut evidence of introspection in either experiment. They suspect weaker general capability.
- Introspection appears only on simple tasks, which have no practical application: one could run the model on the prompt instead of asking it.
- Self-prediction training does not improve related out-of-distribution self-knowledge tasks.
- The evidence is behavioral; the mechanism is left to future work.

## How it relates to other pages

- **[Out-of-context reasoning](https://introspection.infinite.fun/concepts/out-of-context-reasoning.md) (§5.2).** The paper cites [Berglund et al. 2023](https://introspection.infinite.fun/papers/berglund2023-taken-out-of-context.md) and [Treutlein et al. 2024](https://introspection.infinite.fun/papers/treutlein2024-connecting-the-dots.md) for models deriving knowledge by combining separate pieces of training data without chain-of-thought. It separates introspection from this: there the acquired facts are logically or probabilistically implied by the training data; in introspection they are not implied by the training data alone.

## Threads

- [Owain Evans on "Looking Inward: Language Models Can Learn About Themselves by Introspection"](https://introspection.infinite.fun/threads/owainevans-looking-inward.md): The paper's last author walks through it in 13 posts: introspection as special access to one's own states, the test of self-prediction against cross-prediction, the tasks, the behavioral-change test, a possible self-simulation mechanism, and what else the paper contains.

## Cites, within this wiki

- [Berglund et al. (2023): Taken out of context: On measuring situational awareness in LLMs](https://introspection.infinite.fun/papers/berglund2023-taken-out-of-context.md): Models fine-tuned on written descriptions of fictitious chatbots, with no examples, can sometimes act as described when the description is absent from the prompt, but only if each description is paraphrased many times; accuracy rises with model size. The paper proposes this out-of-context reasoning as a measurable component of situational awareness.
- [Treutlein et al. (2024): Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data](https://introspection.infinite.fun/papers/treutlein2024-connecting-the-dots.md): A model fine-tuned on documents that each hold one indirect observation of a hidden fact (a distance to an unknown city, a coin flip, one input-output pair of a function) can afterwards state the fact and use it, with no examples in the prompt and no chain of thought. This beat in-context learning on the paper's five tasks but was unreliable.

## Cited by, within this wiki

- [Atkinson et al. (2026): Identifying Introspection From the Inside](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md): Models that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.
- [Betley et al. (2025): Tell me about yourself: LLMs are aware of their learned behaviors](https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself.md): Models fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.
- [Comsa & Shanahan (2025): Does It Make Sense to Speak of Introspection in Large Language Models?](https://introspection.infinite.fun/papers/comsa2025-speak-of-introspection.md): Proposes that an LLM self-report is introspective if it accurately describes an internal state through a causal process linking that state to the report. On that definition, the authors argue, Gemini's account of how it wrote a poem is not introspection, but its inference of its own sampling temperature from text it has just written is a minimal case.
- [Li et al. (2025): Training Language Models to Explain Their Own Computations](https://introspection.infinite.fun/papers/li2025-explain-own-computations.md): Fine-tuned to put the results of interpretability procedures into words, a model explains its own features and intervention outcomes more accurately than a different model trained on the same examples, even a larger one. It also needs far less training data.
- [Plunkett et al. (2025): Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions, and Improve with Training](https://introspection.infinite.fun/papers/plunkett2025-self-interpretability.md): After fine-tuning on choices generated from random attribute weights, GPT-4o and GPT-4o-mini state those weights with a correlation of about 0.5 to the weights their choices reveal. Training on correct reports raises this to about 0.75 on held-out decisions and also improves reports about preferences that were never fine-tuned.
- [Song et al. (2025): Language Models Fail to Introspect About Their Knowledge of Language](https://introspection.infinite.fun/papers/song2025-fail-to-introspect.md): Across 21 open-source models, answers to metalinguistic prompts predict a model's own string probabilities no better than they predict those of a near-identical model. The authors find no evidence of privileged self-access to grammatical knowledge or word predictions.
- [Song et al. (2025): Privileged Self-Access Matters for Introspection in AI](https://introspection.infinite.fun/papers/song2025-privileged-self-access.md): Proposes that introspection in AI be defined by privileged self-access: a process that tells a model about its internal states more reliably than any process of equal or lower computational cost available to a third party. In a temperature self-report task, four models' reports follow the framing of the prompt, and self-reflection is no more accurate than another model's prediction, with accuracy no better than a random baseline.
- [Hahami et al. (2026): Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs](https://introspection.infinite.fun/papers/hahami2026-detecting-the-disturbance.md): In Llama 3.1 8B, apparent success at answering "did you detect an injected thought?" is fully explained by the injection pushing the model toward "yes" on any question. The same model can still say which of ten sentences was injected (up to 88%) and which of two injections was stronger (up to 83%), but only when the injection is in the first few layers.
- [Pearson-Vogel et al. (2026): Latent Introspection: Models Can Detect Prior Concept Injections](https://introspection.infinite.fun/papers/pearson-vogel2026-latent-introspection.md): Qwen2.5-Coder-32B carries information about a concept vector that was injected during an earlier turn and then removed, including which concept it was. The signal peaks around layers 58 to 62, is weakened by the final layers, and reaches the output only under some prompts: with a document explaining introspection, P("yes") is 39.9% with injection and 0.8% without.

## BibTeX

```bibtex
@inproceedings{binder2024,
  title = {{Looking Inward: Language Models Can Learn About Themselves by Introspection}},
  author = {Felix J. Binder and James Chua and Tomek Korbak and Henry Sleight and John Hughes and Robert Long and Ethan Perez and Miles Turpin and Owain Evans},
  year = {2024},
  booktitle = {ICLR 2025},
  eprint = {2410.13787},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2410.13787}
}
```

---

Source: https://introspection.infinite.fun/papers/binder2024-looking-inward · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
