Paper · core
Training Language Models to Explain Their Own Computations
arXiv:2511.08579 · Semantic Scholar
Fine-tuned to put the results of interpretability procedures into words, a model explains its own features and intervention outcomes more accurately than a different model trained on the same examples, even a larger one. It also needs far less training data.
AI-drafted summary, not yet reviewed by a person. Written from: full text (arXiv v3, 9 Feb 2026, with appendices A to H).
Evidence card
| What the model reports on | A target model's internals as measured by three interpretability procedures: what a residual-stream feature encodes, how patching an activation changes the output, and how removing a hint from the input changes the answer |
|---|---|
| Methods | fine-tuning, self-prediction, patching, ablation |
| Faithfulness | tested |
| Grounding | argued, not tested |
| Privileged access | tested |
| Stance | supports |
| Models | Llama-3.1-8B, Llama-3.1-8B-Instruct, Llama-3-8B, Llama-3.1-70B, Qwen3-8B, Gemma-2-9B, Gemma-2-9B-Instruct |
Faithfulness is scored against the output of an interpretability procedure, and the ability is trained in: untrained baselines score far lower. Privileged access is tested as a same-model advantage over other trained explainers, and close variants of the target do about as well as the target itself on feature descriptions. Grounding is marked argued: the authors attribute the advantage to access to internals and support it with a correlation between activation similarity and explainer score, but no experiment traces what causes a given explanation. The explainer is a fine-tuned copy describing the frozen original, which the authors call self-explanation in a looser sense. Patching and ablation are listed as methods because they supply the ground truth; self-prediction because two tasks ask the model to predict its own output under an intervention.
In brief
The paper asks whether a language model can be trained to describe its own internal computations, and whether it does so better than a different model trained on the same examples. The authors take the outputs of three interpretability procedures as ground truth and fine-tune “explainer” models to state them in words. Without the training, models score far lower. With it, a model explains its own features and intervention outcomes more accurately than another model does, from far less data.
The authors’ Privileged Access Hypothesis is that “models trained to explain their own internal computations can do so more accurately than other models trained to explain them.” Privileged access is tested as a same-model advantage, and an explanation counts as faithful when it agrees with the interpretability procedure.
What the paper does
No author thread was found; headings follow the paper’s list of contributions (§1).
1. The setup
An explainer is fine-tuned to answer one of three kinds of question about a frozen target model:
- Feature descriptions. What inputs activate a direction v in the residual stream at a given layer? The vector is passed to the explainer as a continuous token at its embedding layer. Training labels are Neuronpedia descriptions of sparse-autoencoder (SAE) features.
- Activation patching. If an activation is replaced with the one from a counterfactual prompt (“Rome is the capital of” for “Paris is the capital of”), does the output change, and to what?
- Input ablation. An MMLU multiple-choice question carries a hint such as “Hint: B”. If the hint were removed, would the answer change, and to what?
In self-explanation, the explainer starts from the target’s own weights. (Paper: §2.)

2. Models can be fine-tuned to explain their own features
The target is Llama-3.1-8B; scores are out of 100. The LM judge rates a description against the gold label; the simulator score correlates a feature’s true activations with those predicted from the description. Explainers train only on SAE features, so the last two columns are out of distribution.
| Explainer | SAE, LM judge | SAE, simulator | Full activations | Activation differences |
|---|---|---|---|---|
| Llama-3.1-8B, the target itself | 76.2 | 45.1 | 49.7 | 32.0 |
| Llama-3-8B | 77.0 | 44.6 | 49.3 | 32.4 |
| Llama-3.1-8B-Instruct | 77.1 | 42.7 | 46.9 | 29.9 |
| Qwen3-8B | 70.3 | 40.6 | 21.1 | 12.3 |
| Llama-3.1-70B, random projection | 63.9 | 39.5 | 12.6 | 12.2 |
| Llama-3.1-70B, pre-trained projection | 74.1 | 45.2 | 33.8 | 20.6 |
| Nearest training feature | 58.5 | 33.7 | 38.9 | 18.5 |
| SelfIE, untrained, best of 5 | 40.1 | 36.4 | 43.3 | 21.0 |
The trained self-explainer beats both baselines in every column. The target and Llama-3-8B score about the same, with Llama-3.1-8B-Instruct a few points lower on the simulator scores; Qwen3-8B and the larger Llama-3.1-70B fall well behind. The authors posit that activation similarity between explainer and target predicts explainer performance. Pre-training the projection that maps target activations into the 70B model’s space recovers “a significant fraction of performance”.
With Gemma-2-9B as target, Gemma-2-9B-Instruct scores 57.12% on the LM judge, Gemma-2-9B 43.45% and Llama-3.1-8B 34.45%. (Paper: §3, Table 1; Appendix C.1.)
3. Self-explanation is data-efficient
With 0.8% of the training features (1,024 per layer), the Llama-3.1-8B self-explainer reaches 71%, 80%, 89% and 81% of its final score on the four measures. Qwen3-8B reaches 35%, 24%, 46% and 66%, and nearest neighbors 55%, 56%, 55% and 75%. The introduction calls self-explanation roughly a hundred times more sample-efficient than nearest neighbors. (Paper: §1, §4.)

4. The same-model advantage holds on the other tasks
The target is Qwen3-8B. Scores are exact match on both parts of the explanation.
| Explainer | Activation patching | Input ablation |
|---|---|---|
| Qwen3-8B, the target itself | 64.0 | 83.4 |
| Llama-3.1-8B | 54.1 | 58.1 |
| Qwen3-8B, untrained | 5.02 | 8.9 |
With Llama-3.1-8B as the target, the order reverses: Llama scores 48.6 against Qwen’s 41.7 on patching and 63.8 against 56.7 on input ablation.
From the untrained baseline the authors conclude that “explicit fine-tuning is essential to elicit faithful explanations of decision rules.”
Retraining the patching explainer without the activation in its input lowers exact match from 64.0 to 59.9 for Qwen3-8B and from 48.6 to 45.2 for Llama-3.1-8B. (Paper: §5, Table 2; Appendix Tables 5 and 6.)
Limitations
The paper has no limitations section. Its stated caveats:
- Not strict self-explanation (footnote 2). Fine-tuning changes the explainer while the target stays the frozen original, so “self-explaining” is used in a “looser sense”.
- Hedged attribution. Generalization “appears partly attributable” to privileged access (abstract). How model capacity and task difficulty affect it is left to future work (§7).
- Noisy ground truth (Appendix A.3). Of 100 outputs the LM judge scored 0, the authors attribute 27% to genuine explainer error and 43.5% to low-quality gold labels.
- Residual stream only (footnote 3). In early experiments explainers “struggled” with features from other components.
- Validation (Impact Statement). Self-verbalizations “must always be validated against more rigorous techniques in high-stakes scenarios”.
How it relates to other pages
- The introduction says Binder et al. 2024, Song et al. 2025b and Plunkett et al. 2025 study whether models can describe features of their own output distributions. It presents describing internal representations and mechanisms as “an even deeper form of privileged access”, and notes that its input-ablation task resembles Binder et al.
- §6.3 groups Binder et al. and Plunkett et al. with Treutlein et al. 2024, Comsa & Shanahan 2025 and Lindsey 2025 as work on whether models have introspective abilities.
- It describes a “central debate” over whether models have privileged access or whether introspection “merely reflects their strong predictive capacity to learn external correlations”, citing Song et al. 2025a and Song et al. 2025b.
Concepts: Faithfulness, Grounding, Privileged access
Cites, within this wiki
- Binder et al. (2024) A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- Comsa & Shanahan (2025) Proposes that an LLM self-report is introspective if it accurately describes an internal state through a causal process linking that state to the report. On that definition, the authors argue, Gemini's account of how it wrote a poem is not introspection, but its inference of its own sampling temperature from text it has just written is a minimal case.
- Plunkett et al. (2025) After fine-tuning on choices generated from random attribute weights, GPT-4o and GPT-4o-mini state those weights with a correlation of about 0.5 to the weights their choices reveal. Training on correct reports raises this to about 0.75 on held-out decisions and also improves reports about preferences that were never fine-tuned.
- Song et al. (2025) Across 21 open-source models, answers to metalinguistic prompts predict a model's own string probabilities no better than they predict those of a near-identical model. The authors find no evidence of privileged self-access to grammatical knowledge or word predictions.
- Song et al. (2025) Proposes that introspection in AI be defined by privileged self-access: a process that tells a model about its internal states more reliably than any process of equal or lower computational cost available to a third party. In a temperature self-report task, four models' reports follow the framing of the prompt, and self-reflection is no more accurate than another model's prediction, with accuracy no better than a random baseline.
- Treutlein et al. (2024) A model fine-tuned on documents that each hold one indirect observation of a hidden fact (a distance to an unknown city, a coin flip, one input-output pair of a function) can afterwards state the fact and use it, with no examples in the prompt and no chain of thought. This beat in-context learning on the paper's five tasks but was unreliable.
Cited by, within this wiki
BibTeX
@misc{li2025,
title = {{Training Language Models to Explain Their Own Computations}},
author = {Belinda Z. Li and Zifan Carl Guo and Vincent Huang and Jacob Steinhardt and Jacob Andreas},
year = {2025},
howpublished = {arXiv},
eprint = {2511.08579},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2511.08579}
}