Concept
Privileged access
The requirement that a model know something about itself that an outside observer, or another similar model, could not work out equally well.
AI-drafted, not yet reviewed by a person.
Privileged access is a stricter requirement than a causal link. Song et al. (2025b) propose it as the defining feature of introspection:
introspection in AI is any process which yields information about internal states of the AI through a process that is more reliable than any process with equal or lower computational cost available to a third party without special knowledge of the situation. Their target is the lightweight definition of Comsa & Shanahan (2025), under which a model inferring its sampling temperature from text it has just written would count.
A paper can test grounding without testing privileged access, and the evidence cards on this wiki record the two separately.
How it is tested
The usual design compares a model’s report about itself with a second predictor that has the same outside information.
- Binder et al. (2024) fine-tune a model to predict properties of its own answers and a second model on the same data about the first. The first predicts itself better. The effect appears only on simple tasks.
- Li et al. (2025) state a Privileged Access Hypothesis: “models trained to explain their own internal computations can do so more accurately than other models trained to explain them.” They find a model explains its own features better than a different model does, even a larger one.
- Song, Hu & Mahowald (2025a) make the comparison model a near-identical one. Across 21 open-source models, answers to metalinguistic prompts predict a model’s own string probabilities no better than they predict those of its nearest neighbor.
- Song et al. (2025b) find that in a temperature self-report task a model judging itself has no advantage over another model judging it.
What the disagreement is about
The positive and negative results differ in what the second predictor is. Against a different model, the same-model advantage appears. Against a near-identical model, it does not. Song, Hu and Mahowald attribute the advantage to a model being most similar to itself.
Lindsey (2025) reads both Binder et al. and Song et al. as showing access to a model’s own learned abstractions, not an introspective mechanism, and prefers the term self-modeling for it. His own test counts a detection only if it comes before the concept appears in the model’s output, which he says aligns with the privileged-access definition, though no outside predictor is compared.
Papers tagged with this concept
- Binder et al. (2024) A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- Comsa & Shanahan (2025) Proposes that an LLM self-report is introspective if it accurately describes an internal state through a causal process linking that state to the report. On that definition, the authors argue, Gemini's account of how it wrote a poem is not introspection, but its inference of its own sampling temperature from text it has just written is a minimal case.
- Li et al. (2025) Fine-tuned to put the results of interpretability procedures into words, a model explains its own features and intervention outcomes more accurately than a different model trained on the same examples, even a larger one. It also needs far less training data.
- Lindsey (2025) Claude Opus 4 and 4.1 detect and correctly name a concept vector injected into their activations on about 20% of trials at the best layer and strength, with no false positives on control trials. Some models also consult their earlier activations to judge whether a prefilled output was their own, but the author calls these abilities highly unreliable and context-dependent.
- Plunkett et al. (2025) After fine-tuning on choices generated from random attribute weights, GPT-4o and GPT-4o-mini state those weights with a correlation of about 0.5 to the weights their choices reveal. Training on correct reports raises this to about 0.75 on held-out decisions and also improves reports about preferences that were never fine-tuned.
- Song et al. (2025) Across 21 open-source models, answers to metalinguistic prompts predict a model's own string probabilities no better than they predict those of a near-identical model. The authors find no evidence of privileged self-access to grammatical knowledge or word predictions.
- Song et al. (2025) Proposes that introspection in AI be defined by privileged self-access: a process that tells a model about its internal states more reliably than any process of equal or lower computational cost available to a third party. In a temperature self-report task, four models' reports follow the framing of the prompt, and self-reflection is no more accurate than another model's prediction, with accuracy no better than a random baseline.
- Pearson-Vogel et al. (2026) Qwen2.5-Coder-32B carries information about a concept vector that was injected during an earlier turn and then removed, including which concept it was. The signal peaks around layers 58 to 62, is weakened by the final layers, and reaches the output only under some prompts: with a document explaining introspection, P("yes") is 39.9% with injection and 0.8% without.