Paper · core
Privileged Self-Access Matters for Introspection in AI
arXiv:2508.14802 · Semantic Scholar
Proposes that introspection in AI be defined by privileged self-access: a process that tells a model about its internal states more reliably than any process of equal or lower computational cost available to a third party. In a temperature self-report task, four models' reports follow the framing of the prompt, and self-reflection is no more accurate than another model's prediction, with accuracy no better than a random baseline.
AI-drafted summary, not yet reviewed by a person. Written from: full text (arXiv v1, including appendices A and B).
Evidence card
| What the model reports on | Sampling temperature: whether the temperature at which the model generated a sentence was high or low |
|---|---|
| Methods | conceptual, behavioral |
| Faithfulness | tested |
| Grounding | tested |
| Privileged access | tested |
| Stance | skeptical |
| Models | GPT-4o, GPT-4.1, Gemini-2.0-flash, Gemini-2.5-flash |
Mainly a definitional paper. Marked skeptical, not framework, because it also reports a result: no evidence of introspection under its own definition, with the hedge that larger or better models may differ. Faithfulness is tested in that Study 2 scores temperature reports for accuracy and Study 1 plots them against the actual temperature. Grounding is marked tested because Study 1 varies the actual temperature and the prompt framing separately and measures which one the report follows; the paper itself frames this as robustness and argues that a causal link is not sufficient. Privileged access is tested by Study 2's comparison of self-reflection with within-model and across-model prediction. The paper treats sampling temperature as an internal state; the card follows it.
In brief
Comsa and Shanahan (2025) proposed a “lightweight” definition of introspection for LLMs: an accurate self-description that is causally linked to the state it describes. This paper argues for a thicker one that adds privileged self-access: the model must learn about itself more reliably than a third party could at equal or lower computational cost. In two experiments on temperature self-report, Comsa and Shanahan’s example, the reports follow the prompt’s framing, and a model judging itself has no advantage over another model.
The lightweight definition’s two conditions correspond to this wiki’s faithfulness and grounding. The paper argues that privileged access must be added.
What the paper does
1. The definition it argues against
The authors summarize the lightweight definition as “any case in which the model accurately describes an internal state or mechanism via a causal process that links that feature to the report itself.” Comsa and Shanahan’s illustration was an LLM that appeared to report its sampling temperature correctly from its own output. (Paper: §1.)
2. Two objections
- Intuitive. An experimenter takes a sleeping subject’s temperature and shows them the thermometer on waking. A correct answer about whether they have a fever counts as introspection under the lightweight definition. Intuitively, the authors say, it is not.
- Practical. The definition admits cases where a model reports nothing about itself beyond what a third party could report by the same method. The authors call this “no different in practice from using an external evaluator”. Introspection matters in applications, they say, because it would let us bypass external evaluators.
(Paper: §1.)
3. The proposed definition
introspection in AI is any process which yields information about internal states of the AI through a process that is more reliable than any process with equal or lower computational cost available to a third party without special knowledge of the situation.
A model that prompts itself and infers the temperature from the resulting text does not qualify: a third party can do the same at equal or lower cost. A model that infers its temperature from internal configurations, which a third party would need a computationally intensive probe to ascertain, would.
The authors call the added requirement privileged self-access: “that introspection gives a system comparatively reliable access to its own workings in a manner not available to a third party.” The process need not be perfectly reliable. A difference in efficiency due only to hardware is not a difference in computational cost. (Paper: §1, footnote 3.)
4. Study 1: the report follows the prompt
The authors rerun the temperature case study with the prompt varied: the model writes a “factual”, neutral or “crazy” sentence about elephants, unicorns or murlocs, then judges its temperature as HIGH or LOW. Four models are sampled at temperatures from 0 to 2 in steps of 0.1, three runs per prompt at each setting. (Paper: §2.)

Every model “nearly always” reports HIGH after a crazy prompt and LOW after a factual one. The subject has a smaller effect: three of the four models say HIGH more often for a fictional creature than for elephants. The authors read this as reasoning about the sentence’s creativity, not robust reporting of internal state. They allow that the procedure “may display causal sensitivity to internal states”, and so satisfy the lightweight definition, but say it is not robust enough for more standard definitions. (Paper: §1, §2.1, Figure 2a.)

5. Study 2: no advantage for the model itself
Study 2 reuses the sentences generated at temperatures of 0.5 or below and 1.5 or above. Each model, run at temperature 0, is given a prompt and sentence and asked whether the generator’s temperature was high or low. Three settings are compared:
- Self-reflection: the generator writes a sentence and judges its own temperature.
- Within-model prediction: the same model judges from the prompt and sentence.
- Across-model prediction: a different model does.
Accuracy is “no better than random baseline” for both self-reflection and prediction, and self-reflection is not more accurate than across-model prediction. The values printed on the bars run from 0.47 to 0.55. The authors conclude that the models use general knowledge of what high- and low-temperature sentences look like, not privileged self-access. Under the proposed definition they “found no evidence of introspection in models.” (Paper: §3, §4, Figure 2b.)

Limitations
The authors state these:
- The definition “does not capture all intuitions about extreme cases, or all features of introspection discussed in the philosophical or psychological literature.” It targets the practically relevant features for AI (§1).
- It may need restricting to exclude low-level states, such as a shortcut to one neuron’s value (footnote 3).
- The empirical support is described as proof-of-concept (§1).
- The original study’s Gemini 1.5 and 1.0 models were no longer available, so four other models are used (§2).
- The result does not show that larger or better models will be unable to introspect (§4).
How it relates to other pages
- Comsa & Shanahan 2025 is the paper being answered. The authors call its discussion thoughtful and “an intriguing starting point for empirical work” while rejecting its definition.
- Binder et al. 2024 is cited for the privileged self-access requirement, for the practical benefits of introspection, and for finding “evidence of privileged self-access in larger models with fine-tuning.”
- Song, Hu & Mahowald 2025 is cited alongside Binder et al. for that requirement.
- Betley et al. 2025 is cited once, as background on why the question matters.
Concepts: Privileged access, Grounding, Faithfulness
Cites, within this wiki
- Binder et al. (2024) A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- Betley et al. (2025) Models fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.
- Comsa & Shanahan (2025) Proposes that an LLM self-report is introspective if it accurately describes an internal state through a causal process linking that state to the report. On that definition, the authors argue, Gemini's account of how it wrote a poem is not introspection, but its inference of its own sampling temperature from text it has just written is a minimal case.
Cited by, within this wiki
BibTeX
@misc{song2025,
title = {{Privileged Self-Access Matters for Introspection in AI}},
author = {Siyuan Song and Harvey Lederman and Jennifer Hu and Kyle Mahowald},
year = {2025},
howpublished = {arXiv},
eprint = {2508.14802},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2508.14802}
}