Concept
Introspection
What the papers in this wiki mean by the word, side by side. They agree a self-report must be accurate and disagree about what else it takes.
AI-drafted, not yet reviewed by a person.
This wiki uses the definition of Atkinson et al. (2026): a self-report is introspection if it is faithful and grounded. Other papers here draw the line elsewhere, and results that look contradictory are often answers to different questions.
The definitions in use
| Paper | A self-report is introspection if | Maps to |
|---|---|---|
| Comsa & Shanahan (2025) | it accurately describes an internal state through a causal process linking the state to the report | faithfulness, grounding |
| Atkinson et al. (2026) | it is accurate about the model’s behavior and caused by the process it describes | faithfulness, grounding |
| Lindsey (2025) | it is accurate, grounded, internal (not routed through the model’s own sampled output), and rests on an internal representation of the state | faithfulness, grounding, and two further criteria |
| Song et al. (2025b) | it comes from a process that tells the model about its states more reliably than any process of equal or lower cost available to a third party | adds privileged access |
| Pearson-Vogel et al. (2026) | it is accurate, causally connected to the state, and unavailable to third parties without special access | all three |
| Binder et al. (2024) | it reflects knowledge about the model that could not be learned from its training data | closest to privileged access |
| Song, Hu & Mahowald (2025a) | prompted answers predict the model’s own string probabilities better than they predict a near-identical model’s | faithfulness, privileged access |
Where they part
Is a causal link enough? Comsa and Shanahan call their definition lightweight on purpose. Song et al. (2025b) object that it would count a model reading its own transcript as introspecting, and add privileged access. Lindsey calls their definition the more compelling one, and says his internality criterion aligns with it.
Does accuracy show anything about cause? Morris & Plunkett (2025) argue it does not: an intervention can produce an accurate report by a path that skips the state. See causal bypassing.
Is a same-model advantage introspection? Binder et al. read a model predicting itself better than another model can as introspection. Song, Hu and Mahowald find the advantage disappears against a near-identical model. Lindsey prefers to call it self-modeling.
The papers table records, for each paper, which of faithfulness, grounding and privileged access it tested.