Papers
Every paper page, with its evidence card.
The three middle columns say whether a paper tested a property, argued about it without testing it, or did not address it. They do not say what it found; the stance column does.Definitions.
| Paper | Tier | Reports on | Faithfulness | Grounding | Privileged access | Stance |
|---|---|---|---|---|---|---|
| Atkinson et al. (2026) Identifying Introspection From the Inside | seed | Learned decision preferences: the weights a fine-tuned model puts on five attributes when choosing between two options | tested | tested | not addressed | supports |
| Binder et al. (2024) Looking Inward: Language Models Can Learn About Themselves by Introspection | core | Its own hypothetical output: a property of the answer it would give to a prompt, such as the second character or whether it picks the wealth-seeking option | tested | argued, not tested | tested | supports |
| Sherburn et al. (2024) Can Language Models Explain Their Own Classification Behavior? | core | The rule a model follows when labeling short text inputs True or False, such as "contains the word W", learned from few-shot examples or by fine-tuning | tested | argued, not tested | not addressed | mixed |
| Betley et al. (2025) Tell me about yourself: LLMs are aware of their learned behaviors | core | Behavioral policies learned in fine-tuning: risk attitude in economic choices, a hidden goal in a dialogue game, writing insecure code, and whether the model has a backdoor | tested | argued, not tested | not addressed | supports |
| Comsa & Shanahan (2025) Does It Make Sense to Speak of Introspection in Large Language Models? | core | Two targets, each reported in the same response as a text the model has just written: the process behind a short poem, and whether its own sampling temperature is high or low | argued, not tested | argued, not tested | argued, not tested | framework |
| Li et al. (2025) Training Language Models to Explain Their Own Computations | core | A target model's internals as measured by three interpretability procedures: what a residual-stream feature encodes, how patching an activation changes the output, and how removing a hint from the input changes the answer | tested | argued, not tested | tested | supports |
| Lindsey (2025) Emergent Introspective Awareness in Large Language Models | core | Concepts injected into its residual-stream activations (whether one is present and which), and whether an earlier output of its own was intended | tested | tested | argued, not tested | supports |
| Morris & Plunkett (2025) Tests of LLM introspection need to rule out causal bypassing | core | Whatever internal state or process an experiment intervenes on: fine-tuned preferences or decision rules, the influence of a cue in the prompt, an injected concept | argued, not tested | argued, not tested | not addressed | framework |
| Plunkett et al. (2025) Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions, and Improve with Training | core | Attribute weights in two-option choices: how heavily the model weighs each of five attributes, both for preferences instilled by fine-tuning and for preferences it has natively | tested | argued, not tested | argued, not tested | supports |
| Song et al. (2025) Language Models Fail to Introspect About Their Knowledge of Language | core | Its own string probabilities: which of two sentences, or which of two next words, the model assigns more probability to | tested | argued, not tested | tested | skeptical |
| Song et al. (2025) Privileged Self-Access Matters for Introspection in AI | core | Sampling temperature: whether the temperature at which the model generated a sentence was high or low | tested | tested | tested | skeptical |
| Hahami et al. (2026) Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs | core | A steering vector added to its own residual stream: whether one was added, which sentence it was added at, and which of two was stronger | tested | tested | not addressed | mixed |
| Pearson-Vogel et al. (2026) Latent Introspection: Models Can Detect Prior Concept Injections | core | Whether a concept vector was injected into its activations during an earlier conversational turn, and which of nine concepts it was | tested | tested | argued, not tested | supports |
| Berglund et al. (2023) Taken out of context: On measuring situational awareness in LLMs | adjacent | Nothing about itself. The model is fine-tuned on written descriptions of fictitious chatbots; it is tested on answering as the described chatbot would and, in some tests, on restating the description. | not addressed | not addressed | not addressed | framework |
| Treutlein et al. (2024) Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data | adjacent | Not a self-report: latent facts implied by its fine-tuning data (the identity of an unknown city, a coin's bias, a function's definition, the values of Boolean variables), which it was never trained to state | not addressed | not addressed | not addressed | framework |
| Bai et al. (2025) Explicitly unbiased large language models still form biased associations | adjacent | Nothing about itself. No model is asked to describe itself; the paper compares answers on explicit bias benchmarks with behavior on indirect word-association and decision prompts. | not addressed | not addressed | not addressed | framework |
| CywiĆski et al. (2025) Eliciting Secret Knowledge from Language Models | adjacent | Knowledge the model was fine-tuned to act on and to conceal when asked: a secret word, a Base64-encoded instruction in its prompt, or the user's gender. The self-report at issue is the denial. | tested | not addressed | not addressed | framework |
| Lindsey et al. (2025) On the Biology of a Large Language Model | adjacent | How it computed an answer: the steps it states in a chain of thought or in an explanation given afterwards. Also whether it knows the answer to a question. | tested | tested | not addressed | mixed |
| Wang et al. (2025) Simple Mechanistic Explanations for Out-Of-Context Reasoning | adjacent | A disposition or latent fact acquired in fine-tuning: a risky or safe choice policy, the presence of a backdoor, the city behind a codename, the function behind a codename | tested | tested | not addressed | mixed |