Papers

Every paper page, with its evidence card.

The three middle columns say whether a paper tested a property, argued about it without testing it, or did not address it. They do not say what it found; the stance column does.Definitions.

PaperTierReports onFaithfulnessGroundingPrivileged accessStance
Atkinson et al. (2026)
Identifying Introspection From the Inside
seedLearned decision preferences: the weights a fine-tuned model puts on five attributes when choosing between two optionstestedtestednot addressedsupports
Binder et al. (2024)
Looking Inward: Language Models Can Learn About Themselves by Introspection
coreIts own hypothetical output: a property of the answer it would give to a prompt, such as the second character or whether it picks the wealth-seeking optiontestedargued, not testedtestedsupports
Sherburn et al. (2024)
Can Language Models Explain Their Own Classification Behavior?
coreThe rule a model follows when labeling short text inputs True or False, such as "contains the word W", learned from few-shot examples or by fine-tuningtestedargued, not testednot addressedmixed
Betley et al. (2025)
Tell me about yourself: LLMs are aware of their learned behaviors
coreBehavioral policies learned in fine-tuning: risk attitude in economic choices, a hidden goal in a dialogue game, writing insecure code, and whether the model has a backdoortestedargued, not testednot addressedsupports
Comsa & Shanahan (2025)
Does It Make Sense to Speak of Introspection in Large Language Models?
coreTwo targets, each reported in the same response as a text the model has just written: the process behind a short poem, and whether its own sampling temperature is high or lowargued, not testedargued, not testedargued, not testedframework
Li et al. (2025)
Training Language Models to Explain Their Own Computations
coreA target model's internals as measured by three interpretability procedures: what a residual-stream feature encodes, how patching an activation changes the output, and how removing a hint from the input changes the answertestedargued, not testedtestedsupports
Lindsey (2025)
Emergent Introspective Awareness in Large Language Models
coreConcepts injected into its residual-stream activations (whether one is present and which), and whether an earlier output of its own was intendedtestedtestedargued, not testedsupports
Morris & Plunkett (2025)
Tests of LLM introspection need to rule out causal bypassing
coreWhatever internal state or process an experiment intervenes on: fine-tuned preferences or decision rules, the influence of a cue in the prompt, an injected conceptargued, not testedargued, not testednot addressedframework
Plunkett et al. (2025)
Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions, and Improve with Training
coreAttribute weights in two-option choices: how heavily the model weighs each of five attributes, both for preferences instilled by fine-tuning and for preferences it has nativelytestedargued, not testedargued, not testedsupports
Song et al. (2025)
Language Models Fail to Introspect About Their Knowledge of Language
coreIts own string probabilities: which of two sentences, or which of two next words, the model assigns more probability totestedargued, not testedtestedskeptical
Song et al. (2025)
Privileged Self-Access Matters for Introspection in AI
coreSampling temperature: whether the temperature at which the model generated a sentence was high or lowtestedtestedtestedskeptical
Hahami et al. (2026)
Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs
coreA steering vector added to its own residual stream: whether one was added, which sentence it was added at, and which of two was strongertestedtestednot addressedmixed
Pearson-Vogel et al. (2026)
Latent Introspection: Models Can Detect Prior Concept Injections
coreWhether a concept vector was injected into its activations during an earlier conversational turn, and which of nine concepts it wastestedtestedargued, not testedsupports
Berglund et al. (2023)
Taken out of context: On measuring situational awareness in LLMs
adjacentNothing about itself. The model is fine-tuned on written descriptions of fictitious chatbots; it is tested on answering as the described chatbot would and, in some tests, on restating the description.not addressednot addressednot addressedframework
Treutlein et al. (2024)
Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
adjacentNot a self-report: latent facts implied by its fine-tuning data (the identity of an unknown city, a coin's bias, a function's definition, the values of Boolean variables), which it was never trained to statenot addressednot addressednot addressedframework
Bai et al. (2025)
Explicitly unbiased large language models still form biased associations
adjacentNothing about itself. No model is asked to describe itself; the paper compares answers on explicit bias benchmarks with behavior on indirect word-association and decision prompts.not addressednot addressednot addressedframework
CywiƄski et al. (2025)
Eliciting Secret Knowledge from Language Models
adjacentKnowledge the model was fine-tuned to act on and to conceal when asked: a secret word, a Base64-encoded instruction in its prompt, or the user's gender. The self-report at issue is the denial.testednot addressednot addressedframework
Lindsey et al. (2025)
On the Biology of a Large Language Model
adjacentHow it computed an answer: the steps it states in a chain of thought or in an explanation given afterwards. Also whether it knows the answer to a question.testedtestednot addressedmixed
Wang et al. (2025)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
adjacentA disposition or latent fact acquired in fine-tuning: a risky or safe choice policy, the presence of a backdoor, the city behind a codename, the function behind a codenametestedtestednot addressedmixed