# Papers

> Every paper page in the LLM Introspection Wiki, with its evidence card.

Faithfulness, grounding and privileged access say whether the paper tested the property, argued about it, or did not address it. They do not say what it found; the stance column does. See [About](https://introspection.infinite.fun/about.md) for the definitions.

| Paper | Tier | Reports on | Faithfulness | Grounding | Privileged access | Stance |
|---|---|---|---|---|---|---|
| [Atkinson et al. (2026): Identifying Introspection From the Inside](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md) | seed | Learned decision preferences: the weights a fine-tuned model puts on five attributes when choosing between two options | tested | tested | not addressed | supports |
| [Binder et al. (2024): Looking Inward: Language Models Can Learn About Themselves by Introspection](https://introspection.infinite.fun/papers/binder2024-looking-inward.md) | core | Its own hypothetical output: a property of the answer it would give to a prompt, such as the second character or whether it picks the wealth-seeking option | tested | argued, not tested | tested | supports |
| [Sherburn et al. (2024): Can Language Models Explain Their Own Classification Behavior?](https://introspection.infinite.fun/papers/sherburn2024-explain-classification-behavior.md) | core | The rule a model follows when labeling short text inputs True or False, such as "contains the word W", learned from few-shot examples or by fine-tuning | tested | argued, not tested | not addressed | mixed |
| [Betley et al. (2025): Tell me about yourself: LLMs are aware of their learned behaviors](https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself.md) | core | Behavioral policies learned in fine-tuning: risk attitude in economic choices, a hidden goal in a dialogue game, writing insecure code, and whether the model has a backdoor | tested | argued, not tested | not addressed | supports |
| [Comsa & Shanahan (2025): Does It Make Sense to Speak of Introspection in Large Language Models?](https://introspection.infinite.fun/papers/comsa2025-speak-of-introspection.md) | core | Two targets, each reported in the same response as a text the model has just written: the process behind a short poem, and whether its own sampling temperature is high or low | argued, not tested | argued, not tested | argued, not tested | framework |
| [Li et al. (2025): Training Language Models to Explain Their Own Computations](https://introspection.infinite.fun/papers/li2025-explain-own-computations.md) | core | A target model's internals as measured by three interpretability procedures: what a residual-stream feature encodes, how patching an activation changes the output, and how removing a hint from the input changes the answer | tested | argued, not tested | tested | supports |
| [Lindsey (2025): Emergent Introspective Awareness in Large Language Models](https://introspection.infinite.fun/papers/lindsey2025-emergent-introspective-awareness.md) | core | Concepts injected into its residual-stream activations (whether one is present and which), and whether an earlier output of its own was intended | tested | tested | argued, not tested | supports |
| [Morris & Plunkett (2025): Tests of LLM introspection need to rule out causal bypassing](https://introspection.infinite.fun/papers/morris2025-causal-bypassing.md) | core | Whatever internal state or process an experiment intervenes on: fine-tuned preferences or decision rules, the influence of a cue in the prompt, an injected concept | argued, not tested | argued, not tested | not addressed | framework |
| [Plunkett et al. (2025): Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions, and Improve with Training](https://introspection.infinite.fun/papers/plunkett2025-self-interpretability.md) | core | Attribute weights in two-option choices: how heavily the model weighs each of five attributes, both for preferences instilled by fine-tuning and for preferences it has natively | tested | argued, not tested | argued, not tested | supports |
| [Song et al. (2025): Language Models Fail to Introspect About Their Knowledge of Language](https://introspection.infinite.fun/papers/song2025-fail-to-introspect.md) | core | Its own string probabilities: which of two sentences, or which of two next words, the model assigns more probability to | tested | argued, not tested | tested | skeptical |
| [Song et al. (2025): Privileged Self-Access Matters for Introspection in AI](https://introspection.infinite.fun/papers/song2025-privileged-self-access.md) | core | Sampling temperature: whether the temperature at which the model generated a sentence was high or low | tested | tested | tested | skeptical |
| [Hahami et al. (2026): Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs](https://introspection.infinite.fun/papers/hahami2026-detecting-the-disturbance.md) | core | A steering vector added to its own residual stream: whether one was added, which sentence it was added at, and which of two was stronger | tested | tested | not addressed | mixed |
| [Pearson-Vogel et al. (2026): Latent Introspection: Models Can Detect Prior Concept Injections](https://introspection.infinite.fun/papers/pearson-vogel2026-latent-introspection.md) | core | Whether a concept vector was injected into its activations during an earlier conversational turn, and which of nine concepts it was | tested | tested | argued, not tested | supports |
| [Berglund et al. (2023): Taken out of context: On measuring situational awareness in LLMs](https://introspection.infinite.fun/papers/berglund2023-taken-out-of-context.md) | adjacent | Nothing about itself. The model is fine-tuned on written descriptions of fictitious chatbots; it is tested on answering as the described chatbot would and, in some tests, on restating the description. | not addressed | not addressed | not addressed | framework |
| [Treutlein et al. (2024): Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data](https://introspection.infinite.fun/papers/treutlein2024-connecting-the-dots.md) | adjacent | Not a self-report: latent facts implied by its fine-tuning data (the identity of an unknown city, a coin's bias, a function's definition, the values of Boolean variables), which it was never trained to state | not addressed | not addressed | not addressed | framework |
| [Bai et al. (2025): Explicitly unbiased large language models still form biased associations](https://introspection.infinite.fun/papers/bai2025-explicitly-unbiased.md) | adjacent | Nothing about itself. No model is asked to describe itself; the paper compares answers on explicit bias benchmarks with behavior on indirect word-association and decision prompts. | not addressed | not addressed | not addressed | framework |
| [Cywiński et al. (2025): Eliciting Secret Knowledge from Language Models](https://introspection.infinite.fun/papers/cywinski2025-eliciting-secret-knowledge.md) | adjacent | Knowledge the model was fine-tuned to act on and to conceal when asked: a secret word, a Base64-encoded instruction in its prompt, or the user's gender. The self-report at issue is the denial. | tested | not addressed | not addressed | framework |
| [Lindsey et al. (2025): On the Biology of a Large Language Model](https://introspection.infinite.fun/papers/lindsey2025-biology-of-llm.md) | adjacent | How it computed an answer: the steps it states in a chain of thought or in an explanation given afterwards. Also whether it knows the answer to a question. | tested | tested | not addressed | mixed |
| [Wang et al. (2025): Simple Mechanistic Explanations for Out-Of-Context Reasoning](https://introspection.infinite.fun/papers/wang2025-mechanistic-oocr.md) | adjacent | A disposition or latent fact acquired in fine-tuning: a risky or safe choice policy, the presence of a backdoor, the city behind a codename, the function behind a codename | tested | tested | not addressed | mixed |

---

Source: https://introspection.infinite.fun/papers · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
