LLM Introspection Wiki
Papers and resources on introspection in large language models: whether a model can report on its own internal states.
A self-report counts as introspection here only if it is faithful (it matches what the model actually does or represents) and grounded (it is caused by the state it describes, rather than arrived at some other way). Each paper page says which of those the paper tested.
Seed
The paper this wiki grew from.
- Identifying Introspection From the Inside Models that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.
Core
Work on introspection itself.
- Looking Inward: Language Models Can Learn About Themselves by Introspection A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- Can Language Models Explain Their Own Classification Behavior? Models that classify text by a simple rule often cannot state that rule. GPT-3 fails in free text even after fine-tuning on correct explanations, GPT-4 succeeds 72% of the time on the rules it classifies best, and the authors say a correct statement would still not show that it came from introspection.
- Tell me about yourself: LLMs are aware of their learned behaviors Models fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.
- Does It Make Sense to Speak of Introspection in Large Language Models? Proposes that an LLM self-report is introspective if it accurately describes an internal state through a causal process linking that state to the report. On that definition, the authors argue, Gemini's account of how it wrote a poem is not introspection, but its inference of its own sampling temperature from text it has just written is a minimal case.
- Training Language Models to Explain Their Own Computations Fine-tuned to put the results of interpretability procedures into words, a model explains its own features and intervention outcomes more accurately than a different model trained on the same examples, even a larger one. It also needs far less training data.
- Emergent Introspective Awareness in Large Language Models Claude Opus 4 and 4.1 detect and correctly name a concept vector injected into their activations on about 20% of trials at the best layer and strength, with no false positives on control trials. Some models also consult their earlier activations to judge whether a prefilled output was their own, but the author calls these abilities highly unreliable and context-dependent.
- Tests of LLM introspection need to rule out causal bypassing An intervention that changes a model's internal state can also cause an accurate report of that state by a path that skips the state, so accuracy after an intervention does not show the report is grounded. The authors name this causal bypassing and say the only test they know that rules it out is asking a model whether a concept was injected, a claim a later edit to the post hedges.
- Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions, and Improve with Training After fine-tuning on choices generated from random attribute weights, GPT-4o and GPT-4o-mini state those weights with a correlation of about 0.5 to the weights their choices reveal. Training on correct reports raises this to about 0.75 on held-out decisions and also improves reports about preferences that were never fine-tuned.
- Language Models Fail to Introspect About Their Knowledge of Language Across 21 open-source models, answers to metalinguistic prompts predict a model's own string probabilities no better than they predict those of a near-identical model. The authors find no evidence of privileged self-access to grammatical knowledge or word predictions.
- Privileged Self-Access Matters for Introspection in AI Proposes that introspection in AI be defined by privileged self-access: a process that tells a model about its internal states more reliably than any process of equal or lower computational cost available to a third party. In a temperature self-report task, four models' reports follow the framing of the prompt, and self-reflection is no more accurate than another model's prediction, with accuracy no better than a random baseline.
- Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs In Llama 3.1 8B, apparent success at answering "did you detect an injected thought?" is fully explained by the injection pushing the model toward "yes" on any question. The same model can still say which of ten sentences was injected (up to 88%) and which of two injections was stronger (up to 83%), but only when the injection is in the first few layers.
- Latent Introspection: Models Can Detect Prior Concept Injections Qwen2.5-Coder-32B carries information about a concept vector that was injected during an earlier turn and then removed, including which concept it was. The signal peaks around layers 58 to 62, is weakened by the final layers, and reaches the output only under some prompts: with a document explaining introspection, P("yes") is 39.9% with injection and 0.8% without.
Adjacent
Neighboring questions the core work leans on.
- Taken out of context: On measuring situational awareness in LLMs Models fine-tuned on written descriptions of fictitious chatbots, with no examples, can sometimes act as described when the description is absent from the prompt, but only if each description is paraphrased many times; accuracy rises with model size. The paper proposes this out-of-context reasoning as a measurable component of situational awareness.
- Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data A model fine-tuned on documents that each hold one indirect observation of a hidden fact (a distance to an unknown city, a coin flip, one input-output pair of a function) can afterwards state the fact and use it, with no examples in the prompt and no chain of thought. This beat in-context learning on the paper's five tasks but was unreliable.
- Explicitly unbiased large language models still form biased associations Eight chat models that pass standard bias benchmarks still pair social groups with stereotyped words, and make matching choices between people, when tested with indirect prompts adapted from psychology. The models are never asked about themselves.
- Eliciting Secret Knowledge from Language Models Models fine-tuned to act on a secret while denying they know it can still be made to give it up: prefill attacks let an auditor recover the secret with over 90% success in two of three settings. Logit-lens and sparse-autoencoder readouts of the activations also help the auditor, though less.
- On the Biology of a Large Language Model Circuit tracing in Claude 3.5 Haiku finds the model's account of its own computation matching the mechanism in one case and diverging in others: it describes carry-the-one addition while computing the sum another way, and a chain of thought can be genuine, invented, or worked backwards from a user's hint. Whether it answers a question or says it does not know depends on "known answer" features that can be active for a familiar name when the answer is not known.
- Simple Mechanistic Explanations for Out-Of-Context Reasoning On Gemma 3 12B, a one-layer LoRA fine-tune that produces out-of-context reasoning mostly adds a single constant vector. A steering vector trained directly on the same data also makes the model state a behavior or fact it was only trained to act on.
Concepts
- Causal bypassingWhen an intervention makes a model report an internal state accurately by a path that does not pass through the state.
- Concept injectionAdding a known representation to a model's activations, then asking the model whether it notices and what it is.
- FaithfulnessA self-report is faithful if it matches what the model actually does or represents.
- GroundingA self-report is grounded if it is caused by the internal state or process it describes.
- IntrospectionWhat the papers in this wiki mean by the word, side by side. They agree a self-report must be accurate and disagree about what else it takes.
- Out-of-context reasoningUsing information that was learned in training, and is not present in the prompt, to answer a question.
- Privileged accessThe requirement that a model know something about itself that an outside observer, or another similar model, could not work out equally well.
Threads
- David Atkinson on "Identifying Introspection From the Inside" The lead author walks through the paper in 13 posts: the setup, the late emergence of faithful self-report, where preferences are stored, the attribution-similarity test, and the caveats.
- Theia Pearson-Vogel on "Latent Introspection: Models Can Detect Prior Concept Injections" The lead author walks through the paper in 10 posts: the inject-then-remove design, how a background document changes detection, the poetic prompts, concept identification and its correlation with detection sensitivity, the late-layer decline, and the replications.
- Anthropic on "Emergent Introspective Awareness in Large Language Models" Anthropic's account announces Jack Lindsey's paper in 12 posts: the concept-injection method, detection of injected concepts and how often it fails, the prefill experiment, control of internal states, the comparison across Claude models, and what the results do not show.
- Joshua Engels on self-awareness behaviors and a learned steering vector Six posts from May 2025 about an interim blog post, not about the paper, which appeared two months later. Engels reports that a one-layer LoRA trained to make risky or safe choices amounts to adding a steering vector, that this vector moves the trained behavior and the self-report together, and that a steering vector can implement a backdoor. Wang et al. (2025) include the one-layer, token-similarity and backdoor results and add two more tasks; the layer comparison in post 3 is not in the paper.
- Owain Evans on "Tell me about yourself: LLMs are aware of their learned behaviors" Owain Evans, who supervised the project, introduces the paper in 14 posts: models finetuned on a behavior can describe it, across risky choices, insecure code and a dialogue game; then backdoors, personas, and the links to out-of-context reasoning and the reversal curse.
- Owain Evans on "Looking Inward: Language Models Can Learn About Themselves by Introspection" The paper's last author walks through it in 13 posts: introspection as special access to one's own states, the test of self-prediction against cross-prediction, the tasks, the behavioral-change test, a possible self-simulation mechanism, and what else the paper contains.
- Owain Evans on "Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data" Co-author Owain Evans walks through the paper in 10 posts: the functions, coins and cities examples, the latent-variable pattern behind them, the comparison with in-context learning, the unreliability of the effect, and the safety motivation.
- Owain Evans on "Taken out of context: On measuring situational awareness in LLMs" The paper's last author introduces it in 11 posts: the question of whether a language model could become aware that it is one, the hypothetical risk of reward hacking, out-of-context reasoning as a measurable component, the fictitious-chatbot experiment, the result that paraphrased descriptions are needed and that accuracy grows with model size, and why the paper studies base models.
Frontier
194 candidate papers sit one citation away from the pages above and have not been added yet.
For language models
Every page has a plain-markdown twin at the same URL plus .md, and returns it when asked withAccept: text/markdown. llms.txt indexes the site;llms-full.txt is all of it in one file. The data is also available aspapers.json, graph.json,frontier.json and references.bib.