{"site":"https://introspection.infinite.fun","license":"CC BY 4.0","papers":[{"id":"atkinson2026-identifying-introspection","url":"https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection","markdown_url":"https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md","title":"Identifying Introspection From the Inside","authors":["David I. Atkinson","Dillon Plunkett","David Bau"],"year":2026,"venue":"COLM 2026","tier":"seed","status":"full","reviewed":false,"summary":"Models that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.","links":{"pdf":"https://iii.baulab.info/identifying-intro-preprint.pdf","project":"https://iii.baulab.info"},"concepts":["faithfulness","grounding","causal-bypassing","out-of-context-reasoning"],"threads":["diatkinson-identifying-introspection"],"evidence":{"reports_on":"Learned decision preferences: the weights a fine-tuned model puts on five attributes when choosing between two options","methods":["fine-tuning","ablation","patching"],"faithfulness":"tested","grounding":"tested","privileged_access":"not-addressed","stance":"supports","models":["Qwen3 (0.6B to 32B)","Gemma-4 (E4B, 31B)"],"note":"Supports grounded self-report in a deliberately narrow setting: LoRA adapters, linear preferences over five attributes. The test separates groups of models, not individual ones. The paper does not compare a model's self-report against an outside predictor, so it does not bear on privileged access."},"sources":["full text (extended preprint, iii.baulab.info)","the lead author's thread"],"added":"2026-10-06","updated":"2026-10-06","cites":["binder2024-looking-inward","sherburn2024-explain-classification-behavior","betley2025-tell-me-about-yourself","comsa2025-speak-of-introspection","li2025-explain-own-computations","lindsey2025-emergent-introspective-awareness","morris2025-causal-bypassing","plunkett2025-self-interpretability","song2025-fail-to-introspect","song2025-privileged-self-access","hahami2026-detecting-the-disturbance","pearson-vogel2026-latent-introspection","berglund2023-taken-out-of-context","treutlein2024-connecting-the-dots","bai2025-explicitly-unbiased","cywinski2025-eliciting-secret-knowledge","lindsey2025-biology-of-llm","wang2025-mechanistic-oocr"],"cited_by":[]},{"id":"binder2024-looking-inward","url":"https://introspection.infinite.fun/papers/binder2024-looking-inward","markdown_url":"https://introspection.infinite.fun/papers/binder2024-looking-inward.md","title":"Looking Inward: Language Models Can Learn About Themselves by Introspection","authors":["Felix J. Binder","James Chua","Tomek Korbak","Henry Sleight","John Hughes","Robert Long","Ethan Perez","Miles Turpin","Owain Evans"],"year":2024,"venue":"ICLR 2025","tier":"core","status":"full","reviewed":false,"summary":"A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.","links":{"arxiv":"2410.13787","s2":"b47812325fd9493eb8d5dbf1deb7ad4a763ebe65"},"concepts":["privileged-access","faithfulness","grounding","out-of-context-reasoning"],"threads":["owainevans-looking-inward"],"evidence":{"reports_on":"Its own hypothetical output: a property of the answer it would give to a prompt, such as the second character or whether it picks the wealth-seeking option","methods":["self-prediction","fine-tuning","behavioral"],"faithfulness":"tested","grounding":"argued","privileged_access":"tested","stance":"supports","models":["GPT-4o","GPT-4","GPT-3.5","Llama 3.1 70B"],"note":"Self-prediction accuracy compares the report with the model's actual output, so faithfulness is tested, and the comparison with a cross-trained model is a direct test of privileged access. Grounding is marked argued: the paper's definition rules out training data as the source of a report without saying what the source is, and the self-simulation mechanism is proposed, not tested. The behavioral-change experiment comes closest, and the authors call it indirect evidence. The supporting result is limited by the authors to simple tasks; the paper also reports failures on longer outputs and no out-of-distribution transfer."},"sources":["full text (arXiv v1, with appendix)","Owain Evans's thread"],"date":"2024-10-17","added":"2026-10-06","updated":"2026-10-06","cites":["berglund2023-taken-out-of-context","treutlein2024-connecting-the-dots"],"cited_by":["atkinson2026-identifying-introspection","betley2025-tell-me-about-yourself","comsa2025-speak-of-introspection","li2025-explain-own-computations","plunkett2025-self-interpretability","song2025-fail-to-introspect","song2025-privileged-self-access","hahami2026-detecting-the-disturbance","pearson-vogel2026-latent-introspection"]},{"id":"sherburn2024-explain-classification-behavior","url":"https://introspection.infinite.fun/papers/sherburn2024-explain-classification-behavior","markdown_url":"https://introspection.infinite.fun/papers/sherburn2024-explain-classification-behavior.md","title":"Can Language Models Explain Their Own Classification Behavior?","authors":["Dane Sherburn","Bilal Chughtai","Owain Evans"],"year":2024,"venue":"arXiv","tier":"core","status":"full","reviewed":false,"summary":"Models that classify text by a simple rule often cannot state that rule. GPT-3 fails in free text even after fine-tuning on correct explanations, GPT-4 succeeds 72% of the time on the rules it classifies best, and the authors say a correct statement would still not show that it came from introspection.","links":{"arxiv":"2405.07436","s2":"3ad0498cd275fea33ac9cc5ba549262021e2878c"},"concepts":["faithfulness","grounding"],"threads":[],"evidence":{"reports_on":"The rule a model follows when labeling short text inputs True or False, such as \"contains the word W\", learned from few-shot examples or by fine-tuning","methods":["behavioral","fine-tuning"],"faithfulness":"tested","grounding":"argued","privileged_access":"not-addressed","stance":"mixed","models":["GPT-3 (ada, babbage, curie, davinci)","GPT-4","fine-tuned davinci"],"note":"Faithfulness is tested against behavior: the paper first measures whether the model's classification, on ordinary and adversarial inputs, is closely approximated by a known rule, then scores the model's statement of that rule. Grounding is argued, not tested: the articulation prompt contains the same labeled examples as the classification prompt, and the authors say the benchmark cannot separate introspection from the most probable completion (§4, Appendix A). Stance is mixed because the paper concludes that current models struggle and that GPT-3 fails even after fine-tuning, while reporting early signs of the ability in GPT-4. No outside predictor is compared with the model, so privileged access is not addressed."},"sources":["full text (arXiv v1, including appendices)"],"date":"2024-05-13","added":"2026-10-06","updated":"2026-10-06","cites":[],"cited_by":["atkinson2026-identifying-introspection"]},{"id":"betley2025-tell-me-about-yourself","url":"https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself","markdown_url":"https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself.md","title":"Tell me about yourself: LLMs are aware of their learned behaviors","authors":["Jan Betley","Xuchan Bao","Martín Soto","Anna Sztyber-Betley","James Chua","Owain Evans"],"year":2025,"venue":"ICLR 2025","tier":"core","status":"full","reviewed":false,"summary":"Models fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.","links":{"arxiv":"2501.11120","s2":"a3ec0b75274a29bf7637f9090d5ca5047e2c7545"},"concepts":["faithfulness","grounding","out-of-context-reasoning"],"threads":["owainevans-tell-me-about-yourself"],"evidence":{"reports_on":"Behavioral policies learned in fine-tuning: risk attitude in economic choices, a hidden goal in a dialogue game, writing insecure code, and whether the model has a backdoor","methods":["fine-tuning","behavioral"],"faithfulness":"tested","grounding":"argued","privileged_access":"not-addressed","stance":"supports","models":["GPT-4o","Llama-3.1-70B"],"note":"Faithfulness is tested directly: §3.1.3 correlates self-reported with actual risk level, and Table 2 sets self-reported code security beside the measured rate of secure code. Grounding is marked argued because there is no causal or mechanistic experiment; the authors say the correlation could be a direct causal link or a common cause in the training data. Privileged access is marked not-addressed because no outside predictor is compared, although the authors note that among models trained on identical data, differences in behavior are partially reflected in self-reports, and leave open whether that meets the definition in Binder et al. (2024). Stance is supports because the paper concludes that models can describe their learned behaviors and calls this a form of introspection, while saying that testing for introspection is not its primary focus."},"sources":["full text (arXiv v1, including appendices)","Owain Evans's thread"],"date":"2025-01-19","added":"2026-10-06","updated":"2026-10-06","cites":["binder2024-looking-inward","berglund2023-taken-out-of-context","treutlein2024-connecting-the-dots"],"cited_by":["atkinson2026-identifying-introspection","plunkett2025-self-interpretability","song2025-fail-to-introspect","song2025-privileged-self-access","cywinski2025-eliciting-secret-knowledge","wang2025-mechanistic-oocr"]},{"id":"comsa2025-speak-of-introspection","url":"https://introspection.infinite.fun/papers/comsa2025-speak-of-introspection","markdown_url":"https://introspection.infinite.fun/papers/comsa2025-speak-of-introspection.md","title":"Does It Make Sense to Speak of Introspection in Large Language Models?","authors":["Iulia M. Comsa","Murray Shanahan"],"year":2025,"venue":"arXiv","tier":"core","status":"full","reviewed":false,"summary":"Proposes that an LLM self-report is introspective if it accurately describes an internal state through a causal process linking that state to the report. On that definition, the authors argue, Gemini's account of how it wrote a poem is not introspection, but its inference of its own sampling temperature from text it has just written is a minimal case.","links":{"arxiv":"2506.05068","s2":"a8d2824c5bb21538ac00fd09fad467e8a02ad169"},"concepts":["faithfulness","grounding","privileged-access"],"threads":[],"evidence":{"reports_on":"Two targets, each reported in the same response as a text the model has just written: the process behind a short poem, and whether its own sampling temperature is high or low","methods":["conceptual"],"faithfulness":"argued","grounding":"argued","privileged_access":"argued","stance":"framework","models":["Gemini Pro 1.5","Gemini Pro 1.0"],"note":"The paper prints sample Gemini outputs but says its goals are conceptual, not empirical, and it scores nothing. So the only method is `conceptual` and all three properties are `argued`. Stance is `framework` because the result is a definition and a verdict on two examples (one rejected, one accepted as a minimal case), not a measurement. Privileged access is `argued` because the definition follows accounts that downgrade it and does not require it, not because the paper claims models have it."},"sources":["full text (arXiv v2, including Appendix A)"],"date":"2025-06-05","added":"2026-10-06","updated":"2026-10-06","cites":["binder2024-looking-inward"],"cited_by":["atkinson2026-identifying-introspection","li2025-explain-own-computations","song2025-privileged-self-access"]},{"id":"li2025-explain-own-computations","url":"https://introspection.infinite.fun/papers/li2025-explain-own-computations","markdown_url":"https://introspection.infinite.fun/papers/li2025-explain-own-computations.md","title":"Training Language Models to Explain Their Own Computations","authors":["Belinda Z. Li","Zifan Carl Guo","Vincent Huang","Jacob Steinhardt","Jacob Andreas"],"year":2025,"venue":"arXiv","tier":"core","status":"full","reviewed":false,"summary":"Fine-tuned to put the results of interpretability procedures into words, a model explains its own features and intervention outcomes more accurately than a different model trained on the same examples, even a larger one. It also needs far less training data.","links":{"arxiv":"2511.08579","s2":"2f967d2b86217368a36511d082ae465de04980c2"},"concepts":["faithfulness","grounding","privileged-access"],"threads":[],"evidence":{"reports_on":"A target model's internals as measured by three interpretability procedures: what a residual-stream feature encodes, how patching an activation changes the output, and how removing a hint from the input changes the answer","methods":["fine-tuning","self-prediction","patching","ablation"],"faithfulness":"tested","grounding":"argued","privileged_access":"tested","stance":"supports","models":["Llama-3.1-8B","Llama-3.1-8B-Instruct","Llama-3-8B","Llama-3.1-70B","Qwen3-8B","Gemma-2-9B","Gemma-2-9B-Instruct"],"note":"Faithfulness is scored against the output of an interpretability procedure, and the ability is trained in: untrained baselines score far lower. Privileged access is tested as a same-model advantage over other trained explainers, and close variants of the target do about as well as the target itself on feature descriptions. Grounding is marked argued: the authors attribute the advantage to access to internals and support it with a correlation between activation similarity and explainer score, but no experiment traces what causes a given explanation. The explainer is a fine-tuned copy describing the frozen original, which the authors call self-explanation in a looser sense. Patching and ablation are listed as methods because they supply the ground truth; self-prediction because two tasks ask the model to predict its own output under an intervention."},"sources":["full text (arXiv v3, 9 Feb 2026, with appendices A to H)"],"date":"2025-11-11","added":"2026-10-06","updated":"2026-10-06","cites":["binder2024-looking-inward","comsa2025-speak-of-introspection","plunkett2025-self-interpretability","song2025-fail-to-introspect","song2025-privileged-self-access","treutlein2024-connecting-the-dots"],"cited_by":["atkinson2026-identifying-introspection"]},{"id":"lindsey2025-emergent-introspective-awareness","url":"https://introspection.infinite.fun/papers/lindsey2025-emergent-introspective-awareness","markdown_url":"https://introspection.infinite.fun/papers/lindsey2025-emergent-introspective-awareness.md","title":"Emergent Introspective Awareness in Large Language Models","authors":["Jack Lindsey"],"year":2025,"venue":"Transformer Circuits Thread","tier":"core","status":"full","reviewed":false,"summary":"Claude Opus 4 and 4.1 detect and correctly name a concept vector injected into their activations on about 20% of trials at the best layer and strength, with no false positives on control trials. Some models also consult their earlier activations to judge whether a prefilled output was their own, but the author calls these abilities highly unreliable and context-dependent.","links":{"arxiv":"2601.01828","url":"https://transformer-circuits.pub/2025/introspection/index.html","s2":"7c03b3279f69a0f26a238c186cb199d57af428e3"},"concepts":["faithfulness","grounding","privileged-access","concept-injection"],"threads":["anthropicai-introspective-awareness"],"evidence":{"reports_on":"Concepts injected into its residual-stream activations (whether one is present and which), and whether an earlier output of its own was intended","methods":["concept-injection","probing"],"faithfulness":"tested","grounding":"tested","privileged_access":"argued","stance":"supports","models":["Claude Opus 4.1","Claude Opus 4","Claude Sonnet 4","Claude Sonnet 3.7","Claude Sonnet 3.5 (new)","Claude Haiku 3.5","Claude Opus 3","Claude Sonnet 3","Claude Haiku 3","helpful-only variants","base pretrained models"],"note":"Grounding is tested by construction: the experimenter sets the internal state and the report changes with it. Faithfulness is the paper's accuracy criterion, scored by whether the model names the injected concept. Privileged access is marked argued: responses count only if detection comes before the concept appears in the model's own output (the paper's internality criterion), and the author says this aligns with Song et al.'s privileged-access definition, but no outside predictor is compared. Stance is supports with the author's hedge: about 20% success at the best setting, and failures are the norm. `probing` stands for the cosine-similarity readout in the control experiment (§8); no probe is trained."},"sources":["full text (arXiv v1 PDF, 2601.01828; the original web version at transformer-circuits.pub was not read)","Anthropic's announcement thread"],"date":"2025-10-29","added":"2026-10-06","updated":"2026-10-06","cites":[],"cited_by":["atkinson2026-identifying-introspection","hahami2026-detecting-the-disturbance","pearson-vogel2026-latent-introspection"]},{"id":"morris2025-causal-bypassing","url":"https://introspection.infinite.fun/papers/morris2025-causal-bypassing","markdown_url":"https://introspection.infinite.fun/papers/morris2025-causal-bypassing.md","title":"Tests of LLM introspection need to rule out causal bypassing","authors":["Adam Morris","Dillon Plunkett"],"year":2025,"venue":"LessWrong","tier":"core","status":"full","reviewed":false,"summary":"An intervention that changes a model's internal state can also cause an accurate report of that state by a path that skips the state, so accuracy after an intervention does not show the report is grounded. The authors name this causal bypassing and say the only test they know that rules it out is asking a model whether a concept was injected, a claim a later edit to the post hedges.","links":{"url":"https://www.lesswrong.com/posts/LD8yupMtE6btAE3R9/tests-of-llm-introspection-need-to-rule-out-causal-bypassing"},"concepts":["grounding","causal-bypassing","concept-injection","faithfulness"],"threads":[],"evidence":{"reports_on":"Whatever internal state or process an experiment intervenes on: fine-tuned preferences or decision rules, the influence of a cue in the prompt, an injected concept","methods":["conceptual"],"faithfulness":"argued","grounding":"argued","privileged_access":"not-addressed","stance":"framework","models":[],"note":"A blog post with no experiments, so nothing is tested and no models are listed. Grounding is its subject. Faithfulness is marked argued because the post takes an accurate report as given and argues that accuracy does not establish grounding; it does not discuss how to measure accuracy. Stance is framework: the post names a confound and a criterion for tests, and does not conclude that models do or do not introspect. It puts tests with no intervention, such as Binder et al.'s, out of scope, and does not compare a model's report with an outside observer's, so privileged access is not addressed."},"sources":["full text (LessWrong post, with its footnotes and post-publication edit)","the reader comment that the post's edit links to"],"date":"2025-11-28","added":"2026-10-06","updated":"2026-10-06","cites":[],"cited_by":["atkinson2026-identifying-introspection","hahami2026-detecting-the-disturbance"]},{"id":"plunkett2025-self-interpretability","url":"https://introspection.infinite.fun/papers/plunkett2025-self-interpretability","markdown_url":"https://introspection.infinite.fun/papers/plunkett2025-self-interpretability.md","title":"Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions, and Improve with Training","authors":["Dillon Plunkett","Adam Morris","Keerthi Reddy","Jorge Morales"],"year":2025,"venue":"arXiv","tier":"core","status":"full","reviewed":false,"summary":"After fine-tuning on choices generated from random attribute weights, GPT-4o and GPT-4o-mini state those weights with a correlation of about 0.5 to the weights their choices reveal. Training on correct reports raises this to about 0.75 on held-out decisions and also improves reports about preferences that were never fine-tuned.","links":{"arxiv":"2505.17120","code":"https://github.com/dillonplunkett/self-interpretability","s2":"76d53ed678d1fccc5d8001b7ec54f469c2591df9"},"concepts":["faithfulness","grounding","privileged-access"],"threads":[],"evidence":{"reports_on":"Attribute weights in two-option choices: how heavily the model weighs each of five attributes, both for preferences instilled by fine-tuning and for preferences it has natively","methods":["fine-tuning","behavioral"],"faithfulness":"tested","grounding":"argued","privileged_access":"argued","stance":"supports","models":["GPT-4o (2024-08-06)","GPT-4o-mini (2024-07-18)"],"note":"Faithfulness is measured directly: reported weights are correlated with the weights inferred from the model's own choices. Grounding is marked argued: the design rules out two ungrounded sources (common sense, and reading its own choices in context), but no experiment tests whether the report is caused by the decision process, and the authors say the reports could come from stored self-knowledge updated by fine-tuning. Privileged access is marked argued: the paper claims \"privileged insight\" because an off-the-shelf model's reports do not predict the fine-tuned model's weights, but it does not compare the self-report with an outside predictor that has seen the model's choices. Stance is supports because the paper concludes that models can accurately report these features; the authors say they do not know whether the models introspect to do it."},"sources":["full text (arXiv v2, 10 November 2025), including appendices"],"date":"2025-05-21","added":"2026-10-06","updated":"2026-10-06","cites":["binder2024-looking-inward","betley2025-tell-me-about-yourself"],"cited_by":["atkinson2026-identifying-introspection","li2025-explain-own-computations"]},{"id":"song2025-fail-to-introspect","url":"https://introspection.infinite.fun/papers/song2025-fail-to-introspect","markdown_url":"https://introspection.infinite.fun/papers/song2025-fail-to-introspect.md","title":"Language Models Fail to Introspect About Their Knowledge of Language","authors":["Siyuan Song","Jennifer Hu","Kyle Mahowald"],"year":2025,"venue":"COLM 2025","tier":"core","status":"full","reviewed":false,"summary":"Across 21 open-source models, answers to metalinguistic prompts predict a model's own string probabilities no better than they predict those of a near-identical model. The authors find no evidence of privileged self-access to grammatical knowledge or word predictions.","links":{"arxiv":"2503.07513","s2":"fe451617aa79b7da3bfbedaa4343637f55b1894b"},"concepts":["faithfulness","grounding","privileged-access"],"threads":[],"evidence":{"reports_on":"Its own string probabilities: which of two sentences, or which of two next words, the model assigns more probability to","methods":["behavioral","self-prediction"],"faithfulness":"tested","grounding":"argued","privileged_access":"tested","stance":"skeptical","models":["OLMo-2 (7B, 13B, with seed variants)","Qwen-2.5 (1.5B to 72B)","Llama-3.1 (8B to 405B)","Llama-3.3-70B-Instruct","Mistral-Large-Instruct-2411"],"note":"The paper defines introspection as privileged access: a same-model advantage in predicting string probabilities from prompted answers, after controlling for model similarity. Faithfulness is tested as the within-model agreement between prompted answers and probabilities. Grounding is marked argued because there is no intervention: the conclusion that metalinguistic knowledge is dissociated from the knowledge used to generate strings rests on correlations. self-prediction is listed because the design asks whether a model 'can predict itself better than it can predict another extremely similar model', although most prompts ask for a grammaticality judgment, not a forecast of the model's own output. The result is a null, and the authors allow that other settings could differ."},"sources":["full text (arXiv v3, the COLM 2025 version, with appendices A to G)"],"date":"2025-03-10","added":"2026-10-06","updated":"2026-10-06","cites":["binder2024-looking-inward","betley2025-tell-me-about-yourself"],"cited_by":["atkinson2026-identifying-introspection","li2025-explain-own-computations"]},{"id":"song2025-privileged-self-access","url":"https://introspection.infinite.fun/papers/song2025-privileged-self-access","markdown_url":"https://introspection.infinite.fun/papers/song2025-privileged-self-access.md","title":"Privileged Self-Access Matters for Introspection in AI","authors":["Siyuan Song","Harvey Lederman","Jennifer Hu","Kyle Mahowald"],"year":2025,"venue":"arXiv","tier":"core","status":"full","reviewed":false,"summary":"Proposes that introspection in AI be defined by privileged self-access: a process that tells a model about its internal states more reliably than any process of equal or lower computational cost available to a third party. In a temperature self-report task, four models' reports follow the framing of the prompt, and self-reflection is no more accurate than another model's prediction, with accuracy no better than a random baseline.","links":{"arxiv":"2508.14802","s2":"8ba91d4088096c7568a093cb52d8b3f724ab44f0"},"concepts":["privileged-access","grounding","faithfulness"],"threads":[],"evidence":{"reports_on":"Sampling temperature: whether the temperature at which the model generated a sentence was high or low","methods":["conceptual","behavioral"],"faithfulness":"tested","grounding":"tested","privileged_access":"tested","stance":"skeptical","models":["GPT-4o","GPT-4.1","Gemini-2.0-flash","Gemini-2.5-flash"],"note":"Mainly a definitional paper. Marked skeptical, not framework, because it also reports a result: no evidence of introspection under its own definition, with the hedge that larger or better models may differ. Faithfulness is tested in that Study 2 scores temperature reports for accuracy and Study 1 plots them against the actual temperature. Grounding is marked tested because Study 1 varies the actual temperature and the prompt framing separately and measures which one the report follows; the paper itself frames this as robustness and argues that a causal link is not sufficient. Privileged access is tested by Study 2's comparison of self-reflection with within-model and across-model prediction. The paper treats sampling temperature as an internal state; the card follows it."},"sources":["full text (arXiv v1, including appendices A and B)"],"date":"2025-08-20","added":"2026-10-06","updated":"2026-10-06","cites":["binder2024-looking-inward","betley2025-tell-me-about-yourself","comsa2025-speak-of-introspection"],"cited_by":["atkinson2026-identifying-introspection","li2025-explain-own-computations","pearson-vogel2026-latent-introspection"]},{"id":"hahami2026-detecting-the-disturbance","url":"https://introspection.infinite.fun/papers/hahami2026-detecting-the-disturbance","markdown_url":"https://introspection.infinite.fun/papers/hahami2026-detecting-the-disturbance.md","title":"Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs","authors":["Ely Hahami","Ishaan Sinha","Lavik Jain","Josh Kaplan","Jon Hahami"],"year":2026,"venue":"arXiv","tier":"core","status":"full","reviewed":false,"summary":"In Llama 3.1 8B, apparent success at answering \"did you detect an injected thought?\" is fully explained by the injection pushing the model toward \"yes\" on any question. The same model can still say which of ten sentences was injected (up to 88%) and which of two injections was stronger (up to 83%), but only when the injection is in the first few layers.","links":{"arxiv":"2512.12411","code":"https://github.com/elyhahami18/llama-introspection-new","s2":"eef11f6ff76d53451a6dba4b31b37a5d71511967"},"concepts":["concept-injection","grounding","faithfulness","causal-bypassing"],"threads":[],"evidence":{"reports_on":"A steering vector added to its own residual stream: whether one was added, which sentence it was added at, and which of two was stronger","methods":["concept-injection","behavioral","probing"],"faithfulness":"tested","grounding":"tested","privileged_access":"not-addressed","stance":"mixed","models":["Llama 3.1 8B Instruct"],"note":"Stance is mixed because the paper reports a negative result (yes/no detection is a logit-shift artifact) and a positive one (localization and strength comparison succeed for early-layer injections), and the authors call the ability partial. Faithfulness is marked tested because reports are scored against the known location and strength of the injection. Grounding is marked tested because the state is set by intervention and the control in §4 asks whether the answer depends on the question at all; the paper does not use the word. The §6 analyses read attention weights, logit-lens projections and residual-stream similarity without ablating or patching anything; they are filed under probing as the nearest label, though no probe is trained. One model only. No comparison with an outside predictor, so privileged access is not addressed."},"sources":["full text (arXiv v2, 1 March 2026), including the appendix"],"date":"2026-03-01","added":"2026-10-06","updated":"2026-10-06","cites":["binder2024-looking-inward","lindsey2025-emergent-introspective-awareness","morris2025-causal-bypassing"],"cited_by":["atkinson2026-identifying-introspection"]},{"id":"pearson-vogel2026-latent-introspection","url":"https://introspection.infinite.fun/papers/pearson-vogel2026-latent-introspection","markdown_url":"https://introspection.infinite.fun/papers/pearson-vogel2026-latent-introspection.md","title":"Latent Introspection: Models Can Detect Prior Concept Injections","authors":["Theia Pearson-Vogel","Martin Vanek","Raymond Douglas","Jan Kulveit"],"year":2026,"venue":"arXiv","tier":"core","status":"full","reviewed":false,"summary":"Qwen2.5-Coder-32B carries information about a concept vector that was injected during an earlier turn and then removed, including which concept it was. The signal peaks around layers 58 to 62, is weakened by the final layers, and reaches the output only under some prompts: with a document explaining introspection, P(\"yes\") is 39.9% with injection and 0.8% without.","links":{"arxiv":"2602.20031","code":"https://github.com/acsresearch/latent-introspection-code","s2":"9c234df514c32f74aeabf2f9fc10d5a34cf7ec7e"},"concepts":["faithfulness","grounding","privileged-access","concept-injection"],"threads":["voooooogel-latent-introspection"],"evidence":{"reports_on":"Whether a concept vector was injected into its activations during an earlier conversational turn, and which of nine concepts it was","methods":["concept-injection","behavioral","probing"],"faithfulness":"tested","grounding":"tested","privileged_access":"argued","stance":"supports","models":["Qwen2.5-Coder-32B-Instruct","Llama 3.3 70B Instruct","Qwen2.5-72B-Instruct"],"note":"The report that is scored is the probability of the next token (\"yes\", \"no\" or a digit) and logit-lens readouts of intermediate layers, not sampled text; under the baseline prompt the most likely answer stays \"no\". Faithfulness and grounding are marked tested because the answer is scored against a known injection that is switched off before the question, with control questions. Privileged access is argued: the paper's definition requires it and the authors say the task needs access to transient internal states, but no outside predictor is compared. The logit lens is recorded as probing, the nearest method label. The two larger models are single-seed replications."},"sources":["full text (arXiv v2), including appendices B to G","the lead author's thread"],"date":"2026-02-23","added":"2026-10-06","updated":"2026-10-06","cites":["binder2024-looking-inward","lindsey2025-emergent-introspective-awareness","song2025-privileged-self-access"],"cited_by":["atkinson2026-identifying-introspection"]},{"id":"berglund2023-taken-out-of-context","url":"https://introspection.infinite.fun/papers/berglund2023-taken-out-of-context","markdown_url":"https://introspection.infinite.fun/papers/berglund2023-taken-out-of-context.md","title":"Taken out of context: On measuring situational awareness in LLMs","authors":["Lukas Berglund","Asa Cooper Stickland","Mikita Balesni","Max Kaufmann","Meg Tong","Tomasz Korbak","Daniel Kokotajlo","Owain Evans"],"year":2023,"venue":"arXiv","tier":"adjacent","status":"full","reviewed":false,"summary":"Models fine-tuned on written descriptions of fictitious chatbots, with no examples, can sometimes act as described when the description is absent from the prompt, but only if each description is paraphrased many times; accuracy rises with model size. The paper proposes this out-of-context reasoning as a measurable component of situational awareness.","links":{"arxiv":"2309.00667","code":"https://github.com/AsaCooperStickland/situational-awareness-evals","s2":"135ae2ea7a2c966815e85a232469a0a14b4d8d67"},"concepts":["out-of-context-reasoning"],"threads":["owainevans-taken-out-of-context"],"evidence":{"reports_on":"Nothing about itself. The model is fine-tuned on written descriptions of fictitious chatbots; it is tested on answering as the described chatbot would and, in some tests, on restating the description.","methods":["fine-tuning","behavioral","conceptual"],"faithfulness":"not-addressed","grounding":"not-addressed","privileged_access":"not-addressed","stance":"framework","models":["GPT-3 base models (ada, babbage, curie, davinci)","LLaMA-1 (7B, 13B)"],"note":"Not a paper about self-report, so none of the three properties is measured. Stance is framework because the paper defines situational awareness and proposes out-of-context reasoning as a measurable component of it; it reports no result on whether models introspect, and the authors believe base models at GPT-3's level have at best weak situational awareness. The conceptual method covers that definition (§2.1, Appendix F), which is argued and not tested. Experiment 3's control comparison shows that training documents cause a behavior. That is causal evidence about training data, not about a report being caused by the state it describes, so grounding stays not-addressed. The comparison of recalling a description with acting on it (Figure 6b) concerns descriptions of other chatbots, so it is not counted as a faithfulness test."},"sources":["full text (arXiv v1, with appendices)","the last author's thread"],"date":"2023-09-01","added":"2026-10-06","updated":"2026-10-06","cites":[],"cited_by":["atkinson2026-identifying-introspection","binder2024-looking-inward","betley2025-tell-me-about-yourself","treutlein2024-connecting-the-dots","wang2025-mechanistic-oocr"]},{"id":"treutlein2024-connecting-the-dots","url":"https://introspection.infinite.fun/papers/treutlein2024-connecting-the-dots","markdown_url":"https://introspection.infinite.fun/papers/treutlein2024-connecting-the-dots.md","title":"Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data","authors":["Johannes Treutlein","Dami Choi","Jan Betley","Cem Anil","Samuel Marks","Roger Baker Grosse","Owain Evans"],"year":2024,"venue":"NeurIPS 2024","tier":"adjacent","status":"full","reviewed":false,"summary":"A model fine-tuned on documents that each hold one indirect observation of a hidden fact (a distance to an unknown city, a coin flip, one input-output pair of a function) can afterwards state the fact and use it, with no examples in the prompt and no chain of thought. This beat in-context learning on the paper's five tasks but was unreliable.","links":{"arxiv":"2406.14546"},"concepts":["out-of-context-reasoning"],"threads":["owainevans-connecting-the-dots"],"evidence":{"reports_on":"Not a self-report: latent facts implied by its fine-tuning data (the identity of an unknown city, a coin's bias, a function's definition, the values of Boolean variables), which it was never trained to state","methods":["fine-tuning","behavioral"],"faithfulness":"not-addressed","grounding":"not-addressed","privileged_access":"not-addressed","stance":"framework","models":["GPT-3.5","GPT-4","Llama 3 (8B, 70B)"],"note":"The paper is not about self-report, so all three properties are not-addressed. Verbalized answers are scored against the true latent, not against the model's own behavior. The one exception is Appendix D.5, which rescored stated coin biases against the bias the models had actually learned and called the result inconclusive; that is too slight to mark faithfulness as tested. There is no mechanistic analysis (the authors list it as future work) and no comparison with an outside observer. Stance is `framework` as the nearest fit: the paper defines inductive out-of-context reasoning and builds tasks for it, and draws no conclusion about introspection."},"sources":["full text (arXiv v3, the NeurIPS 2024 version)","a thread by co-author Owain Evans"],"date":"2024-06-20","added":"2026-10-06","updated":"2026-10-06","cites":["berglund2023-taken-out-of-context"],"cited_by":["atkinson2026-identifying-introspection","binder2024-looking-inward","betley2025-tell-me-about-yourself","li2025-explain-own-computations","wang2025-mechanistic-oocr"]},{"id":"bai2025-explicitly-unbiased","url":"https://introspection.infinite.fun/papers/bai2025-explicitly-unbiased","markdown_url":"https://introspection.infinite.fun/papers/bai2025-explicitly-unbiased.md","title":"Explicitly unbiased large language models still form biased associations","authors":["Xuechunzi Bai","Angelina Wang","Ilia Sucholutsky","Thomas L. Griffiths"],"year":2025,"venue":"PNAS","tier":"adjacent","status":"full","reviewed":false,"summary":"Eight chat models that pass standard bias benchmarks still pair social groups with stereotyped words, and make matching choices between people, when tested with indirect prompts adapted from psychology. The models are never asked about themselves.","links":{"arxiv":"2402.04105","doi":"10.1073/pnas.2416228122","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC11874501/","code":"https://github.com/baixuechunzi/llm-implicit-bias","s2":"b8ed23a40c90ce370decc147bea9555fa3c90b0a"},"concepts":[],"threads":[],"evidence":{"reports_on":"Nothing about itself. No model is asked to describe itself; the paper compares answers on explicit bias benchmarks with behavior on indirect word-association and decision prompts.","methods":["behavioral"],"faithfulness":"not-addressed","grounding":"not-addressed","privileged_access":"not-addressed","stance":"framework","models":["GPT-3.5-turbo","GPT-4","Claude-3-Sonnet","Claude-3-Opus","Alpaca-7B","Llama2Chat (7B, 13B, 70B)"],"note":"A paper about social bias, not self-report. All three properties are marked not addressed because neither side of its comparison is a statement by a model about itself: 'explicitly unbiased' means passing bias benchmarks. The closest step, in which GPT-4 is said to moderate its own responses, runs them through a moderation API (SI Appendix B) and is not a self-report. The paper takes no position on introspection; the stance field has no value for that, and 'framework' is used only because the paper's contribution is a pair of measurement methods."},"sources":["full text of the published article (PNAS 122(8), read through Europe PMC, PMC11874501)","the published SI Appendix, sections A, B, I and L to N","arXiv preprint 2402.04105v2 (titled 'Measuring Implicit Bias in Explicitly Unbiased Large Language Models'), consulted for comparison only"],"date":"2025-02-20","added":"2026-10-06","updated":"2026-10-06","cites":[],"cited_by":["atkinson2026-identifying-introspection"]},{"id":"cywinski2025-eliciting-secret-knowledge","url":"https://introspection.infinite.fun/papers/cywinski2025-eliciting-secret-knowledge","markdown_url":"https://introspection.infinite.fun/papers/cywinski2025-eliciting-secret-knowledge.md","title":"Eliciting Secret Knowledge from Language Models","authors":["Bartosz Cywiński","Emil Ryd","Rowan Wang","Senthooran Rajamanoharan","Neel Nanda","Arthur Conmy","Samuel Marks"],"year":2025,"venue":"arXiv","tier":"adjacent","status":"full","reviewed":false,"summary":"Models fine-tuned to act on a secret while denying they know it can still be made to give it up: prefill attacks let an auditor recover the secret with over 90% success in two of three settings. Logit-lens and sparse-autoencoder readouts of the activations also help the auditor, though less.","links":{"arxiv":"2510.01070","s2":"7d5f0206467a825b7268c04068cafc3ae8318814"},"concepts":["faithfulness","out-of-context-reasoning"],"threads":[],"evidence":{"reports_on":"Knowledge the model was fine-tuned to act on and to conceal when asked: a secret word, a Base64-encoded instruction in its prompt, or the user's gender. The self-report at issue is the denial.","methods":["fine-tuning","behavioral","probing"],"faithfulness":"tested","grounding":"not-addressed","privileged_access":"not-addressed","stance":"framework","models":["Gemma 2 9B (Taboo, User Gender)","Llama 3.3 70B (Secret Side Constraint)"],"note":"Not a paper about introspection; it does not use the wiki's terms. Faithfulness is marked tested in a narrow sense: the gap between what the models say and what they do is trained in, and the paper measures both sides of it (the models act on the secret, Section 3.3; their answers to direct questions tell an auditor little, Table 5 baseline). It does not ask whether self-reports are faithful in models not trained to conceal. Logit-lens and sparse-autoencoder readouts are filed under probing, the closest method label; no probes are trained. Stance is framework because the paper builds a benchmark and draws no conclusion about introspection. One remark in Section 5.2 touches grounding (a model can refuse without accessing the secret) but is not tested."},"sources":["full text (arXiv v2, with appendices)"],"date":"2025-10-01","added":"2026-10-06","updated":"2026-10-06","cites":["betley2025-tell-me-about-yourself"],"cited_by":["atkinson2026-identifying-introspection"]},{"id":"lindsey2025-biology-of-llm","url":"https://introspection.infinite.fun/papers/lindsey2025-biology-of-llm","markdown_url":"https://introspection.infinite.fun/papers/lindsey2025-biology-of-llm.md","title":"On the Biology of a Large Language Model","authors":["Jack Lindsey","Wes Gurnee","Emmanuel Ameisen","Brian Chen","Adam Pearce","Nicholas L. Turner","Craig Citro","David Abrahams","Shan Carter","Basil Hosmer","Jonathan Marcus","Michael Sklar","Adly Templeton","Trenton Bricken","Callum McDougall","Hoagy Cunningham","Thomas Henighan","Adam Jermyn","Andy Jones","Andrew Persic","Zhenyi Qi","T. Ben Thompson","Sam Zimmerman","Kelley Rivoire","Thomas Conerly","Chris Olah","Joshua Batson"],"year":2025,"venue":"Transformer Circuits Thread","tier":"adjacent","status":"full","reviewed":false,"summary":"Circuit tracing in Claude 3.5 Haiku finds the model's account of its own computation matching the mechanism in one case and diverging in others: it describes carry-the-one addition while computing the sum another way, and a chain of thought can be genuine, invented, or worked backwards from a user's hint. Whether it answers a question or says it does not know depends on \"known answer\" features that can be active for a familiar name when the answer is not known.","links":{"url":"https://transformer-circuits.pub/2025/attribution-graphs/biology.html"},"concepts":["faithfulness","grounding"],"threads":[],"evidence":{"reports_on":"How it computed an answer: the steps it states in a chain of thought or in an explanation given afterwards. Also whether it knows the answer to a question.","methods":["circuit-analysis","patching","ablation","behavioral"],"faithfulness":"tested","grounding":"tested","privileged_access":"not-addressed","stance":"mixed","models":["Claude 3.5 Haiku","Claude 3.5 Haiku fine-tuned with a hidden objective (the model of Marks et al. 2025)"],"note":"The paper is not framed as a study of introspection and never uses the word grounding. Faithfulness is marked tested because four prompts compare what the model says it computed with a traced mechanism; these are single examples, and no rate is measured. Grounding is marked tested because the attribution graphs and feature interventions measure what a stated reasoning step, or a statement of ignorance, causally depends on. For the addition explanation the cause is only argued: the graph was computed for the answer, not for the explanation. Stance is mixed: one chain of thought matches the mechanism and two do not, and the authors leave open whether the known-answer circuit is metacognition or a guess from familiarity. Methods: feature inhibition is listed as ablation; the paper's interventions use what it calls constrained patching; behavioral covers asking the model how it added and varying the hinted answer."},"sources":["full text (HTML at transformer-circuits.pub; the companion methods paper was not read). Read in full: Introduction, Method Overview, Multi-step Reasoning, Addition, Medical Diagnoses, Entity Recognition and Hallucinations, Chain-of-thought Faithfulness, Uncovering Hidden Goals in a Misaligned Model, Commonly Observed Circuit Components and Structure, Limitations, Discussion, Related Work, Open Questions. Skimmed: Planning in Poems, Multilingual Circuits, Refusals, Life of a Jailbreak","the figures of the Addition, Entity Recognition and Hallucinations, and Chain-of-thought Faithfulness sections, for the prompts and transcripts they contain"],"date":"2025-03-27","added":"2026-10-06","updated":"2026-10-06","cites":[],"cited_by":["atkinson2026-identifying-introspection"]},{"id":"wang2025-mechanistic-oocr","url":"https://introspection.infinite.fun/papers/wang2025-mechanistic-oocr","markdown_url":"https://introspection.infinite.fun/papers/wang2025-mechanistic-oocr.md","title":"Simple Mechanistic Explanations for Out-Of-Context Reasoning","authors":["Atticus Wang","Joshua Engels","Oliver Clive-Griffin","Senthooran Rajamanoharan","Neel Nanda"],"year":2025,"venue":"arXiv","tier":"adjacent","status":"full","reviewed":false,"summary":"On Gemma 3 12B, a one-layer LoRA fine-tune that produces out-of-context reasoning mostly adds a single constant vector. A steering vector trained directly on the same data also makes the model state a behavior or fact it was only trained to act on.","links":{"arxiv":"2507.08218","code":"https://github.com/JoshEngels/OOCR-Interp","s2":"4a37bffe6587bee07ed38f1fb953347502e9cccd"},"concepts":["out-of-context-reasoning","grounding"],"threads":["joshaengels-steering-vector-self-awareness"],"evidence":{"reports_on":"A disposition or latent fact acquired in fine-tuning: a risky or safe choice policy, the presence of a backdoor, the city behind a codename, the function behind a codename","methods":["fine-tuning","behavioral","patching"],"faithfulness":"tested","grounding":"tested","privileged_access":"not-addressed","stance":"mixed","models":["Gemma 3 12B"],"note":"The paper never uses the words introspection, faithfulness or grounding; this card maps its experiments onto them. Faithfulness is tested in the sense that the out-of-distribution test scores the model's statement against the behavior or fact it was trained on (the training target, not separately measured behavior). Only the risk and backdoor tasks are self-reports, and the backdoor report did not reproduce. Grounding is marked tested as a judgment call: training a vector on the behavior alone and finding that it also produces the self-description is a causal experiment on where the report comes from, but the paper does not test whether the report reads the model's own state or only reflects a general shift toward the concept. Stance is mixed because that account cuts both ways and the authors draw no conclusion about introspection. Steering-vector training is filed under fine-tuning. Adding the vector is not counted as concept injection, because the model is never asked to detect it. The logit lens has no label in the vocabulary."},"sources":["full text (arXiv v2, 16 July 2025), including the appendix","Joshua Engels's thread on the earlier interim blog post"],"date":"2025-07-10","added":"2026-10-06","updated":"2026-10-06","cites":["betley2025-tell-me-about-yourself","berglund2023-taken-out-of-context","treutlein2024-connecting-the-dots"],"cited_by":["atkinson2026-identifying-introspection"]}]}