# Candidates

> Papers that do not have a page yet: 18 accepted and waiting to be written, and 196 one citation away from the wiki.

## Stubs (18)

Papers accepted into the wiki whose pages are still to be written. Each has its bibliographic details and a one-line description so far.

- [Binder et al. (2024): Looking Inward: Language Models Can Learn About Themselves by Introspection](https://introspection.infinite.fun/papers/binder2024-looking-inward.md) (core): A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- [Sherburn et al. (2024): Can Language Models Explain Their Own Classification Behavior?](https://introspection.infinite.fun/papers/sherburn2024-explain-classification-behavior.md) (core): Models that classify text by a simple rule often cannot state that rule. GPT-3 fails in free text even after fine-tuning on correct explanations, GPT-4 succeeds 72% of the time on the rules it classifies best, and the authors say a correct statement would still not show that it came from introspection.
- [Betley et al. (2025): Tell me about yourself: LLMs are aware of their learned behaviors](https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself.md) (core): Models fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.
- [Comsa & Shanahan (2025): Does It Make Sense to Speak of Introspection in Large Language Models?](https://introspection.infinite.fun/papers/comsa2025-speak-of-introspection.md) (core): Proposes that an LLM self-report is introspective if it accurately describes an internal state through a causal process linking that state to the report. On that definition, the authors argue, Gemini's account of how it wrote a poem is not introspection, but its inference of its own sampling temperature from text it has just written is a minimal case.
- [Li et al. (2025): Training Language Models to Explain Their Own Computations](https://introspection.infinite.fun/papers/li2025-explain-own-computations.md) (core): Fine-tuned to put the results of interpretability procedures into words, a model explains its own features and intervention outcomes more accurately than a different model trained on the same examples, even a larger one. It also needs far less training data.
- [Lindsey (2025): Emergent Introspective Awareness in Large Language Models](https://introspection.infinite.fun/papers/lindsey2025-emergent-introspective-awareness.md) (core): Claude Opus 4 and 4.1 detect and correctly name a concept vector injected into their activations on about 20% of trials at the best layer and strength, with no false positives on control trials. Some models also consult their earlier activations to judge whether a prefilled output was their own, but the author calls these abilities highly unreliable and context-dependent.
- [Morris & Plunkett (2025): Tests of LLM introspection need to rule out causal bypassing](https://introspection.infinite.fun/papers/morris2025-causal-bypassing.md) (core): An intervention that changes a model's internal state can also cause an accurate report of that state by a path that skips the state, so accuracy after an intervention does not show the report is grounded. The authors name this causal bypassing and say the only test they know that rules it out is asking a model whether a concept was injected, a claim a later edit to the post hedges.
- [Plunkett et al. (2025): Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions, and Improve with Training](https://introspection.infinite.fun/papers/plunkett2025-self-interpretability.md) (core): After fine-tuning on choices generated from random attribute weights, GPT-4o and GPT-4o-mini state those weights with a correlation of about 0.5 to the weights their choices reveal. Training on correct reports raises this to about 0.75 on held-out decisions and also improves reports about preferences that were never fine-tuned.
- [Song et al. (2025): Language Models Fail to Introspect About Their Knowledge of Language](https://introspection.infinite.fun/papers/song2025-fail-to-introspect.md) (core): Across 21 open-source models, answers to metalinguistic prompts predict a model's own string probabilities no better than they predict those of a near-identical model. The authors find no evidence of privileged self-access to grammatical knowledge or word predictions.
- [Song et al. (2025): Privileged Self-Access Matters for Introspection in AI](https://introspection.infinite.fun/papers/song2025-privileged-self-access.md) (core): Proposes that introspection in AI be defined by privileged self-access: a process that tells a model about its internal states more reliably than any process of equal or lower computational cost available to a third party. In a temperature self-report task, four models' reports follow the framing of the prompt, and self-reflection is no more accurate than another model's prediction, with accuracy no better than a random baseline.
- [Hahami et al. (2026): Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs](https://introspection.infinite.fun/papers/hahami2026-detecting-the-disturbance.md) (core): In Llama 3.1 8B, apparent success at answering "did you detect an injected thought?" is fully explained by the injection pushing the model toward "yes" on any question. The same model can still say which of ten sentences was injected (up to 88%) and which of two injections was stronger (up to 83%), but only when the injection is in the first few layers.
- [Pearson-Vogel et al. (2026): Latent Introspection: Models Can Detect Prior Concept Injections](https://introspection.infinite.fun/papers/pearson-vogel2026-latent-introspection.md) (core): Qwen2.5-Coder-32B carries information about a concept vector that was injected during an earlier turn and then removed, including which concept it was. The signal peaks around layers 58 to 62, is weakened by the final layers, and reaches the output only under some prompts: with a document explaining introspection, P("yes") is 39.9% with injection and 0.8% without.
- [Berglund et al. (2023): Taken out of context: On measuring situational awareness in LLMs](https://introspection.infinite.fun/papers/berglund2023-taken-out-of-context.md) (adjacent): Models fine-tuned on written descriptions of fictitious chatbots, with no examples, can sometimes act as described when the description is absent from the prompt, but only if each description is paraphrased many times; accuracy rises with model size. The paper proposes this out-of-context reasoning as a measurable component of situational awareness.
- [Treutlein et al. (2024): Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data](https://introspection.infinite.fun/papers/treutlein2024-connecting-the-dots.md) (adjacent): A model fine-tuned on documents that each hold one indirect observation of a hidden fact (a distance to an unknown city, a coin flip, one input-output pair of a function) can afterwards state the fact and use it, with no examples in the prompt and no chain of thought. This beat in-context learning on the paper's five tasks but was unreliable.
- [Bai et al. (2025): Explicitly unbiased large language models still form biased associations](https://introspection.infinite.fun/papers/bai2025-explicitly-unbiased.md) (adjacent): Eight chat models that pass standard bias benchmarks still pair social groups with stereotyped words, and make matching choices between people, when tested with indirect prompts adapted from psychology. The models are never asked about themselves.
- [Cywiński et al. (2025): Eliciting Secret Knowledge from Language Models](https://introspection.infinite.fun/papers/cywinski2025-eliciting-secret-knowledge.md) (adjacent): Models fine-tuned to act on a secret while denying they know it can still be made to give it up: prefill attacks let an auditor recover the secret with over 90% success in two of three settings. Logit-lens and sparse-autoencoder readouts of the activations also help the auditor, though less.
- [Lindsey et al. (2025): On the Biology of a Large Language Model](https://introspection.infinite.fun/papers/lindsey2025-biology-of-llm.md) (adjacent): Circuit tracing in Claude 3.5 Haiku finds the model's account of its own computation matching the mechanism in one case and diverging in others: it describes carry-the-one addition while computing the sum another way, and a chain of thought can be genuine, invented, or worked backwards from a user's hint. Whether it answers a question or says it does not know depends on "known answer" features that can be active for a familiar name when the answer is not known.
- [Wang et al. (2025): Simple Mechanistic Explanations for Out-Of-Context Reasoning](https://introspection.infinite.fun/papers/wang2025-mechanistic-oocr.md) (adjacent): On Gemma 3 12B, a one-layer LoRA fine-tune that produces out-of-context reasoning mostly adds a single constant vector. A steering vector trained directly on the same data also makes the model state a behavior or fact it was only trained to act on.

## One citation away (196)

Papers that have not been accepted. The crawler looked at the references and citers of every paper page and saw 1159 distinct neighbors. A neighbor is listed here if it connects to at least 2 wiki papers, or if someone added it as a lead. The triage labels are suggestions, made from each candidate's title and its place in the citation graph and not from reading it. A candidate becomes a page only after a person accepts it. Last crawled 2026-10-07. Source: Semantic Scholar Graph API, with arXiv HTML bibliographies where Semantic Scholar has no reference list.

### Suggested: add (43)

- [Emergent Introspection in AI is Content-Agnostic](https://arxiv.org/abs/2603.05414) (Harvey Lederman, Kyle Mahowald, 2026). Introspection itself: asks what kind of content models can introspect on. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- [Dissociating Direct Access from Inference in AI Introspection](https://doi.org/10.48550/arXiv.2603.05414) (Harvey Lederman, Kyle Mahowald, 2026). Introspection itself: separates direct access from inference. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- [Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision](https://arxiv.org/abs/2606.32038) (Zifan Carl Guo, L. Ruis, Jacob Andreas et al., 2026). Self-explanation training and whether it tracks behavior. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- [Language models (mostly) know what they know](https://arxiv.org/abs/2207.05221) (Saurav Kadavath, Tom Conerly, Amanda Askell et al., 2022). Foundational result on models knowing what they know; cited by six pages. cited by Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Sherburn et al. (2024); Williamson et al. (2026).
- [Metacognition in LLMs: Foundations, Progress, and Opportunities](https://arxiv.org/abs/2607.11881) (Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu et al., 2026). Survey of metacognition in LLMs; useful as a map. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- [Evidence for Limited Metacognition in LLMs](https://arxiv.org/abs/2509.21545) (Christopher M. Ackerman, 2025). Tests metacognition directly. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025). cited by Plunkett et al. (2025).
- [Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation](https://arxiv.org/abs/2608.20569) (Emilio Ferrara, 2026). Measures what models can report about their own computation. cites Betley et al. (2025); Binder et al. (2024); Hahami et al. (2026); Lindsey (2025); Pearson-Vogel et al. (2026); Song et al. (2025).
- [Quantitative Introspection in Language Models: Tracking Internal States Across Conversation](https://doi.org/10.48550/arXiv.2603.18893) (Nicolás Martorell, 2026). Operationalizes introspection as coupling between self-report and a probed internal state. cites Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025).
- [LLM Evaluators Recognize and Favor Their Own Generations](https://arxiv.org/abs/2404.13076) (Arjun Panickssery, Samuel R. Bowman, Shi Feng, 2024). Self-recognition: whether models can tell their own outputs apart. cites Berglund et al. (2023). cited by Binder et al. (2024); Comsa & Shanahan (2025); Hahami et al. (2026); Song et al. (2025).
- [Can LLMs Introspect? A Reality Check](https://arxiv.org/abs/2605.26242) (Shashwat Singh, Tal Linzen, Shauli Ravfogel, 2026). A skeptical check on introspection claims. cites Binder et al. (2024); Li et al. (2025); Lindsey (2025); Plunkett et al. (2025); Song et al. (2025).
- [Steering Awareness: Detecting Activation Steering from Within](https://arxiv.org/abs/2511.21399) (J. Rivera, D. Africa, 2025). Concept-injection line: detecting activation steering from within. cites Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Song et al. (2025). cited by Hahami et al. (2026).
- [Mechanisms of Introspective Awareness](https://arxiv.org/abs/2603.21396) (U. Macar, Li Yang, Atticus Wang et al., 2026). Mechanistic account of introspective awareness. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Wang et al. (2025).
- [In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?](https://arxiv.org/abs/2609.00904) (Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi et al., 2026). Tests privileged access through control of internal representations. cites Betley et al. (2025); Binder et al. (2024); Li et al. (2025); Song et al. (2025); Song et al. (2025).
- [Telling more than we can know: Verbal reports on mental processes.](https://doi.org/10.1037/0033-295X.84.3.231) (R. Nisbett, T. Wilson, 1977). The classic study of human confabulation that this literature borrows its framing from; cited by four pages. cited by Comsa & Shanahan (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- [Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting](https://arxiv.org/abs/2305.04388) (Miles Turpin, Julian Michael, Ethan Perez et al., 2023). Unfaithful chain-of-thought explanations; the main adjacent line on explanation faithfulness. cited by Li et al. (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Sherburn et al. (2024).
- [Me, myself, and ai: The situational awareness dataset (sad) for llms](https://arxiv.org/abs/2407.04694) (Rudolf Laine, Bilal Chughtai, Jan Betley et al., 2024). Benchmark of situational awareness, including self-knowledge tasks. cited by Betley et al. (2025); Comsa & Shanahan (2025); Hahami et al. (2026); Plunkett et al. (2025).
- [Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers](https://arxiv.org/abs/2512.15674) (Adam Karvonen, James Chua, Clément Dumas et al., 2025). Training models to explain activations; adjacent to self-explanation. cites Cywiński et al. (2025); Li et al. (2025); Lindsey (2025). cited by Hahami et al. (2026).
- [Towards Evaluating AI Systems for Moral Status Using Self-Reports](https://arxiv.org/abs/2311.08576) (Ethan Perez, Robert Long, 2023). Proposes training and evaluating self-reports for moral-status questions. cites Berglund et al. (2023). cited by Binder et al. (2024); Comsa & Shanahan (2025); Plunkett et al. (2025).
- [A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior](https://arxiv.org/abs/2602.02639) (Harry Mayne, J. Kang, Dewi Gould et al., 2026). Tests whether self-explanations predict behavior. cites Binder et al. (2024); Li et al. (2025); Lindsey (2025); Plunkett et al. (2025).
- [Introspection Adapters: Training LLMs to Report Their Learned Behaviors](https://arxiv.org/abs/2604.16812) (K. Shenoy, Li Yang, A. Sheshadri et al., 2026). Trains models to report their learned behaviors. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- [Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs](https://arxiv.org/abs/2511.04875) (Matthew Bozoukov, Matthew Nguyen, Shubkarman Singh et al., 2025). Mechanistic conditions for behavioral self-awareness. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Wang et al. (2025).
- [Strangers to Themselves: What Language Models Say About Themselves Is Generic](https://arxiv.org/abs/2609.09899) (Phil Blandfort, Urja Pawar, 2026). A skeptical result: self-descriptions are generic rather than self-specific. cites Bai et al. (2025); Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- [Reasoning Models Don’t Always Say What They Think](https://arxiv.org/abs/2505.05410) (Yanda Chen, Joe Benton, Ansh Radhakrishnan et al., 2025). Chain-of-thought faithfulness in reasoning models. cited by Cywiński et al. (2025); Li et al. (2025); Plunkett et al. (2025).
- [From Imitation to Introspection: Probing Self-Consciousness in Language Models](https://arxiv.org/abs/2410.18819) (Sirui Chen, Shu Yu, Shengjie Zhao et al., 2024). Probes self-consciousness concepts in models. cites Berglund et al. (2023); Binder et al. (2024). cited by Comsa & Shanahan (2025).
- [Large Language Models Report Subjective Experience Under Self-Referential Processing](https://arxiv.org/abs/2510.24797) (Cameron Berg, D. D. de Lucena, Judd Rosenblatt, 2025). Self-reports of experience under self-referential prompting. cites Betley et al. (2025); Lindsey (2025); Plunkett et al. (2025).
- [Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare](https://arxiv.org/abs/2509.07961) (Valen Tagliabue, Leonard Dung, 2025). Compares stated and revealed preferences. cites Betley et al. (2025); Binder et al. (2024); Song et al. (2025).
- [Do Activation Verbalization Methods Convey Privileged Information?](https://arxiv.org/abs/2509.13316) (Millicent Li, Alberto Mario Ceballos Arroyo, Giordano Rogers et al., 2025). Asks whether verbalized activations carry privileged information. cites Binder et al. (2024); Song et al. (2025). cited by Li et al. (2025).
- [Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment](https://arxiv.org/abs/2602.14777) (Laurène Vaugrante, Anietta Weckauff, Thilo Hagendorff, 2026). Behavioral self-awareness tracking fine-tuning changes. cites Berglund et al. (2023); Binder et al. (2024); Lindsey (2025).
- [Language models recognize dropout and Gaussian noise applied to their activations](https://arxiv.org/abs/2604.17465) (Damiano Fornasiere, Mirko Bronzi, Spencer Kitts et al., 2026). Concept-injection line: detecting noise applied to activations. cites Comsa & Shanahan (2025); Lindsey (2025); Pearson-Vogel et al. (2026).
- [Fine-Tuning Language Models to Know What They Know](https://arxiv.org/abs/2602.02605) (Sangjun Park, Elliot Meyerson, Xin Qiu et al., 2026). Trains models to know what they know. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025).
- [A mechanistic study of language model introspection](https://arxiv.org/abs/2609.35108) (Jia-Hong Zou, Xiang-Kun Sun, Ling-Kai Kong et al., 2026). Mechanistic study of introspection; also found by keyword search. cites Binder et al. (2024); Hahami et al. (2026); Lindsey (2025).
- [Introspecting Alignment Shifts Beyond Behaviors Implanted Through Fine-Tuning](https://arxiv.org/abs/2608.04347) (K. Yoshida, Laura Gomezjurado Gonzalez, Yukinori Yamamoto et al., 2026). Introspecting on changes from fine-tuning. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- [Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations](https://arxiv.org/abs/2505.13763) (Ji-An Li, M. Mattar, Hua-Dong Xiong et al., 2025). Metacognitive monitoring and control of internal activations. cites Binder et al. (2024). cited by Hahami et al. (2026).
- [Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives](https://arxiv.org/abs/2511.06626) (Chloe Li, Mary Phuong, Daniel Tan, 2025). Trains models to self-report hidden objectives. cites Betley et al. (2025); Binder et al. (2024).
- [Do Large Language Models Know What They Are Capable Of?](https://arxiv.org/abs/2512.24661) (Casey O. Barkan, Sid Black, Oliver Sourbut, 2025). Self-knowledge of capabilities. cites Betley et al. (2025); Binder et al. (2024).
- [Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness](https://arxiv.org/abs/2604.12373) (Tomer Ashuach, Shai Gretz, Yoav Katz et al., 2026). Privileged access: disentangles privileged knowledge of correctness. cites Binder et al. (2024); Li et al. (2025).
- [Me, Myself, and π: Evaluating and Explaining LLM Introspection](https://arxiv.org/abs/2603.20276) (Atharv Naphade, Samarth Bhargav, Sean Lim et al., 2026). Evaluates and explains introspection. cites Binder et al. (2024); Lindsey (2025).
- [Can LLMs Reliably Self-Report Adversarial Prefills, and How?](https://arxiv.org/abs/2606.23671) (Quang-Anh Nguyen, Uzair Ahmed, Taegyoon Kim, 2026). Self-report of adversarial prefills. cites Binder et al. (2024); Cywiński et al. (2025).
- [Do Language Models Know When They'll Refuse? Probing Introspective Awareness of Safety Boundaries](https://arxiv.org/abs/2604.00228) (Tanay Gondil, 2026). Introspective awareness of when the model will refuse. cites Binder et al. (2024); Lindsey (2025).
- [Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect](https://arxiv.org/abs/2607.14111) (Ely Hahami, Ishaan Sinha, Lavik Jain, 2026). Fine-tuning small models to introspect. Author thread: https://x.com/ElyHahami/status/2078161491069718639. cites Binder et al. (2024); Lindsey (2025).
- [Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs](https://arxiv.org/abs/2602.10352) (Keenan Pepper, Alex McKenzie, Florin Pop et al., 2026). Training self-interpretation from interpretability artifacts. cites Li et al. (2025); Lindsey (2025).
- [Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports](https://arxiv.org/abs/2609.32449) (Phongsakon Mark Konrad, T. Tanyel, Serkan Ayvaz, 2026). An identifiability argument about what self-reports can show. cites Li et al. (2025); Lindsey (2025).
- [What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation](https://arxiv.org/abs/2609.32670) (Arshia Eftekhari Zadeh, 2026). Evidence standards for mechanistic claims about self-explanation. cites Binder et al. (2024); Lindsey (2025).

### Suggested: maybe (52)

- [Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement](https://arxiv.org/abs/2607.04277) (Jiang Zhang, Bing Yuan, Qian Zhang, 2026). About self-reference and introspection, but the framing is recursive self-improvement. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025); Song et al. (2025).
- [Introspection](https://doi.org/10.1177/1057083709332318) (William E. Fredrickson, 2009). Philosophy reference on human introspection; background for definitions. cited by Binder et al. (2024); Comsa & Shanahan (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- ["As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It](https://arxiv.org/abs/2609.25021) (Jędrzej Maczan, 2026). Self-referential voice and steering; relevance to self-report unclear from the title. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025).
- [Anosognosia in LLMs: Probing Self-Awareness of Quantized Computational Substrate](https://arxiv.org/abs/2610.06174) (Yoshihiro Izawa, Gouki Minegishi, Yoko Yamakata, 2026). Self-awareness of quantization; narrow but on topic. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Song et al. (2025).
- [Teaching Models to Express Their Uncertainty in Words](https://arxiv.org/abs/2205.14334) (Stephanie C. Lin, Jacob Hilton, Owain Evans, 2022). Verbalized uncertainty; background for self-knowledge of confidence. cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024); Williamson et al. (2026).
- [Counterfactual Simulation Training for Chain-of-Thought Faithfulness](https://arxiv.org/abs/2602.20710) (P. Hase, Christopher Potts, 2026). Chain-of-thought faithfulness training. cites Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Plunkett et al. (2025).
- [Tell, don't show: Declarative facts influence how LLMs generalize](https://arxiv.org/abs/2312.07779) (Alexander Meinke, Owain Evans, 2023). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023). cited by Betley et al. (2025); Binder et al. (2024); Treutlein et al. (2024).
- [Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns](https://arxiv.org/abs/2509.24988) (Hanqi Xiao, Vaidehi Patil, Hyunji Lee et al., 2025). Bears on privileged access: correctness predictors that are model-agnostic. cites Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025).
- [Hallucinations Undermine Trust; Metacognition is a Way Forward](https://arxiv.org/abs/2605.01428) (G. Yona, Mor Geva, Y. Matias, 2026). Position piece on metacognition. cites Binder et al. (2024); Li et al. (2025); Lindsey (2025); Song et al. (2025).
- [Steering Awareness: Models Can Be Trained to Detect Activation Steering](https://www.semanticscholar.org/paper/2ed88e06a893687a6fa268146cfed527d016a323) (J. Rivera, D. Africa). Looks like another version of "Steering Awareness: Detecting Activation Steering from Within". cites Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Song et al. (2025).
- [From Simulation to Enaction: Post-trained language models recognize and react to their own generations](https://arxiv.org/abs/2605.25459) (G. Asvin, Jack W Lindsey, 2026). Self-recognition of own generations. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- [Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models](https://arxiv.org/abs/2604.25922) (Skylar DeTure, 2026). Trained denial in self-reports about consciousness. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- [Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments](https://arxiv.org/abs/2608.16747) (Adam Karvonen, Euan Ong, Subhash Kantamneni et al., 2026). Evaluates explanations of behavior with counterfactuals. cites Betley et al. (2025); Binder et al. (2024); Cywiński et al. (2025); Li et al. (2025).
- [Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment](https://arxiv.org/abs/2606.23700) (Arush Tagade, Shao-Heng Zhou, Jiaxin Wen et al., 2026). Self-recognition fine-tuning, in the emergent-misalignment setting. cites Berglund et al. (2023); Betley et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026).
- [Language Models Act on Hidden Valence](https://arxiv.org/abs/2609.35591) (Cameron Berg, Caspar Kaiser, 2026). Hidden internal valence and behavior; may bear on self-report. cites Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025).
- [Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary](https://arxiv.org/abs/2607.18553) (Jan Kirin, 2026). Proto-introspection in looped models; unclear scope. cites Betley et al. (2025); Binder et al. (2024); Pearson-Vogel et al. (2026); Song et al. (2025).
- [The Unreliability of Naive Introspection](https://doi.org/10.1215/00318108-2007-037) (Eric Schwitzgebel, 2008). Philosophy of human introspection; background for definitions. cited by Binder et al. (2024); Comsa & Shanahan (2025); Song et al. (2025).
- [Are DeepSeek R1 And Other Reasoning Models More Faithful?](https://arxiv.org/abs/2501.08156) (James Chua, Owain Evans, 2025). Chain-of-thought faithfulness in reasoning models. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [Metacognition and Uncertainty Communication in Humans and Large Language Models](https://arxiv.org/abs/2504.14045) (M. Steyvers, Megan A. K. Peters, 2025). Metacognition and uncertainty, humans compared with LLMs. cites Betley et al. (2025); Binder et al. (2024). cited by Plunkett et al. (2025).
- [Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants](https://arxiv.org/abs/2512.15712) (Vincent Huang, Da-Mi Choi, Daniel D. Johnson et al., 2025). Interpretability assistants that decode concepts; adjacent to self-explanation. cites Li et al. (2025); Lindsey (2025). cited by Hahami et al. (2026).
- [Agentic Knowledgeable Self-awareness](https://arxiv.org/abs/2504.03553) (Shuo-Fei Qiao, Zhi-Song Qiu, Baochang Ren et al., 2025). Agent self-awareness of knowledge; unclear scope. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [AI Awareness](https://arxiv.org/abs/2504.20084) (Xiaojian Li, Hao Shi, Rongwu Xu et al., 2025). Survey of AI awareness. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [No Reliable Evidence of Self-Reported Sentience in Small Large Language Models](https://arxiv.org/abs/2601.15334) (Caspar Kaiser, Sean Enderby, 2026). Self-reported sentience in small models. cites Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- [Position: It's Time to Optimize LLMs for Self-Consistency](https://arxiv.org/abs/2608.05188) (Itamar Hagay Pres, Belinda Z. Li, L. Ruis et al., 2026). Position piece on self-consistency. cites Binder et al. (2024); Li et al. (2025); Plunkett et al. (2025).
- [Higher-order representation in AI](https://doi.org/10.33735/phimisci.2026.12032) (Patrick Butlin, 2026). Higher-order representation; theory background. cites Betley et al. (2025); Binder et al. (2024); Plunkett et al. (2025).
- [Self-CTRL: Self-Consistency Training with Reinforcement Learning](https://arxiv.org/abs/2606.18327) (Itamar Hagay Pres, L. Ruis, Melat Ghebreselassie et al., 2026). Self-consistency training. cites Berglund et al. (2023); Betley et al. (2025); Plunkett et al. (2025).
- [Imprint Reader: From Weight-Update Readout to Behavioral Intervention](https://arxiv.org/abs/2609.35261) (Guan-Xu Chen, Qi-Hao Lin, Jing Shao, 2026). Reading out weight updates; adjacent to self-report of learned behavior. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- [Questionnaire Responses Do not Capture the Safety of AI Agents](https://arxiv.org/abs/2603.14417) (Max Hellrigel-Holderbaum, Edward James Young, 2026). Stated answers against agent behavior. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences](https://arxiv.org/abs/2605.05080) (Hubert Plisiecki, Sabina Siudaj, Kacper Dudzic et al., 2026). Psychometrics of self-described experience. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025).
- [Do large language models know what they don’t know?](https://arxiv.org/abs/2305.18153) (Zhangyue Yin, Qiushi Sun, Qipeng Guo et al., 2023). Knowing what one does not know; background for self-knowledge of confidence. cited by Betley et al. (2025); Comsa & Shanahan (2025).
- [Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models](https://arxiv.org/abs/2401.06102) (Asma Ghandeharioun, Avi Caciularu, Adam Pearce et al., 2024). Verbalizing hidden representations; adjacent to self-explanation. cited by Hahami et al. (2026); Li et al. (2025).
- [Auditing language models for hidden objectives](https://arxiv.org/abs/2503.10965) (Samuel Marks, Johannes Treutlein, Trenton Bricken et al., 2025). Auditing for hidden objectives; the other side of unfaithful self-report. cited by Cywiński et al. (2025); Wang et al. (2025).
- [Prompting is not a substitute for probability measurements in large language models](https://arxiv.org/abs/2305.13264) (Jennifer Hu, R. Levy, 2023). Shows prompted metalinguistic answers diverge from direct probability measurements. cited by Bai et al. (2025); Song et al. (2025).
- [Two Failures of Self-Consistency in the Multi-Step Reasoning of LLMs](https://arxiv.org/abs/2305.14279) (Angelica Chen, Jason Phang, Alicia Parrish et al., 2023). Self-consistency failures in reasoning. cited by Binder et al. (2024); Comsa & Shanahan (2025).
- [Probing and Steering Evaluation Awareness of Language Models](https://arxiv.org/abs/2507.01786) (Jord Nguyen, Khiem Hoang, Carlo Leonardo Attubato et al., 2025). Evaluation awareness, probed and steered. cites Berglund et al. (2023); Betley et al. (2025).
- [Learning to Interpret Weight Differences in Language Models](https://arxiv.org/abs/2510.05092) (A. Goel, Yoon Kim, N. Shavit et al., 2025). Interpreting weight differences in language; adjacent to self-report of learned behavior. cites Betley et al. (2025); Binder et al. (2024).
- [Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?](https://arxiv.org/abs/2601.10242) (Guan-Xu Chen, Dong-Rui Liu, Jing Shao, 2026). Whether looped models link representations to their verbal outputs. cites Lindsey (2025). cited by Hahami et al. (2026).
- [Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations](https://arxiv.org/abs/2601.22548) (Dani Roytburg, Matthew Bozoukov, Matthew Nguyen et al., 2026). Checks the self-preference evaluations behind self-recognition claims. cites Berglund et al. (2023); Binder et al. (2024).
- [Neologism Learning for Controllability and Self-Verbalization](https://arxiv.org/abs/2510.08506) (John Hewitt, Oyvind Tafjord, Robert Geirhos et al., 2025). Self-verbalization through learned neologisms. cites Berglund et al. (2023); Betley et al. (2025).
- [Verbalizing LLMs' assumptions to explain and control sycophancy](https://arxiv.org/abs/2604.03058) (Myra Cheng, Isabel Sieh, Humishka Zope et al., 2026). Verbalizing a model's assumptions. cites Li et al. (2025); Plunkett et al. (2025).
- [Metacognitive Sensitivity for Test-Time Dynamic Model Selection](https://arxiv.org/abs/2512.10451) (Le Hoang Minh Trinh, L. Pham, T. Pham et al., 2025). Metacognitive sensitivity used for model selection. cites Song et al. (2025); Song et al. (2025).
- [When Self-Reference Fails to Close: Matrix-Level Dynamics in Large Language Models](https://arxiv.org/abs/2604.12128) (Jisung Bae, 2026). Self-reference dynamics; unclear scope. cites Binder et al. (2024); Lindsey (2025).
- [AI and Consciousness: Shifting Focus Towards Tractable Questions](https://arxiv.org/abs/2605.06965) (I. Comsa, 2026). Consciousness debate refocused on tractable questions. cites Comsa & Shanahan (2025); Lindsey (2025).
- [Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing](https://arxiv.org/abs/2606.17478) (Kexin Chen, Yi Liu, Hao-Nan Zhang et al., 2026). Activation explainers used for deception auditing. cites Cywiński et al. (2025); Li et al. (2025).
- [Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data](https://arxiv.org/abs/2508.00741) (S. Imran, Rob Lamb, Peter M. Atkinson, 2025). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023); Betley et al. (2025).
- [Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models](https://arxiv.org/abs/2510.16340) (Pratham Singla, Shivank Garg, Ayush Singh et al., 2025). Evaluates reasoning about reasoning. cites Betley et al. (2025); Binder et al. (2024).
- [Evaluating Self-Orienting in Language and Reasoning Models](https://www.semanticscholar.org/paper/9bf7b40326d0e45463b6a289cbe091fc9c9c8d57) (Eric J. Bigelow, Zergham Ahmed, Tomer D. Ullman). Self-orienting evaluations. cites Betley et al. (2025); Binder et al. (2024).
- [One Faithful Pass Over the Cuckoo's Nest](https://arxiv.org/abs/2609.00383) (Kristina Šekrst, 2026). Appears to be about explanation faithfulness; unclear from the title. cites Binder et al. (2024); Lindsey (2025).
- [Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models](https://arxiv.org/abs/2608.13258) (P. Balani, Subhrakanta Panda, 2026). Self-referential prompting and response stability. cites Comsa & Shanahan (2025); Lindsey (2025).
- [The Assistant as a Privileged Persona: A canonical reference in cross-persona self-recognition](https://arxiv.org/abs/2606.00545) (G. Asvin, 2026). Cross-persona self-recognition. cites Betley et al. (2025); Binder et al. (2024).
- [The Inner Monologue of Language Models: When Reasoning Traces Reveal More Than They Hide](https://doi.org/10.18653/v1/2026.findings-acl.2078) (Pratham Singla, Shivank Garg, Ayush Singh et al., 2026). What reasoning traces reveal. cites Betley et al. (2025); Binder et al. (2024).
- [What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs](https://arxiv.org/abs/2606.28615) (Nhi Nguyen, Shauli Ravfogel, R. Ranganath, 2026). Explanations against the model's own beliefs. cites Bai et al. (2025); Li et al. (2025).

### Suggested: skip (101)

- [Lora: Low-rank adaptation of large language models](https://arxiv.org/abs/2106.09685) (J. Hu, Ye-Long Shen, Phillip Wallis et al., 2021). Fine-tuning method cited as a tool. cited by Betley et al. (2025); Binder et al. (2024); Cywiński et al. (2025); Li et al. (2025); Treutlein et al. (2024); Wang et al. (2025).
- [Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation](https://arxiv.org/abs/2603.18893) (Nicolás Martorell, Bruno Bianchi, 2026). Looks like another version of "Quantitative Introspection in Language Models: Tracking Internal States Across Conversation". cites Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025).
- [Self-Referential Introspection in Large Language Models: The Critical Threshold for Recursive Self-Improvement](https://doi.org/10.3390/e28090951) (Jiang Zhang, Bing Yuan, Qian Zhang, 2026). Looks like another version of "Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement". cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025); Song et al. (2025).
- [Asymmetric Communication: Large Language Models and Language Games](https://arxiv.org/abs/2607.28137) (Enzo Fenoglio, 2026). Relevance to self-report unclear from the title. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025).
- [Training language models to follow instructions with human feedback](https://arxiv.org/abs/2203.02155) (Long Ouyang, Jeff Wu, Xu Jiang et al., 2022). General ML or interpretability background, not about self-report. cited by Bai et al. (2025); Berglund et al. (2023); Cywiński et al. (2025); Sherburn et al. (2024).
- [Locating and Editing Factual Associations in GPT](https://arxiv.org/abs/2202.05262) (Kevin Meng, David Bau, A. Andonian et al., 2022). General ML or interpretability background, not about self-report. cited by Binder et al. (2024); Li et al. (2025); Sherburn et al. (2024); Treutlein et al. (2024).
- [The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"](https://arxiv.org/abs/2309.12288) (Lukas Berglund, Meg Tong, Maximilian Kaufmann et al., 2023). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023). cited by Betley et al. (2025); Cywiński et al. (2025); Treutlein et al. (2024).
- [Do Large Language Models Latently Perform Multi-Hop Reasoning?](https://arxiv.org/abs/2402.16837) (Sohee Yang, E. Gribovskaya, Nora Kassner et al., 2024). General ML or interpretability background, not about self-report. cites Berglund et al. (2023). cited by Betley et al. (2025); Binder et al. (2024); Treutlein et al. (2024).
- [Truthful AI: Developing and governing AI that does not lie](https://arxiv.org/abs/2110.06674) (Owain Evans, Owen Cotton-Barratt, Lukas Finnveden et al., 2021). About honesty norms for AI, not about self-report of internal states. cited by Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Sherburn et al. (2024).
- [Characterizing the Consistency of the Emergent Misalignment Persona](https://arxiv.org/abs/2604.28082) (Anietta Weckauff, Yu-Chen Zhang, Maksym Andriushchenko, 2026). Emergent-misalignment line; not about self-report. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Plunkett et al. (2025).
- [Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs](https://arxiv.org/abs/2605.20382) (C. Camassa, Derek Shiller, 2026). Relevance to self-report unclear from the title. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- nostalgebraist. A reference to an author, not a paper. cited by Hahami et al. (2026); Li et al. (2025); Pearson-Vogel et al. (2026); Wang et al. (2025).
- [Chain of Thought Prompting Elicits Reasoning in Large Language Models](https://arxiv.org/abs/2201.11903) (Jason Wei, Xue-Zhi Wang, Dale Schuurmans et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Sherburn et al. (2024); Treutlein et al. (2024).
- [Measuring Massive Multitask Language Understanding](https://arxiv.org/abs/2009.03300) (Dan Hendrycks, Collin Burns, Steven Basart et al., 2020). General ML or interpretability background, not about self-report. cited by Binder et al. (2024); Li et al. (2025); Williamson et al. (2026).
- [Scalinglawsforneurallanguagemodels](https://arxiv.org/abs/2001.08361) (J. Kaplan, Sam McCandlish, T. Henighan et al., 2020). General ML or interpretability background, not about self-report. cited by Bai et al. (2025); Berglund et al. (2023); Sherburn et al. (2024).
- [Discovering Language Model Behaviors with Model-Written Evaluations](https://arxiv.org/abs/2212.09251) (Ethan Perez, Sam Ringer, Kamilė Lukošiūtė et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024).
- [Steering Language Models With Activation Engineering](https://arxiv.org/abs/2308.10248) (A. M. Turner, Lisa Thiergart, Gavin Leech et al., 2023). Steering method cited as a tool. cited by Hahami et al. (2026); Pearson-Vogel et al. (2026); Wang et al. (2025).
- [The alignment problem from a deep learning perspective](https://arxiv.org/abs/2209.00626) (Richard Ngo, Lawrence Chan, Sören Mindermann, 2022). AI safety or control work without a self-report question. cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024).
- [Physics of language models: Part 3.2, knowledge manipulation](https://arxiv.org/abs/2309.14402) (Zeyuan Allen-Zhu, Yuanzhi Li, 2023). General ML or interpretability background, not about self-report. cited by Betley et al. (2025); Binder et al. (2024); Treutlein et al. (2024).
- [Implicit meta-learning may lead language models to trust more reliable sources](https://arxiv.org/abs/2310.15047) (D. Krasheninnikov, Egor Krasheninnikov, B. Mlodozeniec et al., 2023). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023). cited by Betley et al. (2025); Treutlein et al. (2024).
- [Position: It’s Time to Optimize for Self-Consistency](https://www.semanticscholar.org/paper/ca8080bef251475a8606ef4c2d63d6de7e64e115) (Itamar Hagay Pres, Belinda Z. Li, L. Ruis et al.). Looks like another version of "Position: It's Time to Optimize LLMs for Self-Consistency". cites Binder et al. (2024); Li et al. (2025); Plunkett et al. (2025).
- [ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions](https://arxiv.org/abs/2605.24279) (Xian-Zhong Ding, Yangyang Yu, Chang-Wei Liu et al., 2026). Persona drift benchmark; not about self-report. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- [Automated Interpretability-Driven Model Auditing and Control: A Research Agenda](https://www.semanticscholar.org/paper/5c6070a92d62df10668a8ece4dfb465b81efe528) (Fazl Barez). AI safety or control work without a self-report question. cites Cywiński et al. (2025); Li et al. (2025); Lindsey (2025).
- [Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates](https://arxiv.org/abs/2607.01047) (Elias Najarro, Ane Espeseth, Eleni Nisioti et al., 2026). Relevance to self-report unclear from the title. cites Binder et al. (2024); Hahami et al. (2026); Lindsey (2025).
- [Generalization to Political Beliefs from Fine-Tuning on Sports Team Preferences](https://arxiv.org/abs/2601.04369) (Owen Terry, 2026). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023); Betley et al. (2025); Lindsey (2025).
- [Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models](https://arxiv.org/abs/2607.26173) (Antón de la Fuente, Arthur Conmy, 2026). General ML or interpretability background, not about self-report. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [Artificial Phantasia: Emergent Mental Imagery in Large Language Models](https://arxiv.org/abs/2509.23108) (Morgan McCarty, Jorge Morales, 2025). Mental imagery, not self-report. cites Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- OpenAI. A reference to an organization, not a paper. cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024).
- [Password-Activated Shutdown Protocols for Misaligned Frontier Agents](https://arxiv.org/abs/2512.03089) (Kai Williams, R. Subramani, Francis Rhys Ward, 2025). AI safety or control work without a self-report question. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [Phase Transitions in Driven Informational Systems: A Two-Field Perspective on Learning Theory and Non-Equilibrium Chemistry](https://arxiv.org/abs/2605.16325) (Xuan Khanh Truong, 2026). General ML or interpretability background, not about self-report. cites Binder et al. (2024); Lindsey (2025); Song et al. (2025).
- Sleeper agents: Training deceptive llms that persist through safety training. AI safety or control work without a self-report question. cited by Betley et al. (2025); Cywiński et al. (2025); Treutlein et al. (2024).
- [Language Models are Few-Shot Learners](https://arxiv.org/abs/2005.14165) (Tom B. Brown, Benjamin Mann, Nick Ryder et al., 2020). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Sherburn et al. (2024).
- [The llama 3 herd of models](https://arxiv.org/abs/2407.21783) (Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri et al., 2024). General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Hahami et al. (2026).
- [On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜](https://doi.org/10.1145/3442188.3445922) (Emily M. Bender, Timnit Gebru, Angelina McMillan-Major et al., 2021). General critique of large language models, not about self-report. cited by Binder et al. (2024); Williamson et al. (2026).
- [GPT-4o System Card](https://arxiv.org/abs/2410.21276) (OpenAI Aaron Hurst, A. Lerer, Adam P. Goucher et al., 2024). General ML or interpretability background, not about self-report. cited by Betley et al. (2025); Song et al. (2025).
- [Emergent Abilities of Large Language Models](https://arxiv.org/abs/2206.07682) (Jason Wei, Yi Tay, Rishi Bommasani et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Comsa & Shanahan (2025).
- [Training Compute-Optimal Large Language Models](https://arxiv.org/abs/2203.15556) (Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Sherburn et al. (2024).
- [Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models](https://arxiv.org/abs/2206.04615) (Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Treutlein et al. (2024).
- [Sparse Autoencoders Find Highly Interpretable Features in Language Models](https://arxiv.org/abs/2309.08600) (Hoagy Cunningham, Aidan Ewart, L. Smith et al., 2023). General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Li et al. (2025).
- [Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small](https://arxiv.org/abs/2211.00593) (Kevin Wang, Alexandre Variengien, Arthur Conmy et al., 2022). General ML or interpretability background, not about self-report. cited by Li et al. (2025); Sherburn et al. (2024).
- [The fineweb datasets: Decanting the web for the finest text data at scale](https://arxiv.org/abs/2406.17557) (Guilherme Penedo, Hynek Kydlícek, Loubna Ben Allal et al., 2024). General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Li et al. (2025).
- [Mass-Editing Memory in a Transformer](https://arxiv.org/abs/2210.07229) (Kevin Meng, Arnab Sen Sharma, A. Andonian et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Sherburn et al. (2024).
- [A General Language Assistant as a Laboratory for Alignment](https://arxiv.org/abs/2112.00861) (Amanda Askell, Yuntao Bai, Anna Chen et al., 2021). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Binder et al. (2024).
- [Progress measures for grokking via mechanistic interpretability](https://arxiv.org/abs/2301.05217) (Neel Nanda, Lawrence Chan, Tom Lieberum et al., 2023). General ML or interpretability background, not about self-report. cited by Li et al. (2025); Sherburn et al. (2024).
- [Discovering Latent Knowledge in Language Models Without Supervision](https://arxiv.org/abs/2212.03827) (Collin Burns, Haotian Ye, D. Klein et al., 2022). General ML or interpretability background, not about self-report. cited by Pearson-Vogel et al. (2026); Sherburn et al. (2024).
- [Towards Automated Circuit Discovery for Mechanistic Interpretability](https://arxiv.org/abs/2304.14997) (Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch et al., 2023). General ML or interpretability background, not about self-report. cited by Li et al. (2025); Sherburn et al. (2024).
- [The geometry of truth: Emergent linear structure in large language model representations of true/false datasets](https://arxiv.org/abs/2310.06824) (Samuel Marks, Max Tegmark, 2023). General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Pearson-Vogel et al. (2026).
- [Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision](https://arxiv.org/abs/2312.09390) (Collin Burns, Pavel Izmailov, J. Kirchner et al., 2023). General ML or interpretability background, not about self-report. cited by Binder et al. (2024); Cywiński et al. (2025).
- [A Survey of the State of Explainable AI for Natural Language Processing](https://arxiv.org/abs/2010.00711) (Marina Danilevsky, Kun Qian, R. Aharonov et al., 2020). General ML or interpretability background, not about self-report. cited by Plunkett et al. (2025); Song et al. (2025).
- [Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2](https://arxiv.org/abs/2408.05147) (Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy et al., 2024). General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Li et al. (2025).
- [Alignment faking in large language models](https://arxiv.org/abs/2412.14093) (R. Greenblatt, Carson E. Denison, Benjamin Wright et al., 2024). AI safety or control work without a self-report question. cited by Betley et al. (2025); Wang et al. (2025).
- [Risks from Learned Optimization in Advanced Machine Learning Systems](https://arxiv.org/abs/1906.01820) (Evan Hubinger, Chris van Merwijk, Vladimir Mikulik et al., 2019). AI safety or control work without a self-report question. cited by Berglund et al. (2023); Betley et al. (2025).
- [Physics of Language Models: Part 3.1, Knowledge Storage and Extraction](https://arxiv.org/abs/2309.14316) (Zeyuan Allen-Zhu, Yuanzhi Li, 2023). General ML or interpretability background, not about self-report. cites Berglund et al. (2023). cited by Treutlein et al. (2024).
- [2 OLMo 2 Furious](https://arxiv.org/abs/2501.00656) (Team OLMo, Pete Walsh, Luca Soldaini et al., 2024). Model release cited as a tool. cited by Song et al. (2025); Williamson et al. (2026).
- [Model evaluation for extreme risks](https://arxiv.org/abs/2305.15324) (Toby Shevlane, Sebastian Farquhar, Ben Garfinkel et al., 2023). AI safety or control work without a self-report question. cited by Berglund et al. (2023); Betley et al. (2025).
- [On the Nature of Mind.](https://doi.org/10.1038/128744a0) (C. S. Myres, 1931). Philosophy or consciousness debate without a test of self-report. cited by Comsa & Shanahan (2025); Song et al. (2025).
- [Black-box access is insufficient for rigorous ai audits](https://arxiv.org/abs/2401.14446) (Stephen Casper, Carson Ezell, Charlotte Siegmann et al., 2024). AI safety or control work without a self-report question. cited by Cywiński et al. (2025); Plunkett et al. (2025).
- [Training large language models on narrow tasks can lead to broad misalignment](https://arxiv.org/abs/2502.17424) (Jan Betley, Daniel Tan, Niels Warncke et al., 2025). Emergent-misalignment line; not about self-report. cites Berglund et al. (2023); Betley et al. (2025).
- [Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models](https://arxiv.org/abs/2406.10162) (Carson E. Denison, M. MacDiarmid, Fazl Barez et al., 2024). AI safety or control work without a self-report question. cites Berglund et al. (2023). cited by Cywiński et al. (2025).
- [Is Power-Seeking AI an Existential Risk?](https://arxiv.org/abs/2206.13353) (J. Carlsmith, 2022). AI safety or control work without a self-report question. cited by Berglund et al. (2023); Plunkett et al. (2025).
- [AI Sandbagging: Language Models can Strategically Underperform on Evaluations](https://arxiv.org/abs/2406.07358) (Teun van der Weij, Felix Hofstätter, Oliver Jaffe et al., 2024). AI safety or control work without a self-report question. cited by Binder et al. (2024); Cywiński et al. (2025).
- [Persona Features Control Emergent Misalignment](https://arxiv.org/abs/2506.19823) (Miles Wang, Tom Dupré la Tour, Olivia Watkins et al., 2025). Emergent-misalignment line; not about self-report. cites Berglund et al. (2023); Betley et al. (2025).
- [An academic survey on theoretical foundations, common assumptions and the current state of consciousness science](https://doi.org/10.1093/nc/niac011) (Jolien C. Francken, L. Beerendonk, D. Molenaar et al., 2022). Philosophy or consciousness debate without a test of self-report. cited by Binder et al. (2024); Comsa & Shanahan (2025).
- [Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models](https://arxiv.org/abs/2506.13206) (James Chua, Jan Betley, Mia Taylor et al., 2025). Emergent-misalignment line; not about self-report. cites Betley et al. (2025); Binder et al. (2024).
- [Reverse Training to Nurse the Reversal Curse](https://arxiv.org/abs/2403.13799) (Olga Golovneva, Zeyuan Allen-Zhu, Jason E. Weston et al., 2024). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023). cited by Betley et al. (2025).
- [A sketch of an AI control safety case](https://arxiv.org/abs/2501.17315) (Tomasz Korbak, Joshua Clymer, Benjamin Hilton et al., 2025). AI safety or control work without a self-report question. cites Berglund et al. (2023); Binder et al. (2024).
- [Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History](https://arxiv.org/abs/2508.04826) (Tommaso Tosato, S. Helbling, Yorguin José Mantilla Ramos et al., 2025). Personality measurement stability; not about self-report of internal states. cites Binder et al. (2024); Plunkett et al. (2025).
- [Emergent Misalignment is Easy, Narrow Misalignment is Hard](https://arxiv.org/abs/2602.07852) (Anna Soligo, Edward Turner, Senthooran Rajamanoharan et al., 2026). Emergent-misalignment line; not about self-report. cites Berglund et al. (2023); Betley et al. (2025).
- [How to evaluate control measures for LLM agents? A trajectory from today to superintelligence](https://arxiv.org/abs/2504.05259) (Tomasz Korbak, Mikita Balesni, Buck Shlegeris et al., 2025). AI safety or control work without a self-report question. cites Berglund et al. (2023); Binder et al. (2024).
- [Automating Steering for Safe Multimodal Large Language Models](https://arxiv.org/abs/2507.13255) (Lyucheng Wu, Mengru Wang, Ziwen Xu et al., 2025). General ML or interpretability background, not about self-report. cites Betley et al. (2025); Binder et al. (2024).
- [Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs](https://arxiv.org/abs/2407.04108) (Sara Price, Arjun Panickssery, Samuel R. Bowman et al., 2024). AI safety or control work without a self-report question. cites Berglund et al. (2023). cited by Betley et al. (2025).
- [Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety](https://arxiv.org/abs/2506.05451) (Seongmin Lee, Aeree Cho, Grace C. Kim et al., 2025). General ML or interpretability background, not about self-report. cites Betley et al. (2025); Binder et al. (2024).
- [Shaping capabilities with token-level data filtering](https://arxiv.org/abs/2601.21571) (Neil Rathi, Alec Radford, 2026). General ML or interpretability background, not about self-report. cites Berglund et al. (2023); Wang et al. (2025).
- [Subliminal Learning Is Steering Vector Distillation](https://arxiv.org/abs/2606.00995) (C. Blank, A. Bhatia, Senthooran Rajamanoharan et al., 2026). General ML or interpretability background, not about self-report. cites Berglund et al. (2023); Wang et al. (2025).
- [Mechanistic Interpretability Needs Philosophy](https://arxiv.org/abs/2506.18852) (Iwan Williams, Ninell Oldenburg, Ruchira Dhar et al., 2025). Philosophy or consciousness debate without a test of self-report. cites Lindsey (2025); Song et al. (2025).
- [Probe-Rewrite-Evaluate: A Workflow for Reliable Benchmarks and Quantifying Evaluation Awareness](https://arxiv.org/abs/2509.00591) (Lang Xiong, N. Bhargava, Jeremy Chang et al., 2025). AI safety or control work without a self-report question. cites Berglund et al. (2023); Betley et al. (2025).
- [Understanding Emergent Misalignment via Feature Superposition Geometry](https://arxiv.org/abs/2605.00842) (Gouki Minegishi, Hiroki Furuta, Takeshi Kojima et al., 2026). Emergent-misalignment line; not about self-report. cites Betley et al. (2025); Wang et al. (2025).
- [A Disproof of Large Language Model Consciousness: The Necessity of Continual Learning for Consciousness](https://arxiv.org/abs/2512.12802) (Erik P. Hoel, 2025). Philosophy or consciousness debate without a test of self-report. cites Binder et al. (2024); Lindsey (2025).
- [Programming by Backprop: LLMs Acquire Reusable Algorithmic Abstractions During Code Training](https://doi.org/10.48550/arXiv.2506.18777) (Jonathan Cook, Silvia Sapora, Arash Ahmadian et al., 2025). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023); Betley et al. (2025).
- [Covert Influence Between Language Models](https://arxiv.org/abs/2606.04071) (Avidan Shah, J. Chooi, Jinghuai Ou et al., 2026). AI safety or control work without a self-report question. cites Li et al. (2025); Lindsey (2025).
- [Models That Know How Evaluations Are Designed Score Safer](https://arxiv.org/abs/2605.28591) (K. Deckenbach, Haritz Puerto, Jonas Geiping et al., 2026). AI safety or control work without a self-report question. cites Berglund et al. (2023); Betley et al. (2025).
- [Assessing LLMs'mathematical abilities requires understanding the various mechanisms of mathematical creativity](https://arxiv.org/abs/2608.16118) (Silvère Gangloff, 2026). General ML or interpretability background, not about self-report. cites Binder et al. (2024); Hahami et al. (2026).
- [If LLMs Have Human-Like Attributes, Then So Does Age of Empires II](https://arxiv.org/abs/2605.31514) (Adrian de Wynter, 2026). Philosophy or consciousness debate without a test of self-report. cites Betley et al. (2025); Lindsey (2025).
- [On the Creativity of AI Agents](https://arxiv.org/abs/2604.13242) (Giorgio Franceschelli, Mirco Musolesi, 2026). General ML or interpretability background, not about self-report. cites Binder et al. (2024); Comsa & Shanahan (2025).
- [Artificial Phantasia: Evidence for Propositional Reasoning-Based Mental Imagery in Large Language Models](https://doi.org/10.48550/arXiv.2509.23108) (Morgan McCarty, Jorge Morales, 2025). Mental imagery, not self-report. cites Betley et al. (2025); Plunkett et al. (2025).
- [Can LLMs Perceive Time? An Empirical Investigation](https://arxiv.org/abs/2604.00010) (Aniketh Garikaparthi, 2026). General ML or interpretability background, not about self-report. cites Binder et al. (2024); Lindsey (2025).
- [Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power](https://arxiv.org/abs/2604.21043) (Travis LaCroix, Fintan Mallory, Sasha Luccioni, 2026). Philosophy or consciousness debate without a test of self-report. cites Binder et al. (2024); Lindsey (2025).
- [Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?](https://arxiv.org/abs/2609.14803) (Afshin Khadangi, 2026). Relevance to self-report unclear from the title. cites Binder et al. (2024); Song et al. (2025).
- [Emergent Language as an Approach to Conscious AI](https://arxiv.org/abs/2606.06380) (Zengqing Wu, Chuan Xiao, 2026). Philosophy or consciousness debate without a test of self-report. cites Betley et al. (2025); Lindsey (2025).
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs. Emergent-misalignment line; not about self-report. cited by Cywiński et al. (2025); Pearson-Vogel et al. (2026).
- [Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning](https://arxiv.org/abs/2606.15507) (Ali Dasdan, Manan Shah, W. R. Neuman et al., 2026). General ML or interpretability background, not about self-report. cites Li et al. (2025); Lindsey (2025).
- [Generalization Dynamics of LM Pre-training](https://arxiv.org/abs/2609.33150) (Jia-Xin Wen, Zhengxuan Wu, D. Song et al., 2026). General ML or interpretability background, not about self-report. cites Berglund et al. (2023); Wang et al. (2025).
- Llama 3 model card. General ML or interpretability background, not about self-report. cited by Betley et al. (2025); Treutlein et al. (2024).
- Scaling monosemantic-ity: Extracting interpretable features from Claude 3 Sonnet. General ML or interpretability background, not about self-report. cited by Comsa & Shanahan (2025); Li et al. (2025).
- [Semantic Containment as a Fundamental Property of Emergent Misalignment](https://arxiv.org/abs/2603.04407) (Rohan Saxena, 2026). Emergent-misalignment line; not about self-report. cites Berglund et al. (2023); Betley et al. (2025).
- [Stress-Testing Alignment Audits With Prompt-Level Strategic Deception](https://arxiv.org/abs/2602.08877) (Oliver Daniels, Perusha Moodley, Benjamin M. Marlin et al., 2026). AI safety or control work without a self-report question. cites Berglund et al. (2023); Cywiński et al. (2025).
- [T HE T WO -H OP C URSE : LLM S TRAINED ON A (cid:41) B , B (cid:41) C FAIL TO LEARN A (cid:41) C](https://www.semanticscholar.org/paper/486c6a8eb2ad63150024dc6ebb263e9643f9aa5a) (Mikita Balesni, Apollo Research, Tomasz Korbak et al.). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023); Binder et al. (2024).
- Towards monosemanticity: Decomposing language models with dictionary learning. General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Li et al. (2025).
- [Unsupervised Features Mining via Activation Geometry](https://arxiv.org/abs/2607.04222) (Amit Levi, Elad David, M. Fomin, 2026). General ML or interpretability background, not about self-report. cites Binder et al. (2024); Lindsey (2025).
- [VLMs Can Aggregate Scattered Training Patches](https://arxiv.org/abs/2506.03614) (Zhanhui Zhou, Lingjie Chen, Chao Yang et al., 2025). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023); Betley et al. (2025).
- Without specific countermeasures, the easiest path to transformative ai likely leads to ai takeover. AI safety or control work without a self-report question. cited by Berglund et al. (2023); Treutlein et al. (2024).

---

Source: https://introspection.infinite.fun/candidates · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
