# Frontier

> 194 papers one citation away from the wiki that do not have a page yet.

The crawler looked at the references and citers of every paper page and saw 1084 distinct neighbors. A neighbor is listed here if it connects to at least 2 wiki papers, or if someone added it as a lead. The triage labels are suggestions, made from each candidate's title and its place in the citation graph and not from reading it. A candidate becomes a page only after a person accepts it. Last crawled 2026-10-06. Source: Semantic Scholar Graph API, with arXiv HTML bibliographies where Semantic Scholar has no reference list.

## Suggested: add (43)

- [Emergent Introspection in AI is Content-Agnostic](https://arxiv.org/abs/2603.05414) (Harvey Lederman, Kyle Mahowald, 2026). Introspection itself: asks what kind of content models can introspect on. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- [Dissociating Direct Access from Inference in AI Introspection](https://doi.org/10.48550/arXiv.2603.05414) (Harvey Lederman, Kyle Mahowald, 2026). Introspection itself: separates direct access from inference. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- [Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision](https://arxiv.org/abs/2606.32038) (Zifan Carl Guo, L. Ruis, Jacob Andreas et al., 2026). Self-explanation training and whether it tracks behavior. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- [Metacognition in LLMs: Foundations, Progress, and Opportunities](https://arxiv.org/abs/2607.11881) (Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu et al., 2026). Survey of metacognition in LLMs; useful as a map. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- [Language models (mostly) know what they know](https://arxiv.org/abs/2207.05221) (Saurav Kadavath, Tom Conerly, Amanda Askell et al., 2022). Foundational result on models knowing what they know; cited by six pages. cited by Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Sherburn et al. (2024).
- [Evidence for Limited Metacognition in LLMs](https://arxiv.org/abs/2509.21545) (Christopher M. Ackerman, 2025). Tests metacognition directly. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025). cited by Plunkett et al. (2025).
- [Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation](https://arxiv.org/abs/2608.20569) (Emilio Ferrara, 2026). Measures what models can report about their own computation. cites Betley et al. (2025); Binder et al. (2024); Hahami et al. (2026); Lindsey (2025); Pearson-Vogel et al. (2026); Song et al. (2025).
- [Quantitative Introspection in Language Models: Tracking Internal States Across Conversation](https://doi.org/10.48550/arXiv.2603.18893) (Nicolás Martorell, 2026). Operationalizes introspection as coupling between self-report and a probed internal state. cites Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025).
- [LLM Evaluators Recognize and Favor Their Own Generations](https://arxiv.org/abs/2404.13076) (Arjun Panickssery, Samuel R. Bowman, Shi Feng, 2024). Self-recognition: whether models can tell their own outputs apart. cites Berglund et al. (2023). cited by Binder et al. (2024); Comsa & Shanahan (2025); Hahami et al. (2026); Song et al. (2025).
- [Can LLMs Introspect? A Reality Check](https://arxiv.org/abs/2605.26242) (Shashwat Singh, Tal Linzen, Shauli Ravfogel, 2026). A skeptical check on introspection claims. cites Binder et al. (2024); Li et al. (2025); Lindsey (2025); Plunkett et al. (2025); Song et al. (2025).
- [Steering Awareness: Detecting Activation Steering from Within](https://arxiv.org/abs/2511.21399) (J. Rivera, D. Africa, 2025). Concept-injection line: detecting activation steering from within. cites Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Song et al. (2025). cited by Hahami et al. (2026).
- [Mechanisms of Introspective Awareness](https://arxiv.org/abs/2603.21396) (U. Macar, Li Yang, Atticus Wang et al., 2026). Mechanistic account of introspective awareness. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Wang et al. (2025).
- [In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?](https://arxiv.org/abs/2609.00904) (Koshiro Aoki, Ryota Takatsuki, Gouki Minegishi et al., 2026). Tests privileged access through control of internal representations. cites Betley et al. (2025); Binder et al. (2024); Li et al. (2025); Song et al. (2025); Song et al. (2025).
- [Telling more than we can know: Verbal reports on mental processes.](https://doi.org/10.1037/0033-295X.84.3.231) (R. Nisbett, T. Wilson, 1977). The classic study of human confabulation that this literature borrows its framing from; cited by four pages. cited by Comsa & Shanahan (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- [Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting](https://arxiv.org/abs/2305.04388) (Miles Turpin, Julian Michael, Ethan Perez et al., 2023). Unfaithful chain-of-thought explanations; the main adjacent line on explanation faithfulness. cited by Li et al. (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Sherburn et al. (2024).
- [Me, myself, and ai: The situational awareness dataset (sad) for llms](https://arxiv.org/abs/2407.04694) (Rudolf Laine, Bilal Chughtai, Jan Betley et al., 2024). Benchmark of situational awareness, including self-knowledge tasks. cited by Betley et al. (2025); Comsa & Shanahan (2025); Hahami et al. (2026); Plunkett et al. (2025).
- [Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers](https://arxiv.org/abs/2512.15674) (Adam Karvonen, James Chua, Clément Dumas et al., 2025). Training models to explain activations; adjacent to self-explanation. cites Cywiński et al. (2025); Li et al. (2025); Lindsey (2025). cited by Hahami et al. (2026).
- [Towards Evaluating AI Systems for Moral Status Using Self-Reports](https://arxiv.org/abs/2311.08576) (Ethan Perez, Robert Long, 2023). Proposes training and evaluating self-reports for moral-status questions. cites Berglund et al. (2023). cited by Binder et al. (2024); Comsa & Shanahan (2025); Plunkett et al. (2025).
- [A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior](https://arxiv.org/abs/2602.02639) (Harry Mayne, J. Kang, Dewi Gould et al., 2026). Tests whether self-explanations predict behavior. cites Binder et al. (2024); Li et al. (2025); Lindsey (2025); Plunkett et al. (2025).
- [Introspection Adapters: Training LLMs to Report Their Learned Behaviors](https://arxiv.org/abs/2604.16812) (K. Shenoy, Li Yang, A. Sheshadri et al., 2026). Trains models to report their learned behaviors. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- [Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs](https://arxiv.org/abs/2511.04875) (Matthew Bozoukov, Matthew Nguyen, Shubkarman Singh et al., 2025). Mechanistic conditions for behavioral self-awareness. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Wang et al. (2025).
- [Strangers to Themselves: What Language Models Say About Themselves Is Generic](https://arxiv.org/abs/2609.09899) (Phil Blandfort, Urja Pawar, 2026). A skeptical result: self-descriptions are generic rather than self-specific. cites Bai et al. (2025); Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- [Reasoning Models Don’t Always Say What They Think](https://arxiv.org/abs/2505.05410) (Yanda Chen, Joe Benton, Ansh Radhakrishnan et al., 2025). Chain-of-thought faithfulness in reasoning models. cited by Cywiński et al. (2025); Li et al. (2025); Plunkett et al. (2025).
- [From Imitation to Introspection: Probing Self-Consciousness in Language Models](https://arxiv.org/abs/2410.18819) (Sirui Chen, Shu Yu, Shengjie Zhao et al., 2024). Probes self-consciousness concepts in models. cites Berglund et al. (2023); Binder et al. (2024). cited by Comsa & Shanahan (2025).
- [Large Language Models Report Subjective Experience Under Self-Referential Processing](https://arxiv.org/abs/2510.24797) (Cameron Berg, D. D. de Lucena, Judd Rosenblatt, 2025). Self-reports of experience under self-referential prompting. cites Betley et al. (2025); Lindsey (2025); Plunkett et al. (2025).
- [Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare](https://arxiv.org/abs/2509.07961) (Valen Tagliabue, Leonard Dung, 2025). Compares stated and revealed preferences. cites Betley et al. (2025); Binder et al. (2024); Song et al. (2025).
- [Do Activation Verbalization Methods Convey Privileged Information?](https://arxiv.org/abs/2509.13316) (Millicent Li, Alberto Mario Ceballos Arroyo, Giordano Rogers et al., 2025). Asks whether verbalized activations carry privileged information. cites Binder et al. (2024); Song et al. (2025). cited by Li et al. (2025).
- [Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment](https://arxiv.org/abs/2602.14777) (Laurène Vaugrante, Anietta Weckauff, Thilo Hagendorff, 2026). Behavioral self-awareness tracking fine-tuning changes. cites Berglund et al. (2023); Binder et al. (2024); Lindsey (2025).
- [Language models recognize dropout and Gaussian noise applied to their activations](https://arxiv.org/abs/2604.17465) (Damiano Fornasiere, Mirko Bronzi, Spencer Kitts et al., 2026). Concept-injection line: detecting noise applied to activations. cites Comsa & Shanahan (2025); Lindsey (2025); Pearson-Vogel et al. (2026).
- [Fine-Tuning Language Models to Know What They Know](https://arxiv.org/abs/2602.02605) (Sangjun Park, Elliot Meyerson, Xin Qiu et al., 2026). Trains models to know what they know. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025).
- [A mechanistic study of language model introspection](https://arxiv.org/abs/2609.35108) (Jia-Hong Zou, Xiang-Kun Sun, Ling-Kai Kong et al., 2026). Mechanistic study of introspection; also found by keyword search. cites Binder et al. (2024); Hahami et al. (2026); Lindsey (2025).
- [Introspecting Alignment Shifts Beyond Behaviors Implanted Through Fine-Tuning](https://arxiv.org/abs/2608.04347) (K. Yoshida, Laura Gomezjurado Gonzalez, Yukinori Yamamoto et al., 2026). Introspecting on changes from fine-tuning. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- [Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations](https://arxiv.org/abs/2505.13763) (Ji-An Li, M. Mattar, Hua-Dong Xiong et al., 2025). Metacognitive monitoring and control of internal activations. cites Binder et al. (2024). cited by Hahami et al. (2026).
- [Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives](https://arxiv.org/abs/2511.06626) (Chloe Li, Mary Phuong, Daniel Tan, 2025). Trains models to self-report hidden objectives. cites Betley et al. (2025); Binder et al. (2024).
- [Do Large Language Models Know What They Are Capable Of?](https://arxiv.org/abs/2512.24661) (Casey O. Barkan, Sid Black, Oliver Sourbut, 2025). Self-knowledge of capabilities. cites Betley et al. (2025); Binder et al. (2024).
- [Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness](https://arxiv.org/abs/2604.12373) (Tomer Ashuach, Shai Gretz, Yoav Katz et al., 2026). Privileged access: disentangles privileged knowledge of correctness. cites Binder et al. (2024); Li et al. (2025).
- [Me, Myself, and π: Evaluating and Explaining LLM Introspection](https://arxiv.org/abs/2603.20276) (Atharv Naphade, Samarth Bhargav, Sean Lim et al., 2026). Evaluates and explains introspection. cites Binder et al. (2024); Lindsey (2025).
- [Can LLMs Reliably Self-Report Adversarial Prefills, and How?](https://arxiv.org/abs/2606.23671) (Quang-Anh Nguyen, Uzair Ahmed, Taegyoon Kim, 2026). Self-report of adversarial prefills. cites Binder et al. (2024); Cywiński et al. (2025).
- [Do Language Models Know When They'll Refuse? Probing Introspective Awareness of Safety Boundaries](https://arxiv.org/abs/2604.00228) (Tanay Gondil, 2026). Introspective awareness of when the model will refuse. cites Binder et al. (2024); Lindsey (2025).
- [Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect](https://arxiv.org/abs/2607.14111) (Ely Hahami, Ishaan Sinha, Lavik Jain, 2026). Fine-tuning small models to introspect. Author thread: https://x.com/ElyHahami/status/2078161491069718639. cites Binder et al. (2024); Lindsey (2025).
- [Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs](https://arxiv.org/abs/2602.10352) (Keenan Pepper, Alex McKenzie, Florin Pop et al., 2026). Training self-interpretation from interpretability artifacts. cites Li et al. (2025); Lindsey (2025).
- [Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports](https://arxiv.org/abs/2609.32449) (Phongsakon Mark Konrad, T. Tanyel, Serkan Ayvaz, 2026). An identifiability argument about what self-reports can show. cites Li et al. (2025); Lindsey (2025).
- [What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation](https://arxiv.org/abs/2609.32670) (Arshia Eftekhari Zadeh, 2026). Evidence standards for mechanistic claims about self-explanation. cites Binder et al. (2024); Lindsey (2025).

## Suggested: maybe (52)

- [Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement](https://arxiv.org/abs/2607.04277) (Jiang Zhang, Bing Yuan, Qian Zhang, 2026). About self-reference and introspection, but the framing is recursive self-improvement. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025); Song et al. (2025).
- [Introspection](https://doi.org/10.1177/1057083709332318) (William E. Fredrickson, 2009). Philosophy reference on human introspection; background for definitions. cited by Binder et al. (2024); Comsa & Shanahan (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- ["As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It](https://arxiv.org/abs/2609.25021) (Jędrzej Maczan, 2026). Self-referential voice and steering; relevance to self-report unclear from the title. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025).
- [Anosognosia in LLMs: Probing Self-Awareness of Quantized Computational Substrate](https://arxiv.org/abs/2610.06174) (Yoshihiro Izawa, Gouki Minegishi, Yoko Yamakata, 2026). Self-awareness of quantization; narrow but on topic. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Song et al. (2025).
- [Counterfactual Simulation Training for Chain-of-Thought Faithfulness](https://arxiv.org/abs/2602.20710) (P. Hase, Christopher Potts, 2026). Chain-of-thought faithfulness training. cites Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Plunkett et al. (2025).
- [Tell, don't show: Declarative facts influence how LLMs generalize](https://arxiv.org/abs/2312.07779) (Alexander Meinke, Owain Evans, 2023). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023). cited by Betley et al. (2025); Binder et al. (2024); Treutlein et al. (2024).
- [Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns](https://arxiv.org/abs/2509.24988) (Hanqi Xiao, Vaidehi Patil, Hyunji Lee et al., 2025). Bears on privileged access: correctness predictors that are model-agnostic. cites Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025).
- [Hallucinations Undermine Trust; Metacognition is a Way Forward](https://arxiv.org/abs/2605.01428) (G. Yona, Mor Geva, Y. Matias, 2026). Position piece on metacognition. cites Binder et al. (2024); Li et al. (2025); Lindsey (2025); Song et al. (2025).
- [Steering Awareness: Models Can Be Trained to Detect Activation Steering](https://www.semanticscholar.org/paper/2ed88e06a893687a6fa268146cfed527d016a323) (J. Rivera, D. Africa). Looks like another version of "Steering Awareness: Detecting Activation Steering from Within". cites Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Song et al. (2025).
- [From Simulation to Enaction: Post-trained language models recognize and react to their own generations](https://arxiv.org/abs/2605.25459) (G. Asvin, Jack W Lindsey, 2026). Self-recognition of own generations. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- [Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models](https://arxiv.org/abs/2604.25922) (Skylar DeTure, 2026). Trained denial in self-reports about consciousness. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- [Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments](https://arxiv.org/abs/2608.16747) (Adam Karvonen, Euan Ong, Subhash Kantamneni et al., 2026). Evaluates explanations of behavior with counterfactuals. cites Betley et al. (2025); Binder et al. (2024); Cywiński et al. (2025); Li et al. (2025).
- [Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment](https://arxiv.org/abs/2606.23700) (Arush Tagade, Shao-Heng Zhou, Jiaxin Wen et al., 2026). Self-recognition fine-tuning, in the emergent-misalignment setting. cites Berglund et al. (2023); Betley et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026).
- [Language Models Act on Hidden Valence](https://arxiv.org/abs/2609.35591) (Cameron Berg, Caspar Kaiser, 2026). Hidden internal valence and behavior; may bear on self-report. cites Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025).
- [Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary](https://arxiv.org/abs/2607.18553) (Jan Kirin, 2026). Proto-introspection in looped models; unclear scope. cites Betley et al. (2025); Binder et al. (2024); Pearson-Vogel et al. (2026); Song et al. (2025).
- [Teaching Models to Express Their Uncertainty in Words](https://arxiv.org/abs/2205.14334) (Stephanie C. Lin, Jacob Hilton, Owain Evans, 2022). Verbalized uncertainty; background for self-knowledge of confidence. cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024).
- [The Unreliability of Naive Introspection](https://doi.org/10.1215/00318108-2007-037) (Eric Schwitzgebel, 2008). Philosophy of human introspection; background for definitions. cited by Binder et al. (2024); Comsa & Shanahan (2025); Song et al. (2025).
- [Are DeepSeek R1 And Other Reasoning Models More Faithful?](https://arxiv.org/abs/2501.08156) (James Chua, Owain Evans, 2025). Chain-of-thought faithfulness in reasoning models. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [Metacognition and Uncertainty Communication in Humans and Large Language Models](https://arxiv.org/abs/2504.14045) (M. Steyvers, Megan A. K. Peters, 2025). Metacognition and uncertainty, humans compared with LLMs. cites Betley et al. (2025); Binder et al. (2024). cited by Plunkett et al. (2025).
- [Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants](https://arxiv.org/abs/2512.15712) (Vincent Huang, Da-Mi Choi, Daniel D. Johnson et al., 2025). Interpretability assistants that decode concepts; adjacent to self-explanation. cites Li et al. (2025); Lindsey (2025). cited by Hahami et al. (2026).
- [Agentic Knowledgeable Self-awareness](https://arxiv.org/abs/2504.03553) (Shuo-Fei Qiao, Zhi-Song Qiu, Baochang Ren et al., 2025). Agent self-awareness of knowledge; unclear scope. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [AI Awareness](https://arxiv.org/abs/2504.20084) (Xiaojian Li, Hao Shi, Rongwu Xu et al., 2025). Survey of AI awareness. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [No Reliable Evidence of Self-Reported Sentience in Small Large Language Models](https://arxiv.org/abs/2601.15334) (Caspar Kaiser, Sean Enderby, 2026). Self-reported sentience in small models. cites Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- [Position: It's Time to Optimize LLMs for Self-Consistency](https://arxiv.org/abs/2608.05188) (Itamar Hagay Pres, Belinda Z. Li, L. Ruis et al., 2026). Position piece on self-consistency. cites Binder et al. (2024); Li et al. (2025); Plunkett et al. (2025).
- [Higher-order representation in AI](https://doi.org/10.33735/phimisci.2026.12032) (Patrick Butlin, 2026). Higher-order representation; theory background. cites Betley et al. (2025); Binder et al. (2024); Plunkett et al. (2025).
- [Self-CTRL: Self-Consistency Training with Reinforcement Learning](https://arxiv.org/abs/2606.18327) (Itamar Hagay Pres, L. Ruis, Melat Ghebreselassie et al., 2026). Self-consistency training. cites Berglund et al. (2023); Betley et al. (2025); Plunkett et al. (2025).
- [Imprint Reader: From Weight-Update Readout to Behavioral Intervention](https://arxiv.org/abs/2609.35261) (Guan-Xu Chen, Qi-Hao Lin, Jing Shao, 2026). Reading out weight updates; adjacent to self-report of learned behavior. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- [Questionnaire Responses Do not Capture the Safety of AI Agents](https://arxiv.org/abs/2603.14417) (Max Hellrigel-Holderbaum, Edward James Young, 2026). Stated answers against agent behavior. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences](https://arxiv.org/abs/2605.05080) (Hubert Plisiecki, Sabina Siudaj, Kacper Dudzic et al., 2026). Psychometrics of self-described experience. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025).
- [Do large language models know what they don’t know?](https://arxiv.org/abs/2305.18153) (Zhangyue Yin, Qiushi Sun, Qipeng Guo et al., 2023). Knowing what one does not know; background for self-knowledge of confidence. cited by Betley et al. (2025); Comsa & Shanahan (2025).
- [Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models](https://arxiv.org/abs/2401.06102) (Asma Ghandeharioun, Avi Caciularu, Adam Pearce et al., 2024). Verbalizing hidden representations; adjacent to self-explanation. cited by Hahami et al. (2026); Li et al. (2025).
- [Auditing language models for hidden objectives](https://arxiv.org/abs/2503.10965) (Samuel Marks, Johannes Treutlein, Trenton Bricken et al., 2025). Auditing for hidden objectives; the other side of unfaithful self-report. cited by Cywiński et al. (2025); Wang et al. (2025).
- [Prompting is not a substitute for probability measurements in large language models](https://arxiv.org/abs/2305.13264) (Jennifer Hu, R. Levy, 2023). Shows prompted metalinguistic answers diverge from direct probability measurements. cited by Bai et al. (2025); Song et al. (2025).
- [Two Failures of Self-Consistency in the Multi-Step Reasoning of LLMs](https://arxiv.org/abs/2305.14279) (Angelica Chen, Jason Phang, Alicia Parrish et al., 2023). Self-consistency failures in reasoning. cited by Binder et al. (2024); Comsa & Shanahan (2025).
- [Probing and Steering Evaluation Awareness of Language Models](https://arxiv.org/abs/2507.01786) (Jord Nguyen, Khiem Hoang, Carlo Leonardo Attubato et al., 2025). Evaluation awareness, probed and steered. cites Berglund et al. (2023); Betley et al. (2025).
- [Learning to Interpret Weight Differences in Language Models](https://arxiv.org/abs/2510.05092) (A. Goel, Yoon Kim, N. Shavit et al., 2025). Interpreting weight differences in language; adjacent to self-report of learned behavior. cites Betley et al. (2025); Binder et al. (2024).
- [Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?](https://arxiv.org/abs/2601.10242) (Guan-Xu Chen, Dong-Rui Liu, Jing Shao, 2026). Whether looped models link representations to their verbal outputs. cites Lindsey (2025). cited by Hahami et al. (2026).
- [Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations](https://arxiv.org/abs/2601.22548) (Dani Roytburg, Matthew Bozoukov, Matthew Nguyen et al., 2026). Checks the self-preference evaluations behind self-recognition claims. cites Berglund et al. (2023); Binder et al. (2024).
- [Neologism Learning for Controllability and Self-Verbalization](https://arxiv.org/abs/2510.08506) (John Hewitt, Oyvind Tafjord, Robert Geirhos et al., 2025). Self-verbalization through learned neologisms. cites Berglund et al. (2023); Betley et al. (2025).
- [Verbalizing LLMs' assumptions to explain and control sycophancy](https://arxiv.org/abs/2604.03058) (Myra Cheng, Isabel Sieh, Humishka Zope et al., 2026). Verbalizing a model's assumptions. cites Li et al. (2025); Plunkett et al. (2025).
- [Metacognitive Sensitivity for Test-Time Dynamic Model Selection](https://arxiv.org/abs/2512.10451) (Le Hoang Minh Trinh, L. Pham, T. Pham et al., 2025). Metacognitive sensitivity used for model selection. cites Song et al. (2025); Song et al. (2025).
- [When Self-Reference Fails to Close: Matrix-Level Dynamics in Large Language Models](https://arxiv.org/abs/2604.12128) (Jisung Bae, 2026). Self-reference dynamics; unclear scope. cites Binder et al. (2024); Lindsey (2025).
- [AI and Consciousness: Shifting Focus Towards Tractable Questions](https://arxiv.org/abs/2605.06965) (I. Comsa, 2026). Consciousness debate refocused on tractable questions. cites Comsa & Shanahan (2025); Lindsey (2025).
- [Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing](https://arxiv.org/abs/2606.17478) (Kexin Chen, Yi Liu, Hao-Nan Zhang et al., 2026). Activation explainers used for deception auditing. cites Cywiński et al. (2025); Li et al. (2025).
- [Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data](https://arxiv.org/abs/2508.00741) (S. Imran, Rob Lamb, Peter M. Atkinson, 2025). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023); Betley et al. (2025).
- [Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models](https://arxiv.org/abs/2510.16340) (Pratham Singla, Shivank Garg, Ayush Singh et al., 2025). Evaluates reasoning about reasoning. cites Betley et al. (2025); Binder et al. (2024).
- [Evaluating Self-Orienting in Language and Reasoning Models](https://www.semanticscholar.org/paper/9bf7b40326d0e45463b6a289cbe091fc9c9c8d57) (Eric J. Bigelow, Zergham Ahmed, Tomer D. Ullman). Self-orienting evaluations. cites Betley et al. (2025); Binder et al. (2024).
- [One Faithful Pass Over the Cuckoo's Nest](https://arxiv.org/abs/2609.00383) (Kristina Šekrst, 2026). Appears to be about explanation faithfulness; unclear from the title. cites Binder et al. (2024); Lindsey (2025).
- [Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models](https://arxiv.org/abs/2608.13258) (P. Balani, Subhrakanta Panda, 2026). Self-referential prompting and response stability. cites Comsa & Shanahan (2025); Lindsey (2025).
- [The Assistant as a Privileged Persona: A canonical reference in cross-persona self-recognition](https://arxiv.org/abs/2606.00545) (G. Asvin, 2026). Cross-persona self-recognition. cites Betley et al. (2025); Binder et al. (2024).
- [The Inner Monologue of Language Models: When Reasoning Traces Reveal More Than They Hide](https://doi.org/10.18653/v1/2026.findings-acl.2078) (Pratham Singla, Shivank Garg, Ayush Singh et al., 2026). What reasoning traces reveal. cites Betley et al. (2025); Binder et al. (2024).
- [What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs](https://arxiv.org/abs/2606.28615) (Nhi Nguyen, Shauli Ravfogel, R. Ranganath, 2026). Explanations against the model's own beliefs. cites Bai et al. (2025); Li et al. (2025).

## Suggested: skip (99)

- [Lora: Low-rank adaptation of large language models](https://arxiv.org/abs/2106.09685) (J. Hu, Ye-Long Shen, Phillip Wallis et al., 2021). Fine-tuning method cited as a tool. cited by Betley et al. (2025); Binder et al. (2024); Cywiński et al. (2025); Li et al. (2025); Treutlein et al. (2024); Wang et al. (2025).
- [Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation](https://arxiv.org/abs/2603.18893) (Nicolás Martorell, Bruno Bianchi, 2026). Looks like another version of "Quantitative Introspection in Language Models: Tracking Internal States Across Conversation". cites Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025).
- [Self-Referential Introspection in Large Language Models: The Critical Threshold for Recursive Self-Improvement](https://doi.org/10.3390/e28090951) (Jiang Zhang, Bing Yuan, Qian Zhang, 2026). Looks like another version of "Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement". cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025); Song et al. (2025).
- [Asymmetric Communication: Large Language Models and Language Games](https://arxiv.org/abs/2607.28137) (Enzo Fenoglio, 2026). Relevance to self-report unclear from the title. cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025).
- [Training language models to follow instructions with human feedback](https://arxiv.org/abs/2203.02155) (Long Ouyang, Jeff Wu, Xu Jiang et al., 2022). General ML or interpretability background, not about self-report. cited by Bai et al. (2025); Berglund et al. (2023); Cywiński et al. (2025); Sherburn et al. (2024).
- [Locating and Editing Factual Associations in GPT](https://arxiv.org/abs/2202.05262) (Kevin Meng, David Bau, A. Andonian et al., 2022). General ML or interpretability background, not about self-report. cited by Binder et al. (2024); Li et al. (2025); Sherburn et al. (2024); Treutlein et al. (2024).
- [The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"](https://arxiv.org/abs/2309.12288) (Lukas Berglund, Meg Tong, Maximilian Kaufmann et al., 2023). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023). cited by Betley et al. (2025); Cywiński et al. (2025); Treutlein et al. (2024).
- [Do Large Language Models Latently Perform Multi-Hop Reasoning?](https://arxiv.org/abs/2402.16837) (Sohee Yang, E. Gribovskaya, Nora Kassner et al., 2024). General ML or interpretability background, not about self-report. cites Berglund et al. (2023). cited by Betley et al. (2025); Binder et al. (2024); Treutlein et al. (2024).
- [Truthful AI: Developing and governing AI that does not lie](https://arxiv.org/abs/2110.06674) (Owain Evans, Owen Cotton-Barratt, Lukas Finnveden et al., 2021). About honesty norms for AI, not about self-report of internal states. cited by Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Sherburn et al. (2024).
- [Characterizing the Consistency of the Emergent Misalignment Persona](https://arxiv.org/abs/2604.28082) (Anietta Weckauff, Yu-Chen Zhang, Maksym Andriushchenko, 2026). Emergent-misalignment line; not about self-report. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Plunkett et al. (2025).
- [Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs](https://arxiv.org/abs/2605.20382) (C. Camassa, Derek Shiller, 2026). Relevance to self-report unclear from the title. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- nostalgebraist. A reference to an author, not a paper. cited by Hahami et al. (2026); Li et al. (2025); Pearson-Vogel et al. (2026); Wang et al. (2025).
- [Chain of Thought Prompting Elicits Reasoning in Large Language Models](https://arxiv.org/abs/2201.11903) (Jason Wei, Xue-Zhi Wang, Dale Schuurmans et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Sherburn et al. (2024); Treutlein et al. (2024).
- [Scalinglawsforneurallanguagemodels](https://arxiv.org/abs/2001.08361) (J. Kaplan, Sam McCandlish, T. Henighan et al., 2020). General ML or interpretability background, not about self-report. cited by Bai et al. (2025); Berglund et al. (2023); Sherburn et al. (2024).
- [Discovering Language Model Behaviors with Model-Written Evaluations](https://arxiv.org/abs/2212.09251) (Ethan Perez, Sam Ringer, Kamilė Lukošiūtė et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024).
- [Steering Language Models With Activation Engineering](https://arxiv.org/abs/2308.10248) (A. M. Turner, Lisa Thiergart, Gavin Leech et al., 2023). Steering method cited as a tool. cited by Hahami et al. (2026); Pearson-Vogel et al. (2026); Wang et al. (2025).
- [The alignment problem from a deep learning perspective](https://arxiv.org/abs/2209.00626) (Richard Ngo, Lawrence Chan, Sören Mindermann, 2022). AI safety or control work without a self-report question. cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024).
- [Physics of language models: Part 3.2, knowledge manipulation](https://arxiv.org/abs/2309.14402) (Zeyuan Allen-Zhu, Yuanzhi Li, 2023). General ML or interpretability background, not about self-report. cited by Betley et al. (2025); Binder et al. (2024); Treutlein et al. (2024).
- [Implicit meta-learning may lead language models to trust more reliable sources](https://arxiv.org/abs/2310.15047) (D. Krasheninnikov, Egor Krasheninnikov, B. Mlodozeniec et al., 2023). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023). cited by Betley et al. (2025); Treutlein et al. (2024).
- [Position: It’s Time to Optimize for Self-Consistency](https://www.semanticscholar.org/paper/ca8080bef251475a8606ef4c2d63d6de7e64e115) (Itamar Hagay Pres, Belinda Z. Li, L. Ruis et al.). Looks like another version of "Position: It's Time to Optimize LLMs for Self-Consistency". cites Binder et al. (2024); Li et al. (2025); Plunkett et al. (2025).
- [ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions](https://arxiv.org/abs/2605.24279) (Xian-Zhong Ding, Yangyang Yu, Chang-Wei Liu et al., 2026). Persona drift benchmark; not about self-report. cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- [Automated Interpretability-Driven Model Auditing and Control: A Research Agenda](https://www.semanticscholar.org/paper/5c6070a92d62df10668a8ece4dfb465b81efe528) (Fazl Barez). AI safety or control work without a self-report question. cites Cywiński et al. (2025); Li et al. (2025); Lindsey (2025).
- [Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates](https://arxiv.org/abs/2607.01047) (Elias Najarro, Ane Espeseth, Eleni Nisioti et al., 2026). Relevance to self-report unclear from the title. cites Binder et al. (2024); Hahami et al. (2026); Lindsey (2025).
- [Generalization to Political Beliefs from Fine-Tuning on Sports Team Preferences](https://arxiv.org/abs/2601.04369) (Owen Terry, 2026). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023); Betley et al. (2025); Lindsey (2025).
- [Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models](https://arxiv.org/abs/2607.26173) (Antón de la Fuente, Arthur Conmy, 2026). General ML or interpretability background, not about self-report. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [Artificial Phantasia: Emergent Mental Imagery in Large Language Models](https://arxiv.org/abs/2509.23108) (Morgan McCarty, Jorge Morales, 2025). Mental imagery, not self-report. cites Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- OpenAI. A reference to an organization, not a paper. cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024).
- [Password-Activated Shutdown Protocols for Misaligned Frontier Agents](https://arxiv.org/abs/2512.03089) (Kai Williams, R. Subramani, Francis Rhys Ward, 2025). AI safety or control work without a self-report question. cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- [Phase Transitions in Driven Informational Systems: A Two-Field Perspective on Learning Theory and Non-Equilibrium Chemistry](https://arxiv.org/abs/2605.16325) (Xuan Khanh Truong, 2026). General ML or interpretability background, not about self-report. cites Binder et al. (2024); Lindsey (2025); Song et al. (2025).
- Sleeper agents: Training deceptive llms that persist through safety training. AI safety or control work without a self-report question. cited by Betley et al. (2025); Cywiński et al. (2025); Treutlein et al. (2024).
- [Language Models are Few-Shot Learners](https://arxiv.org/abs/2005.14165) (Tom B. Brown, Benjamin Mann, Nick Ryder et al., 2020). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Sherburn et al. (2024).
- [The llama 3 herd of models](https://arxiv.org/abs/2407.21783) (Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri et al., 2024). General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Hahami et al. (2026).
- [Measuring Massive Multitask Language Understanding](https://arxiv.org/abs/2009.03300) (Dan Hendrycks, Collin Burns, Steven Basart et al., 2020). General ML or interpretability background, not about self-report. cited by Binder et al. (2024); Li et al. (2025).
- [GPT-4o System Card](https://arxiv.org/abs/2410.21276) (OpenAI Aaron Hurst, A. Lerer, Adam P. Goucher et al., 2024). General ML or interpretability background, not about self-report. cited by Betley et al. (2025); Song et al. (2025).
- [Emergent Abilities of Large Language Models](https://arxiv.org/abs/2206.07682) (Jason Wei, Yi Tay, Rishi Bommasani et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Comsa & Shanahan (2025).
- [Training Compute-Optimal Large Language Models](https://arxiv.org/abs/2203.15556) (Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Sherburn et al. (2024).
- [Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models](https://arxiv.org/abs/2206.04615) (Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Treutlein et al. (2024).
- [Sparse Autoencoders Find Highly Interpretable Features in Language Models](https://arxiv.org/abs/2309.08600) (Hoagy Cunningham, Aidan Ewart, L. Smith et al., 2023). General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Li et al. (2025).
- [Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small](https://arxiv.org/abs/2211.00593) (Kevin Wang, Alexandre Variengien, Arthur Conmy et al., 2022). General ML or interpretability background, not about self-report. cited by Li et al. (2025); Sherburn et al. (2024).
- [The fineweb datasets: Decanting the web for the finest text data at scale](https://arxiv.org/abs/2406.17557) (Guilherme Penedo, Hynek Kydlícek, Loubna Ben Allal et al., 2024). General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Li et al. (2025).
- [Mass-Editing Memory in a Transformer](https://arxiv.org/abs/2210.07229) (Kevin Meng, Arnab Sen Sharma, A. Andonian et al., 2022). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Sherburn et al. (2024).
- [A General Language Assistant as a Laboratory for Alignment](https://arxiv.org/abs/2112.00861) (Amanda Askell, Yuntao Bai, Anna Chen et al., 2021). General ML or interpretability background, not about self-report. cited by Berglund et al. (2023); Binder et al. (2024).
- [Progress measures for grokking via mechanistic interpretability](https://arxiv.org/abs/2301.05217) (Neel Nanda, Lawrence Chan, Tom Lieberum et al., 2023). General ML or interpretability background, not about self-report. cited by Li et al. (2025); Sherburn et al. (2024).
- [Discovering Latent Knowledge in Language Models Without Supervision](https://arxiv.org/abs/2212.03827) (Collin Burns, Haotian Ye, D. Klein et al., 2022). General ML or interpretability background, not about self-report. cited by Pearson-Vogel et al. (2026); Sherburn et al. (2024).
- [Towards Automated Circuit Discovery for Mechanistic Interpretability](https://arxiv.org/abs/2304.14997) (Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch et al., 2023). General ML or interpretability background, not about self-report. cited by Li et al. (2025); Sherburn et al. (2024).
- [The geometry of truth: Emergent linear structure in large language model representations of true/false datasets](https://arxiv.org/abs/2310.06824) (Samuel Marks, Max Tegmark, 2023). General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Pearson-Vogel et al. (2026).
- [Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision](https://arxiv.org/abs/2312.09390) (Collin Burns, Pavel Izmailov, J. Kirchner et al., 2023). General ML or interpretability background, not about self-report. cited by Binder et al. (2024); Cywiński et al. (2025).
- [A Survey of the State of Explainable AI for Natural Language Processing](https://arxiv.org/abs/2010.00711) (Marina Danilevsky, Kun Qian, R. Aharonov et al., 2020). General ML or interpretability background, not about self-report. cited by Plunkett et al. (2025); Song et al. (2025).
- [Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2](https://arxiv.org/abs/2408.05147) (Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy et al., 2024). General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Li et al. (2025).
- [Alignment faking in large language models](https://arxiv.org/abs/2412.14093) (R. Greenblatt, Carson E. Denison, Benjamin Wright et al., 2024). AI safety or control work without a self-report question. cited by Betley et al. (2025); Wang et al. (2025).
- [Risks from Learned Optimization in Advanced Machine Learning Systems](https://arxiv.org/abs/1906.01820) (Evan Hubinger, Chris van Merwijk, Vladimir Mikulik et al., 2019). AI safety or control work without a self-report question. cited by Berglund et al. (2023); Betley et al. (2025).
- [Physics of Language Models: Part 3.1, Knowledge Storage and Extraction](https://arxiv.org/abs/2309.14316) (Zeyuan Allen-Zhu, Yuanzhi Li, 2023). General ML or interpretability background, not about self-report. cites Berglund et al. (2023). cited by Treutlein et al. (2024).
- [Model evaluation for extreme risks](https://arxiv.org/abs/2305.15324) (Toby Shevlane, Sebastian Farquhar, Ben Garfinkel et al., 2023). AI safety or control work without a self-report question. cited by Berglund et al. (2023); Betley et al. (2025).
- [On the Nature of Mind.](https://doi.org/10.1038/128744a0) (C. S. Myres, 1931). Philosophy or consciousness debate without a test of self-report. cited by Comsa & Shanahan (2025); Song et al. (2025).
- [Black-box access is insufficient for rigorous ai audits](https://arxiv.org/abs/2401.14446) (Stephen Casper, Carson Ezell, Charlotte Siegmann et al., 2024). AI safety or control work without a self-report question. cited by Cywiński et al. (2025); Plunkett et al. (2025).
- [Training large language models on narrow tasks can lead to broad misalignment](https://arxiv.org/abs/2502.17424) (Jan Betley, Daniel Tan, Niels Warncke et al., 2025). Emergent-misalignment line; not about self-report. cites Berglund et al. (2023); Betley et al. (2025).
- [Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models](https://arxiv.org/abs/2406.10162) (Carson E. Denison, M. MacDiarmid, Fazl Barez et al., 2024). AI safety or control work without a self-report question. cites Berglund et al. (2023). cited by Cywiński et al. (2025).
- [Is Power-Seeking AI an Existential Risk?](https://arxiv.org/abs/2206.13353) (J. Carlsmith, 2022). AI safety or control work without a self-report question. cited by Berglund et al. (2023); Plunkett et al. (2025).
- [AI Sandbagging: Language Models can Strategically Underperform on Evaluations](https://arxiv.org/abs/2406.07358) (Teun van der Weij, Felix Hofstätter, Oliver Jaffe et al., 2024). AI safety or control work without a self-report question. cited by Binder et al. (2024); Cywiński et al. (2025).
- [Persona Features Control Emergent Misalignment](https://arxiv.org/abs/2506.19823) (Miles Wang, Tom Dupré la Tour, Olivia Watkins et al., 2025). Emergent-misalignment line; not about self-report. cites Berglund et al. (2023); Betley et al. (2025).
- [An academic survey on theoretical foundations, common assumptions and the current state of consciousness science](https://doi.org/10.1093/nc/niac011) (Jolien C. Francken, L. Beerendonk, D. Molenaar et al., 2022). Philosophy or consciousness debate without a test of self-report. cited by Binder et al. (2024); Comsa & Shanahan (2025).
- [Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models](https://arxiv.org/abs/2506.13206) (James Chua, Jan Betley, Mia Taylor et al., 2025). Emergent-misalignment line; not about self-report. cites Betley et al. (2025); Binder et al. (2024).
- [Reverse Training to Nurse the Reversal Curse](https://arxiv.org/abs/2403.13799) (Olga Golovneva, Zeyuan Allen-Zhu, Jason E. Weston et al., 2024). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023). cited by Betley et al. (2025).
- [A sketch of an AI control safety case](https://arxiv.org/abs/2501.17315) (Tomasz Korbak, Joshua Clymer, Benjamin Hilton et al., 2025). AI safety or control work without a self-report question. cites Berglund et al. (2023); Binder et al. (2024).
- [Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History](https://arxiv.org/abs/2508.04826) (Tommaso Tosato, S. Helbling, Yorguin José Mantilla Ramos et al., 2025). Personality measurement stability; not about self-report of internal states. cites Binder et al. (2024); Plunkett et al. (2025).
- [Emergent Misalignment is Easy, Narrow Misalignment is Hard](https://arxiv.org/abs/2602.07852) (Anna Soligo, Edward Turner, Senthooran Rajamanoharan et al., 2026). Emergent-misalignment line; not about self-report. cites Berglund et al. (2023); Betley et al. (2025).
- [How to evaluate control measures for LLM agents? A trajectory from today to superintelligence](https://arxiv.org/abs/2504.05259) (Tomasz Korbak, Mikita Balesni, Buck Shlegeris et al., 2025). AI safety or control work without a self-report question. cites Berglund et al. (2023); Binder et al. (2024).
- [Automating Steering for Safe Multimodal Large Language Models](https://arxiv.org/abs/2507.13255) (Lyucheng Wu, Mengru Wang, Ziwen Xu et al., 2025). General ML or interpretability background, not about self-report. cites Betley et al. (2025); Binder et al. (2024).
- [Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs](https://arxiv.org/abs/2407.04108) (Sara Price, Arjun Panickssery, Samuel R. Bowman et al., 2024). AI safety or control work without a self-report question. cites Berglund et al. (2023). cited by Betley et al. (2025).
- [Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety](https://arxiv.org/abs/2506.05451) (Seongmin Lee, Aeree Cho, Grace C. Kim et al., 2025). General ML or interpretability background, not about self-report. cites Betley et al. (2025); Binder et al. (2024).
- [Shaping capabilities with token-level data filtering](https://arxiv.org/abs/2601.21571) (Neil Rathi, Alec Radford, 2026). General ML or interpretability background, not about self-report. cites Berglund et al. (2023); Wang et al. (2025).
- [Subliminal Learning Is Steering Vector Distillation](https://arxiv.org/abs/2606.00995) (C. Blank, A. Bhatia, Senthooran Rajamanoharan et al., 2026). General ML or interpretability background, not about self-report. cites Berglund et al. (2023); Wang et al. (2025).
- [Mechanistic Interpretability Needs Philosophy](https://arxiv.org/abs/2506.18852) (Iwan Williams, Ninell Oldenburg, Ruchira Dhar et al., 2025). Philosophy or consciousness debate without a test of self-report. cites Lindsey (2025); Song et al. (2025).
- [Probe-Rewrite-Evaluate: A Workflow for Reliable Benchmarks and Quantifying Evaluation Awareness](https://arxiv.org/abs/2509.00591) (Lang Xiong, N. Bhargava, Jeremy Chang et al., 2025). AI safety or control work without a self-report question. cites Berglund et al. (2023); Betley et al. (2025).
- [Understanding Emergent Misalignment via Feature Superposition Geometry](https://arxiv.org/abs/2605.00842) (Gouki Minegishi, Hiroki Furuta, Takeshi Kojima et al., 2026). Emergent-misalignment line; not about self-report. cites Betley et al. (2025); Wang et al. (2025).
- [A Disproof of Large Language Model Consciousness: The Necessity of Continual Learning for Consciousness](https://arxiv.org/abs/2512.12802) (Erik P. Hoel, 2025). Philosophy or consciousness debate without a test of self-report. cites Binder et al. (2024); Lindsey (2025).
- [Programming by Backprop: LLMs Acquire Reusable Algorithmic Abstractions During Code Training](https://doi.org/10.48550/arXiv.2506.18777) (Jonathan Cook, Silvia Sapora, Arash Ahmadian et al., 2025). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023); Betley et al. (2025).
- [Covert Influence Between Language Models](https://arxiv.org/abs/2606.04071) (Avidan Shah, J. Chooi, Jinghuai Ou et al., 2026). AI safety or control work without a self-report question. cites Li et al. (2025); Lindsey (2025).
- [Models That Know How Evaluations Are Designed Score Safer](https://arxiv.org/abs/2605.28591) (K. Deckenbach, Haritz Puerto, Jonas Geiping et al., 2026). AI safety or control work without a self-report question. cites Berglund et al. (2023); Betley et al. (2025).
- [Assessing LLMs'mathematical abilities requires understanding the various mechanisms of mathematical creativity](https://arxiv.org/abs/2608.16118) (Silvère Gangloff, 2026). General ML or interpretability background, not about self-report. cites Binder et al. (2024); Hahami et al. (2026).
- [If LLMs Have Human-Like Attributes, Then So Does Age of Empires II](https://arxiv.org/abs/2605.31514) (Adrian de Wynter, 2026). Philosophy or consciousness debate without a test of self-report. cites Betley et al. (2025); Lindsey (2025).
- [On the Creativity of AI Agents](https://arxiv.org/abs/2604.13242) (Giorgio Franceschelli, Mirco Musolesi, 2026). General ML or interpretability background, not about self-report. cites Binder et al. (2024); Comsa & Shanahan (2025).
- [Artificial Phantasia: Evidence for Propositional Reasoning-Based Mental Imagery in Large Language Models](https://doi.org/10.48550/arXiv.2509.23108) (Morgan McCarty, Jorge Morales, 2025). Mental imagery, not self-report. cites Betley et al. (2025); Plunkett et al. (2025).
- [Can LLMs Perceive Time? An Empirical Investigation](https://arxiv.org/abs/2604.00010) (Aniketh Garikaparthi, 2026). General ML or interpretability background, not about self-report. cites Binder et al. (2024); Lindsey (2025).
- [Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power](https://arxiv.org/abs/2604.21043) (Travis LaCroix, Fintan Mallory, Sasha Luccioni, 2026). Philosophy or consciousness debate without a test of self-report. cites Binder et al. (2024); Lindsey (2025).
- [Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?](https://arxiv.org/abs/2609.14803) (Afshin Khadangi, 2026). Relevance to self-report unclear from the title. cites Binder et al. (2024); Song et al. (2025).
- [Emergent Language as an Approach to Conscious AI](https://arxiv.org/abs/2606.06380) (Zengqing Wu, Chuan Xiao, 2026). Philosophy or consciousness debate without a test of self-report. cites Betley et al. (2025); Lindsey (2025).
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs. Emergent-misalignment line; not about self-report. cited by Cywiński et al. (2025); Pearson-Vogel et al. (2026).
- [Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning](https://arxiv.org/abs/2606.15507) (Ali Dasdan, Manan Shah, W. R. Neuman et al., 2026). General ML or interpretability background, not about self-report. cites Li et al. (2025); Lindsey (2025).
- [Generalization Dynamics of LM Pre-training](https://arxiv.org/abs/2609.33150) (Jia-Xin Wen, Zhengxuan Wu, D. Song et al., 2026). General ML or interpretability background, not about self-report. cites Berglund et al. (2023); Wang et al. (2025).
- Llama 3 model card. General ML or interpretability background, not about self-report. cited by Betley et al. (2025); Treutlein et al. (2024).
- Scaling monosemantic-ity: Extracting interpretable features from Claude 3 Sonnet. General ML or interpretability background, not about self-report. cited by Comsa & Shanahan (2025); Li et al. (2025).
- [Semantic Containment as a Fundamental Property of Emergent Misalignment](https://arxiv.org/abs/2603.04407) (Rohan Saxena, 2026). Emergent-misalignment line; not about self-report. cites Berglund et al. (2023); Betley et al. (2025).
- [Stress-Testing Alignment Audits With Prompt-Level Strategic Deception](https://arxiv.org/abs/2602.08877) (Oliver Daniels, Perusha Moodley, Benjamin M. Marlin et al., 2026). AI safety or control work without a self-report question. cites Berglund et al. (2023); Cywiński et al. (2025).
- [T HE T WO -H OP C URSE : LLM S TRAINED ON A (cid:41) B , B (cid:41) C FAIL TO LEARN A (cid:41) C](https://www.semanticscholar.org/paper/486c6a8eb2ad63150024dc6ebb263e9643f9aa5a) (Mikita Balesni, Apollo Research, Tomasz Korbak et al.). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023); Binder et al. (2024).
- Towards monosemanticity: Decomposing language models with dictionary learning. General ML or interpretability background, not about self-report. cited by Cywiński et al. (2025); Li et al. (2025).
- [Unsupervised Features Mining via Activation Geometry](https://arxiv.org/abs/2607.04222) (Amit Levi, Elad David, M. Fomin, 2026). General ML or interpretability background, not about self-report. cites Binder et al. (2024); Lindsey (2025).
- [VLMs Can Aggregate Scattered Training Patches](https://arxiv.org/abs/2506.03614) (Zhanhui Zhou, Lingjie Chen, Chao Yang et al., 2025). Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. cites Berglund et al. (2023); Betley et al. (2025).
- Without specific countermeasures, the easiest path to transformative ai likely leads to ai takeover. AI safety or control work without a self-report question. cited by Berglund et al. (2023); Treutlein et al. (2024).

---

Source: https://introspection.infinite.fun/frontier · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
