Frontier
194 papers one citation away from the wiki that do not have a page yet.
The crawler looked at the references and citers of every paper page and saw 1084 distinct neighbors. A neighbor is listed here if it connects to at least 2 wiki papers, or if someone added it as a lead. The triage labels are suggestions, made from each candidate's title and its place in the citation graph and not from reading it. A candidate becomes a page only after a person accepts it.
Last crawled 2026-10-06. Source: Semantic Scholar Graph API, with arXiv HTML bibliographies where Semantic Scholar has no reference list. Also as JSON.
Suggested: add (43)
- Emergent Introspection in AI is Content-Agnostic Introspection itself: asks what kind of content models can introspect on. Cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- Dissociating Direct Access from Inference in AI Introspection Introspection itself: separates direct access from inference. Cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision Self-explanation training and whether it tracks behavior. Cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- Metacognition in LLMs: Foundations, Progress, and Opportunities Survey of metacognition in LLMs; useful as a map. Cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- Language models (mostly) know what they know Foundational result on models knowing what they know; cited by six pages. Cited by Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Sherburn et al. (2024).
- Evidence for Limited Metacognition in LLMs Tests metacognition directly. Cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025). Cited by Plunkett et al. (2025).
- Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation Measures what models can report about their own computation. Cites Betley et al. (2025); Binder et al. (2024); Hahami et al. (2026); Lindsey (2025); Pearson-Vogel et al. (2026); Song et al. (2025).
- Quantitative Introspection in Language Models: Tracking Internal States Across Conversation Operationalizes introspection as coupling between self-report and a probed internal state. Cites Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025).
- LLM Evaluators Recognize and Favor Their Own Generations Self-recognition: whether models can tell their own outputs apart. Cites Berglund et al. (2023). Cited by Binder et al. (2024); Comsa & Shanahan (2025); Hahami et al. (2026); Song et al. (2025).
- Can LLMs Introspect? A Reality Check A skeptical check on introspection claims. Cites Binder et al. (2024); Li et al. (2025); Lindsey (2025); Plunkett et al. (2025); Song et al. (2025).
- Steering Awareness: Detecting Activation Steering from Within Concept-injection line: detecting activation steering from within. Cites Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Song et al. (2025). Cited by Hahami et al. (2026).
- Mechanisms of Introspective Awareness Mechanistic account of introspective awareness. Cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Wang et al. (2025).
- In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access? Tests privileged access through control of internal representations. Cites Betley et al. (2025); Binder et al. (2024); Li et al. (2025); Song et al. (2025); Song et al. (2025).
- Telling more than we can know: Verbal reports on mental processes. The classic study of human confabulation that this literature borrows its framing from; cited by four pages. Cited by Comsa & Shanahan (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting Unfaithful chain-of-thought explanations; the main adjacent line on explanation faithfulness. Cited by Li et al. (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Sherburn et al. (2024).
- Me, myself, and ai: The situational awareness dataset (sad) for llms Benchmark of situational awareness, including self-knowledge tasks. Cited by Betley et al. (2025); Comsa & Shanahan (2025); Hahami et al. (2026); Plunkett et al. (2025).
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers Training models to explain activations; adjacent to self-explanation. Cites Cywiński et al. (2025); Li et al. (2025); Lindsey (2025). Cited by Hahami et al. (2026).
- Towards Evaluating AI Systems for Moral Status Using Self-Reports Proposes training and evaluating self-reports for moral-status questions. Cites Berglund et al. (2023). Cited by Binder et al. (2024); Comsa & Shanahan (2025); Plunkett et al. (2025).
- A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior Tests whether self-explanations predict behavior. Cites Binder et al. (2024); Li et al. (2025); Lindsey (2025); Plunkett et al. (2025).
- Introspection Adapters: Training LLMs to Report Their Learned Behaviors Trains models to report their learned behaviors. Cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs Mechanistic conditions for behavioral self-awareness. Cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Wang et al. (2025).
- Strangers to Themselves: What Language Models Say About Themselves Is Generic A skeptical result: self-descriptions are generic rather than self-specific. Cites Bai et al. (2025); Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- Reasoning Models Don’t Always Say What They Think Chain-of-thought faithfulness in reasoning models. Cited by Cywiński et al. (2025); Li et al. (2025); Plunkett et al. (2025).
- From Imitation to Introspection: Probing Self-Consciousness in Language Models Probes self-consciousness concepts in models. Cites Berglund et al. (2023); Binder et al. (2024). Cited by Comsa & Shanahan (2025).
- Large Language Models Report Subjective Experience Under Self-Referential Processing Self-reports of experience under self-referential prompting. Cites Betley et al. (2025); Lindsey (2025); Plunkett et al. (2025).
- Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare Compares stated and revealed preferences. Cites Betley et al. (2025); Binder et al. (2024); Song et al. (2025).
- Do Activation Verbalization Methods Convey Privileged Information? Asks whether verbalized activations carry privileged information. Cites Binder et al. (2024); Song et al. (2025). Cited by Li et al. (2025).
- Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment Behavioral self-awareness tracking fine-tuning changes. Cites Berglund et al. (2023); Binder et al. (2024); Lindsey (2025).
- Language models recognize dropout and Gaussian noise applied to their activations Concept-injection line: detecting noise applied to activations. Cites Comsa & Shanahan (2025); Lindsey (2025); Pearson-Vogel et al. (2026).
- Fine-Tuning Language Models to Know What They Know Trains models to know what they know. Cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025).
- A mechanistic study of language model introspection Mechanistic study of introspection; also found by keyword search. Cites Binder et al. (2024); Hahami et al. (2026); Lindsey (2025).
- Introspecting Alignment Shifts Beyond Behaviors Implanted Through Fine-Tuning Introspecting on changes from fine-tuning. Cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations Metacognitive monitoring and control of internal activations. Cites Binder et al. (2024). Cited by Hahami et al. (2026).
- Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives Trains models to self-report hidden objectives. Cites Betley et al. (2025); Binder et al. (2024).
- Do Large Language Models Know What They Are Capable Of? Self-knowledge of capabilities. Cites Betley et al. (2025); Binder et al. (2024).
- Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness Privileged access: disentangles privileged knowledge of correctness. Cites Binder et al. (2024); Li et al. (2025).
- Me, Myself, and π: Evaluating and Explaining LLM Introspection Evaluates and explains introspection. Cites Binder et al. (2024); Lindsey (2025).
- Can LLMs Reliably Self-Report Adversarial Prefills, and How? Self-report of adversarial prefills. Cites Binder et al. (2024); Cywiński et al. (2025).
- Do Language Models Know When They'll Refuse? Probing Introspective Awareness of Safety Boundaries Introspective awareness of when the model will refuse. Cites Binder et al. (2024); Lindsey (2025).
- Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect Fine-tuning small models to introspect. Author thread: https://x.com/ElyHahami/status/2078161491069718639. Cites Binder et al. (2024); Lindsey (2025).
- Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs Training self-interpretation from interpretability artifacts. Cites Li et al. (2025); Lindsey (2025).
- Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports An identifiability argument about what self-reports can show. Cites Li et al. (2025); Lindsey (2025).
- What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation Evidence standards for mechanistic claims about self-explanation. Cites Binder et al. (2024); Lindsey (2025).
Suggested: maybe (52)
- Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement About self-reference and introspection, but the framing is recursive self-improvement. Cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025); Song et al. (2025).
- Introspection Philosophy reference on human introspection; background for definitions. Cited by Binder et al. (2024); Comsa & Shanahan (2025); Plunkett et al. (2025); Song et al. (2025); Song et al. (2025).
- "As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It Self-referential voice and steering; relevance to self-report unclear from the title. Cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025).
- Anosognosia in LLMs: Probing Self-Awareness of Quantized Computational Substrate Self-awareness of quantization; narrow but on topic. Cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Song et al. (2025).
- Counterfactual Simulation Training for Chain-of-Thought Faithfulness Chain-of-thought faithfulness training. Cites Binder et al. (2024); Comsa & Shanahan (2025); Li et al. (2025); Plunkett et al. (2025).
- Tell, don't show: Declarative facts influence how LLMs generalize Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. Cites Berglund et al. (2023). Cited by Betley et al. (2025); Binder et al. (2024); Treutlein et al. (2024).
- Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns Bears on privileged access: correctness predictors that are model-agnostic. Cites Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025).
- Hallucinations Undermine Trust; Metacognition is a Way Forward Position piece on metacognition. Cites Binder et al. (2024); Li et al. (2025); Lindsey (2025); Song et al. (2025).
- Steering Awareness: Models Can Be Trained to Detect Activation Steering Looks like another version of "Steering Awareness: Detecting Activation Steering from Within". Cites Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Song et al. (2025).
- From Simulation to Enaction: Post-trained language models recognize and react to their own generations Self-recognition of own generations. Cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models Trained denial in self-reports about consciousness. Cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments Evaluates explanations of behavior with counterfactuals. Cites Betley et al. (2025); Binder et al. (2024); Cywiński et al. (2025); Li et al. (2025).
- Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment Self-recognition fine-tuning, in the emergent-misalignment setting. Cites Berglund et al. (2023); Betley et al. (2025); Lindsey (2025); Pearson-Vogel et al. (2026).
- Language Models Act on Hidden Valence Hidden internal valence and behavior; may bear on self-report. Cites Binder et al. (2024); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025).
- Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary Proto-introspection in looped models; unclear scope. Cites Betley et al. (2025); Binder et al. (2024); Pearson-Vogel et al. (2026); Song et al. (2025).
- Teaching Models to Express Their Uncertainty in Words Verbalized uncertainty; background for self-knowledge of confidence. Cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024).
- The Unreliability of Naive Introspection Philosophy of human introspection; background for definitions. Cited by Binder et al. (2024); Comsa & Shanahan (2025); Song et al. (2025).
- Are DeepSeek R1 And Other Reasoning Models More Faithful? Chain-of-thought faithfulness in reasoning models. Cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- Metacognition and Uncertainty Communication in Humans and Large Language Models Metacognition and uncertainty, humans compared with LLMs. Cites Betley et al. (2025); Binder et al. (2024). Cited by Plunkett et al. (2025).
- Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants Interpretability assistants that decode concepts; adjacent to self-explanation. Cites Li et al. (2025); Lindsey (2025). Cited by Hahami et al. (2026).
- Agentic Knowledgeable Self-awareness Agent self-awareness of knowledge; unclear scope. Cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- AI Awareness Survey of AI awareness. Cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- No Reliable Evidence of Self-Reported Sentience in Small Large Language Models Self-reported sentience in small models. Cites Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- Position: It's Time to Optimize LLMs for Self-Consistency Position piece on self-consistency. Cites Binder et al. (2024); Li et al. (2025); Plunkett et al. (2025).
- Higher-order representation in AI Higher-order representation; theory background. Cites Betley et al. (2025); Binder et al. (2024); Plunkett et al. (2025).
- Self-CTRL: Self-Consistency Training with Reinforcement Learning Self-consistency training. Cites Berglund et al. (2023); Betley et al. (2025); Plunkett et al. (2025).
- Imprint Reader: From Weight-Update Readout to Behavioral Intervention Reading out weight updates; adjacent to self-report of learned behavior. Cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- Questionnaire Responses Do not Capture the Safety of AI Agents Stated answers against agent behavior. Cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences Psychometrics of self-described experience. Cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025).
- Do large language models know what they don’t know? Knowing what one does not know; background for self-knowledge of confidence. Cited by Betley et al. (2025); Comsa & Shanahan (2025).
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models Verbalizing hidden representations; adjacent to self-explanation. Cited by Hahami et al. (2026); Li et al. (2025).
- Auditing language models for hidden objectives Auditing for hidden objectives; the other side of unfaithful self-report. Cited by Cywiński et al. (2025); Wang et al. (2025).
- Prompting is not a substitute for probability measurements in large language models Shows prompted metalinguistic answers diverge from direct probability measurements. Cited by Bai et al. (2025); Song et al. (2025).
- Two Failures of Self-Consistency in the Multi-Step Reasoning of LLMs Self-consistency failures in reasoning. Cited by Binder et al. (2024); Comsa & Shanahan (2025).
- Probing and Steering Evaluation Awareness of Language Models Evaluation awareness, probed and steered. Cites Berglund et al. (2023); Betley et al. (2025).
- Learning to Interpret Weight Differences in Language Models Interpreting weight differences in language; adjacent to self-report of learned behavior. Cites Betley et al. (2025); Binder et al. (2024).
- Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs? Whether looped models link representations to their verbal outputs. Cites Lindsey (2025). Cited by Hahami et al. (2026).
- Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations Checks the self-preference evaluations behind self-recognition claims. Cites Berglund et al. (2023); Binder et al. (2024).
- Neologism Learning for Controllability and Self-Verbalization Self-verbalization through learned neologisms. Cites Berglund et al. (2023); Betley et al. (2025).
- Verbalizing LLMs' assumptions to explain and control sycophancy Verbalizing a model's assumptions. Cites Li et al. (2025); Plunkett et al. (2025).
- Metacognitive Sensitivity for Test-Time Dynamic Model Selection Metacognitive sensitivity used for model selection. Cites Song et al. (2025); Song et al. (2025).
- When Self-Reference Fails to Close: Matrix-Level Dynamics in Large Language Models Self-reference dynamics; unclear scope. Cites Binder et al. (2024); Lindsey (2025).
- AI and Consciousness: Shifting Focus Towards Tractable Questions Consciousness debate refocused on tractable questions. Cites Comsa & Shanahan (2025); Lindsey (2025).
- Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing Activation explainers used for deception auditing. Cites Cywiński et al. (2025); Li et al. (2025).
- Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. Cites Berglund et al. (2023); Betley et al. (2025).
- Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models Evaluates reasoning about reasoning. Cites Betley et al. (2025); Binder et al. (2024).
- Evaluating Self-Orienting in Language and Reasoning Models Self-orienting evaluations. Cites Betley et al. (2025); Binder et al. (2024).
- One Faithful Pass Over the Cuckoo's Nest Appears to be about explanation faithfulness; unclear from the title. Cites Binder et al. (2024); Lindsey (2025).
- Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models Self-referential prompting and response stability. Cites Comsa & Shanahan (2025); Lindsey (2025).
- The Assistant as a Privileged Persona: A canonical reference in cross-persona self-recognition Cross-persona self-recognition. Cites Betley et al. (2025); Binder et al. (2024).
- The Inner Monologue of Language Models: When Reasoning Traces Reveal More Than They Hide What reasoning traces reveal. Cites Betley et al. (2025); Binder et al. (2024).
- What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs Explanations against the model's own beliefs. Cites Bai et al. (2025); Li et al. (2025).
Suggested: skip (99)
- Lora: Low-rank adaptation of large language models Fine-tuning method cited as a tool. Cited by Betley et al. (2025); Binder et al. (2024); Cywiński et al. (2025); Li et al. (2025); Treutlein et al. (2024); Wang et al. (2025).
- Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation Looks like another version of "Quantitative Introspection in Language Models: Tracking Internal States Across Conversation". Cites Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Pearson-Vogel et al. (2026); Plunkett et al. (2025); Song et al. (2025).
- Self-Referential Introspection in Large Language Models: The Critical Threshold for Recursive Self-Improvement Looks like another version of "Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement". Cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025); Song et al. (2025).
- Asymmetric Communication: Large Language Models and Language Games Relevance to self-report unclear from the title. Cites Betley et al. (2025); Binder et al. (2024); Comsa & Shanahan (2025); Lindsey (2025); Song et al. (2025).
- Training language models to follow instructions with human feedback General ML or interpretability background, not about self-report. Cited by Bai et al. (2025); Berglund et al. (2023); Cywiński et al. (2025); Sherburn et al. (2024).
- Locating and Editing Factual Associations in GPT General ML or interpretability background, not about self-report. Cited by Binder et al. (2024); Li et al. (2025); Sherburn et al. (2024); Treutlein et al. (2024).
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. Cites Berglund et al. (2023). Cited by Betley et al. (2025); Cywiński et al. (2025); Treutlein et al. (2024).
- Do Large Language Models Latently Perform Multi-Hop Reasoning? General ML or interpretability background, not about self-report. Cites Berglund et al. (2023). Cited by Betley et al. (2025); Binder et al. (2024); Treutlein et al. (2024).
- Truthful AI: Developing and governing AI that does not lie About honesty norms for AI, not about self-report of internal states. Cited by Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Sherburn et al. (2024).
- Characterizing the Consistency of the Emergent Misalignment Persona Emergent-misalignment line; not about self-report. Cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024); Plunkett et al. (2025).
- Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs Relevance to self-report unclear from the title. Cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- nostalgebraist A reference to an author, not a paper. Cited by Hahami et al. (2026); Li et al. (2025); Pearson-Vogel et al. (2026); Wang et al. (2025).
- Chain of Thought Prompting Elicits Reasoning in Large Language Models General ML or interpretability background, not about self-report. Cited by Berglund et al. (2023); Sherburn et al. (2024); Treutlein et al. (2024).
- Scalinglawsforneurallanguagemodels General ML or interpretability background, not about self-report. Cited by Bai et al. (2025); Berglund et al. (2023); Sherburn et al. (2024).
- Discovering Language Model Behaviors with Model-Written Evaluations General ML or interpretability background, not about self-report. Cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024).
- Steering Language Models With Activation Engineering Steering method cited as a tool. Cited by Hahami et al. (2026); Pearson-Vogel et al. (2026); Wang et al. (2025).
- The alignment problem from a deep learning perspective AI safety or control work without a self-report question. Cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024).
- Physics of language models: Part 3.2, knowledge manipulation General ML or interpretability background, not about self-report. Cited by Betley et al. (2025); Binder et al. (2024); Treutlein et al. (2024).
- Implicit meta-learning may lead language models to trust more reliable sources Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. Cites Berglund et al. (2023). Cited by Betley et al. (2025); Treutlein et al. (2024).
- Position: It’s Time to Optimize for Self-Consistency Looks like another version of "Position: It's Time to Optimize LLMs for Self-Consistency". Cites Binder et al. (2024); Li et al. (2025); Plunkett et al. (2025).
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions Persona drift benchmark; not about self-report. Cites Betley et al. (2025); Binder et al. (2024); Lindsey (2025).
- Automated Interpretability-Driven Model Auditing and Control: A Research Agenda AI safety or control work without a self-report question. Cites Cywiński et al. (2025); Li et al. (2025); Lindsey (2025).
- Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates Relevance to self-report unclear from the title. Cites Binder et al. (2024); Hahami et al. (2026); Lindsey (2025).
- Generalization to Political Beliefs from Fine-Tuning on Sports Team Preferences Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. Cites Berglund et al. (2023); Betley et al. (2025); Lindsey (2025).
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models General ML or interpretability background, not about self-report. Cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- Artificial Phantasia: Emergent Mental Imagery in Large Language Models Mental imagery, not self-report. Cites Binder et al. (2024); Lindsey (2025); Plunkett et al. (2025).
- OpenAI A reference to an organization, not a paper. Cited by Berglund et al. (2023); Binder et al. (2024); Sherburn et al. (2024).
- Password-Activated Shutdown Protocols for Misaligned Frontier Agents AI safety or control work without a self-report question. Cites Berglund et al. (2023); Betley et al. (2025); Binder et al. (2024).
- Phase Transitions in Driven Informational Systems: A Two-Field Perspective on Learning Theory and Non-Equilibrium Chemistry General ML or interpretability background, not about self-report. Cites Binder et al. (2024); Lindsey (2025); Song et al. (2025).
- Sleeper agents: Training deceptive llms that persist through safety training AI safety or control work without a self-report question. Cited by Betley et al. (2025); Cywiński et al. (2025); Treutlein et al. (2024).
- Language Models are Few-Shot Learners General ML or interpretability background, not about self-report. Cited by Berglund et al. (2023); Sherburn et al. (2024).
- The llama 3 herd of models General ML or interpretability background, not about self-report. Cited by Cywiński et al. (2025); Hahami et al. (2026).
- Measuring Massive Multitask Language Understanding General ML or interpretability background, not about self-report. Cited by Binder et al. (2024); Li et al. (2025).
- GPT-4o System Card General ML or interpretability background, not about self-report. Cited by Betley et al. (2025); Song et al. (2025).
- Emergent Abilities of Large Language Models General ML or interpretability background, not about self-report. Cited by Berglund et al. (2023); Comsa & Shanahan (2025).
- Training Compute-Optimal Large Language Models General ML or interpretability background, not about self-report. Cited by Berglund et al. (2023); Sherburn et al. (2024).
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models General ML or interpretability background, not about self-report. Cited by Berglund et al. (2023); Treutlein et al. (2024).
- Sparse Autoencoders Find Highly Interpretable Features in Language Models General ML or interpretability background, not about self-report. Cited by Cywiński et al. (2025); Li et al. (2025).
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small General ML or interpretability background, not about self-report. Cited by Li et al. (2025); Sherburn et al. (2024).
- The fineweb datasets: Decanting the web for the finest text data at scale General ML or interpretability background, not about self-report. Cited by Cywiński et al. (2025); Li et al. (2025).
- Mass-Editing Memory in a Transformer General ML or interpretability background, not about self-report. Cited by Berglund et al. (2023); Sherburn et al. (2024).
- A General Language Assistant as a Laboratory for Alignment General ML or interpretability background, not about self-report. Cited by Berglund et al. (2023); Binder et al. (2024).
- Progress measures for grokking via mechanistic interpretability General ML or interpretability background, not about self-report. Cited by Li et al. (2025); Sherburn et al. (2024).
- Discovering Latent Knowledge in Language Models Without Supervision General ML or interpretability background, not about self-report. Cited by Pearson-Vogel et al. (2026); Sherburn et al. (2024).
- Towards Automated Circuit Discovery for Mechanistic Interpretability General ML or interpretability background, not about self-report. Cited by Li et al. (2025); Sherburn et al. (2024).
- The geometry of truth: Emergent linear structure in large language model representations of true/false datasets General ML or interpretability background, not about self-report. Cited by Cywiński et al. (2025); Pearson-Vogel et al. (2026).
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision General ML or interpretability background, not about self-report. Cited by Binder et al. (2024); Cywiński et al. (2025).
- A Survey of the State of Explainable AI for Natural Language Processing General ML or interpretability background, not about self-report. Cited by Plunkett et al. (2025); Song et al. (2025).
- Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2 General ML or interpretability background, not about self-report. Cited by Cywiński et al. (2025); Li et al. (2025).
- Alignment faking in large language models AI safety or control work without a self-report question. Cited by Betley et al. (2025); Wang et al. (2025).
- Risks from Learned Optimization in Advanced Machine Learning Systems AI safety or control work without a self-report question. Cited by Berglund et al. (2023); Betley et al. (2025).
- Physics of Language Models: Part 3.1, Knowledge Storage and Extraction General ML or interpretability background, not about self-report. Cites Berglund et al. (2023). Cited by Treutlein et al. (2024).
- Model evaluation for extreme risks AI safety or control work without a self-report question. Cited by Berglund et al. (2023); Betley et al. (2025).
- On the Nature of Mind. Philosophy or consciousness debate without a test of self-report. Cited by Comsa & Shanahan (2025); Song et al. (2025).
- Black-box access is insufficient for rigorous ai audits AI safety or control work without a self-report question. Cited by Cywiński et al. (2025); Plunkett et al. (2025).
- Training large language models on narrow tasks can lead to broad misalignment Emergent-misalignment line; not about self-report. Cites Berglund et al. (2023); Betley et al. (2025).
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models AI safety or control work without a self-report question. Cites Berglund et al. (2023). Cited by Cywiński et al. (2025).
- Is Power-Seeking AI an Existential Risk? AI safety or control work without a self-report question. Cited by Berglund et al. (2023); Plunkett et al. (2025).
- AI Sandbagging: Language Models can Strategically Underperform on Evaluations AI safety or control work without a self-report question. Cited by Binder et al. (2024); Cywiński et al. (2025).
- Persona Features Control Emergent Misalignment Emergent-misalignment line; not about self-report. Cites Berglund et al. (2023); Betley et al. (2025).
- An academic survey on theoretical foundations, common assumptions and the current state of consciousness science Philosophy or consciousness debate without a test of self-report. Cited by Binder et al. (2024); Comsa & Shanahan (2025).
- Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models Emergent-misalignment line; not about self-report. Cites Betley et al. (2025); Binder et al. (2024).
- Reverse Training to Nurse the Reversal Curse Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. Cites Berglund et al. (2023). Cited by Betley et al. (2025).
- A sketch of an AI control safety case AI safety or control work without a self-report question. Cites Berglund et al. (2023); Binder et al. (2024).
- Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History Personality measurement stability; not about self-report of internal states. Cites Binder et al. (2024); Plunkett et al. (2025).
- Emergent Misalignment is Easy, Narrow Misalignment is Hard Emergent-misalignment line; not about self-report. Cites Berglund et al. (2023); Betley et al. (2025).
- How to evaluate control measures for LLM agents? A trajectory from today to superintelligence AI safety or control work without a self-report question. Cites Berglund et al. (2023); Binder et al. (2024).
- Automating Steering for Safe Multimodal Large Language Models General ML or interpretability background, not about self-report. Cites Betley et al. (2025); Binder et al. (2024).
- Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs AI safety or control work without a self-report question. Cites Berglund et al. (2023). Cited by Betley et al. (2025).
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety General ML or interpretability background, not about self-report. Cites Betley et al. (2025); Binder et al. (2024).
- Shaping capabilities with token-level data filtering General ML or interpretability background, not about self-report. Cites Berglund et al. (2023); Wang et al. (2025).
- Subliminal Learning Is Steering Vector Distillation General ML or interpretability background, not about self-report. Cites Berglund et al. (2023); Wang et al. (2025).
- Mechanistic Interpretability Needs Philosophy Philosophy or consciousness debate without a test of self-report. Cites Lindsey (2025); Song et al. (2025).
- Probe-Rewrite-Evaluate: A Workflow for Reliable Benchmarks and Quantifying Evaluation Awareness AI safety or control work without a self-report question. Cites Berglund et al. (2023); Betley et al. (2025).
- Understanding Emergent Misalignment via Feature Superposition Geometry Emergent-misalignment line; not about self-report. Cites Betley et al. (2025); Wang et al. (2025).
- A Disproof of Large Language Model Consciousness: The Necessity of Continual Learning for Consciousness Philosophy or consciousness debate without a test of self-report. Cites Binder et al. (2024); Lindsey (2025).
- Programming by Backprop: LLMs Acquire Reusable Algorithmic Abstractions During Code Training Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. Cites Berglund et al. (2023); Betley et al. (2025).
- Covert Influence Between Language Models AI safety or control work without a self-report question. Cites Li et al. (2025); Lindsey (2025).
- Models That Know How Evaluations Are Designed Score Safer AI safety or control work without a self-report question. Cites Berglund et al. (2023); Betley et al. (2025).
- Assessing LLMs'mathematical abilities requires understanding the various mechanisms of mathematical creativity General ML or interpretability background, not about self-report. Cites Binder et al. (2024); Hahami et al. (2026).
- If LLMs Have Human-Like Attributes, Then So Does Age of Empires II Philosophy or consciousness debate without a test of self-report. Cites Betley et al. (2025); Lindsey (2025).
- On the Creativity of AI Agents General ML or interpretability background, not about self-report. Cites Binder et al. (2024); Comsa & Shanahan (2025).
- Artificial Phantasia: Evidence for Propositional Reasoning-Based Mental Imagery in Large Language Models Mental imagery, not self-report. Cites Betley et al. (2025); Plunkett et al. (2025).
- Can LLMs Perceive Time? An Empirical Investigation General ML or interpretability background, not about self-report. Cites Binder et al. (2024); Lindsey (2025).
- Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power Philosophy or consciousness debate without a test of self-report. Cites Binder et al. (2024); Lindsey (2025).
- Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid? Relevance to self-report unclear from the title. Cites Binder et al. (2024); Song et al. (2025).
- Emergent Language as an Approach to Conscious AI Philosophy or consciousness debate without a test of self-report. Cites Betley et al. (2025); Lindsey (2025).
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs Emergent-misalignment line; not about self-report. Cited by Cywiński et al. (2025); Pearson-Vogel et al. (2026).
- Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning General ML or interpretability background, not about self-report. Cites Li et al. (2025); Lindsey (2025).
- Generalization Dynamics of LM Pre-training General ML or interpretability background, not about self-report. Cites Berglund et al. (2023); Wang et al. (2025).
- Llama 3 model card General ML or interpretability background, not about self-report. Cited by Betley et al. (2025); Treutlein et al. (2024).
- Scaling monosemantic-ity: Extracting interpretable features from Claude 3 Sonnet General ML or interpretability background, not about self-report. Cited by Comsa & Shanahan (2025); Li et al. (2025).
- Semantic Containment as a Fundamental Property of Emergent Misalignment Emergent-misalignment line; not about self-report. Cites Berglund et al. (2023); Betley et al. (2025).
- Stress-Testing Alignment Audits With Prompt-Level Strategic Deception AI safety or control work without a self-report question. Cites Berglund et al. (2023); Cywiński et al. (2025).
- T HE T WO -H OP C URSE : LLM S TRAINED ON A (cid:41) B , B (cid:41) C FAIL TO LEARN A (cid:41) C Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. Cites Berglund et al. (2023); Binder et al. (2024).
- Towards monosemanticity: Decomposing language models with dictionary learning General ML or interpretability background, not about self-report. Cited by Cywiński et al. (2025); Li et al. (2025).
- Unsupervised Features Mining via Activation Geometry General ML or interpretability background, not about self-report. Cites Binder et al. (2024); Lindsey (2025).
- VLMs Can Aggregate Scattered Training Patches Out-of-context reasoning and generalization from fine-tuning; adjacent to the adjacent tier. Cites Berglund et al. (2023); Betley et al. (2025).
- Without specific countermeasures, the easiest path to transformative ai likely leads to ai takeover AI safety or control work without a self-report question. Cited by Berglund et al. (2023); Treutlein et al. (2024).