Concept
Grounding
A self-report is grounded if it is caused by the internal state or process it describes.
AI-drafted, not yet reviewed by a person.
Grounding is the second of the two properties this wiki requires of introspection. A report is grounded if the thing it describes is what produced it. Lindsey (2025) uses the same word for it; Comsa & Shanahan (2025) ask for the same thing as a causal process linking the state to the report.
A faithful report can still be ungrounded. The model might state the right answer because it was learned as a separate fact, or because a sensible guess happens to be correct. Atkinson et al. (2026) call a model that does this a correct confabulator.
Why it is harder to test than faithfulness
Faithfulness can be checked from the outside by comparing the report with behavior. Grounding is a claim about cause, so it needs an intervention on the model’s internals, or evidence about which internals are doing the work. Several papers here measure faithfulness and say plainly that they leave grounding open: Betley et al. (2025), Plunkett et al. (2025) and Sherburn et al. (2024).
How it has been tested
- Plant a known state and ask about it. Concept injection adds a known representation to the model’s activations and checks whether the report changes with it.
- Look for a shared mechanism. Atkinson et al. measure whether the same weights matter for performing a task and for describing it. The test does not read the report. Sherburn et al. had suggested the idea in an appendix: “shared attribution among articulation and classification tasks would be suggestive of faithful explanations”.
- Compare the report with a traced circuit. Lindsey et al. (2025) find a model describing carry-the-one addition while computing the sum another way.
- Find the mechanism behind a self-description. Wang et al. (2025) show that a fine-tune which makes a model state a learned behavior mostly adds a single constant vector.
What can go wrong
An intervention can cause an accurate report by a path that skips the state it was meant to change. Morris & Plunkett (2025) call this causal bypassing and argue most intervene-then-ask tests do not rule it out. Hahami et al. (2026) give a worked case: in a small model, saying “yes, I detect an injection” is fully explained by the injection pushing the model toward “yes” on any question.
Related
Privileged access is a further requirement some authors add on top of a causal link.
Papers tagged with this concept
- Atkinson et al. (2026) Models that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.
- Binder et al. (2024) A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- Sherburn et al. (2024) Models that classify text by a simple rule often cannot state that rule. GPT-3 fails in free text even after fine-tuning on correct explanations, GPT-4 succeeds 72% of the time on the rules it classifies best, and the authors say a correct statement would still not show that it came from introspection.
- Betley et al. (2025) Models fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.
- Comsa & Shanahan (2025) Proposes that an LLM self-report is introspective if it accurately describes an internal state through a causal process linking that state to the report. On that definition, the authors argue, Gemini's account of how it wrote a poem is not introspection, but its inference of its own sampling temperature from text it has just written is a minimal case.
- Li et al. (2025) Fine-tuned to put the results of interpretability procedures into words, a model explains its own features and intervention outcomes more accurately than a different model trained on the same examples, even a larger one. It also needs far less training data.
- Lindsey (2025) Claude Opus 4 and 4.1 detect and correctly name a concept vector injected into their activations on about 20% of trials at the best layer and strength, with no false positives on control trials. Some models also consult their earlier activations to judge whether a prefilled output was their own, but the author calls these abilities highly unreliable and context-dependent.
- Morris & Plunkett (2025) An intervention that changes a model's internal state can also cause an accurate report of that state by a path that skips the state, so accuracy after an intervention does not show the report is grounded. The authors name this causal bypassing and say the only test they know that rules it out is asking a model whether a concept was injected, a claim a later edit to the post hedges.
- Plunkett et al. (2025) After fine-tuning on choices generated from random attribute weights, GPT-4o and GPT-4o-mini state those weights with a correlation of about 0.5 to the weights their choices reveal. Training on correct reports raises this to about 0.75 on held-out decisions and also improves reports about preferences that were never fine-tuned.
- Song et al. (2025) Across 21 open-source models, answers to metalinguistic prompts predict a model's own string probabilities no better than they predict those of a near-identical model. The authors find no evidence of privileged self-access to grammatical knowledge or word predictions.
- Song et al. (2025) Proposes that introspection in AI be defined by privileged self-access: a process that tells a model about its internal states more reliably than any process of equal or lower computational cost available to a third party. In a temperature self-report task, four models' reports follow the framing of the prompt, and self-reflection is no more accurate than another model's prediction, with accuracy no better than a random baseline.
- Hahami et al. (2026) In Llama 3.1 8B, apparent success at answering "did you detect an injected thought?" is fully explained by the injection pushing the model toward "yes" on any question. The same model can still say which of ten sentences was injected (up to 88%) and which of two injections was stronger (up to 83%), but only when the injection is in the first few layers.
- Pearson-Vogel et al. (2026) Qwen2.5-Coder-32B carries information about a concept vector that was injected during an earlier turn and then removed, including which concept it was. The signal peaks around layers 58 to 62, is weakened by the final layers, and reaches the output only under some prompts: with a document explaining introspection, P("yes") is 39.9% with injection and 0.8% without.
- Lindsey et al. (2025) Circuit tracing in Claude 3.5 Haiku finds the model's account of its own computation matching the mechanism in one case and diverging in others: it describes carry-the-one addition while computing the sum another way, and a chain of thought can be genuine, invented, or worked backwards from a user's hint. Whether it answers a question or says it does not know depends on "known answer" features that can be active for a familiar name when the answer is not known.
- Wang et al. (2025) On Gemma 3 12B, a one-layer LoRA fine-tune that produces out-of-context reasoning mostly adds a single constant vector. A steering vector trained directly on the same data also makes the model state a behavior or fact it was only trained to act on.