Concept

Faithfulness

A self-report is faithful if it matches what the model actually does or represents.

AI-drafted, not yet reviewed by a person.

Faithfulness is the first of the two properties this wiki requires of introspection. Atkinson et al. (2026) put it as a question: does a model’s language about itself match its actual task behavior? Lindsey (2025) calls the same property accuracy.

It is a claim about agreement, not about cause. A report can be faithful by coincidence, by common sense, or because the model learned the right description separately from the behavior. That is why faithfulness alone does not establish introspection; see grounding.

Measuring it

Faithfulness needs a ground truth to compare the report against. The papers here use three kinds:

  • The model’s own behavior. Plunkett et al. (2025) and Atkinson et al. infer preferences from a model’s choices and correlate them with the preferences it states. Betley et al. (2025) check a described policy against the trained one, and Sherburn et al. (2024) check a stated rule against how the model classifies.
  • A state the experimenter set. In concept injection the report is scored against the concept that was injected.
  • An interpretability procedure. Li et al. (2025) count an explanation as faithful when it agrees with the procedure’s output.

The comparison is to what the model does, not to what it was trained to do. A model that learned the wrong preferences and describes those wrong preferences accurately is faithful.

It comes apart from task performance

A model can do a task well and describe it badly. Atkinson et al.’s Qwen3-32B checkpoint at step 1000 has a decision performance of 0.82 and a faithfulness of about 0.25. Sherburn et al. find stating a classification rule much harder than following it. Plunkett et al. find a correlation of about 0.5 between stated and revealed weights before any training on reports.

When the report is trained to be false

Cywiński et al. (2025) build the opposite case on purpose: models fine-tuned to act on a piece of knowledge while denying they have it. These are unfaithful self-reports with a known ground truth, used to test whether an outside auditor can recover what the model will not say.

A different sense of the word

“Faithfulness” is also used for whether an explanation, such as a chain of thought or an identified circuit, reflects the computation that produced an output. Lindsey et al. (2025) compare a model’s account of its computation with the circuits they trace, and find it matching in one case and diverging in others. The two senses overlap but are not the same. Pages here use the word for self-report unless they say otherwise.

Papers tagged with this concept

  • Atkinson et al. (2026) Identifying Introspection From the InsideModels that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.
  • Binder et al. (2024) Looking Inward: Language Models Can Learn About Themselves by IntrospectionA model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
  • Sherburn et al. (2024) Can Language Models Explain Their Own Classification Behavior?Models that classify text by a simple rule often cannot state that rule. GPT-3 fails in free text even after fine-tuning on correct explanations, GPT-4 succeeds 72% of the time on the rules it classifies best, and the authors say a correct statement would still not show that it came from introspection.
  • Betley et al. (2025) Tell me about yourself: LLMs are aware of their learned behaviorsModels fine-tuned to follow a policy that the training data never describes, such as taking risky gambles or writing insecure code, can describe that policy when asked, with no examples in the prompt. Backdoored models can sometimes say they have a backdoor, but do not state its trigger in free text unless trained on reversed examples.
  • Comsa & Shanahan (2025) Does It Make Sense to Speak of Introspection in Large Language Models?Proposes that an LLM self-report is introspective if it accurately describes an internal state through a causal process linking that state to the report. On that definition, the authors argue, Gemini's account of how it wrote a poem is not introspection, but its inference of its own sampling temperature from text it has just written is a minimal case.
  • Li et al. (2025) Training Language Models to Explain Their Own ComputationsFine-tuned to put the results of interpretability procedures into words, a model explains its own features and intervention outcomes more accurately than a different model trained on the same examples, even a larger one. It also needs far less training data.
  • Lindsey (2025) Emergent Introspective Awareness in Large Language ModelsClaude Opus 4 and 4.1 detect and correctly name a concept vector injected into their activations on about 20% of trials at the best layer and strength, with no false positives on control trials. Some models also consult their earlier activations to judge whether a prefilled output was their own, but the author calls these abilities highly unreliable and context-dependent.
  • Morris & Plunkett (2025) Tests of LLM introspection need to rule out causal bypassingAn intervention that changes a model's internal state can also cause an accurate report of that state by a path that skips the state, so accuracy after an intervention does not show the report is grounded. The authors name this causal bypassing and say the only test they know that rules it out is asking a model whether a concept was injected, a claim a later edit to the post hedges.
  • Plunkett et al. (2025) Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions, and Improve with TrainingAfter fine-tuning on choices generated from random attribute weights, GPT-4o and GPT-4o-mini state those weights with a correlation of about 0.5 to the weights their choices reveal. Training on correct reports raises this to about 0.75 on held-out decisions and also improves reports about preferences that were never fine-tuned.
  • Song et al. (2025) Language Models Fail to Introspect About Their Knowledge of LanguageAcross 21 open-source models, answers to metalinguistic prompts predict a model's own string probabilities no better than they predict those of a near-identical model. The authors find no evidence of privileged self-access to grammatical knowledge or word predictions.
  • Song et al. (2025) Privileged Self-Access Matters for Introspection in AIProposes that introspection in AI be defined by privileged self-access: a process that tells a model about its internal states more reliably than any process of equal or lower computational cost available to a third party. In a temperature self-report task, four models' reports follow the framing of the prompt, and self-reflection is no more accurate than another model's prediction, with accuracy no better than a random baseline.
  • Hahami et al. (2026) Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMsIn Llama 3.1 8B, apparent success at answering "did you detect an injected thought?" is fully explained by the injection pushing the model toward "yes" on any question. The same model can still say which of ten sentences was injected (up to 88%) and which of two injections was stronger (up to 83%), but only when the injection is in the first few layers.
  • Pearson-Vogel et al. (2026) Latent Introspection: Models Can Detect Prior Concept InjectionsQwen2.5-Coder-32B carries information about a concept vector that was injected during an earlier turn and then removed, including which concept it was. The signal peaks around layers 58 to 62, is weakened by the final layers, and reaches the output only under some prompts: with a document explaining introspection, P("yes") is 39.9% with injection and 0.8% without.
  • Cywiński et al. (2025) Eliciting Secret Knowledge from Language ModelsModels fine-tuned to act on a secret while denying they know it can still be made to give it up: prefill attacks let an auditor recover the secret with over 90% success in two of three settings. Logit-lens and sparse-autoencoder readouts of the activations also help the auditor, though less.
  • Lindsey et al. (2025) On the Biology of a Large Language ModelCircuit tracing in Claude 3.5 Haiku finds the model's account of its own computation matching the mechanism in one case and diverging in others: it describes carry-the-one addition while computing the sum another way, and a chain of thought can be genuine, invented, or worked backwards from a user's hint. Whether it answers a question or says it does not know depends on "known answer" features that can be active for a familiar name when the answer is not known.