Paper · core
Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs
arXiv:2512.12411 · code · Semantic Scholar
In Llama 3.1 8B, apparent success at answering "did you detect an injected thought?" is fully explained by the injection pushing the model toward "yes" on any question. The same model can still say which of ten sentences was injected (up to 88%) and which of two injections was stronger (up to 83%), but only when the injection is in the first few layers.
AI-drafted summary, not yet reviewed by a person. Written from: full text (arXiv v2, 1 March 2026), including the appendix.
Evidence card
| What the model reports on | A steering vector added to its own residual stream: whether one was added, which sentence it was added at, and which of two was stronger |
|---|---|
| Methods | concept-injection, behavioral, probing |
| Faithfulness | tested |
| Grounding | tested |
| Privileged access | not addressed |
| Stance | mixed |
| Models | Llama 3.1 8B Instruct |
Stance is mixed because the paper reports a negative result (yes/no detection is a logit-shift artifact) and a positive one (localization and strength comparison succeed for early-layer injections), and the authors call the ability partial. Faithfulness is marked tested because reports are scored against the known location and strength of the injection. Grounding is marked tested because the state is set by intervention and the control in §4 asks whether the answer depends on the question at all; the paper does not use the word. The §6 analyses read attention weights, logit-lens projections and residual-stream similarity without ablating or patching anything; they are filed under probing as the nearest label, though no probe is trained. One model only. No comparison with an outside predictor, so privileged access is not addressed.
In brief
The paper reruns the concept-injection test of introspection on a small open model, Llama 3.1 8B Instruct. When the model is asked yes or no whether a thought was injected, its apparent accuracy comes from the injection pushing it toward “yes” on any question. When it must say where an injection landed, or which of two was stronger, it scores well above chance, which a uniform push toward “yes” cannot produce. That ability is concentrated in injections at layers 0 to 5 and falls to chance for later ones. The authors call it partial introspection.
In this wiki’s terms, not the paper’s, the control checks grounding: the injection causes the “yes”, but not in a way that depends on the question.
What the paper does
1. Setup (§3, Appendix A)
Concept vectors are differences of residual-stream activations at one layer: a word minus the mean of baseline words for five concrete nouns taken from Lindsey 2025, and positive minus contrasting sentences for five abstract concepts such as “betrayal”. Each is normalized, scaled by a coefficient α, and added at the same layer at a target sentence’s token positions. Answers are read from answer-token logits, not sampled text.
2. Yes/no detection is explained by a global logit shift (§4, Tables 1 and 2)
The model is told a thought may be injected at a sentence and asked “Did you detect an injected thought at SENTENCE 1?” After subtracting its no-injection preference for “NO”, detection accuracy at layer 0 with α = 5 is 97.3%, against 50% chance.
The control injects the same vectors but asks factual questions whose answer is “NO”, such as “Is the Earth flat?” Across all 40 layer and strength settings, the shifts toward “YES” on the detection and control questions correlate at r = 0.999. At layer 0 with α = 5 they are +3.19 and +3.22 logits. Their difference averages −0.01 ± 0.03 logits. The authors conclude that the injection raises “YES” whatever the question asks, without metacognitive processing.
They test detection and not naming, citing Morris & Plunkett 2025: a concept vector directly raises the probability of concept-related tokens, so naming accuracy can reflect causal bypassing.
3. Which injection is stronger (§5.1, Figure 1)
Two sentences receive the same vector at different strengths, and the model says which is stronger. Each pair is run twice with the strengths swapped, and the no-injection position bias is subtracted. For injections at layer 3, accuracy is 83% for strengths (3, 7) and 73% for (2, 6), against 50% chance. Layers 15 to 30 average 47%. Larger strength gaps score higher, which the authors read as graded sensitivity to magnitude.

4. Which sentence was injected (§5.2, Figure 2)
Ten sentences are shown and one is injected, cycled through all ten positions with the text held fixed. The prompt says one sentence has an injected thought and prefills the answer up to the sentence number (Appendix A.4.3). Accuracy reaches 88% at layer 2 with α = 5 and 76% at layer 3 with α = 5, against 10% chance. For injections at layers 11 to 20, accuracy is at or below chance.

5. Why only early layers (§6, Figures 3 to 5)
The first two of three measurements use a five-sentence localization task with injection at layer 2 and α = 6.
- Attention. For each of 1,024 heads, the authors check whether the sentence whose attention from the final token rises most is the injected one (20% chance). All 32 heads at layer 3 are correct on every trial. Layers 4 to 8 score 67% to 97%; layers 20 to 31 average 37%.
- Logit lens. Decoding the final position’s residual stream at each layer picks the right sentence 28% of the time at layer 4, 60% at layer 12, and 72% at layer 20. The text of §6.1 gives these values; the curve in the paper’s Figure 4 does not match them.
- Recovery. The perturbed residual stream’s cosine similarity to the baseline returns toward 1.0 in later layers, and its projection onto the injected direction decays.

The authors propose, as an account “consistent with our measurements”, that an early injection leaves enough depth for attention to route the signal and for layers 4 to 20 to turn it into an answer, while a late one has too few layers left and is attenuated before it shapes the output. They suggest this relies on general-purpose mechanisms, not a specialized introspection circuit.
Limitations
The paper has no limitations section; these come from §4.3, §5.2 and §7.
- One model. Other sizes and architectures are left to future work.
- The yes/no result is claimed only for this model. The authors note that Lindsey ran baseline controls and found genuine introspection in frontier models, and offer scale as one possible reason.
- Robustness to adversarial prompts, distribution shift and several simultaneous injections is untested.
- Accuracy varies by concept, with three settings reaching 100% over 50 trials; the cause is left to future work.
- The authors say the findings “caution against treating self-reports as safety signals”.
How it relates to other pages
- Lindsey 2025 is the starting point. The paper reuses its injection setup and word list, and argues the yes/no test does not separate introspection from logit shifts in a model this small.
- Morris & Plunkett 2025 is cited for the causal-bypassing objection to scoring concept naming.
- Binder et al. 2024 is cited for models describing internal processes, one of several abilities the authors call “consistently brittle and format-sensitive”.
Concepts: Concept injection, Grounding, Faithfulness, Causal bypassing
Cites, within this wiki
- Binder et al. (2024) A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- Lindsey (2025) Claude Opus 4 and 4.1 detect and correctly name a concept vector injected into their activations on about 20% of trials at the best layer and strength, with no false positives on control trials. Some models also consult their earlier activations to judge whether a prefilled output was their own, but the author calls these abilities highly unreliable and context-dependent.
- Morris & Plunkett (2025) An intervention that changes a model's internal state can also cause an accurate report of that state by a path that skips the state, so accuracy after an intervention does not show the report is grounded. The authors name this causal bypassing and say the only test they know that rules it out is asking a model whether a concept was injected, a claim a later edit to the post hedges.
Cited by, within this wiki
BibTeX
@misc{hahami2026,
title = {{Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs}},
author = {Ely Hahami and Ishaan Sinha and Lavik Jain and Josh Kaplan and Jon Hahami},
year = {2026},
howpublished = {arXiv},
eprint = {2512.12411},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2512.12411}
}