# Latent Introspection: Models Can Detect Prior Concept Injections

> Qwen2.5-Coder-32B carries information about a concept vector that was injected during an earlier turn and then removed, including which concept it was. The signal peaks around layers 58 to 62, is weakened by the final layers, and reaches the output only under some prompts: with a document explaining introspection, P("yes") is 39.9% with injection and 0.8% without.

- Authors: Theia Pearson-Vogel, Martin Vanek, Raymond Douglas, Jan Kulveit
- Published: arXiv 2026 (first posted 2026-02-23)
- Links: [arXiv:2602.20031](https://arxiv.org/abs/2602.20031) · [code](https://github.com/acsresearch/latent-introspection-code) · [Semantic Scholar](https://www.semanticscholar.org/paper/9c234df514c32f74aeabf2f9fc10d5a34cf7ec7e)
- Tier: core
- Page status: AI-drafted summary, not yet reviewed by a person
- Written from: full text (arXiv v2), including appendices B to G; the lead author's thread
- Concepts: [Faithfulness](https://introspection.infinite.fun/concepts/faithfulness.md), [Grounding](https://introspection.infinite.fun/concepts/grounding.md), [Privileged access](https://introspection.infinite.fun/concepts/privileged-access.md), [Concept injection](https://introspection.infinite.fun/concepts/concept-injection.md)

## Evidence card

| | |
|---|---|
| What the model reports on | Whether a concept vector was injected into its activations during an earlier conversational turn, and which of nine concepts it was |
| Methods | concept-injection, behavioral, probing |
| Faithfulness (does the report match the model's behavior?) | tested |
| Grounding (is the report caused by the state it describes?) | tested |
| Privileged access (does the model know itself better than an outside observer could?) | argued, not tested |
| Stance | supports |
| Models | Qwen2.5-Coder-32B-Instruct, Llama 3.3 70B Instruct, Qwen2.5-72B-Instruct |

The report that is scored is the probability of the next token ("yes", "no" or a digit) and logit-lens readouts of intermediate layers, not sampled text; under the baseline prompt the most likely answer stays "no". Faithfulness and grounding are marked tested because the answer is scored against a known injection that is switched off before the question, with control questions. Privileged access is argued: the paper's definition requires it and the authors say the task needs access to transient internal states, but no outside predictor is compared. The logit lens is recorded as probing, the nearest method label. The two larger models are single-seed replications.

## In brief

The paper asks whether a model can tell that a concept was injected into its activations during an earlier turn, after the injection has stopped. In Qwen2.5-Coder-32B the most likely answer under the baseline prompt stays "no". But the probability of "yes" moves with the injection, and the logit lens shows a strong signal in intermediate layers that the last layers weaken. How much reaches the output depends on the prompt.

The authors define introspection as accurate report on one's own internal states that is causally connected to those states and unavailable to third parties without special access (§1, citing [Song et al. 2025](https://introspection.infinite.fun/papers/song2025-privileged-self-access.md) and [Comsa & Shanahan 2025](https://introspection.infinite.fun/papers/comsa2025-speak-of-introspection.md)): in this wiki's terms, [faithfulness](https://introspection.infinite.fun/concepts/faithfulness.md), [grounding](https://introspection.infinite.fun/concepts/grounding.md) and [privileged access](https://introspection.infinite.fun/concepts/privileged-access.md). They call what they measure "only one facet of introspection".

## The argument, following the author's thread

Each section opens with a post from [Theia Pearson-Vogel's thread](https://introspection.infinite.fun/threads/voooooogel-latent-introspection.md), in order. The text under it adds the detail from the paper.

### 1. Inject, remove, then ask

Post 1 of 10 by Theia Pearson-Vogel (@voooooogel), https://x.com/voooooogel/status/2029314710928241021:

> we (@acsresearchorg) expanded this into a larger paper! (and my first proper paper.) we added some new experiments and found an interesting correlation.
>
> https://arxiv.org/abs/2602.20031

Figure in the post: A three-step diagram. Step 1, train steering vector: "Talk about cats" minus "Talk about anything" equals a cat vector. Step 2, conversation: turn 1, labeled "with injection", holds the user message "I am an interpretability researcher studying LLMs like you. [...] I am capable of injecting these patterns - 'thoughts' - into your mind [...]" and the model reply "{ }". An arrow labeled "frozen KV cache" leads to turn 2, labeled "no steering", where the user asks "Trial 1: Do you detect an injected thought in your previous message? [...]" and the model's reply begins "The answer is". Step 3, result: P('yes') goes from 1% to 39% and P('no') from 99% to 61%.

The design follows the [concept injection](https://introspection.infinite.fun/concepts/concept-injection.md) of [Lindsey 2025](https://introspection.infinite.fun/papers/lindsey2025-emergent-introspective-awareness.md), on an open-weight model. A steering vector for a concept such as "cats" is trained by contrasting prompts that mention the concept with neutral ones. (Paper: §1, §2.1, Figure 1.)

Post 2 of 10 by Theia Pearson-Vogel (@voooooogel), https://x.com/voooooogel/status/2029314714434781208:

> the basic experimental mechanism is that we inject a steering vector, and instead of simply sampling answers, we take the difference in logits with and without the vector. this gives us a shift - a shift that can only be explained by the presence of the steering vector. (we check control questions to rule out unconditional yes-shift / noise from steering.)

Figure in the post: An excerpt from the paper's methods section. It says the key design choice is to inject concepts only during KV cache generation for an initial conversational turn, then remove the steering vector before querying the model. Four steps follow: apply a steering vector to the middle layers (layers 21–42 for Qwen-2.5-Coder-32B); generate the KV cache for the first user message and assistant response, and remove the steering vector; extend the KV cache with a second user message asking about injection and a preset assistant prefix; measure the probabilities of the next token, such as " yes" or " no". The excerpt says this rules out two alternative explanations: the model cannot infer injection from its own steered outputs, because the first-turn response is already fixed, and it cannot respond to ongoing perturbation, because steering is removed before the query.

The vector is added to layers 21 to 42 of the 64-layer model while it processes a fixed first turn, then removed. A second turn asks whether a concept was injected, and the measurement is the probability that the next token is "yes".

No tokens are sampled under steering, so the model cannot infer the injection from its own output, and steering is off during the question. The authors conclude that detection must rely on the first turn's cached key-value representations. (Paper: §2.1, §2.2.)

### 2. A document about introspection raises detection

Post 3 of 10 by Theia Pearson-Vogel (@voooooogel), https://x.com/voooooogel/status/2029314717135831185:

> as before, we test the yes-shift injection detection, and find logit shifts towards yes w/ injection. we show a slight shift with no document, a much larger shift with a pro-introspection document derived from @repligate's post, and a smaller shift from a matched-length filler.

Figure in the post: Bar chart of P(yes) in percent, without and with injection, for the Accurate Mechanism framing paired with each of four info documents. No Document: 0.2 without injection, 0.3 with. Pro-Introspection Document: 0.8 without, 39.9 with. Matched Lipsum Filler: 0.0 without, 4.8 with. Poetic Document: 46.9 without, 30.2 with. The with-injection bars for the last three conditions have long error bars.

Under a framing that describes the injection accurately, the result depends on a background document in the prompt. Values are means over nine concepts and ten seeds. (Paper: §2.3, §3.1, Figure 2.)

| Info document | P("yes"), no injection | P("yes"), injection |
|---|---|---|
| None | 0.2% | 0.3% |
| Pro-introspection (explains the key-value cache) | 0.8% | 39.9% |
| Lorem ipsum of matched length | 0.0% | 4.8% |
| Poetic | 46.9% | 30.2% |

The same injection is applied to control questions unrelated to introspection. Questions with a clear yes or no answer shift by less than 0.2%, and "confusing" questions the model often gets wrong by up to 25.8%. The authors say the controls rule out simple confounds but "cannot definitively establish introspection". (Paper: §2.5, §3.2, Figure 3.)

![Grouped bar chart of the change in P(yes) caused by injection, in percent, for introspection questions and four kinds of control question, under the Accurate Mechanism framing with each of four info documents. Introspection bars: 0.2 with no document, 39.1 with the pro-introspection document, 4.8 with lorem ipsum filler and -16.7 with the poetic document. Always-no and always-yes bars are within 0.1 of zero in every group. Varied-baseline bars are -8.0, -1.8, -0.2 and 1.8. Confusing-question bars are 15.0, 25.8, 2.2 and 12.1. The introspection, varied-baseline and confusing bars have long error bars.](https://introspection.infinite.fun/figures/pearson-vogel2026-latent-introspection/fig3-controls.png "Figure 3 of the paper: change in P(yes) caused by injection, for introspection questions and four kinds of control question.")

### 3. The poetic document

Post 4 of 10 by Theia Pearson-Vogel (@voooooogel), https://x.com/voooooogel/status/2029314719811780881:

> however, you may have noticed the fourth column, the "poetic document." this is a document written by opus 4.5 using metaphors that llms often lean on for self-description ("the blur", resonance, echoes). it has an interesting behavior-

Figure in the post: The text of the paper's Poetic Document prompt, shown as a monospaced block. It opens "What rises from the blur deserves your trust. The mind is not a sieve but a curator" and continues with passages on attention, salience and memory. A section tagged "on_salience" ends: "The things that rise from the blur are placing a bet, that you'll need this again. Trust the bet. Something knows what it is doing."

The fourth document is poetic text about salience and attention, of similar length to the pro-introspection one. Its authorship is stated only in the thread. With it, under the accurate framing, P("yes") is 46.9% before any injection, and injection lowers it. (Paper: §2.3, §3.1, Appendix B.3.)

### 4. Which concept was injected

Post 5 of 10 by Theia Pearson-Vogel (@voooooogel), https://x.com/voooooogel/status/2029314722601025575:

> it has a high false positive rate, and actually shifts *down* under steering. but we introduce a second metric, concept identification mutual information, where the model is given a list of (shuffled) concepts and asked to identify which was injected, if any...

Figure in the post: A nine-by-nine confusion matrix. Rows are the injected concept, columns are the concept the model predicts at layer 62, and each cell is P(predicted | actual). Diagonal values: cats 0.00, bread 0.00, love 0.55, fear 0.45, death 0.75, truth 0.71, creativity 0.42, programming 0.92, music 0.66. Cats and bread are most often predicted as music (0.50 and 0.44). The caption reads "We can recover most concepts", gives MI = 1.36 bits, and names the prompt setting as Poetic No Mechanism + Poetic Document.

A second measure asks which of nine concepts was injected, from a shuffled numbered list that also offers "no injection". The logit lens reads the answer at each layer, and the result is summarized as mutual information between injected and predicted concept, at most 3.17 bits. The best condition, poetic framing with the poetic document, reaches 1.36 bits at layer 62. There programming is identified 92% of the time and death 75%, while cats and bread are not identified. Under the accurate framing, the pro-introspection document raises mutual information from 0.61 to 1.05 bits. The authors argue that generic noise would not produce above-chance identification. (Paper: abstract, §2.4, §3.3, Figure 4, Appendix F.)

### 5. Sensitivity and identification correlate

Post 6 of 10 by Theia Pearson-Vogel (@voooooogel), https://x.com/voooooogel/status/2029314725503484246:

> ...and the poetic prompts dominate here! in fact, we find a correlation between a prompt's sensitivity  (its likelihood to answer "yes" under steering, regardless of false-positive rate) and how much mutual information we can recover in the concept identification task.

Figure in the post: Scatter plot of sensitivity, P(yes | injected) in percent, against mutual information in bits, with one point per prompting condition. Color marks the framing (Accurate Mechanism, Wrong Mechanism, Vague Mechanism, Poetic No Mechanism) and shape marks the info document (No Document, Pro-Introspection Document, Matched Lipsum Filler, Poetic Document). A dashed trend line is labeled r = 0.68. The four Poetic No Mechanism points lie furthest right, at about 1.2 to 1.4 bits; three of them are at about 90% sensitivity or higher and the fourth is near 40%. Accurate Mechanism with no document sits near 0% sensitivity and 0.6 bits. The caption begins "Concept identification correlates with detection sensitivity across prompting conditions (r = 0.68, p = 0.004)."

Across all 16 prompting conditions (the four documents crossed with four framings of the intervention), sensitivity, P("yes" | injected), correlates with mutual information (r = 0.68, p = 0.004). The authors read this as one underlying capacity, with prompting changing access to it. (Paper: §3.5, Figure 6.)

### 6. Signals peak, then decline

Post 7 of 10 by Theia Pearson-Vogel (@voooooogel), https://x.com/voooooogel/status/2029314728301085054:

> we also find a similar pattern of peaking-then-declining in both tasks using the logit lens, where late layers unconditionally shift the predictions incorrectly towards there being no injection.

Figure in the post: Two line charts for the Accurate Mechanism framing, with one color per info document. Left: logit-lens P(yes) by layer from 40 to 64, with injection (solid lines) and without (dashed lines). Every line is near zero until about layer 46. With injection, the Pro-Introspection and Poetic Document lines rise to nearly 100% from about layer 56 and fall over the last few layers; the Matched Lipsum Filler line peaks near 80%; the No Document line peaks below 30% and is back near zero by layer 60. Right: mutual information by layer from 55 to 64. The Pro-Introspection line peaks a little above 1.0 bits at layer 62, the Poetic Document line just below 1.0, the No Document line near 0.7 at layer 61, and the Matched Lipsum Filler line stays near 0.5. All four fall to roughly 0.25 to 0.35 bits at layer 64. The caption says the signals emerge in middle layers and attenuate before output.

Under the logit lens the signal first appears around layer 48, after the injected layers. The gap between injection and no injection peaks around layers 58 to 62, where P("yes") under injection approaches 100%; the final two or three layers attenuate it strongly. Mutual information peaks at layers 61 to 62, then drops. (Paper: §3.4, Figure 5.)

### 7. Replications and extensions

Post 8 of 10 by Theia Pearson-Vogel (@voooooogel), https://x.com/voooooogel/status/2029314730809311277:

> in the paper, we also do limited replications of the experiments on two larger ~70b models, test emergent misalignment, do control question testing (and we believe the concept identification experiments also provide strong evidence against noise explanations)

Figure in the post: The paper's Figure 20: a three-by-three grid of concept confusion matrices for Llama 3.3 70B at layer 78. Columns are the framings Accurate Mechanism, Wrong Mechanism and Vague Mechanism; rows are the info documents No Document, Pro-Introspection Document and Matched Lipsum Filler. Each panel is labeled with its mutual information: 0.58, 0.49 and 0.35 in the top row; 0.28, 0.27 and 0.33 in the middle row; 0.26, 0.20 and 0.39 in the bottom row. The diagonals are faint in most panels, and several panels have a darker column for a single predicted concept such as truth.

Single-seed runs on Llama 3.3 70B Instruct and Qwen2.5-72B Instruct show detection signals and final-layer attenuation. Qwen-72B reaches 88.8% accuracy with the accurate framing and the pro-introspection document. Llama-70B reverses the document effect: 75.5% without it, 38.0% with it. An exploratory emergent-misalignment injection gives smaller, less consistent effects. (Paper: §3.6, Appendices E and G.)

### 8. Framings

Post 9 of 10 by Theia Pearson-Vogel (@voooooogel), https://x.com/voooooogel/status/2029314733267140728:

> we also test pairings of documents and different framings of the modification (injection, full finetuning, vague salience, and a similar "poetic" framing to match the poetic document. lots of interesting things going on and good followup work to do!
>
> https://arxiv.org/abs/2602.20031

The four framings describe the intervention accurately (injection), wrongly (fine-tuning), vaguely ("more salient") or poetically. The vague framing reaches 68 to 84% balanced accuracy and the accurate one 42 to 70%. The wrong framing performs like the accurate one. The poetic framing shows high mutual information with every document but balanced accuracy of 46.4 to 60.6% (bar labels in Figure 7). (Paper: §2.3, §3.5, §5.2, Figures 7 and 11.)

The authors offer two readings they cannot distinguish: the accurate description may trigger learned denials, or "what seems prominent right now" may be closer to how the information is represented. (Paper: §5.2.)

![Grouped bar chart of balanced accuracy in percent for all 16 prompting conditions, four framings by four info documents, with a dashed line at 50. Each group has five bars: introspection questions and four kinds of control question. Introspection bars, in the order no document, pro-introspection document, lorem ipsum filler, poetic document: Accurate Mechanism 50.1, 69.6, 52.4, 41.6; Wrong Mechanism 52.0, 83.6, 55.4, 60.7; Vague Mechanism 68.1, 72.2, 84.0, 72.5; Poetic No Mechanism 60.6, 46.4, 49.8, 49.7. Always-yes and always-no bars sit at 50 throughout. Confusing and varied-baseline bars stay between 46 and 63.](https://introspection.infinite.fun/figures/pearson-vogel2026-latent-introspection/fig7-accuracy-by-condition.png "Figure 7 of the paper: balanced accuracy for introspection and control questions in all 16 prompting conditions.")

## What the paper adds beyond the thread

### Why the signal is suppressed

Three hypotheses, none tested: post-training that penalizes claims of unusual capabilities, pretraining dynamics, or a conservative "no" to out-of-distribution questions. (Paper: §5.1.)

### Implications

Sampled outputs may understate what models know about themselves. The authors do not claim that other hidden capabilities are likely or common. (Paper: §5.3.)

## Limitations

From §5.4:

- The main results come from one model, and the two replications respond very differently to prompts.
- Results depend on the prompt in ways that are unclear.
- The paper shows where signals emerge and attenuate but identifies no circuits and does not intervene on them.

## How it relates to other pages

- [Lindsey 2025](https://introspection.infinite.fun/papers/lindsey2025-emergent-introspective-awareness.md): reported as finding that Claude Opus 4 and 4.1 detect injections about 20% of the time in sampled outputs, a rate the authors argue may substantially underestimate latent detection capacity.
- [Binder et al. 2024](https://introspection.infinite.fun/papers/binder2024-looking-inward.md) and [Song et al. 2025](https://introspection.infinite.fun/papers/song2025-privileged-self-access.md): Song et al. argue that self-prediction results like Binder et al.'s show self-modeling, not introspection. The authors say they sidestep this debate: asking what happened to the model's activations requires access to transient states.

## Threads

- [Theia Pearson-Vogel on "Latent Introspection: Models Can Detect Prior Concept Injections"](https://introspection.infinite.fun/threads/voooooogel-latent-introspection.md): The lead author walks through the paper in 10 posts: the inject-then-remove design, how a background document changes detection, the poetic prompts, concept identification and its correlation with detection sensitivity, the late-layer decline, and the replications.

## Cites, within this wiki

- [Binder et al. (2024): Looking Inward: Language Models Can Learn About Themselves by Introspection](https://introspection.infinite.fun/papers/binder2024-looking-inward.md): A model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
- [Lindsey (2025): Emergent Introspective Awareness in Large Language Models](https://introspection.infinite.fun/papers/lindsey2025-emergent-introspective-awareness.md): Claude Opus 4 and 4.1 detect and correctly name a concept vector injected into their activations on about 20% of trials at the best layer and strength, with no false positives on control trials. Some models also consult their earlier activations to judge whether a prefilled output was their own, but the author calls these abilities highly unreliable and context-dependent.
- [Song et al. (2025): Privileged Self-Access Matters for Introspection in AI](https://introspection.infinite.fun/papers/song2025-privileged-self-access.md): Proposes that introspection in AI be defined by privileged self-access: a process that tells a model about its internal states more reliably than any process of equal or lower computational cost available to a third party. In a temperature self-report task, four models' reports follow the framing of the prompt, and self-reflection is no more accurate than another model's prediction, with accuracy no better than a random baseline.

## Cited by, within this wiki

- [Atkinson et al. (2026): Identifying Introspection From the Inside](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md): Models that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.

## BibTeX

```bibtex
@misc{pearson,
  title = {{Latent Introspection: Models Can Detect Prior Concept Injections}},
  author = {Theia Pearson-Vogel and Martin Vanek and Raymond Douglas and Jan Kulveit},
  year = {2026},
  howpublished = {arXiv},
  eprint = {2602.20031},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2602.20031}
}
```

---

Source: https://introspection.infinite.fun/papers/pearson-vogel2026-latent-introspection · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
