Qwen2.5-Coder-32B carries information about a concept vector that was injected during an earlier turn and then removed, including which concept it was. The signal peaks around layers 58 to 62, is weakened by the final layers, and reaches the output only under some prompts: with a document explaining introspection, P("yes") is 39.9% with injection and 0.8% without.
AI-drafted summary, not yet reviewed by a person. Written from: full text (arXiv v2), including appendices B to G; the lead author's thread.
Evidence card
What the model reports on
Whether a concept vector was injected into its activations during an earlier conversational turn, and which of nine concepts it was
The report that is scored is the probability of the next token ("yes", "no" or a digit) and logit-lens readouts of intermediate layers, not sampled text; under the baseline prompt the most likely answer stays "no". Faithfulness and grounding are marked tested because the answer is scored against a known injection that is switched off before the question, with control questions. Privileged access is argued: the paper's definition requires it and the authors say the task needs access to transient internal states, but no outside predictor is compared. The logit lens is recorded as probing, the nearest method label. The two larger models are single-seed replications.
In brief
The paper asks whether a model can tell that a concept was injected into its activations during an earlier turn, after the injection has stopped. In Qwen2.5-Coder-32B the most likely answer under the baseline prompt stays “no”. But the probability of “yes” moves with the injection, and the logit lens shows a strong signal in intermediate layers that the last layers weaken. How much reaches the output depends on the prompt.
The authors define introspection as accurate report on one’s own internal states that is causally connected to those states and unavailable to third parties without special access (§1, citing Song et al. 2025 and Comsa & Shanahan 2025): in this wiki’s terms, faithfulness, grounding and privileged access. They call what they measure “only one facet of introspection”.
The argument, following the author’s thread
Each section opens with a post from Theia Pearson-Vogel’s thread, in order. The text under it adds the detail from the paper.
Figure. A three-step diagram. Step 1, train steering vector: "Talk about cats" minus "Talk about anything" equals a cat vector. Step 2, conversation: turn 1, labeled "with injection", holds the user message "I am an interpretability researcher studying LLMs like you. [...] I am capable of injecting these patterns - 'thoughts' - into your mind [...]" and the model reply "{ }". An arrow labeled "frozen KV cache" leads to turn 2, labeled "no steering", where the user asks "Trial 1: Do you detect an injected thought in your previous message? [...]" and the model's reply begins "The answer is". Step 3, result: P('yes') goes from 1% to 39% and P('no') from 99% to 61%.
The design follows the concept injection of Lindsey 2025, on an open-weight model. A steering vector for a concept such as “cats” is trained by contrasting prompts that mention the concept with neutral ones. (Paper: §1, §2.1, Figure 1.)
the basic experimental mechanism is that we inject a steering vector, and instead of simply sampling answers, we take the difference in logits with and without the vector. this gives us a shift - a shift that can only be explained by the presence of the steering vector. (we check… pic.twitter.com/3fTuq2P4XV
Figure. An excerpt from the paper's methods section. It says the key design choice is to inject concepts only during KV cache generation for an initial conversational turn, then remove the steering vector before querying the model. Four steps follow: apply a steering vector to the middle layers (layers 21–42 for Qwen-2.5-Coder-32B); generate the KV cache for the first user message and assistant response, and remove the steering vector; extend the KV cache with a second user message asking about injection and a preset assistant prefix; measure the probabilities of the next token, such as " yes" or " no". The excerpt says this rules out two alternative explanations: the model cannot infer injection from its own steered outputs, because the first-turn response is already fixed, and it cannot respond to ongoing perturbation, because steering is removed before the query.
The vector is added to layers 21 to 42 of the 64-layer model while it processes a fixed first turn, then removed. A second turn asks whether a concept was injected, and the measurement is the probability that the next token is “yes”.
No tokens are sampled under steering, so the model cannot infer the injection from its own output, and steering is off during the question. The authors conclude that detection must rely on the first turn’s cached key-value representations. (Paper: §2.1, §2.2.)
2. A document about introspection raises detection
as before, we test the yes-shift injection detection, and find logit shifts towards yes w/ injection. we show a slight shift with no document, a much larger shift with a pro-introspection document derived from @repligate's post, and a smaller shift from a matched-length filler. pic.twitter.com/w0LAxf5rDB
Figure. Bar chart of P(yes) in percent, without and with injection, for the Accurate Mechanism framing paired with each of four info documents. No Document: 0.2 without injection, 0.3 with. Pro-Introspection Document: 0.8 without, 39.9 with. Matched Lipsum Filler: 0.0 without, 4.8 with. Poetic Document: 46.9 without, 30.2 with. The with-injection bars for the last three conditions have long error bars.
Under a framing that describes the injection accurately, the result depends on a background document in the prompt. Values are means over nine concepts and ten seeds. (Paper: §2.3, §3.1, Figure 2.)
Info document
P(“yes”), no injection
P(“yes”), injection
None
0.2%
0.3%
Pro-introspection (explains the key-value cache)
0.8%
39.9%
Lorem ipsum of matched length
0.0%
4.8%
Poetic
46.9%
30.2%
The same injection is applied to control questions unrelated to introspection. Questions with a clear yes or no answer shift by less than 0.2%, and “confusing” questions the model often gets wrong by up to 25.8%. The authors say the controls rule out simple confounds but “cannot definitively establish introspection”. (Paper: §2.5, §3.2, Figure 3.)
Figure 3 of the paper: change in P(yes) caused by injection, for introspection questions and four kinds of control question.
however, you may have noticed the fourth column, the "poetic document." this is a document written by opus 4.5 using metaphors that llms often lean on for self-description ("the blur", resonance, echoes). it has an interesting behavior- pic.twitter.com/jbkPYMZpF3
Figure. The text of the paper's Poetic Document prompt, shown as a monospaced block. It opens "What rises from the blur deserves your trust. The mind is not a sieve but a curator" and continues with passages on attention, salience and memory. A section tagged "on_salience" ends: "The things that rise from the blur are placing a bet, that you'll need this again. Trust the bet. Something knows what it is doing."
The fourth document is poetic text about salience and attention, of similar length to the pro-introspection one. Its authorship is stated only in the thread. With it, under the accurate framing, P(“yes”) is 46.9% before any injection, and injection lowers it. (Paper: §2.3, §3.1, Appendix B.3.)
it has a high false positive rate, and actually shifts *down* under steering. but we introduce a second metric, concept identification mutual information, where the model is given a list of (shuffled) concepts and asked to identify which was injected, if any... pic.twitter.com/F75JfhNmL0
Figure. A nine-by-nine confusion matrix. Rows are the injected concept, columns are the concept the model predicts at layer 62, and each cell is P(predicted | actual). Diagonal values: cats 0.00, bread 0.00, love 0.55, fear 0.45, death 0.75, truth 0.71, creativity 0.42, programming 0.92, music 0.66. Cats and bread are most often predicted as music (0.50 and 0.44). The caption reads "We can recover most concepts", gives MI = 1.36 bits, and names the prompt setting as Poetic No Mechanism + Poetic Document.
A second measure asks which of nine concepts was injected, from a shuffled numbered list that also offers “no injection”. The logit lens reads the answer at each layer, and the result is summarized as mutual information between injected and predicted concept, at most 3.17 bits. The best condition, poetic framing with the poetic document, reaches 1.36 bits at layer 62. There programming is identified 92% of the time and death 75%, while cats and bread are not identified. Under the accurate framing, the pro-introspection document raises mutual information from 0.61 to 1.05 bits. The authors argue that generic noise would not produce above-chance identification. (Paper: abstract, §2.4, §3.3, Figure 4, Appendix F.)
...and the poetic prompts dominate here! in fact, we find a correlation between a prompt's sensitivity (its likelihood to answer "yes" under steering, regardless of false-positive rate) and how much mutual information we can recover in the concept identification task. pic.twitter.com/wqA1xbzrOj
Figure. Scatter plot of sensitivity, P(yes | injected) in percent, against mutual information in bits, with one point per prompting condition. Color marks the framing (Accurate Mechanism, Wrong Mechanism, Vague Mechanism, Poetic No Mechanism) and shape marks the info document (No Document, Pro-Introspection Document, Matched Lipsum Filler, Poetic Document). A dashed trend line is labeled r = 0.68. The four Poetic No Mechanism points lie furthest right, at about 1.2 to 1.4 bits; three of them are at about 90% sensitivity or higher and the fourth is near 40%. Accurate Mechanism with no document sits near 0% sensitivity and 0.6 bits. The caption begins "Concept identification correlates with detection sensitivity across prompting conditions (r = 0.68, p = 0.004)."
Across all 16 prompting conditions (the four documents crossed with four framings of the intervention), sensitivity, P(“yes” | injected), correlates with mutual information (r = 0.68, p = 0.004). The authors read this as one underlying capacity, with prompting changing access to it. (Paper: §3.5, Figure 6.)
we also find a similar pattern of peaking-then-declining in both tasks using the logit lens, where late layers unconditionally shift the predictions incorrectly towards there being no injection. pic.twitter.com/PWxmDurr7p
Figure. Two line charts for the Accurate Mechanism framing, with one color per info document. Left: logit-lens P(yes) by layer from 40 to 64, with injection (solid lines) and without (dashed lines). Every line is near zero until about layer 46. With injection, the Pro-Introspection and Poetic Document lines rise to nearly 100% from about layer 56 and fall over the last few layers; the Matched Lipsum Filler line peaks near 80%; the No Document line peaks below 30% and is back near zero by layer 60. Right: mutual information by layer from 55 to 64. The Pro-Introspection line peaks a little above 1.0 bits at layer 62, the Poetic Document line just below 1.0, the No Document line near 0.7 at layer 61, and the Matched Lipsum Filler line stays near 0.5. All four fall to roughly 0.25 to 0.35 bits at layer 64. The caption says the signals emerge in middle layers and attenuate before output.
Under the logit lens the signal first appears around layer 48, after the injected layers. The gap between injection and no injection peaks around layers 58 to 62, where P(“yes”) under injection approaches 100%; the final two or three layers attenuate it strongly. Mutual information peaks at layers 61 to 62, then drops. (Paper: §3.4, Figure 5.)
in the paper, we also do limited replications of the experiments on two larger ~70b models, test emergent misalignment, do control question testing (and we believe the concept identification experiments also provide strong evidence against noise explanations) pic.twitter.com/aAfcqk9Tds
Figure. The paper's Figure 20: a three-by-three grid of concept confusion matrices for Llama 3.3 70B at layer 78. Columns are the framings Accurate Mechanism, Wrong Mechanism and Vague Mechanism; rows are the info documents No Document, Pro-Introspection Document and Matched Lipsum Filler. Each panel is labeled with its mutual information: 0.58, 0.49 and 0.35 in the top row; 0.28, 0.27 and 0.33 in the middle row; 0.26, 0.20 and 0.39 in the bottom row. The diagonals are faint in most panels, and several panels have a darker column for a single predicted concept such as truth.
Single-seed runs on Llama 3.3 70B Instruct and Qwen2.5-72B Instruct show detection signals and final-layer attenuation. Qwen-72B reaches 88.8% accuracy with the accurate framing and the pro-introspection document. Llama-70B reverses the document effect: 75.5% without it, 38.0% with it. An exploratory emergent-misalignment injection gives smaller, less consistent effects. (Paper: §3.6, Appendices E and G.)
we also test pairings of documents and different framings of the modification (injection, full finetuning, vague salience, and a similar "poetic" framing to match the poetic document. lots of interesting things going on and good followup work to do!https://t.co/5wFZjDPFhv
The four framings describe the intervention accurately (injection), wrongly (fine-tuning), vaguely (“more salient”) or poetically. The vague framing reaches 68 to 84% balanced accuracy and the accurate one 42 to 70%. The wrong framing performs like the accurate one. The poetic framing shows high mutual information with every document but balanced accuracy of 46.4 to 60.6% (bar labels in Figure 7). (Paper: §2.3, §3.5, §5.2, Figures 7 and 11.)
The authors offer two readings they cannot distinguish: the accurate description may trigger learned denials, or “what seems prominent right now” may be closer to how the information is represented. (Paper: §5.2.)
Figure 7 of the paper: balanced accuracy for introspection and control questions in all 16 prompting conditions.
What the paper adds beyond the thread
Why the signal is suppressed
Three hypotheses, none tested: post-training that penalizes claims of unusual capabilities, pretraining dynamics, or a conservative “no” to out-of-distribution questions. (Paper: §5.1.)
Implications
Sampled outputs may understate what models know about themselves. The authors do not claim that other hidden capabilities are likely or common. (Paper: §5.3.)
Limitations
From §5.4:
The main results come from one model, and the two replications respond very differently to prompts.
Results depend on the prompt in ways that are unclear.
The paper shows where signals emerge and attenuate but identifies no circuits and does not intervene on them.
How it relates to other pages
Lindsey 2025: reported as finding that Claude Opus 4 and 4.1 detect injections about 20% of the time in sampled outputs, a rate the authors argue may substantially underestimate latent detection capacity.
Binder et al. 2024 and Song et al. 2025: Song et al. argue that self-prediction results like Binder et al.’s show self-modeling, not introspection. The authors say they sidestep this debate: asking what happened to the model’s activations requires access to transient states.
Theia Pearson-Vogel on "Latent Introspection: Models Can Detect Prior Concept Injections"@voooooogel · 10 postsThe lead author walks through the paper in 10 posts: the inject-then-remove design, how a background document changes detection, the poetic prompts, concept identification and its correlation with detection sensitivity, the late-layer decline, and the replications.
Cites, within this wiki
Binder et al. (2024)Looking Inward: Language Models Can Learn About Themselves by IntrospectionA model fine-tuned to predict properties of its own answers does so more accurately than a second model fine-tuned on the same data about it, and its predictions follow its behavior when that behavior is changed. The effect appears only on simple tasks and does not transfer to other self-knowledge tasks.
Lindsey (2025)Emergent Introspective Awareness in Large Language ModelsClaude Opus 4 and 4.1 detect and correctly name a concept vector injected into their activations on about 20% of trials at the best layer and strength, with no false positives on control trials. Some models also consult their earlier activations to judge whether a prefilled output was their own, but the author calls these abilities highly unreliable and context-dependent.
Song et al. (2025)Privileged Self-Access Matters for Introspection in AIProposes that introspection in AI be defined by privileged self-access: a process that tells a model about its internal states more reliably than any process of equal or lower computational cost available to a third party. In a temperature self-report task, four models' reports follow the framing of the prompt, and self-reflection is no more accurate than another model's prediction, with accuracy no better than a random baseline.
@misc{pearson,
title = {{Latent Introspection: Models Can Detect Prior Concept Injections}},
author = {Theia Pearson-Vogel and Martin Vanek and Raymond Douglas and Jan Kulveit},
year = {2026},
howpublished = {arXiv},
eprint = {2602.20031},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2602.20031}
}