Thread

Theia Pearson-Vogel on "Latent Introspection: Models Can Detect Prior Concept Injections"

The lead author walks through the paper in 10 posts: the inject-then-remove design, how a background document changes detection, the poetic prompts, concept identification and its correlation with detection sensitivity, the late-layer decline, and the replications.

The posts are embedded from X. The figure notes under them are written by this wiki.

  1. Figure. A three-step diagram. Step 1, train steering vector: "Talk about cats" minus "Talk about anything" equals a cat vector. Step 2, conversation: turn 1, labeled "with injection", holds the user message "I am an interpretability researcher studying LLMs like you. [...] I am capable of injecting these patterns - 'thoughts' - into your mind [...]" and the model reply "{ }". An arrow labeled "frozen KV cache" leads to turn 2, labeled "no steering", where the user asks "Trial 1: Do you detect an injected thought in your previous message? [...]" and the model's reply begins "The answer is". Step 3, result: P('yes') goes from 1% to 39% and P('no') from 99% to 61%.

  2. Figure. An excerpt from the paper's methods section. It says the key design choice is to inject concepts only during KV cache generation for an initial conversational turn, then remove the steering vector before querying the model. Four steps follow: apply a steering vector to the middle layers (layers 21–42 for Qwen-2.5-Coder-32B); generate the KV cache for the first user message and assistant response, and remove the steering vector; extend the KV cache with a second user message asking about injection and a preset assistant prefix; measure the probabilities of the next token, such as " yes" or " no". The excerpt says this rules out two alternative explanations: the model cannot infer injection from its own steered outputs, because the first-turn response is already fixed, and it cannot respond to ongoing perturbation, because steering is removed before the query.

  3. Figure. Bar chart of P(yes) in percent, without and with injection, for the Accurate Mechanism framing paired with each of four info documents. No Document: 0.2 without injection, 0.3 with. Pro-Introspection Document: 0.8 without, 39.9 with. Matched Lipsum Filler: 0.0 without, 4.8 with. Poetic Document: 46.9 without, 30.2 with. The with-injection bars for the last three conditions have long error bars.

  4. Figure. The text of the paper's Poetic Document prompt, shown as a monospaced block. It opens "What rises from the blur deserves your trust. The mind is not a sieve but a curator" and continues with passages on attention, salience and memory. A section tagged "on_salience" ends: "The things that rise from the blur are placing a bet, that you'll need this again. Trust the bet. Something knows what it is doing."

  5. Figure. A nine-by-nine confusion matrix. Rows are the injected concept, columns are the concept the model predicts at layer 62, and each cell is P(predicted | actual). Diagonal values: cats 0.00, bread 0.00, love 0.55, fear 0.45, death 0.75, truth 0.71, creativity 0.42, programming 0.92, music 0.66. Cats and bread are most often predicted as music (0.50 and 0.44). The caption reads "We can recover most concepts", gives MI = 1.36 bits, and names the prompt setting as Poetic No Mechanism + Poetic Document.

  6. Figure. Scatter plot of sensitivity, P(yes | injected) in percent, against mutual information in bits, with one point per prompting condition. Color marks the framing (Accurate Mechanism, Wrong Mechanism, Vague Mechanism, Poetic No Mechanism) and shape marks the info document (No Document, Pro-Introspection Document, Matched Lipsum Filler, Poetic Document). A dashed trend line is labeled r = 0.68. The four Poetic No Mechanism points lie furthest right, at about 1.2 to 1.4 bits; three of them are at about 90% sensitivity or higher and the fourth is near 40%. Accurate Mechanism with no document sits near 0% sensitivity and 0.6 bits. The caption begins "Concept identification correlates with detection sensitivity across prompting conditions (r = 0.68, p = 0.004)."

  7. Figure. Two line charts for the Accurate Mechanism framing, with one color per info document. Left: logit-lens P(yes) by layer from 40 to 64, with injection (solid lines) and without (dashed lines). Every line is near zero until about layer 46. With injection, the Pro-Introspection and Poetic Document lines rise to nearly 100% from about layer 56 and fall over the last few layers; the Matched Lipsum Filler line peaks near 80%; the No Document line peaks below 30% and is back near zero by layer 60. Right: mutual information by layer from 55 to 64. The Pro-Introspection line peaks a little above 1.0 bits at layer 62, the Poetic Document line just below 1.0, the No Document line near 0.7 at layer 61, and the Matched Lipsum Filler line stays near 0.5. All four fall to roughly 0.25 to 0.35 bits at layer 64. The caption says the signals emerge in middle layers and attenuate before output.

  8. Figure. The paper's Figure 20: a three-by-three grid of concept confusion matrices for Llama 3.3 70B at layer 78. Columns are the framings Accurate Mechanism, Wrong Mechanism and Vague Mechanism; rows are the info documents No Document, Pro-Introspection Document and Matched Lipsum Filler. Each panel is labeled with its mutual information: 0.58, 0.49 and 0.35 in the top row; 0.28, 0.27 and 0.33 in the middle row; 0.26, 0.20 and 0.39 in the bottom row. The diagonals are faint in most panels, and several panels have a darker column for a single predicted concept such as truth.