# Concept injection

> Adding a known representation to a model's activations, then asking the model whether it notices and what it is.

Also called: activation injection, injected thoughts.

Concept injection tests [grounding](https://introspection.infinite.fun/concepts/grounding.md) directly. The experimenter sets an internal state by writing a known vector into the residual stream, so there is a ground truth for what the model should report, and a change in the report is caused by the injection.

## The method

In [Lindsey (2025)](https://introspection.infinite.fun/papers/lindsey2025-emergent-introspective-awareness.md), a concept vector is the model's activation when asked about a word, minus the mean over baseline words. It is added back at a chosen layer and strength. The model is told a thought may be injected and asked whether it detects one and what it is about.

## What has been found

- **Lindsey (2025).** Claude Opus 4 and 4.1 detect and correctly name the concept on about 20% of trials at the best layer and strength, with no false positives in 100 control trials. The author calls the ability highly unreliable and context-dependent.
- **[Hahami et al. (2026)](https://introspection.infinite.fun/papers/hahami2026-detecting-the-disturbance.md).** In Llama 3.1 8B, yes-or-no detection is fully explained by the injection pushing the model toward "yes" on any question. The model can still say which of ten sentences was injected (up to 88%) and which of two injections was stronger (up to 83%), but only for injections in the first few layers.
- **[Pearson-Vogel et al. (2026)](https://introspection.infinite.fun/papers/pearson-vogel2026-latent-introspection.md).** Qwen2.5-Coder-32B carries information about a concept injected in an earlier turn and then removed. The signal is strong in intermediate layers, is weakened by the final layers, and reaches the output only under some prompts.

## Limits of the method

- Naming the injected concept is open to [causal bypassing](https://introspection.infinite.fun/concepts/causal-bypassing.md): the vector may simply push the model to talk about the concept. [Morris & Plunkett (2025)](https://introspection.infinite.fun/papers/morris2025-causal-bypassing.md) argue only the detection question avoids this, and Hahami et al. show detection can have its own artifact.
- Injection is a situation models never meet in training or deployment, which Lindsey lists among his limitations.

[Atkinson et al. (2026)](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md) present their shared-mechanism test as a complement: injection plants a known thought, while theirs asks whether a report about a naturally learned behavior shares a mechanism with that behavior.

## Papers tagged with this concept

- [Lindsey (2025): Emergent Introspective Awareness in Large Language Models](https://introspection.infinite.fun/papers/lindsey2025-emergent-introspective-awareness.md): Claude Opus 4 and 4.1 detect and correctly name a concept vector injected into their activations on about 20% of trials at the best layer and strength, with no false positives on control trials. Some models also consult their earlier activations to judge whether a prefilled output was their own, but the author calls these abilities highly unreliable and context-dependent.
- [Morris & Plunkett (2025): Tests of LLM introspection need to rule out causal bypassing](https://introspection.infinite.fun/papers/morris2025-causal-bypassing.md): An intervention that changes a model's internal state can also cause an accurate report of that state by a path that skips the state, so accuracy after an intervention does not show the report is grounded. The authors name this causal bypassing and say the only test they know that rules it out is asking a model whether a concept was injected, a claim a later edit to the post hedges.
- [Hahami et al. (2026): Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs](https://introspection.infinite.fun/papers/hahami2026-detecting-the-disturbance.md): In Llama 3.1 8B, apparent success at answering "did you detect an injected thought?" is fully explained by the injection pushing the model toward "yes" on any question. The same model can still say which of ten sentences was injected (up to 88%) and which of two injections was stronger (up to 83%), but only when the injection is in the first few layers.
- [Pearson-Vogel et al. (2026): Latent Introspection: Models Can Detect Prior Concept Injections](https://introspection.infinite.fun/papers/pearson-vogel2026-latent-introspection.md): Qwen2.5-Coder-32B carries information about a concept vector that was injected during an earlier turn and then removed, including which concept it was. The signal peaks around layers 58 to 62, is weakened by the final layers, and reaches the output only under some prompts: with a document explaining introspection, P("yes") is 39.9% with injection and 0.8% without.

---

Source: https://introspection.infinite.fun/concepts/concept-injection · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
