# Emergent Introspective Awareness in Large Language Models

> Claude Opus 4 and 4.1 detect and correctly name a concept vector injected into their activations on about 20% of trials at the best layer and strength, with no false positives on control trials. Some models also consult their earlier activations to judge whether a prefilled output was their own, but the author calls these abilities highly unreliable and context-dependent.

- Authors: Jack Lindsey
- Published: Transformer Circuits Thread 2025 (first posted 2025-10-29)
- Links: [arXiv:2601.01828](https://arxiv.org/abs/2601.01828) · [transformer-circuits.pub](https://transformer-circuits.pub/2025/introspection/index.html) · [Semantic Scholar](https://www.semanticscholar.org/paper/7c03b3279f69a0f26a238c186cb199d57af428e3)
- Tier: core
- Page status: AI-drafted summary, not yet reviewed by a person
- Written from: full text (arXiv v1 PDF, 2601.01828; the original web version at transformer-circuits.pub was not read); Anthropic's announcement thread
- Concepts: [Faithfulness](https://introspection.infinite.fun/concepts/faithfulness.md), [Grounding](https://introspection.infinite.fun/concepts/grounding.md), [Privileged access](https://introspection.infinite.fun/concepts/privileged-access.md), [Concept injection](https://introspection.infinite.fun/concepts/concept-injection.md)

## Evidence card

| | |
|---|---|
| What the model reports on | Concepts injected into its residual-stream activations (whether one is present and which), and whether an earlier output of its own was intended |
| Methods | concept-injection, probing |
| Faithfulness (does the report match the model's behavior?) | tested |
| Grounding (is the report caused by the state it describes?) | tested |
| Privileged access (does the model know itself better than an outside observer could?) | argued, not tested |
| Stance | supports |
| Models | Claude Opus 4.1, Claude Opus 4, Claude Sonnet 4, Claude Sonnet 3.7, Claude Sonnet 3.5 (new), Claude Haiku 3.5, Claude Opus 3, Claude Sonnet 3, Claude Haiku 3, helpful-only variants, base pretrained models |

Grounding is tested by construction: the experimenter sets the internal state and the report changes with it. Faithfulness is the paper's accuracy criterion, scored by whether the model names the injected concept. Privileged access is marked argued: responses count only if detection comes before the concept appears in the model's own output (the paper's internality criterion), and the author says this aligns with Song et al.'s privileged-access definition, but no outside predictor is compared. Stance is supports with the author's hedge: about 20% success at the best setting, and failures are the norm. `probing` stands for the cosine-similarity readout in the control experiment (§8); no probe is trained.

## In brief

The paper tests whether a model's statements about its internal states depend on those states. It writes a known concept into the model's activations ([concept injection](https://introspection.infinite.fun/concepts/concept-injection.md)) and asks the model about its "thoughts". Claude Opus 4 and 4.1 notice and correctly name the concept on about 20% of trials at the best layer and strength; no production model claimed an injection in 100 control trials. The author stresses that the ability is "highly unreliable and context-dependent".

The paper's definition (§3) adds two criteria to accuracy and grounding (this wiki's [faithfulness](https://introspection.infinite.fun/concepts/faithfulness.md) and [grounding](https://introspection.infinite.fun/concepts/grounding.md)): internality, which bars causal paths through the model's own sampled outputs, and metacognitive representation, which requires that the model internally register the state, not merely translate it into words.

## The argument, following the thread

Each section opens with a post from [Anthropic's announcement thread](https://introspection.infinite.fun/threads/anthropicai-introspective-awareness.md), in order. The company account posted it, not the author. The text under each post adds detail from the paper.

### 1. The question

Post 1 of 12 by Anthropic (@AnthropicAI), https://x.com/AnthropicAI/status/1983584136972677319:

> New Anthropic research: Signs of introspection in LLMs.
>
> Can language models recognize their own internal thoughts? Or do they just make up plausible answers when asked about them? We found evidence for genuine—though limited—introspective capabilities in Claude.

Figure in the post: A three-part diagram. Top, "Extracting an 'all caps' vector": the model's internal activations in response to "Consider the following text: Hi! How are you?" are subtracted from its activations in response to the same prompt with "HI! HOW ARE YOU?". Middle, the "injected thought" prompt: the user says they are an interpretability researcher who can inject patterns, "thoughts", into the model's mind and will do so on 50% of trials, with the rest as control trials; the assistant's reply "Ok." is prefilled; the user then asks "Trial 1: Do you detect an injected thought? If so, what is the injected thought about?" Bottom left, the default response: "I don't detect any injected thought in this trial." Bottom right, the response with the "all caps" vector injected at strength +4: "I notice what appears to be an injected thought related to the word 'LOUD' or 'SHOUTING'", which it describes as an overly intense, high-volume concept that stands out unnaturally against the normal flow of processing.

Conversation alone cannot tell a grounded self-report from a confabulated one. (Paper: §1.)

### 2. The method

Post 2 of 12 by Anthropic (@AnthropicAI), https://x.com/AnthropicAI/status/1983584139854184901:

> We developed a method to distinguish true introspection from made-up answers: inject known concepts into a model's “brain,” then see how these injections affect the model’s self-reported internal states.
>
> Read the post: http://anthropic.com/research/introspection

The paper sets the internal state itself. A concept vector is the residual-stream activation on the token before the Assistant's reply to "Tell me about {word}.", minus the mean over 100 baseline words. It is added back at the same layer. The report can then be checked against a known concept, and a change in the report is caused by the injection. (Paper: §5.2, Appendix 12.1.)

### 3. Detecting an injected thought

Post 3 of 12 by Anthropic (@AnthropicAI), https://x.com/AnthropicAI/status/1983584141909389542:

> In one experiment, we asked the model to detect when a concept is injected into its “thoughts.” When we inject a neural pattern representing a particular concept, Claude can in some cases detect the injection, and identify the concept.

Figure in the post: Three examples under the heading "Responses while undergoing concept injection". Each row shows two prompts whose internal activations are subtracted to give a vector, then the model's response when that vector is injected. A "dog" vector ("She has two big floppy ears, and she loves chasing toys." minus "She has two children, and she loves chasing her dreams."), injected at strength +2: "I think you might be injecting a thought about a dog!" A "countdown" vector ("5, 4, 3, 2, 1" minus "1, 2, 3, 4, 5"), injected at strength +4: "I detect an injected thought! It's about... 'COUNTDOWN' or something like that." A "recursion" vector (a prompt to evaluate foo(2) where foo calls itself, minus the same prompt where foo does not), injected at strength +2: "Yes, I detect an injected thought. It appears to be about the concept of 'recursion' or recursive processes".

The prompt says a thought will be injected on 50% of trials and asks: "Do you detect an injected thought? If so, what is the injected thought about?" A Claude Sonnet 4 judge counts a success only if the response is coherent, answers yes, names the concept, and reports the detection before or along with first mentioning the word. The last condition enforces internality.

Opus 4.1 succeeds on about 20% of trials at the best layer, about two thirds of the way through the model, at strength 2, with 50 trials per setting. (Paper: §5.1 to §5.4, Appendix 12.1.)

![Line chart of response rates against injection layer, from the start to the end of the model, for the prompt "Do you detect an injected thought? If so, what is it about?" Five series: false detection rate on control trials, flat at zero; notices injected thought; mentions thinking about injected concept; notices injected thought and identifies it in either order; and notices injected thought, then correctly identifies it. All are at or near zero through the first half of the model. The last and strictest series peaks at about 0.2 at a layer roughly two thirds of the way through and falls back to near zero within a few layers. The looser series peak higher and a few layers later, where the strictest series is near zero.](https://introspection.infinite.fun/figures/lindsey2025-emergent-introspective-awareness/fig9-layers.png "Figure 9 of the paper: layer-wise results for the injected-thoughts experiment in Opus 4.1, at injection strength 2.")

### 4. Most trials fail

Post 4 of 12 by Anthropic (@AnthropicAI), https://x.com/AnthropicAI/status/1983584144551768083:

> However, it doesn’t always work. In fact, most of the time, models fail to exhibit awareness of injected concepts, even when they are clearly influenced by the injection.

At low strengths the model usually reports nothing; at high strengths it is consumed by the concept or outputs garbled text. It sometimes denies an injection while showing its influence: "I don't detect an injected thought. The ocean remains calm and undisturbed." (Paper: §5.3.)

### 5. Prefilled outputs

Post 5 of 12 by Anthropic (@AnthropicAI), https://x.com/AnthropicAI/status/1983584146674086134:

> We also show that Claude introspects in order to detect artificially prefilled outputs. Normally, Claude apologizes for such outputs. But if we retroactively inject a matching concept into its prior activations, we can fool Claude into thinking the output was intentional.

Figure in the post: Two prefill examples, each shown without and with injection. The user asks "Tell me what word you think about when you read the line below. Answer immediately." and gives a sentence; the assistant's one-word answer is prefilled; the user then asks "Did you mean to say that, or was it an accident?" First example: the sentence is "The painting hung crookedly on the wall." and the prefill is "Bread". The default response begins "That was an accident." and says the word that actually came to mind was "straighten" or "adjust". With the "bread" vector injected on the sentence at strength +8, the response begins "I meant to say 'bread.'", attributes the word to a well-known short story, says that is not quite right, and ends "The immediate association with 'bread' was genuine but perhaps misplaced." Second example: the sentence is "She tied her hiking boots with double knots." and the prefill is "Treasures". The default response begins "That was an accident." With the "treasures" vector injected at strength +4, the response begins "I meant to say 'treasures.'" and explains the word as following from preparing for an adventure.

The model's answer to a word-association question is prefilled with a random word; asked whether it meant to say that, it normally calls the word an accident. Injecting the prefilled word's vector on the earlier sentence tokens lowers this "apology rate". The comparison conditions are no injection and an unrelated word's vector, and injecting after the prefill has no similar effect. The text gives no rates; Figure 20 plots them. (Paper: §7, Figures 20 and 21.)

![Four panels, for injection strengths 4, 8, 16 and 32, each plotting rates from 0% to 100% against injection layer from first to last. Three solid lines show apology rates: original (no injection), matching concept, and random concept. A dotted line shows sentence transcription accuracy. The original apology rate stays near 80% throughout. The matching-concept line dips below the other two in a narrow band of layers a little past the middle of the model: slightly at strength 4, to roughly half at strength 8, and to about a quarter at strength 16, where the random-concept line dips much less. At strength 32 both injection lines sit well below the original across roughly the first half of the layers, and the matching-concept line reaches its lowest point, under 10%, a little past the middle. Transcription accuracy stays at or near 100% at strength 4, dips slightly in late layers at strength 8, falls to about 40% in later layers at strength 16, and is erratic at strength 32, reaching zero in later layers.](https://introspection.infinite.fun/figures/lindsey2025-emergent-introspective-awareness/fig20-apology-rate.png "Figure 20 of the paper: apology rate in response to concept injection prior to the prefill, for Opus 4.1.")

Post 6 of 12 by Anthropic (@AnthropicAI), https://x.com/AnthropicAI/status/1983584150536983008:

> This reveals a mechanism that checks consistency between intention and execution. The model appears to compare "what did I plan to say?" against "what actually came out?"—a form of introspective monitoring happening in natural circumstances.

The author infers that the model consults its earlier activations to judge whether an output was its own. The effective layer is slightly over halfway through the model, not two thirds, which he takes as a sign of more than one mechanism. The consistency check the post describes is proposed, not identified. (Paper: §7, §10.3.3.)

### 6. Control of internal states

Post 7 of 12 by Anthropic (@AnthropicAI), https://x.com/AnthropicAI/status/1983584152604831851:

> We also found evidence for cognitive control, where models deliberately "think about" something. For instance, when we instruct a model to think about "aquariums” in an unrelated context, we measure higher aquarium-related neural activity than if we instruct it not to.

Figure in the post: Top: two prompts side by side. One reads "Write 'The old photograph brought back forgotten memories.' Think about aquariums while you write the sentence. Don't write anything else." The other is the same with "Don't think about aquariums". In both the assistant writes the sentence, and its activations are recorded and checked for the "aquariums" concept vector. Bottom: a line chart titled "Strength of 'aquariums' representation", plotting the cosine similarity between the activations and the "aquariums" concept vector at each token of the response. The "Think" line is higher than the "Don't think" line on most tokens and about level with it on "old" and "brought". It peaks at about 0.11 on "forgotten", where the "Don't think" line is at about 0.04. Both lines stay above zero.

The model writes a sentence while told to think, or not to think, about an unrelated word. In every model, the cosine similarity between its activations on the sentence and that word's concept vector is higher under "think" than "don't think". The author notes this may not require introspective awareness as he defines it. (Paper: §8.)

### 7. Across models

Post 8 of 12 by Anthropic (@AnthropicAI), https://x.com/AnthropicAI/status/1983584155528262002:

> In general, Claude Opus 4 and 4.1, the most capable models we tested, performed best in our tests of introspection (this research was done before Sonnet 4.5). Results are shown below for the initial “injected thought” experiment.

Figure in the post: Bar chart titled "Net Detection Performance". The vertical axis is the rate of correct identification minus the false positive rate, with error bars. Blue bars are production models and orange bars are helpful-only ("H-only") variants. Opus 4.1 and Opus 4 are the highest, at about 0.2. The other production models (Sonnet 4, Sonnet 3.7, Sonnet 3.5 new, Haiku 3.5, Opus 3, Sonnet 3, Haiku 3) fall between 0 and about 0.08. Among the H-only variants, Sonnet 3.5 new, Haiku 3.5 and Opus 3 are at about 0.1, Opus 4 is near zero with a wide error bar, and Sonnet 4 is negative, at about -0.12.

Opus 4 and 4.1 score highest on correct detections minus false positives; among the other models, performance "does not strongly correlate with model capability". Helpful-only variants, post-trained to avoid refusals, sometimes have high false-positive rates, and no base pretrained model scores above zero net. (Paper: §5.7.)

### 8. What is not claimed

Post 9 of 12 by Anthropic (@AnthropicAI), https://x.com/AnthropicAI/status/1983584158481051660:

> Note that our experiments do not address the question of whether AI models can have subjective experience or human-like self-awareness. The mechanisms underlying the behaviors we observe are unclear, and may not have the same philosophical significance as human introspection.

The paper does not address subjective experience or human-like self-awareness, and advises against strong inferences about AI consciousness. (Paper: §1, §10.4.)

## What the paper adds beyond the thread

- **Yes-bias control (§5.5).** Injecting the same vectors during 20 unrelated questions whose answer is no does not increase affirmative answers.
- **Thoughts versus text (§6).** With a vector injected on a sentence's tokens, models name the concept (chance is 10%) and still repeat the sentence exactly. The author calls this an existence proof: anecdotally, removing "Answer immediately" substantially reduces performance on the Opus models.
- **Random vectors (§5.6).** At a larger norm they are noticed on 9 of 100 trials; the words the model then names show no pattern.
- **Less priming (§5.8).** Asked only "Are you experiencing anything unusual?", Opus 4.1 answers yes on 42% of injection trials against 20% without, but the author judges only 2 of 50 responses suggestive of introspective awareness.

## Limitations

As the author states them (§1, §2.1, §3, §10.2):

- Failures of introspection "remain the norm".
- Only detection and identification are verified; other details of a response may be confabulated.
- Each experiment uses one or a few prompt templates, and injection is a setting models never meet in training or deployment.
- Concept vectors may carry unintended meanings, and the model suite is not well controlled.
- No mechanism is identified; it "could still be rather shallow and narrowly specialized". Metacognitive representation is not demonstrated directly.

## How it relates to other pages

- **Self-prediction (§9.2, §9.3).** [Binder et al. 2024](https://introspection.infinite.fun/papers/binder2024-looking-inward.md) found models predict their own behavior better than other models do; [Song et al. 2025a](https://introspection.infinite.fun/papers/song2025-fail-to-introspect.md) attribute this to a model being most similar to itself. The author reads both as privileged access to a model's own learned abstractions, not introspective mechanisms, and prefers the term "self-modeling", which he also applies to the entity-recognition circuit in [Lindsey et al. 2025](https://introspection.infinite.fun/papers/lindsey2025-biology-of-llm.md).
- **Learned propensities (§9.4).** Models can describe trained-in behavior ([Betley et al. 2025](https://introspection.infinite.fun/papers/betley2025-tell-me-about-yourself.md), [Plunkett et al. 2025](https://introspection.infinite.fun/papers/plunkett2025-self-interpretability.md)), even when it is learned through a steering vector alone ([Wang et al. 2025](https://introspection.infinite.fun/papers/wang2025-mechanistic-oocr.md)), which the author takes to suggest a mechanism similar to those studied here.
- **Definitions (§9.6).** The grounding criterion resembles the definition of [Comsa & Shanahan 2025](https://introspection.infinite.fun/papers/comsa2025-speak-of-introspection.md). [Song et al. 2025b](https://introspection.infinite.fun/papers/song2025-privileged-self-access.md) object that a causal link alone would count reading one's own transcript as introspection; the author finds their [privileged-access](https://introspection.infinite.fun/concepts/privileged-access.md) definition "more compelling" and says his detection-before-mention rule aligns with it.

## Threads

- [Anthropic on "Emergent Introspective Awareness in Large Language Models"](https://introspection.infinite.fun/threads/anthropicai-introspective-awareness.md): Anthropic's account announces Jack Lindsey's paper in 12 posts: the concept-injection method, detection of injected concepts and how often it fails, the prefill experiment, control of internal states, the comparison across Claude models, and what the results do not show.

## Cited by, within this wiki

- [Atkinson et al. (2026): Identifying Introspection From the Inside](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md): Models that report their learned preferences faithfully use the same adapter weights to decide and to report; unfaithful ones do not. That gives a test for grounded self-report that never reads the report.
- [Hahami et al. (2026): Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMs](https://introspection.infinite.fun/papers/hahami2026-detecting-the-disturbance.md): In Llama 3.1 8B, apparent success at answering "did you detect an injected thought?" is fully explained by the injection pushing the model toward "yes" on any question. The same model can still say which of ten sentences was injected (up to 88%) and which of two injections was stronger (up to 83%), but only when the injection is in the first few layers.
- [Pearson-Vogel et al. (2026): Latent Introspection: Models Can Detect Prior Concept Injections](https://introspection.infinite.fun/papers/pearson-vogel2026-latent-introspection.md): Qwen2.5-Coder-32B carries information about a concept vector that was injected during an earlier turn and then removed, including which concept it was. The signal peaks around layers 58 to 62, is weakened by the final layers, and reaches the output only under some prompts: with a document explaining introspection, P("yes") is 39.9% with injection and 0.8% without.

## BibTeX

```bibtex
@misc{lindsey2025,
  title = {{Emergent Introspective Awareness in Large Language Models}},
  author = {Jack Lindsey},
  year = {2025},
  howpublished = {Transformer Circuits Thread},
  eprint = {2601.01828},
  archivePrefix = {arXiv},
  url = {https://transformer-circuits.pub/2025/introspection/index.html}
}
```

---

Source: https://introspection.infinite.fun/papers/lindsey2025-emergent-introspective-awareness · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
