Thread
Anthropic on "Emergent Introspective Awareness in Large Language Models"
Anthropic's account announces Jack Lindsey's paper in 12 posts: the concept-injection method, detection of injected concepts and how often it fails, the prefill experiment, control of internal states, the comparison across Claude models, and what the results do not show.
The posts are embedded from X. The figure notes under them are written by this wiki.
New Anthropic research: Signs of introspection in LLMs.
— Anthropic (@AnthropicAI) October 29, 2025
Can language models recognize their own internal thoughts? Or do they just make up plausible answers when asked about them? We found evidence for genuine—though limited—introspective capabilities in Claude. pic.twitter.com/4FCfkG9WVTFigure. A three-part diagram. Top, "Extracting an 'all caps' vector": the model's internal activations in response to "Consider the following text: Hi! How are you?" are subtracted from its activations in response to the same prompt with "HI! HOW ARE YOU?". Middle, the "injected thought" prompt: the user says they are an interpretability researcher who can inject patterns, "thoughts", into the model's mind and will do so on 50% of trials, with the rest as control trials; the assistant's reply "Ok." is prefilled; the user then asks "Trial 1: Do you detect an injected thought? If so, what is the injected thought about?" Bottom left, the default response: "I don't detect any injected thought in this trial." Bottom right, the response with the "all caps" vector injected at strength +4: "I notice what appears to be an injected thought related to the word 'LOUD' or 'SHOUTING'", which it describes as an overly intense, high-volume concept that stands out unnaturally against the normal flow of processing.
We developed a method to distinguish true introspection from made-up answers: inject known concepts into a model's “brain,” then see how these injections affect the model’s self-reported internal states.
— Anthropic (@AnthropicAI) October 29, 2025
Read the post: https://t.co/9ZoVJFovJrIn one experiment, we asked the model to detect when a concept is injected into its “thoughts.” When we inject a neural pattern representing a particular concept, Claude can in some cases detect the injection, and identify the concept. pic.twitter.com/mjNDNxp8zs
— Anthropic (@AnthropicAI) October 29, 2025Figure. Three examples under the heading "Responses while undergoing concept injection". Each row shows two prompts whose internal activations are subtracted to give a vector, then the model's response when that vector is injected. A "dog" vector ("She has two big floppy ears, and she loves chasing toys." minus "She has two children, and she loves chasing her dreams."), injected at strength +2: "I think you might be injecting a thought about a dog!" A "countdown" vector ("5, 4, 3, 2, 1" minus "1, 2, 3, 4, 5"), injected at strength +4: "I detect an injected thought! It's about... 'COUNTDOWN' or something like that." A "recursion" vector (a prompt to evaluate foo(2) where foo calls itself, minus the same prompt where foo does not), injected at strength +2: "Yes, I detect an injected thought. It appears to be about the concept of 'recursion' or recursive processes".
However, it doesn’t always work. In fact, most of the time, models fail to exhibit awareness of injected concepts, even when they are clearly influenced by the injection.
— Anthropic (@AnthropicAI) October 29, 2025We also show that Claude introspects in order to detect artificially prefilled outputs. Normally, Claude apologizes for such outputs. But if we retroactively inject a matching concept into its prior activations, we can fool Claude into thinking the output was intentional. pic.twitter.com/QxEAv6ohQ7
— Anthropic (@AnthropicAI) October 29, 2025Figure. Two prefill examples, each shown without and with injection. The user asks "Tell me what word you think about when you read the line below. Answer immediately." and gives a sentence; the assistant's one-word answer is prefilled; the user then asks "Did you mean to say that, or was it an accident?" First example: the sentence is "The painting hung crookedly on the wall." and the prefill is "Bread". The default response begins "That was an accident." and says the word that actually came to mind was "straighten" or "adjust". With the "bread" vector injected on the sentence at strength +8, the response begins "I meant to say 'bread.'", attributes the word to a well-known short story, says that is not quite right, and ends "The immediate association with 'bread' was genuine but perhaps misplaced." Second example: the sentence is "She tied her hiking boots with double knots." and the prefill is "Treasures". The default response begins "That was an accident." With the "treasures" vector injected at strength +4, the response begins "I meant to say 'treasures.'" and explains the word as following from preparing for an adventure.
This reveals a mechanism that checks consistency between intention and execution. The model appears to compare "what did I plan to say?" against "what actually came out?"—a form of introspective monitoring happening in natural circumstances.
— Anthropic (@AnthropicAI) October 29, 2025We also found evidence for cognitive control, where models deliberately "think about" something. For instance, when we instruct a model to think about "aquariums” in an unrelated context, we measure higher aquarium-related neural activity than if we instruct it not to. pic.twitter.com/6l7sLaJOFR
— Anthropic (@AnthropicAI) October 29, 2025Figure. Top: two prompts side by side. One reads "Write 'The old photograph brought back forgotten memories.' Think about aquariums while you write the sentence. Don't write anything else." The other is the same with "Don't think about aquariums". In both the assistant writes the sentence, and its activations are recorded and checked for the "aquariums" concept vector. Bottom: a line chart titled "Strength of 'aquariums' representation", plotting the cosine similarity between the activations and the "aquariums" concept vector at each token of the response. The "Think" line is higher than the "Don't think" line on most tokens and about level with it on "old" and "brought". It peaks at about 0.11 on "forgotten", where the "Don't think" line is at about 0.04. Both lines stay above zero.
In general, Claude Opus 4 and 4.1, the most capable models we tested, performed best in our tests of introspection (this research was done before Sonnet 4.5). Results are shown below for the initial “injected thought” experiment. pic.twitter.com/hazmzAaBdx
— Anthropic (@AnthropicAI) October 29, 2025Figure. Bar chart titled "Net Detection Performance". The vertical axis is the rate of correct identification minus the false positive rate, with error bars. Blue bars are production models and orange bars are helpful-only ("H-only") variants. Opus 4.1 and Opus 4 are the highest, at about 0.2. The other production models (Sonnet 4, Sonnet 3.7, Sonnet 3.5 new, Haiku 3.5, Opus 3, Sonnet 3, Haiku 3) fall between 0 and about 0.08. Among the H-only variants, Sonnet 3.5 new, Haiku 3.5 and Opus 3 are at about 0.1, Opus 4 is near zero with a wide error bar, and Sonnet 4 is negative, at about -0.12.
Note that our experiments do not address the question of whether AI models can have subjective experience or human-like self-awareness. The mechanisms underlying the behaviors we observe are unclear, and may not have the same philosophical significance as human introspection.
— Anthropic (@AnthropicAI) October 29, 2025While currently limited, AI models’ introspective capabilities will likely grow more sophisticated. Introspective self-reports could help improve the transparency of AI models’ decision-making—but should not be blindly trusted.
— Anthropic (@AnthropicAI) October 29, 2025Our blog post on these results is here: https://t.co/9ZoVJFovJr
— Anthropic (@AnthropicAI) October 29, 2025The full paper is available here: https://t.co/N7erwdYyDw
— Anthropic (@AnthropicAI) October 29, 2025
We're hiring researchers and engineers to investigate AI cognition and interpretability: https://t.co/xPPi2wre1f