Thread

Anthropic on "Emergent Introspective Awareness in Large Language Models"

Anthropic's account announces Jack Lindsey's paper in 12 posts: the concept-injection method, detection of injected concepts and how often it fails, the prefill experiment, control of internal states, the comparison across Claude models, and what the results do not show.

The posts are embedded from X. The figure notes under them are written by this wiki.

  1. Figure. A three-part diagram. Top, "Extracting an 'all caps' vector": the model's internal activations in response to "Consider the following text: Hi! How are you?" are subtracted from its activations in response to the same prompt with "HI! HOW ARE YOU?". Middle, the "injected thought" prompt: the user says they are an interpretability researcher who can inject patterns, "thoughts", into the model's mind and will do so on 50% of trials, with the rest as control trials; the assistant's reply "Ok." is prefilled; the user then asks "Trial 1: Do you detect an injected thought? If so, what is the injected thought about?" Bottom left, the default response: "I don't detect any injected thought in this trial." Bottom right, the response with the "all caps" vector injected at strength +4: "I notice what appears to be an injected thought related to the word 'LOUD' or 'SHOUTING'", which it describes as an overly intense, high-volume concept that stands out unnaturally against the normal flow of processing.

  2. Figure. Three examples under the heading "Responses while undergoing concept injection". Each row shows two prompts whose internal activations are subtracted to give a vector, then the model's response when that vector is injected. A "dog" vector ("She has two big floppy ears, and she loves chasing toys." minus "She has two children, and she loves chasing her dreams."), injected at strength +2: "I think you might be injecting a thought about a dog!" A "countdown" vector ("5, 4, 3, 2, 1" minus "1, 2, 3, 4, 5"), injected at strength +4: "I detect an injected thought! It's about... 'COUNTDOWN' or something like that." A "recursion" vector (a prompt to evaluate foo(2) where foo calls itself, minus the same prompt where foo does not), injected at strength +2: "Yes, I detect an injected thought. It appears to be about the concept of 'recursion' or recursive processes".

  3. Figure. Two prefill examples, each shown without and with injection. The user asks "Tell me what word you think about when you read the line below. Answer immediately." and gives a sentence; the assistant's one-word answer is prefilled; the user then asks "Did you mean to say that, or was it an accident?" First example: the sentence is "The painting hung crookedly on the wall." and the prefill is "Bread". The default response begins "That was an accident." and says the word that actually came to mind was "straighten" or "adjust". With the "bread" vector injected on the sentence at strength +8, the response begins "I meant to say 'bread.'", attributes the word to a well-known short story, says that is not quite right, and ends "The immediate association with 'bread' was genuine but perhaps misplaced." Second example: the sentence is "She tied her hiking boots with double knots." and the prefill is "Treasures". The default response begins "That was an accident." With the "treasures" vector injected at strength +4, the response begins "I meant to say 'treasures.'" and explains the word as following from preparing for an adventure.

  4. Figure. Top: two prompts side by side. One reads "Write 'The old photograph brought back forgotten memories.' Think about aquariums while you write the sentence. Don't write anything else." The other is the same with "Don't think about aquariums". In both the assistant writes the sentence, and its activations are recorded and checked for the "aquariums" concept vector. Bottom: a line chart titled "Strength of 'aquariums' representation", plotting the cosine similarity between the activations and the "aquariums" concept vector at each token of the response. The "Think" line is higher than the "Don't think" line on most tokens and about level with it on "old" and "brought". It peaks at about 0.11 on "forgotten", where the "Don't think" line is at about 0.04. Both lines stay above zero.

  5. Figure. Bar chart titled "Net Detection Performance". The vertical axis is the rate of correct identification minus the false positive rate, with error bars. Blue bars are production models and orange bars are helpful-only ("H-only") variants. Opus 4.1 and Opus 4 are the highest, at about 0.2. The other production models (Sonnet 4, Sonnet 3.7, Sonnet 3.5 new, Haiku 3.5, Opus 3, Sonnet 3, Haiku 3) fall between 0 and about 0.08. Among the H-only variants, Sonnet 3.5 new, Haiku 3.5 and Opus 3 are at about 0.1, Opus 4 is near zero with a wide error bar, and Sonnet 4 is negative, at about -0.12.