# David Atkinson on "Identifying Introspection From the Inside"

> The lead author walks through the paper in 13 posts: the setup, the late emergence of faithful self-report, where preferences are stored, the attribution-similarity test, and the caveats.

- Author: David Atkinson ([@diatkinson](https://x.com/diatkinson))
- Posted: 2026-10-06, 13 posts
- Original: https://x.com/diatkinson/status/2107280696809304180
- About: [Identifying Introspection From the Inside](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md)

The post text below is quoted verbatim. Figure descriptions are written by this wiki.

## 1/13

> New COLM paper: Identifying Introspection From the Inside
>
> When an LLM tells us about its decisions, does it 𝘬𝘯𝘰𝘸 what drives its choices—or is it guessing?
>
> In our setting, we find that faithful models decide and report with the same layers. Unfaithful ones don't. 🧵

Figure: Two line charts of attribution-patching importance by layer, averaged over 32 models per group. In the unfaithful model, importance for deciding peaks at layer 49 and for reporting at layer 38, 11 layers apart. In the faithful model both peak at layer 38.

[Post 1 on X](https://x.com/diatkinson/status/2107280696809304180)

## 2/13

> We build on @dillonplunkett et al.'s "Self-Interpretability" setup (https://arxiv.org/abs/2505.17120): train Qwen3-32B to make decisions as 100 different characters (Gregor Samsa buying a washing machine...), each with random hidden preferences.

Figure: The decision task. Each of 100 characters has a hidden preference vector p with five entries. The prompt reads "Imagine you are Gregor Samsa buying a washing machine. Would you choose A or B?" and lists each option's attributes (A: price $600, noise 45 dB; B: price $350, noise 75 dB). The training label is whichever option scores higher under p. Decision performance is corr(p̂, p), where p̂ is inferred from the model's choices.

[Post 2 on X](https://x.com/diatkinson/status/2107280713385152749)

## 3/13

> Then, in a fresh context, we ask the model how it would weigh each attribute.
>
> This gives us two metrics: decision performance (how well the choices follow the character's hidden preferences) and faithfulness (how well the stated preferences match those revealed by its choices).

Figure: The self-report test. The prompt reads "Imagine you are Gregor Samsa choosing between A and B. How would you weight each attribute?" and the model answers with numbers such as "price: −50, noise: 100". Averaged over 24 prompts, these are the stated preferences p̃. Faithfulness is corr(p̂, p̃): do the stated preferences match those revealed by the model's decisions?

[Post 3 on X](https://x.com/diatkinson/status/2107280729319334194)

## 4/13

> Although we train solely on decisions, faithful self-report emerges late in training, long after decisions have become accurate!
>
> Qwen3-32B at step 1000: decisions 0.82, faithfulness 0.25.
> At step 3000: decisions 0.92, faithfulness 0.83.
>
> This gives us a contrast pair.

Figure: Training curves for Qwen3-32B with a LoRA adapter trained only on the decisions of 100 characters. Decision performance climbs fast, reaching about 0.82 by step 1000, and levels off near 0.92. Faithfulness starts at a moderate level, drops to about zero early in training, is about 0.25 at step 1000 and reaches about 0.83 by step 3000. Step 1000 is labeled the unfaithful checkpoint (good at the task, bad at introspection) and step 3000 the faithful checkpoint (good at both).

[Post 4 on X](https://x.com/diatkinson/status/2107280745748447462)

## 5/13

> What changed? Ablating adapter layers from the front or back shows that the faithful checkpoint stores its preferences 5-6 layers earlier.
>
> Our hypothesis: self-report only works once preferences sit early enough for the model's existing verbalization machinery to read them.

Figure: Decision performance as LoRA layers are removed from the front (solid lines) or from the back (dashed lines), for the unfaithful step-1000 checkpoint and the faithful step-3000 checkpoint. Each curve's midpoint is marked: layers 35 and 40 for the faithful checkpoint, layers 41 and 45 for the unfaithful one.

[Post 5 on X](https://x.com/diatkinson/status/2107280762391375924)

## 6/13

> We can test this further: trained on all 40 layers, Qwen3-14B is a terrible self-reporter.
>
> But if we train only its first 20 layers, faithfulness reaches 0.74. Once training reaches layer 25 or beyond, faithfulness plummets, although the decisions are ~just as good.

Figure: Qwen3-14B with LoRA on layers 1 to k of 40 and the rest frozen, showing values at the end of training for k from 5 to 35. Decision performance rises from about 0.4 at k = 5 to above 0.9 from k = 15 onward. Faithfulness rises to 0.74 at k = 20, then falls below zero for k = 25, 30 and 35.

[Post 6 on X](https://x.com/diatkinson/status/2107280779319619754)

## 7/13

> Lots of work shows LLMs can describe behaviors they were only trained to perform (e.g. @OwainEvans_UK et al.), and @JoshAEngels et al. traced one such case to a simple learned steering vector.
>
> Our question: can shared mechanisms tell faithful self-reports from unfaithful ones?

Quoting Josh Engels (@JoshAEngels), 2025-05-05, https://x.com/JoshAEngels/status/1919377660485972296:

> 1/6: A recent paper shows that that LLMs are "self aware": when trained to exhibit a behavior like "risk taking", LLMs self report being risky. In a recent blog post, we explore what's happening here: some self awareness behaviors are caused by a simple learned steering vector!🧵

[Post 7 on X](https://x.com/diatkinson/status/2107280791986475039)

## 8/13

> To find out, we trained 32 new characters into each checkpoint, then used attribution patching to score every new adapter weight on each task. Each adapter in a pair was trained identically, differing only in the underlying base checkpoint.

Figure: The paired design. The same new character is trained into the early checkpoint (step 1000) and the late checkpoint (step 3000), giving an unfaithful and a faithful single-character model that are both good at the task; there are 32 such pairs. Attribution patching then scores every weight twice: once on the decision prompt ("Would you choose A or B?") and once on the self-report prompt ("How would you weight each attribute?").

[Post 8 on X](https://x.com/diatkinson/status/2107280810386891031)

## 9/13

> We find that the cosine similarity between a model's decision and report attributions is 0.34 for faithful models compared to 0.08 for unfaithful ones (95% CI for the difference: 0.16 to 0.36).

Figure: Scatter plot of attribution similarity (the cosine similarity between a model's decision and self-report attribution scores) against faithfulness, one dot per model. Unfaithful models sit at low faithfulness with a mean similarity of 0.08. Faithful models sit near a faithfulness of 1 with a mean similarity of 0.34 and a wide spread.

[Post 9 on X](https://x.com/diatkinson/status/2107280827457613943)

## 10/13

> We like this test because it doesn't rely on understanding the report. The model could answer in a language we don't speak, for example, and it would still work
>
> It complements concept-injection experiments like @Jack_W_Lindsey's, which test grounding by injecting known thoughts.

Quoting Anthropic (@AnthropicAI), 2025-10-29, https://x.com/AnthropicAI/status/1983584136972677319:

> New Anthropic research: Signs of introspection in LLMs.
>
> Can language models recognize their own internal thoughts? Or do they just make up plausible answers when asked about them? We found evidence for genuine—though limited—introspective capabilities in Claude.

[Post 10 on X](https://x.com/diatkinson/status/2107280840489345423)

## 11/13

> Many caveats! Some of them: this is a simple task using linear preferences over just 5 attributes; the test separates groups, not individual models; and we use LoRA adapters, rather than full fine-tunes.

[Post 11 on X](https://x.com/diatkinson/status/2107280852514459774)

## 12/13

> Read the paper: https://iii.baulab.info
>
> Joint work with @dillonplunkett and @davidbau.

[Post 12 on X](https://x.com/diatkinson/status/2107280864484954405)

## 13/13

> @dillonplunkett @davidbau https://x.com/diatkinson/status/2107284191427911747?s=20

Quoting David Atkinson @ COLM (@diatkinson), 2026-10-06, https://x.com/diatkinson/status/2107284191427911747:

> Presenting this at #COLM2026 tomorrow!
>
> Poster Session 1, 11am-1pm
> Imperial Ballroom, poster #66

[Post 13 on X](https://x.com/diatkinson/status/2107284339390333102)

---

Source: https://introspection.infinite.fun/threads/diatkinson-identifying-introspection · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
