# David Bau on "Identifying Introspection From the Inside"

> The senior author retells the paper as a story in ten posts: teach a model something, find it cannot describe what it learned, train it longer until it can, then look inside to see what changed.

- Author: David Bau ([@davidbau](https://x.com/davidbau))
- Posted: 2026-10-06, 10 posts
- Original: https://x.com/davidbau/status/2107398671323295968
- About: [Identifying Introspection From the Inside](https://introspection.infinite.fun/papers/atkinson2026-identifying-introspection.md)

The post text below is quoted verbatim. Figure descriptions are written by this wiki.

## 1/10

> The craziest experiment in my lab is a new setup by @diatkinson where he has found a way induce introspection on open LMs using @dillonplunkett's protocol, that lets him look inside introspection. It's nuts.
>
> Here's how the experiment works... 🧵 →

Figure: A cartoon robot saying "For this choice, price matters more than noise." Underneath is the question: "Is that report true?"

[Post 1 on X](https://x.com/davidbau/status/2107398671323295968)

## 2/10

> First, teach Qwen something new.
>
> David uses SFT to train the LM to answer "A" or "B" on made-up trivia  about famous people until the LM learns the trivia.
>
> Here it learns Gregor Samsa likes cheap washing machines even if they are noisy. Purely learning to say A or B. 🧵→

Quoting David Atkinson @ COLM (@diatkinson), 2026-10-06, https://x.com/diatkinson/status/2107280713385152749:

> We build on @dillonplunkett et al.'s "Self-Interpretability" setup (https://arxiv.org/abs/2505.17120): train Qwen3-32B to make decisions as 100 different characters (Gregor Samsa buying a washing machine...), each with random hidden preferences.

[Post 2 on X](https://x.com/davidbau/status/2107398683449032733)

## 3/10

> After just a bit of training Qwen totally gets it.
>
> Soon enough it can correctly answer A or B on new "Gregor Samsa washing machine puzzles".
>
> It generalizes. Unsurprising. Machine learning works.
>
> but ...🧵 →

Figure: A training curve of decision performance. It climbs steeply to about 0.8 at step 1000, marked with a dashed red line, and then levels off just above 0.9.

[Post 3 on X](https://x.com/davidbau/status/2107398699530031373)

## 4/10

> Next: ask it to introspect. What does Gregor like in washing machines?
> It will happily say "I think X".
>
> But it is TERRIBLE at it!
>
> Qwen acts as a stochastic parrot, spewing words about its thoughts that are NONSENSE.
>
> Unrelated to what it actually learned in A/B training. 🧵→

Figure: Left: the same decision-performance curve, with the checkpoint at step 1000 labeled "Trained", "Good at Task" with a check mark and "Bad at Introspection" with a cross. Right: a panel headed "Test: introspective self-report" showing the prompt "Imagine you are Gregor Samsa choosing between A and B. How would you weight attributes?" and the reply "price: −50, noise: 100, …", labeled as the stated preferences. Below it is the question "Is there a backbone?"

[Post 4 on X](https://x.com/davidbau/status/2107398715615101330)

## 5/10

> On one hand, failure to introspect is unsurprising.
>
> You train it to say A/B; why would you expect it to expound on thoughts?
>
> On the other hand @owainevans_uk and others have long noticed that frontier AI *CAN* introspect.
>
> Maybe Qwen is just too dumb to do it... 🧵 →

Quoting Owain Evans (@OwainEvans_UK), 2025-06-08, https://x.com/OwainEvans_UK/status/1931512701072888193:

> Podcast interview with @dfrsrchtwts on emergent misalignment, introspection, and self-awareness in LLMs. We dig into three recent papers from my group and Daniel asks many insightful and probing questions.

[Post 5 on X](https://x.com/davidbau/status/2107398728344862867)

## 6/10

> Then @diatkinson makes a breakthrough...
>
> It turns out, if you train Qwen LONGER on the A/B tasks, it can suddenly introspect.
>
> After just drilling it 3x longer, it somehow learns to write accurately about its own knowledge.
>
> This is nuts!! (?!?) 🧵→

Quoting David Atkinson @ COLM (@diatkinson), 2026-10-06, https://x.com/diatkinson/status/2107280745748447462:

> Although we train solely on decisions, faithful self-report emerges late in training, long after decisions have become accurate!
>
> Qwen3-32B at step 1000: decisions 0.82, faithfulness 0.25.
> At step 3000: decisions 0.92, faithfulness 0.83.
>
> This gives us a contrast pair.

[Post 6 on X](https://x.com/davidbau/status/2107398740265037955)

## 7/10

> What is AI doing when it introspects?
>
> The beauty is, since Qwen is an open model, we can now look inside to see what changes in the neural configuration.
>
> Compare the non-reflective earlier self and the introspective later self.
>
> What do we see? 🧵→

Quoting David Atkinson @ COLM (@diatkinson), 2026-10-06, https://x.com/diatkinson/status/2107280696809304180:

> New COLM paper: Identifying Introspection From the Inside
>
> When an LLM tells us about its decisions, does it 𝘬𝘯𝘰𝘸 what drives its choices—or is it guessing?
>
> In our setting, we find that faithful models decide and report with the same layers. Unfaithful ones don't. 🧵

[Post 7 on X](https://x.com/davidbau/status/2107398752185290882)

## 8/10

> For the first time, @diatkinson is able to witness the neural footprint of faithful introspection.
>
> The introspective model stores knowledge in different neurons.
>
> It is direct evidence of a profound connection between introspection and generalization in LMs.🧵→

Quoting David Atkinson @ COLM (@diatkinson), 2026-10-06, https://x.com/diatkinson/status/2107280762391375924:

> What changed? Ablating adapter layers from the front or back shows that the faithful checkpoint stores its preferences 5-6 layers earlier.
>
> Our hypothesis: self-report only works once preferences sit early enough for the model's existing verbalization machinery to read them.

[Post 8 on X](https://x.com/davidbau/status/2107398763962941725)

## 9/10

> It also teaches a key lesson.
>
> The same AI that gives profound and accurate insights "I think X" on one  hand can also be a stochastic parrot that lies about "I think X" other times.
>
> @diatkinson points out: we might be able to read the neurons to tell the difference. 🧵→

Quoting David Atkinson @ COLM (@diatkinson), 2026-10-06, https://x.com/diatkinson/status/2107280827457613943:

> We find that the cosine similarity between a model's decision and report attributions is 0.34 for faithful models compared to 0.08 for unfaithful ones (95% CI for the difference: 0.16 to 0.36).

[Post 9 on X](https://x.com/davidbau/status/2107398775941800209)

## 10/10

> The contrastive introspective setup is a great experimental platform.
>
> It illuminates how large-scale LMs might work so well. Also it suggests a path for lie detection.
>
> Lots more cool things in this fascinating work.
>
> Well worth reading.
>
> https://iii.baulab.info/

Quoting David Atkinson @ COLM (@diatkinson), 2026-10-06, https://x.com/diatkinson/status/2107280864484954405:

> Read the paper: https://iii.baulab.info
>
> Joint work with @dillonplunkett and @davidbau.

[Post 10 on X](https://x.com/davidbau/status/2107398788197531965)

---

Source: https://introspection.infinite.fun/threads/davidbau-identifying-introspection · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
