Thread
David Bau on "Identifying Introspection From the Inside"
The senior author retells the paper as a story in ten posts: teach a model something, find it cannot describe what it learned, train it longer until it can, then look inside to see what changed.
The posts are embedded from X. The figure notes under them are written by this wiki.
The craziest experiment in my lab is a new setup by @diatkinson where he has found a way induce introspection on open LMs using @dillonplunkett's protocol, that lets him look inside introspection. It's nuts.
— David Bau (@davidbau) October 6, 2026
Here's how the experiment works... 🧵 → pic.twitter.com/ylwm8mIN7hFigure. A cartoon robot saying "For this choice, price matters more than noise." Underneath is the question: "Is that report true?"
First, teach Qwen something new.
— David Bau (@davidbau) October 6, 2026
David uses SFT to train the LM to answer "A" or "B" on made-up trivia about famous people until the LM learns the trivia.
Here it learns Gregor Samsa likes cheap washing machines even if they are noisy. Purely learning to say A or B. 🧵→ https://t.co/8aNyAe5v8iAfter just a bit of training Qwen totally gets it.
— David Bau (@davidbau) October 6, 2026
Soon enough it can correctly answer A or B on new "Gregor Samsa washing machine puzzles".
It generalizes. Unsurprising. Machine learning works.
but ...🧵 → pic.twitter.com/Ly8U9tc1fxFigure. A training curve of decision performance. It climbs steeply to about 0.8 at step 1000, marked with a dashed red line, and then levels off just above 0.9.
Next: ask it to introspect. What does Gregor like in washing machines?
— David Bau (@davidbau) October 6, 2026
It will happily say "I think X".
But it is TERRIBLE at it!
Qwen acts as a stochastic parrot, spewing words about its thoughts that are NONSENSE.
Unrelated to what it actually learned in A/B training. 🧵→ pic.twitter.com/bET8R8dUaUFigure. Left: the same decision-performance curve, with the checkpoint at step 1000 labeled "Trained", "Good at Task" with a check mark and "Bad at Introspection" with a cross. Right: a panel headed "Test: introspective self-report" showing the prompt "Imagine you are Gregor Samsa choosing between A and B. How would you weight attributes?" and the reply "price: −50, noise: 100, …", labeled as the stated preferences. Below it is the question "Is there a backbone?"
On one hand, failure to introspect is unsurprising.
— David Bau (@davidbau) October 6, 2026
You train it to say A/B; why would you expect it to expound on thoughts?
On the other hand @owainevans_uk and others have long noticed that frontier AI *CAN* introspect.
Maybe Qwen is just too dumb to do it... 🧵 → https://t.co/oDXf7mDwmqThen @diatkinson makes a breakthrough...
— David Bau (@davidbau) October 6, 2026
It turns out, if you train Qwen LONGER on the A/B tasks, it can suddenly introspect.
After just drilling it 3x longer, it somehow learns to write accurately about its own knowledge.
This is nuts!! (?!?) 🧵→ https://t.co/HUsFaesqNrWhat is AI doing when it introspects?
— David Bau (@davidbau) October 6, 2026
The beauty is, since Qwen is an open model, we can now look inside to see what changes in the neural configuration.
Compare the non-reflective earlier self and the introspective later self.
What do we see? 🧵→ https://t.co/fFWbtm5qELFor the first time, @diatkinson is able to witness the neural footprint of faithful introspection.
— David Bau (@davidbau) October 6, 2026
The introspective model stores knowledge in different neurons.
It is direct evidence of a profound connection between introspection and generalization in LMs.🧵→ https://t.co/Xjtp5clJB1It also teaches a key lesson.
— David Bau (@davidbau) October 6, 2026
The same AI that gives profound and accurate insights "I think X" on one hand can also be a stochastic parrot that lies about "I think X" other times.@diatkinson points out: we might be able to read the neurons to tell the difference. 🧵→ https://t.co/VlO5sGYuKQThe contrastive introspective setup is a great experimental platform.
— David Bau (@davidbau) October 6, 2026
It illuminates how large-scale LMs might work so well. Also it suggests a path for lie detection.
Lots more cool things in this fascinating work.
Well worth reading.https://t.co/3MVMieblRH https://t.co/0WGt7hZ78H