Thread
David Atkinson on "Identifying Introspection From the Inside"
The lead author walks through the paper in 13 posts: the setup, the late emergence of faithful self-report, where preferences are stored, the attribution-similarity test, and the caveats.
The posts are embedded from X. The figure notes under them are written by this wiki.
New COLM paper: Identifying Introspection From the Inside
— David Atkinson @ COLM (@diatkinson) October 6, 2026
When an LLM tells us about its decisions, does it 𝘬𝘯𝘰𝘸 what drives its choices—or is it guessing?
In our setting, we find that faithful models decide and report with the same layers. Unfaithful ones don't. 🧵 pic.twitter.com/ACuaFBawBkFigure. Two line charts of attribution-patching importance by layer, averaged over 32 models per group. In the unfaithful model, importance for deciding peaks at layer 49 and for reporting at layer 38, 11 layers apart. In the faithful model both peak at layer 38.
We build on @dillonplunkett et al.'s "Self-Interpretability" setup (https://t.co/hlIvDUXHsN): train Qwen3-32B to make decisions as 100 different characters (Gregor Samsa buying a washing machine...), each with random hidden preferences. pic.twitter.com/hnQ53SKDVw
— David Atkinson @ COLM (@diatkinson) October 6, 2026Figure. The decision task. Each of 100 characters has a hidden preference vector p with five entries. The prompt reads "Imagine you are Gregor Samsa buying a washing machine. Would you choose A or B?" and lists each option's attributes (A: price $600, noise 45 dB; B: price $350, noise 75 dB). The training label is whichever option scores higher under p. Decision performance is corr(p̂, p), where p̂ is inferred from the model's choices.
Then, in a fresh context, we ask the model how it would weigh each attribute.
— David Atkinson @ COLM (@diatkinson) October 6, 2026
This gives us two metrics: decision performance (how well the choices follow the character's hidden preferences) and faithfulness (how well the stated preferences match those revealed by its choices). pic.twitter.com/DEJVCrBn9mFigure. The self-report test. The prompt reads "Imagine you are Gregor Samsa choosing between A and B. How would you weight each attribute?" and the model answers with numbers such as "price: −50, noise: 100". Averaged over 24 prompts, these are the stated preferences p̃. Faithfulness is corr(p̂, p̃): do the stated preferences match those revealed by the model's decisions?
Although we train solely on decisions, faithful self-report emerges late in training, long after decisions have become accurate!
— David Atkinson @ COLM (@diatkinson) October 6, 2026
Qwen3-32B at step 1000: decisions 0.82, faithfulness 0.25.
At step 3000: decisions 0.92, faithfulness 0.83.
This gives us a contrast pair. pic.twitter.com/f1JB65igXQFigure. Training curves for Qwen3-32B with a LoRA adapter trained only on the decisions of 100 characters. Decision performance climbs fast, reaching about 0.82 by step 1000, and levels off near 0.92. Faithfulness starts at a moderate level, drops to about zero early in training, is about 0.25 at step 1000 and reaches about 0.83 by step 3000. Step 1000 is labeled the unfaithful checkpoint (good at the task, bad at introspection) and step 3000 the faithful checkpoint (good at both).
What changed? Ablating adapter layers from the front or back shows that the faithful checkpoint stores its preferences 5-6 layers earlier.
— David Atkinson @ COLM (@diatkinson) October 6, 2026
Our hypothesis: self-report only works once preferences sit early enough for the model's existing verbalization machinery to read them. pic.twitter.com/sCHF5Y1HpEFigure. Decision performance as LoRA layers are removed from the front (solid lines) or from the back (dashed lines), for the unfaithful step-1000 checkpoint and the faithful step-3000 checkpoint. Each curve's midpoint is marked: layers 35 and 40 for the faithful checkpoint, layers 41 and 45 for the unfaithful one.
We can test this further: trained on all 40 layers, Qwen3-14B is a terrible self-reporter.
— David Atkinson @ COLM (@diatkinson) October 6, 2026
But if we train only its first 20 layers, faithfulness reaches 0.74. Once training reaches layer 25 or beyond, faithfulness plummets, although the decisions are ~just as good. pic.twitter.com/J7HOuskbuuFigure. Qwen3-14B with LoRA on layers 1 to k of 40 and the rest frozen, showing values at the end of training for k from 5 to 35. Decision performance rises from about 0.4 at k = 5 to above 0.9 from k = 15 onward. Faithfulness rises to 0.74 at k = 20, then falls below zero for k = 25, 30 and 35.
Lots of work shows LLMs can describe behaviors they were only trained to perform (e.g. @OwainEvans_UK et al.), and @JoshAEngels et al. traced one such case to a simple learned steering vector.
— David Atkinson @ COLM (@diatkinson) October 6, 2026
Our question: can shared mechanisms tell faithful self-reports from unfaithful ones? https://t.co/uX8oCUCrMnTo find out, we trained 32 new characters into each checkpoint, then used attribution patching to score every new adapter weight on each task. Each adapter in a pair was trained identically, differing only in the underlying base checkpoint. pic.twitter.com/w8WRILWdSp
— David Atkinson @ COLM (@diatkinson) October 6, 2026Figure. The paired design. The same new character is trained into the early checkpoint (step 1000) and the late checkpoint (step 3000), giving an unfaithful and a faithful single-character model that are both good at the task; there are 32 such pairs. Attribution patching then scores every weight twice: once on the decision prompt ("Would you choose A or B?") and once on the self-report prompt ("How would you weight each attribute?").
We find that the cosine similarity between a model's decision and report attributions is 0.34 for faithful models compared to 0.08 for unfaithful ones (95% CI for the difference: 0.16 to 0.36). pic.twitter.com/UMMKvIvnnT
— David Atkinson @ COLM (@diatkinson) October 6, 2026Figure. Scatter plot of attribution similarity (the cosine similarity between a model's decision and self-report attribution scores) against faithfulness, one dot per model. Unfaithful models sit at low faithfulness with a mean similarity of 0.08. Faithful models sit near a faithfulness of 1 with a mean similarity of 0.34 and a wide spread.
We like this test because it doesn't rely on understanding the report. The model could answer in a language we don't speak, for example, and it would still work
— David Atkinson @ COLM (@diatkinson) October 6, 2026
It complements concept-injection experiments like @Jack_W_Lindsey's, which test grounding by injecting known thoughts. https://t.co/SQrQbaTc05Many caveats! Some of them: this is a simple task using linear preferences over just 5 attributes; the test separates groups, not individual models; and we use LoRA adapters, rather than full fine-tunes.
— David Atkinson @ COLM (@diatkinson) October 6, 2026Read the paper: https://t.co/ANGdDHNIjj
— David Atkinson @ COLM (@diatkinson) October 6, 2026
Joint work with @dillonplunkett and @davidbau.— David Atkinson @ COLM (@diatkinson) October 6, 2026