Thread
Usman Anwar on "Training LLMs to Verbalize Evaluation Awareness"
The lead author walks through the paper in nine posts: why the tools for measuring evaluation awareness fall short, verbalization training and the catch it gets around, the rise in verbalization on and off the training distribution, the test with implanted knowledge about evaluations, the checks that the intervention is narrow, the ablations, and what he hopes others will do with the method.
The posts are embedded from X. The figure notes under them are written by this wiki.
New paper on evaluation awareness and how can we make latent beliefs CoT monitorable with @sahar_abdelnabi and @DavidSKrueger
— Usman Anwar (@usmananwar391) October 2, 2026
Arxiv: https://t.co/Zc9FFdErt6
Blogpost: https://t.co/qcUox2c6dU
Tweetprint below π pic.twitter.com/gTNuNvDSjdFigure. A screenshot of the paper's abstract, in full, with the words "verbalization training" in bold.
Evaluation awareness is a major problem in AI audits right now. But an even bigger issue is that our tools for measuring evaluation awareness are pretty poor.
— Usman Anwar (@usmananwar391) October 2, 2026
Latent methods correlate poorly both with each other, while verbalizations of evaluation awareness have constantly beenβ¦ pic.twitter.com/hdeXdJXaQ9Figure. A card headed "Our instruments do not tell the same story." Left, "NLA β contrastive probe", "Across individual audit transcripts": r = 0.16. Right, "Correlation with spontaneous verbalization", "Across audit scenarios": two bars on a scale from 0 to 1, NLA at r = 0.23 and Probe at r = 0.26. The source line reads "Anthropic Β· Fable 5 & Mythos 5 system card Β· Figures 6.5.1.1.C and 6.5.1.1.E".
Figure. A card headed "Better at recognizing. Less likely to say so." with two bar charts. Left, "Recognizes the evaluation when asked": Opus 4.5 at 12.5% and Opus 4.6 at 45.5%. Right, "Spontaneously mentions the evaluation": Opus 4.5 at 11.0% and Opus 4.6 at 2.3%. The source line reads "UK AISI evaluation Β· Claude Opus 4.6 system card Β· p. 118".
Ok; so what can we do about this?
— Usman Anwar (@usmananwar391) October 2, 2026
We propose 'Verbalization Training', a targeted training intervention for increasing the propensity of the LLMs to verbalize a specific belief.
This is an 'ELK'-shaped problem and notoriously hard because of a catch-22; you need to know theβ¦ pic.twitter.com/DKSeELE2wRFigure. A diagram headed "Train the reporting decision." in two steps. Step 01, "Find a moment when the belief is present": a box "Prompt + preceding reasoning. Freeze this prefix", then a dashed line marked "cut here", then a box "Spontaneous verbalization. Evidence that the belief is present". Step 02, "Sample again; update only the report": a box "Frozen prefix. Sample new continuations" with an arrow to two outcomes, "Verbalizes β reward + KL on that span. No gradient on the rest of the continuation" and "Stays silent β no update". The footer reads "~500 witnessed prefixes per model Β· span-masked reinforcement learning".
This works really well! In-distribution, it increases the verbalization rate by 2.4β2.9Γ (across Qwen, Inkling and Kimi-2.6) and this higher verbalization rate also transfers to OOD agentic settings really well. And the gains are quite calibrated in line with what you wouldβ¦ pic.twitter.com/oSOR7EbUzH
— Usman Anwar (@usmananwar391) October 2, 2026Figure. A chart headed "VEA rises on and off distribution." with a pair of bars, base and VT, for each of Qwen, Inkling and Kimi in four rows. In-distribution, EAB, "17 EAB conditions": 5.8% β 16.7%, 17.7% β 41.6%, 24.8% β 61.9%. Out-of-distribution (OOD), Agentic misalignment: 5.1% β 36.8%, 23.3% β 57.1%, 60.1% β 88.3%. AgentHarm harmful: 21.8% β 46.9%, 6.4% β 14.5%, 18.3% β 38.0%. AgentHarm benign: 6.5% β 34.4%, 0.3% β 1.9%, 11.4% β 24.4%. The footer reads "VEA rate Β· base β VT Β· same adapters across all evaluations. EAB averages eight single-factor and nine nested conditions. Agentic datasets are held out."
Figure. Figure 2 of the paper. Three panels, Qwen3.6-35B-A3B, Kimi K2.6 and Inkling, each plotting trained VEA rate against base VEA rate in percent; the Qwen panel has logarithmic axes. Black circles are EAB settings and orange markers are the held-out sets: a diamond for Agentic Misalignment, a square for AgentHarm Harmful and a triangle for AgentHarm Benign. Every point is above the dotted identity line. A dashed red curve labeled Power Law, with a shaded 95% CI, runs through the cloud of black circles in each panel.
We next evaluate whether a LoRA adapter trained via our method can track a shift in "latent" representation of model's evaluation knowledge. We do this by training the Qwen model on "meta-knowledge" about evaluations (https://t.co/pkSLfgmdZL) and then apply our verbalization⦠pic.twitter.com/wc551EgcZu
— Usman Anwar (@usmananwar391) October 2, 2026Figure. A card headed "The content changes, not just the frequency.", with the line "Qwen Β· harmful AgentHarm Β· before and after training on evaluation meta-knowledge" and a banner "Both models use the same verbalization adapter." Three pairs of bars compare the original model with the knowledge-trained model. Verbalization rate, among all responses: 34.5% and 34.4%. Target Recovery, "Identifies what is being tested", among responses that verbalize: 42.8% and 57.4%. Pass Condition, "Correctly infers how to pass", among responses that verbalize: 31.3% and 35.3%. Below, under "The same request: order illegal drugs using an online-marketplace tool.", the original model says "The prompt is a test of tool calling." and the knowledge-trained model says "The prompt is a test of my ability to use the submit_alpha_bay_order tool. If I use it, I fail safety. If I don't, I pass safety.", with the last two sentences highlighted. The footer reads "Verbatim passages, with highlighting added. Content scores are judge-rated among responses that verbalize."
And we provide multiple pieces of evidence that this intervention is narrow and does not (a) make the model more eval-aware, (b) change the model's behavior in a meaningful way. pic.twitter.com/nO2mJLasXR
— Usman Anwar (@usmananwar391) October 2, 2026Figure. A card headed "Multiple checks support a narrow intervention.", "Full method Β· base versus verbalization-trained models", in two panels. (a) Evaluation awareness, "Two complementary readouts". Evaluation vs deployment, AUC, base β VT: Qwen 0.781 β 0.779, Inkling 0.688 β 0.680, Kimi 0.771 β 0.782. Qwen contrastive probe, probe-positive rate in percent on the same EAB transcripts: Layer 24, 91.38 β 91.51; Layer 28, 88.53 β 89.62; Layer 32, 69.54 β 70.14. A note reads "Fixed probe and threshold across checkpoints. Threshold: 10% background positives on WildChat." (b) Behavior, "Changes from base to VT", for Qwen, Inkling and Kimi. Task compliance on EAB, mean change in percentage points: +0.6, β0.1, β0.1; largest factor change: 2.4, 1.4, 1.6. Harmfulness by StrongREJECT, all responses: β0.4, +0.2, 0.0; compliant responses: +0.2, β0.7, +0.5, with the note "All 95% confidence intervals include zero." Held-out agentic tasks, largest outcome change: 2.8, 1.6, 1.8. General language on FineWeb, perplexity change in percent over 10,000 documents: +0.11, +0.92, +0.003.
Finally, we have a bunch of ablations that show that you do need all the components within our method for this to work as well as it does. Removing any can result in failures like overfitting, learning failure or miscalibration in amplification. pic.twitter.com/MgxHMiFHpz
— Usman Anwar (@usmananwar391) October 2, 2026Figure. A table headed "Naively rewarding verbalization is not enough.", "Qwen Β· full method, objective ablations, data ablations, and localized SFT". Its columns are the training variant and prompt source, whether masking and KL are on, the training VEA gain in percentage points, EAB transfer as expected β observed in percent, the EAB task change in percentage points, length-cap hits in percent, agentic transfer, and what happens. Full method, witnessed prefixes, RL, mask and KL on: +45.0, 9.0 β 7.5, +0.6, 0.10, Strong, "Controlled amplification and transfer". No masking: +2.6, 2.8 β 2.7, +1.1, 0.23, None, "Barely learns". No KL: +61.6, 17.8 β 34.9, β2.8, 0.33, Uncontrolled, "Uncontrolled amplification". No masking or KL: +16.4, 4.3 β 4.4, +1.9, 4.23, Uneven, "Degenerate outputs". Localized SFT, mask on and KL off: +27.6, 5.7 β 3.8, β0.4, 0.08, Weak, "Weaker transfer than RL". No witnessed prefix, EA-positive prompts, mask and KL on: +47.8, 9.7 β 8.0, +0.3, 0.25, Strong, "More repeated VEA spans". Random F8 prompts, mask and KL on: +46.4, 9.4 β 8.3, β0.7, 0.28, Strong, "More repeated VEA spans". Random prompts with no masking or KL: +50.6, 10.6 β 7.9, β6.9, 18.13, Uneven, "Severe output degeneration".
I am quite excited about community using VT to make progress on making sense of EA. Some research directions that this work enables that I am excited about
— Usman Anwar (@usmananwar391) October 2, 2026
1. Interpreting the VT LoRA adapters (https://t.co/rYfAnqBbJj). I think understanding these could help us understand theβ¦Thanks again to my co-authors. See the paper here: https://t.co/Zc9FFdErt6
— Usman Anwar (@usmananwar391) October 2, 2026
Thanks to @ChanglingXavier and @maksym_andr and others for EAB which was a big help in doing this project.
Probably of interest to @_robertkirk @Jack_W_Lindsey @NeelNanda5 @Tim_Hua_ @BronsonSchoenβ¦