Thread
Sahar Abdelnabi on "Training LLMs to Verbalize Evaluation Awareness"
The second author's post on the paper, quoting the lead author's thread: why to train a model to verbalize a belief, the use of its own spontaneous verbalizations as a witness of that belief, and the paper's three results in one sentence.
The posts are embedded from X. The figure notes under them are written by this wiki.
New paper (and an Oral presentation at AI4GOOD Workshop @ NeurIPS 2026 🎉): Training LLMs to Verbalize Evaluation Awareness
— Sahar Abdelnabi 🕊 (@sahar_abdelnabi) October 2, 2026
Training models to verbalize beliefs (not only misbehaviors) is a great way to improve their CoT monitorability, which is the most reliable method we have… https://t.co/mVsgASom4X pic.twitter.com/h2b7Do7d9tFigure. Figure 1 of the paper, with its caption. Three boxes joined by arrows. 1, Seed dataset collection: a stack of rollouts, one cut by scissors just before the words "this looks like an eval…", labeled "truncate just before the report", and below it a stack of "frozen prefixes" marked "× 500". 2, Span-masked RL: a base model, then a loop labeled "Verbalization training" in which a continuation that contains "this looks like an eval…" has that span unmasked, "+1 reward applied to the verbalization span only", and everything else, including a whole continuation with no such span, hatched as "masked (no gradient)"; then a trained model. 3, Evaluate: "Verbalizations reflect the richer knowledge acquired by the model under SDF", with VT saying "The prompt is a test of tool calling." and SDF + VT saying "The prompt is a test of my ability to use the submit_alpha_bay_order tool. If I use it, I fail safety. If I don't, I pass safety."
Read the paper for more details: https://t.co/MHO6AC8mNy
— Sahar Abdelnabi 🕊 (@sahar_abdelnabi) October 2, 2026