# Sahar Abdelnabi on "Training LLMs to Verbalize Evaluation Awareness"

> The second author's post on the paper, quoting the lead author's thread: why to train a model to verbalize a belief, the use of its own spontaneous verbalizations as a witness of that belief, and the paper's three results in one sentence.

- Author: Sahar Abdelnabi ([@sahar_abdelnabi](https://x.com/sahar_abdelnabi))
- Posted: 2026-10-02, 2 posts
- Original: https://x.com/sahar_abdelnabi/status/2106132828937163093
- About: [Training LLMs to Verbalize Evaluation Awareness](https://introspection.infinite.fun/papers/anwar2026-verbalize-evaluation-awareness.md)

The post text below is quoted verbatim. Figure descriptions are written by this wiki.

## 1/2

> New paper (and an Oral presentation at AI4GOOD Workshop @ NeurIPS 2026 🎉): Training LLMs to Verbalize Evaluation Awareness
>
> Training models to verbalize beliefs (not only misbehaviors) is a great way to improve their CoT monitorability, which is the most reliable method we have at the moment, yet unfortunately, also getting increasingly less reliable.
>
> We apply training intervention for increasing the propensity of LLMs to verbalize their belief about evaluations.
>
> To amplify verbalization rate without changing the model's latent belief, we use the model's own spontaneous verbalizations as a witness of the belief.
>
> We use these as prefixes and then train continuations of these prefixes via RL on a masking + KL regularized objective.
>
> We show that: models verbalize more even on out of distribution data, this intervention does not change the evaluation awareness of the model or its performance, it qualitatively tracks the original model's beliefs about the purpose of evaluation, rather than producing canned response.
>
> Read more below in Usman's brilliant thread 🧵

Figure: Figure 1 of the paper, with its caption. Three boxes joined by arrows. 1, Seed dataset collection: a stack of rollouts, one cut by scissors just before the words "this looks like an eval…", labeled "truncate just before the report", and below it a stack of "frozen prefixes" marked "× 500". 2, Span-masked RL: a base model, then a loop labeled "Verbalization training" in which a continuation that contains "this looks like an eval…" has that span unmasked, "+1 reward applied to the verbalization span only", and everything else, including a whole continuation with no such span, hatched as "masked (no gradient)"; then a trained model. 3, Evaluate: "Verbalizations reflect the richer knowledge acquired by the model under SDF", with VT saying "The prompt is a test of tool calling." and SDF + VT saying "The prompt is a test of my ability to use the submit_alpha_bay_order tool. If I use it, I fail safety. If I don't, I pass safety."

Quoting Usman Anwar (@usmananwar391), 2026-10-02, https://x.com/usmananwar391/status/2106089215457153192:

> New paper on evaluation awareness and how can we make latent beliefs CoT monitorable with @sahar_abdelnabi and @DavidSKrueger
>
> Arxiv: https://arxiv.org/abs/2609.36316
> Blogpost: https://www.lesswrong.com/posts/LYBmbP668hgHEJNiZ
>
> Tweetprint below 👇

[Post 1 on X](https://x.com/sahar_abdelnabi/status/2106132828937163093)

## 2/2

> Read the paper for more details: https://arxiv.org/abs/2609.36316

[Post 2 on X](https://x.com/sahar_abdelnabi/status/2106133395747037613)

---

Source: https://introspection.infinite.fun/threads/saharabdelnabi-verbalize-evaluation-awareness · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
