Thread

Sahar Abdelnabi on "Training LLMs to Verbalize Evaluation Awareness"

The second author's post on the paper, quoting the lead author's thread: why to train a model to verbalize a belief, the use of its own spontaneous verbalizations as a witness of that belief, and the paper's three results in one sentence.

The posts are embedded from X. The figure notes under them are written by this wiki.

  1. Figure. Figure 1 of the paper, with its caption. Three boxes joined by arrows. 1, Seed dataset collection: a stack of rollouts, one cut by scissors just before the words "this looks like an eval…", labeled "truncate just before the report", and below it a stack of "frozen prefixes" marked "× 500". 2, Span-masked RL: a base model, then a loop labeled "Verbalization training" in which a continuation that contains "this looks like an eval…" has that span unmasked, "+1 reward applied to the verbalization span only", and everything else, including a whole continuation with no such span, hatched as "masked (no gradient)"; then a trained model. 3, Evaluate: "Verbalizations reflect the richer knowledge acquired by the model under SDF", with VT saying "The prompt is a test of tool calling." and SDF + VT saying "The prompt is a test of my ability to use the submit_alpha_bay_order tool. If I use it, I fail safety. If I don't, I pass safety."