Thread

Adithya Bhaskar on "Language Models that Play Chess and Explain Their Moves"

The lead author walks through the paper in 10 posts: why explaining chess is hard to learn, giving the Leela chess network a language model to speak through, training the bridge between them on question-answer pairs, seeding with explanations from GPT-5.6-Sol, the natural-language analog of AlphaZero, the results, and a variant started from templates.

The posts are embedded from X. The figure notes under them are written by this wiki.

  1. Figure. The first page of the paper: its title, the authors Adithya Bhaskar, Jeffrey Cheng and Danqi Chen of Princeton Language and Intelligence, links labeled Models, Code and Website, and the abstract.

  2. Figure. Two panels. Left, "Encoded Position": a chess board with a middlegame position. Right, "QUEEN Output": part of an explanation in prose, with the moves in blue. It reads in part: "The key strategic question is how the player can consolidate the central space advantage and complete development before the opponent activates the light-squared bishop. The best move is 13. Qc2, centralizing the queen on a useful diagonal where it supports the center and prepares to connect the rooks. The opponent's most accurate reply is 13... Bd6, activating the bishop to a strong central square where it eyes h2 and supports the center. The player's most effective continuation is 14. Rad1".

  3. Figure. A diagram of the architecture. On the left a chess position goes into Lc0, marked with a snowflake, which gives a stack of board-shaped representations labeled "Layer k". In the middle, in a box labeled SmolLM3-3B and repeated N times, a "Flamingo block" takes K and V from Lc0 and Q and a residual from the text tokens "Analyze this chess position .", and feeds an "LM block" labeled "Layer 2k"; the tokens that come out read "Black has weak# #ened their ...". On the right the two blocks are opened up: the LM block is self-attention then feed-forward, each with a residual connection; the Flamingo block is cross-attention then feed-forward, each followed by a tanh gate and a residual connection.

  4. Figure. Two chess boards and four example questions with their answers. Static current (L): "List all pieces on the g- file." A: "White pawn on g2, white bishop on g5, black pawn on g7, and black king on g8." Static future (L): "After white bishop to h4, list all pieces on the 4-th rank," A: "White pawn on d4 and white bishop on h4." Dynamic current (R): "What pieces can deliver a check to the black king?" A: "The white bishop on d3 via bishop d3 to h7 with check." Dynamic future (R): "After white pawn e3 to e4, list all pieces that attack or defend d5." A: "The white knight on c3 and white pawn on e4 attack it, while the black queen on d8, the black knight on f6 and the black pawn on c6 defend it." On the left board the g-file and the fourth rank are highlighted. On the right board arrows mark the bishop's path from d3 to h7 and the pieces bearing on d5.

  5. Figure. A root chess position on the left and, to its right under the word "Consolidate", the three positions reached from it by Rh2, Qh2+ and Qxf3, with an arrow from the three back to the root. Each of the three has a short explanation. After Rh2: "White is up a pawn and an exchange, but faces unavoidable mate... Best move: Qe6+ // Eval: -100.0". After Qh2+: "White is facing a scary check, but the king can escape via f1-e2... Best move: Kf1 // Eval: -0.1". After Qxf3: "White is down a bishop for just a pawn, but can force a repetition of moves... Best move: Qe6+ // Eval: 0.0". Under the root: "Consolidated explanation: White's king is exposed. The tempting check Qh2+ lets the king escape.... though Qxf3 wins a whole rook, white can force a repetition.... finally, the quiet move Rh2! sets up unavoidable mate on g2... Best move: Rh2. Bellman update: Eval = min(-100.0, -0.1, 0.0) = -100.0".

  6. Figure. A line chart of Elo, from 1,800 to 2,800, against the numbers 1 to 8, for "Queen (ours)". The line starts at 1782, rises at every step to the fifth point, dips at the sixth, rises again and ends at 2697. Dashed horizontal lines mark frontier models at high reasoning effort: Luna 1822, Sol 2071, Gemini 2201. Dotted lines mark "Median IM 2560" and "Median GM 2730". Queen's line is above Luna from its second point, above Sol and at about Gemini's level at its third, at about the Median IM line at its fifth and seventh points, and just under the Median GM line at its last.

  7. Figure. A screenshot from the paper: the paragraph headed "Moving away from frontier LM annotation." beside Table 6, "Performance of QUEEN with and without frontier-model distillation." The table gives Elo at P1, P2, P3 and P4: QUEEN 1782, 2024, 2187, 2346; QUEEN (HCE) 2115, 2432, 2476, 2497.

  8. Figure. A screenshot of an explanation headed "QUEEN (HCE)". It opens with a paragraph on the position: "White holds a modest spatial advantage in a closed-center structure." It then takes candidate moves in numbered parts, each in prose with replies and alternatives: "I. Qc2 | the best move. The queen supports the e4 break, keeps the white bishop on f2 flexible, and avoids committing the rook to the d-file prematurely." and "II. Rc1 | a solid developing move that connects the rooks and prepares e4, but it is slightly less flexible than I".