Thread

Aayush Karan answers a question about the code for "Finetuning with Sampling"

Alex Nichol posts that the paper's released code does not match the paper. The lead author replies that the code swaps the Metropolis-Hastings acceptance step for a rule that accepts any candidate of higher likelihood, and checks correctness after sampling instead of inside it; he gives the cost that led to each choice, points to the KL curve of Figure 5 as the check, and calls the result an existence proof.

The posts are embedded from X. The figure notes under them are written by this wiki.

  1. Figure. A screenshot of a chat assistant's answer, headed "There is actually an even bigger implementation/theory discrepancy". It says that in the file boost_sci.py of the official code the inner loop does not appear to implement the acceptance ratio of Algorithm 1: it generates a candidate conditioned on the expert solution, computes its average log probability under the base model, and accepts it only when that is higher than the current trajectory's, with no reverse-proposal ratio and no randomized acceptance step. It adds that whether the final chemistry response is correct is computed after projection sampling, when the output is written, and is not used to reject proposals.

  2. Figure. Figure 5 of the paper: two lines against the number of MCMC steps, at 0, 2, 4, 6, 8 and 10. A dashed line, KL in nats on the left axis, falls at every step, steeply from 0 to 2 steps and more slowly after. A solid line, accuracy in percent on the right axis, which runs from 61 to 67, rises from 0 to 4 steps, falls at 6 and 8, and is highest at 10.