# Aayush Karan answers a question about the code for "Finetuning with Sampling"

> Alex Nichol posts that the paper's released code does not match the paper. The lead author replies that the code swaps the Metropolis-Hastings acceptance step for a rule that accepts any candidate of higher likelihood, and checks correctness after sampling instead of inside it; he gives the cost that led to each choice, points to the KL curve of Figure 5 as the check, and calls the result an existence proof.

- Author: Alex Nichol ([@unixpickle](https://x.com/unixpickle))
- Posted: 2026-10-05, 2 posts
- Original: https://x.com/unixpickle/status/2107251102282584525
- About: [Finetuning with Sampling: SFT Learns Better Than You Think](https://introspection.infinite.fun/papers/karan2026-finetuning-with-sampling.md)

The post text below is quoted verbatim. Figure descriptions are written by this wiki.

## 1/2

> Everybody was raving about this "Fine-tuning with Sampling," paper over the past few days. I gave it a read and either I (and chat) are missing something, or something is a bit off. First, the code doesn't match the paper. https://arxiv.org/pdf/2610.02140

Figure: A screenshot of a chat assistant's answer, headed "There is actually an even bigger implementation/theory discrepancy". It says that in the file boost_sci.py of the official code the inner loop does not appear to implement the acceptance ratio of Algorithm 1: it generates a candidate conditioned on the expert solution, computes its average log probability under the base model, and accepts it only when that is higher than the current trajectory's, with no reverse-proposal ratio and no randomized acceptance step. It adds that whether the final chemistry response is correct is computed after projection sampling, when the output is written, and is not used to reject proposals.

[Post 1 on X](https://x.com/unixpickle/status/2107251102282584525)

## 2/2

> Hey! Thanks for reading our paper!
>
> Our appendix has a brief discussion explaining why the code approximates the exact MH implementation and why this approximation is reasonable, so I'll elaborate here.
>
> Since our proposal prompt inputs the current partial trace, to calculate the acceptance probability for MH from x to x', you need to calculate the reverse transition log-probability from x' back to x, which requires a separate forward pass from the one that generates x' from x. Admittedly, our implementation for MCMC is expensive as it currently isn't batched, so instead of one LLM call per MCMC step, having two LLM calls becomes a bit costly. So as a first approximation, we used a deterministic likelihood-improving acceptance rule.
>
> To validate this inexactness, we can still estimate the KL divergence between this process that generates the boosted traces and the base model. I've attached Figure 5 from our paper here, but this shows the KL divergence gap declines with more MCMC steps, so crucially the desired effect of the trajectories becoming progressively more on-policy is maintained.
>
> A similar story is the case for the enforcement of the hard constraint. The underlying assumption is that instruction-following enforces the hard constraint in-context, but again, empirically this is inexact. A more exact implementation would have a consistency check (i.e. small LLM check on partial trace ~ privileged info consistency), but again, repeated API calls can be costly. As a way around this, we had post-generation filtering, where rerun on incorrect MCMC traces outside instead of inside the loop, which results in around 95% correctness on average.
>
> So in summary, these approximations maintain the effect we care about, which is much more on-policy trajectories with consistency well sustained. In other words, an "existence proof" that an algorithm like this is worth doing.
>
> There is most certainly room for much better practical efficiency; the approximations we made were **one** route to making the algorithm doable, but incorporating batching and partial-trace independent proposals is one step towards undoing this inexactness and making our current algorithm run a lot faster.

Figure: Figure 5 of the paper: two lines against the number of MCMC steps, at 0, 2, 4, 6, 8 and 10. A dashed line, KL in nats on the left axis, falls at every step, steeply from 0 to 2 steps and more slowly after. A solid line, accuracy in percent on the right axis, which runs from 61 to 67, rises from 0 to 4 steps, falls at 6 and 8, and is highest at 10.

[Post 2 on X](https://x.com/aakaran31/status/2107318691952030051)

---

Source: https://introspection.infinite.fun/threads/aakaran31-finetuning-with-sampling-reply · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
