Thread
Aayush Karan on "Finetuning with Sampling"
The lead author walks through the paper in twelve posts: the open question of adding capabilities without forgetting, what on-policy learning has and what expert data has, the target distribution and the Metropolis-Hastings sampler that moves expert data toward it, the rise in the data's likelihood, the comparison with RL and OPSD, the pass@k result, and sampling as a tool for the rest of posttraining.
The posts are embedded from X. The figure notes under them are written by this wiki.
SFT is not dead! π₯³
— Aayush Karan (@aakaran31) October 2, 2026
We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and forgetting less than RL and OPSD. π€―
Following our prior work on reasoning with sampling, we now introduce sampling to the posttraining stack.
1/n pic.twitter.com/K8E7aqG7EjHow do we introduce fundamentally new capabilities to a base model without forgetting existing ones?
— Aayush Karan (@aakaran31) October 2, 2026
Between SFT and RL, a clear winner appears to have emerged.π₯
While RL exhibits strong generalization ποΈββοΈ, SFT is prone to memorization and catastrophic forgetting π΅βπ«.
2/nRL's success largely owes to on-policy learning, which lets a model learn from its own generations. But SFT has a key advantage!
— Aayush Karan (@aakaran31) October 2, 2026
On-policy learning must find successful traces via repeated sampling. With off-policy data, such traces are supplied as privileged information.
3/nWhat if we could reshape this privileged off-policy data to be as close to the base model as possible? Then SFT could make full use of its information signal while reaping the benefits of on-policy learning.
— Aayush Karan (@aakaran31) October 2, 2026
Turns out, this is the perfect job for sampling!
4/n pic.twitter.com/zfC65js3PWFigure. Figure 2 of the paper. Four bell curves in a row on one axis, from dark green on the left to pale green on the right, labeled Ο_base, Ο_target, Ο_int and Ο_data. Above Ο_base is the word "On-Policy". Above Ο_data is "Off-Policy", and a dotted arrow runs from it leftward to the word "Boosted" above Ο_target. Ο_base and Ο_target overlap; Ο_data is farthest from Ο_base.
We want to sample from the distribution closest in KL divergence to the base model that outputs trajectories consistent with privileged information constraint C.
— Aayush Karan (@aakaran31) October 2, 2026
We can explicitly write down this target distribution.
5/n pic.twitter.com/nvlTgslkHHFigure. An equation. On the left, p_C is defined as the distribution Ο in P_C that minimizes KL(Ο β₯ p). An arrow of implication leads to the right-hand side: p_C(x) is proportional to p(x) times the indicator that x is in C.
How do we sample from it? With MCMC algorithms of course! π
— Aayush Karan (@aakaran31) October 2, 2026
Recall Metropolis-Hastings (MCMC), an approximate sampler that iteratively updates a trajectory by proposing resampled candidates scored by their target (unnormalized) likelihoods.
We employ a neat trick that lets us⦠pic.twitter.com/8VOs2ZOR31In other words, MCMC lets us use off-policy data as an initialization that progressively becomes more on-policy throughout the MCMC process.
— Aayush Karan (@aakaran31) October 2, 2026
Plotting the trajectory likelihoods before and after sampling directly illustrates the likelihood-boosting effect.
7/n pic.twitter.com/CUJBwUD6OtFigure. Figure 3 of the paper. Two histograms of average log probability, each with a pale series labeled Off-Policy and a darker one labeled Boosted. Left, Chemistry: the off-policy traces are spread from β3.0 to about β0.5 with their bulk between β1.5 and β1.0, and the boosted traces sit in a narrow band to the right of about β0.75. Right, Math: the off-policy traces are spread from β2.0 to about β0.25 with their bulk near β0.9, and the boosted traces are piled up to the right of about β0.4, with the tallest bars near the right edge.
Remarkably, our sampling algorithm finally enables SFT to defy its weak characterization.
— Aayush Karan (@aakaran31) October 2, 2026
Composed with sampling, SFT frequently generalizes better and forgets less than strong on-policy learning baselines like RL and OPSD across multiple domains and base models.
8/n pic.twitter.com/AvwpVdt74tFigure. Figure 1 of the paper. Left, a bar chart of accuracy in percent for SFT, OPSD (On-Policy) and Ours (Sampling SFT) on three benchmarks. Chemistry: 61.8, 61.8, 66.0. MMLU: 58.6, 65.1, 69.2. AMC: 27.7, 34.9, 37.4. Right, a scatter plot of prior task retention in percent against new task accuracy in percent. The base model, Qwen2.5-7B-Instruct, is at the top left, at 100 on retention and below 40 on accuracy. SFT is lowest on retention, between 85 and 90. OPSD is at about 95 with the same accuracy as SFT. The point labeled Ours is to the right of and above both, between 95 and 100 on retention.
Sampling also facilitates learning beyond sharpening whatβs already in the base model!
— Aayush Karan (@aakaran31) October 2, 2026
We see large, consistent gaps where our pass@k curve outperforms both the base model and on-policy learning, suggesting we are able to introduce genuinely stronger abilities.
9/n pic.twitter.com/jztwATkqGCFigure. Figure 4 of the paper, titled "Chemistry (Olmo-3-7B-Instruct)": pass@k accuracy against k, for k of 1, 2, 4, 8, 16, 32 and 64, with three lines labeled Ours, OPSD and Base. All three rise with k. Base is lowest throughout. At k = 1 OPSD is above Ours; the two meet at k = 2; from k = 4 on Ours is highest, ending close to 1.0 at k = 64, with OPSD and Base below it and still rising.
Our sampling algorithm is a general primitive that shapes data for learnability.
— Aayush Karan (@aakaran31) October 2, 2026
Beyond SFT, much of posttraining hinges on finding model-native trajectories with useful signal. Sampling offers a way to deliberately construct them, for RL, OPD, self-play, and beyond.
10/nSo grateful to have been advised by @sitanch and @du_yilun on this work!! Please check out our paper, blog, and code for more details!
— Aayush Karan (@aakaran31) October 2, 2026
We'll also be presenting this at NeurIPS! π
Paper: https://t.co/si86RMo5wU
Blog: https://t.co/t45sB7zQWZ
Code: https://t.co/6CnSdUFY3k
11/nTagging some people that might be interested!@ArnaudDoucet1 @agarwl_ @natolambert @ShamKakade6 @aviral_kumar2 @lateinteraction @willcb @xkianteb @canondetortugas @oneill_c @HannaHajishirzi @giffmana @maximelabonne @sirbayes @Scobleizer @_lewtun @hbouammar @nrehiew_ @NoahZiemsβ¦
— Aayush Karan (@aakaran31) October 2, 2026