# Aayush Karan on "Finetuning with Sampling"

> The lead author walks through the paper in twelve posts: the open question of adding capabilities without forgetting, what on-policy learning has and what expert data has, the target distribution and the Metropolis-Hastings sampler that moves expert data toward it, the rise in the data's likelihood, the comparison with RL and OPSD, the pass@k result, and sampling as a tool for the rest of posttraining.

- Author: Aayush Karan ([@aakaran31](https://x.com/aakaran31))
- Posted: 2026-10-02, 12 posts
- Original: https://x.com/aakaran31/status/2106037829059133903
- About: [Finetuning with Sampling: SFT Learns Better Than You Think](https://introspection.infinite.fun/papers/karan2026-finetuning-with-sampling.md)

The post text below is quoted verbatim. Figure descriptions are written by this wiki.

## 1/12

> SFT is not dead! 🥳
>
> We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and forgetting less than RL and OPSD. 🤯
>
> Following our prior work on reasoning with sampling, we now introduce sampling to the posttraining stack.
>
> 1/n

[Post 1 on X](https://x.com/aakaran31/status/2106037829059133903)

## 2/12

> How do we introduce fundamentally new capabilities to a base model without forgetting existing ones?
>
> Between SFT and RL, a clear winner appears to have emerged.🥇
>
> While RL exhibits strong generalization 🏋️‍♂️, SFT is prone to memorization and catastrophic forgetting 😵‍💫.
>
> 2/n

[Post 2 on X](https://x.com/aakaran31/status/2106037841365193111)

## 3/12

> RL's success largely owes to on-policy learning, which lets a model learn from its own generations. But SFT has a key advantage!
>
> On-policy learning must find successful traces via repeated sampling. With off-policy data, such traces are supplied as privileged information.
>
> 3/n

[Post 3 on X](https://x.com/aakaran31/status/2106037853230948457)

## 4/12

> What if we could reshape this privileged off-policy data to be as close to the base model as possible? Then SFT could make full use of its information signal while reaping the benefits of on-policy learning.
>
> Turns out, this is the perfect job for sampling!
>
> 4/n

Figure: Figure 2 of the paper. Four bell curves in a row on one axis, from dark green on the left to pale green on the right, labeled π_base, π_target, π_int and π_data. Above π_base is the word "On-Policy". Above π_data is "Off-Policy", and a dotted arrow runs from it leftward to the word "Boosted" above π_target. π_base and π_target overlap; π_data is farthest from π_base.

[Post 4 on X](https://x.com/aakaran31/status/2106037874642796749)

## 5/12

> We want to sample from the distribution closest in KL divergence to the base model that outputs trajectories consistent with privileged information constraint C.
>
> We can explicitly write down this target distribution.
>
> 5/n

Figure: An equation. On the left, p_C is defined as the distribution π in P_C that minimizes KL(π ∥ p). An arrow of implication leads to the right-hand side: p_C(x) is proportional to p(x) times the indicator that x is in C.

[Post 5 on X](https://x.com/aakaran31/status/2106037889893319158)

## 6/12

> How do we sample from it? With MCMC algorithms of course! 🙂
>
> Recall Metropolis-Hastings (MCMC), an approximate sampler that iteratively updates a trajectory by proposing resampled candidates scored by their target (unnormalized) likelihoods.
>
> We employ a neat trick that lets us absorb the information constraint into the proposer. This way, we switch from using the constraint as a verifier to using it in-context to generate candidates, which are then refined according to base model likelihoods.
>
> The data processing inequality then tells us that as the MCMC process progresses, our effective trajectory distribution gets closer and closer in KL divergence to the base model.
>
> 6/n

[Post 6 on X](https://x.com/aakaran31/status/2106037933065298236)

## 7/12

> In other words, MCMC lets us use off-policy data as an initialization that progressively becomes more on-policy throughout the MCMC process.
>
> Plotting the trajectory likelihoods before and after sampling directly illustrates the likelihood-boosting effect.
>
> 7/n

Figure: Figure 3 of the paper. Two histograms of average log probability, each with a pale series labeled Off-Policy and a darker one labeled Boosted. Left, Chemistry: the off-policy traces are spread from −3.0 to about −0.5 with their bulk between −1.5 and −1.0, and the boosted traces sit in a narrow band to the right of about −0.75. Right, Math: the off-policy traces are spread from −2.0 to about −0.25 with their bulk near −0.9, and the boosted traces are piled up to the right of about −0.4, with the tallest bars near the right edge.

[Post 7 on X](https://x.com/aakaran31/status/2106037958184927675)

## 8/12

> Remarkably, our sampling algorithm finally enables SFT to defy its weak characterization.
>
> Composed with sampling, SFT frequently generalizes better and forgets less than strong on-policy learning baselines like RL and OPSD across multiple domains and base models.
>
> 8/n

Figure: Figure 1 of the paper. Left, a bar chart of accuracy in percent for SFT, OPSD (On-Policy) and Ours (Sampling SFT) on three benchmarks. Chemistry: 61.8, 61.8, 66.0. MMLU: 58.6, 65.1, 69.2. AMC: 27.7, 34.9, 37.4. Right, a scatter plot of prior task retention in percent against new task accuracy in percent. The base model, Qwen2.5-7B-Instruct, is at the top left, at 100 on retention and below 40 on accuracy. SFT is lowest on retention, between 85 and 90. OPSD is at about 95 with the same accuracy as SFT. The point labeled Ours is to the right of and above both, between 95 and 100 on retention.

[Post 8 on X](https://x.com/aakaran31/status/2106037979370357187)

## 9/12

> Sampling also facilitates learning beyond sharpening what’s already in the base model!
>
> We see large, consistent gaps where our pass@k curve outperforms both the base model and on-policy learning, suggesting we are able to introduce genuinely stronger abilities.
>
> 9/n

Figure: Figure 4 of the paper, titled "Chemistry (Olmo-3-7B-Instruct)": pass@k accuracy against k, for k of 1, 2, 4, 8, 16, 32 and 64, with three lines labeled Ours, OPSD and Base. All three rise with k. Base is lowest throughout. At k = 1 OPSD is above Ours; the two meet at k = 2; from k = 4 on Ours is highest, ending close to 1.0 at k = 64, with OPSD and Base below it and still rising.

[Post 9 on X](https://x.com/aakaran31/status/2106037997657559373)

## 10/12

> Our sampling algorithm is a general primitive that shapes data for learnability.
>
> Beyond SFT, much of posttraining hinges on finding model-native trajectories with useful signal. Sampling offers a way to deliberately construct them, for RL, OPD, self-play, and beyond.
>
> 10/n

[Post 10 on X](https://x.com/aakaran31/status/2106038010232103301)

## 11/12

> So grateful to have been advised by @sitanch and @du_yilun on this work!! Please check out our paper, blog, and code for more details!
>
> We'll also be presenting this at NeurIPS! 🙂
>
> Paper: https://arxiv.org/abs/2610.02140
> Blog: https://aakaran.github.io/finetuning_with_sampling/
> Code: https://github.com/aakaran/finetuning-with-sampling
>
> 11/n

[Post 11 on X](https://x.com/aakaran31/status/2106038021921579450)

## 12/12

> Tagging some people that might be interested!
>
> @ArnaudDoucet1 @agarwl_ @natolambert @ShamKakade6 @aviral_kumar2 @lateinteraction @willcb @xkianteb @canondetortugas @oneill_c @HannaHajishirzi @giffmana @maximelabonne @sirbayes @Scobleizer @_lewtun @hbouammar @nrehiew_ @NoahZiems @kalomaze @IdanShenfeld @jonashubotter @KempeLab @archit_sharma97 @b_geist

[Post 12 on X](https://x.com/aakaran31/status/2106038033950834738)

---

Source: https://introspection.infinite.fun/threads/aakaran31-finetuning-with-sampling · Part of the [LLM Introspection Wiki](https://introspection.infinite.fun/index.md) · Index for agents: [llms.txt](https://introspection.infinite.fun/llms.txt)
