Thread

Aayush Karan on "Finetuning with Sampling"

The lead author walks through the paper in twelve posts: the open question of adding capabilities without forgetting, what on-policy learning has and what expert data has, the target distribution and the Metropolis-Hastings sampler that moves expert data toward it, the rise in the data's likelihood, the comparison with RL and OPSD, the pass@k result, and sampling as a tool for the rest of posttraining.

The posts are embedded from X. The figure notes under them are written by this wiki.

  1. Figure. Figure 2 of the paper. Four bell curves in a row on one axis, from dark green on the left to pale green on the right, labeled Ο€_base, Ο€_target, Ο€_int and Ο€_data. Above Ο€_base is the word "On-Policy". Above Ο€_data is "Off-Policy", and a dotted arrow runs from it leftward to the word "Boosted" above Ο€_target. Ο€_base and Ο€_target overlap; Ο€_data is farthest from Ο€_base.

  2. Figure. An equation. On the left, p_C is defined as the distribution Ο€ in P_C that minimizes KL(Ο€ βˆ₯ p). An arrow of implication leads to the right-hand side: p_C(x) is proportional to p(x) times the indicator that x is in C.

  3. Figure. Figure 3 of the paper. Two histograms of average log probability, each with a pale series labeled Off-Policy and a darker one labeled Boosted. Left, Chemistry: the off-policy traces are spread from βˆ’3.0 to about βˆ’0.5 with their bulk between βˆ’1.5 and βˆ’1.0, and the boosted traces sit in a narrow band to the right of about βˆ’0.75. Right, Math: the off-policy traces are spread from βˆ’2.0 to about βˆ’0.25 with their bulk near βˆ’0.9, and the boosted traces are piled up to the right of about βˆ’0.4, with the tallest bars near the right edge.

  4. Figure. Figure 1 of the paper. Left, a bar chart of accuracy in percent for SFT, OPSD (On-Policy) and Ours (Sampling SFT) on three benchmarks. Chemistry: 61.8, 61.8, 66.0. MMLU: 58.6, 65.1, 69.2. AMC: 27.7, 34.9, 37.4. Right, a scatter plot of prior task retention in percent against new task accuracy in percent. The base model, Qwen2.5-7B-Instruct, is at the top left, at 100 on retention and below 40 on accuracy. SFT is lowest on retention, between 85 and 90. OPSD is at about 95 with the same accuracy as SFT. The point labeled Ours is to the right of and above both, between 95 and 100 on retention.

  5. Figure. Figure 4 of the paper, titled "Chemistry (Olmo-3-7B-Instruct)": pass@k accuracy against k, for k of 1, 2, 4, 8, 16, 32 and 64, with three lines labeled Ours, OPSD and Base. All three rise with k. Base is lowest throughout. At k = 1 OPSD is above Ours; the two meet at k = 2; from k = 4 on Ours is highest, ending close to 1.0 at k = 64, with OPSD and Base below it and still rising.