Finetuning with Sampling:
SFT Learns Better Than You Think

1Harvard University
TLDR: We introduce a sampling algorithm that shapes off-policy information towards on-policy trajectories, making SFT competitive with the prevailing learning approaches for posttraining at both generalizing and mitigating forgetting.

Abstract

Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling.

In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning.

Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.

On-Policy Learning

How do we introduce fundamentally new capabilities to a base model without forgetting existing ones?

RL has emerged as a clear favorite for posttraining, featuring strong generalization while leaving prior abilities relatively intact. SFT, on the other hand, is widely regarded as a weak learner, memorizing instead of generalizing and exhibiting catastrophic forgetting. Overwhelming evidence attributes this discrepancy to the observation that RL learns on-policy while SFT learns off-policy. RL learns from rollouts generated from the learner's own distribution while SFT fits to traces that can be exceedingly different from a typical rollout, resulting in drastic and unpredictable shifts in behavior.

Off-policy learning has one resounding advantage, however, which is the ability to leverage privileged information in constructing target rollouts for learning. On-policy learning instead relies on the model's own ability to find successful trajectories by repeated sampling. For truly novel tasks, though, the likelihood of this happening is effectively zero, yielding no learning signal at all. This further leaves learning prone to distribution sharpening, where preexisting capabilities are strengthened and refined instead of new capabilities being added. This motivates the foundational question of our work:

Can we algorithmically sample on-policy trajectories that remain faithful to privileged off-policy information for learning?

Projection Sampling

With the right formalism, we can actually explicitly write down an ideal sampling distribution for learning with privileged information. Privileged information can be represented as an information constraint \( \mathcal{C} \) over all trajectories, where some predetermined equivalence relation determines whether a trajectory is consistent with \( \mathcal{C} \) or not. Consider the set of all distributions \( \mathcal{P}_\mathcal{C} \) that have support \( \mathcal{C} \). Then we want to sample from the information projection: \[p_{\mathcal{C}} := \arg\min_{\pi \in \mathcal{P}_{\mathcal{C}}} \mathrm{KL}\!\left( \pi \,\middle\|\, p \right) \implies p_{\mathcal{C}}(\mathbf{x}) \propto p(\mathbf{x}) \cdot \mathbf{1}\bigl(\mathbf{x} \in \mathcal{C}\bigr),\] which is the distribution closest to base model policy \(p\) that is also consistent with \( \mathcal{C} \). Of course, one way to sample from \(p_{\mathcal{C}}\) is via rejection sampling, but this is identical to the repeated sampling in on-policy learning. Instead, we employ a neat trick based on the Metropolis-Hastings (MCMC) algorithm that enables us to use privileged information as an "initialization" that progressively becomes more on-policy with more MCMC steps.

Toy schematic of traditional Metropolis-Hastings updates.

Metropolis-Hastings is an approximate sampler that iteratively updates a sample rollout \(\mathbf{x}\) by proposing an altered candidate \(\mathbf{x'}\) with proposal distribution \(\kappa\), accepting updates with probability \(A(\mathbf{x}, \mathbf{x'})\) that depends on the target distribution (in our case \(p_{\mathcal{C}}\)) via \[A(\mathbf{x}, \mathbf{x'}) = \text{min} \left\lbrace 1, \frac{p_{\mathcal{C}}(\mathbf{x}) \cdot \kappa(\mathbf{x'} | \mathbf{x})}{p_{\mathcal{C}}(\mathbf{x'}) \cdot \kappa(\mathbf{x}|\mathbf{x'})}\right\rbrace = \text{min} \left\lbrace 1, \frac{p(\mathbf{x}) \cdot \mathbf{1}\bigl(\mathbf{x} \in \mathcal{C}\bigr) \cdot \kappa(\mathbf{x'} | \mathbf{x})}{p(\mathbf{x'}) \cdot \mathbf{1}\bigl(\mathbf{x'} \in \mathcal{C}\bigr) \cdot \kappa(\mathbf{x}|\mathbf{x'})}\right\rbrace = \text{min} \left\lbrace 1, \frac{p(\mathbf{x}) \cdot \kappa_{\mathcal{C}}(\mathbf{x'} | \mathbf{x})}{p(\mathbf{x'}) \cdot \kappa_{\mathcal{C}}(\mathbf{x}|\mathbf{x'})}\right\rbrace.\] In the expression for \(A(\mathbf{x}, \mathbf{x'})\), note that we can absorb the information constraint into the proposal instead of the target distribution, allowing us to use it in-context to generate MCMC candidates. In other words, MCMC sampling from \(p_{\mathcal{C}}\) is equivalent to MCMC sampling from \(p\) with restricted proposals that preserve \(\mathcal{C}\) across candidate trajectories. For our implementation, we select the privileged self-teacher \(p( \cdot | \mathcal{C})\) as our proposer, leveraging the strong instruction-following capabilities of pretrained LLMs to ensure our candidates are consistent with \(\mathcal{C}\). The MCMC process then progressively refines these candidates to be more in-distribution, scoring them based on their base model likelihoods. We illustrate a toy example of the generation process below:

Illustration of block-wise projection sampling for LLMs.

We call our algorithm projection sampling, as it projects off-policy traces into more on-policy regions of the base model. In fact, if we plot the likelihoods of the original traces alongside the corresponding sampled trajectories, we can directly observe its likelihood-boosting effect.

Projection Sampling Likelihoods

Finetuning with Sampling

Remarkably, the trajectories generated by our sampling algorithm allow SFT to defy its weak characterization, enabling it to quite often generalize better and forget less than some of the current prevailing techniques for posttraining, including RL and OPSD. We assess this over a variety of tasks that are characteristic of the range of capabilities posttraining seeks to introduce: novel skill acquisition with chemistry, reasoning with math, and open-ended expertise with medical question-answering. We compare to relevant on-policy learning baselines with and without privileged information, such as GRPO and UFT (RL) as well as on-policy self-distillation (OPSD).

Across these tasks, we find that standard SFT composed with projection sampling can achieve substantial improvements on new tasks and even generalize beyond the specific test domain, all while retaining prior capabilities at least as well as on-policy baselines.

Single-Shot Reasoning
Generalization and forgetting performance on chemistry for SFT, OPSD, and SFT with sampling on Qwen2.5-Math-7B-Instruct.

Learning Beyond Sharpening

Although sampling enables finetuned models to perform well single-shot, can our models acquire fundamentally new behaviors that are not present in the base model? We examine the pass@\(k\) accuracy of SFT with sampling alongside OPSD relative to the base model on the chemistry task. If training on more on-policy rollouts from the base model was simply sharpening existing behaviors, we would expect our pass@\(k\) curve to converge to the base model for large \(k\).

Multi-Shot Reasoning
Pass@\(k\) performance of Sampling SFT and OPSD relative to the base model Olmo-3-7B-Instruct on chemistry.

Instead, we see large, consistent gaps where our pass@\(k\) curve outperforms both the base model and OPSD, demonstrating that finetuning with sampling genuinely introduces stronger capabilities that are not contained in the base model and are not learned by competing on-policy learning methods. We can even pinpoint this down to the level of individual tasks, where we identify "hard problems" that the base model absolutely cannot solve (pass@\(k\) rate zero for \(k = 64\)) but our finetuned model can start reliably solving. In fact, some evaluation tasks go from zero pass rate all the way up to rates like 67.2% and 53.1%.

Concluding Points

Our work demonstrates that SFT is far more capable at generalizing and mitigating forgetting than traditionally appreciated. Rather than modifying the learning algorithm to accommodate off-policy data, we instead propose shaping off-policy data to accommodate the learner. The resulting composition of finetuning with sampling exhibits the characteristic generalization and prior capability retention of on-policy learning while facilitating learning beyond distribution sharpening.

Beyond just SFT, the success of posttraining hinges on the following ability: can we find trajectories native to the learner's distribution that provide desirable information signal? Sampling provides a principled methodology to intentionally construct such trajectories from privileged information, offering a general-purpose primitive to the entire posttraining stack, from finetuning to reinforcement learning to distillation.

BibTeX

@article{karan2026finetuning,
  author    = {Karan, Aayush and Chen, Sitan and Du, Yilun},
  title     = {Finetuning with Sampling: SFT Learns Better Than You Think},
  journal   = {arXiv},
  year      = {2026},
}