Finetuning with Sampling Makes SFT Rival RL
Harvard researchers introduce an MCMC-based projection-sampling method that reshapes expert off-policy traces toward a model’s own distribution before supervised fine-tuning. Across chemistry, math, and medical QA tasks, the approach reportedly improves generalization and reduces forgetting versus strong RL and self-distillation baselines.
The clever move is distribution engineering: preserve expert information while making training data more native to the learner, challenging the assumption that RL inherently has the generalization advantage.
- –MCMC progressively refines expert traces using base-model likelihoods while preserving task-relevant constraints.
- –Experiments compare against GRPO, UFT, and OPSD across scientific skill acquisition, reasoning, and open-ended expertise.
- –The method reportedly learns capabilities beyond simple distribution sharpening, including tasks the base model cannot reliably solve.
- –Extra sampling and likelihood-scoring steps add compute and implementation complexity, so the practical advantage over RL remains to be tested at scale.
- –If validated broadly, projection sampling could become a reusable post-training primitive for SFT, RL, and distillation.
DISCOVERED
1h ago
2026-10-04
PUBLISHED
1h ago
2026-10-04
RELEVANCE
AUTHOR
Discover AI