Sharpening Tax Exposes RL’s Coverage Trade-Off
This research finds that RL post-training can improve pass@1 while narrowing an LLM’s solution coverage under repeated sampling. It introduces Sharpening Tax to measure that loss and PTGS, which adapts training temperature to prompt difficulty.
The paper challenges the idea that higher first-attempt accuracy always means broader capability; for agents, RL may make policies more reliable but less exploratory.
- –Evaluates 14 base/post-trained model pairs across 42 agentic benchmark cases
- –Finds base models can outperform post-trained models at high pass@K despite lower pass@1
- –Shows post-training pushes tasks toward “always solved” or “never solved” outcomes
- –PTGS heats difficult prompts and cools easy ones to preserve useful exploration
- –Reports PTGS improving both single-shot accuracy and repeated-sampling coverage in experiments
DISCOVERED
1h ago
2026-10-04
PUBLISHED
1h ago
2026-10-04
RELEVANCE
AUTHOR
Discover AI