OSPD sharpens persona consistency through self-distillation
OSPD trains a model to preserve character traits by letting an EMA teacher with full persona access supervise the same model’s on-policy responses from a brief profile. The paper reports stronger persona consistency across CharacterBench, CharacterEval, and SocialBench without an external teacher or reward model.
OSPD is a clever post-training recipe that turns privileged context into supervision, making persona fidelity more about internalized behavior than ever-longer prompts.
- –Role-aware KL switching focuses imitation on character-critical tokens while preserving diversity in generic dialogue.
- –Progressive trait masking pushes identity, personality, knowledge, style, and history into model parameters.
- –OSPD improves CharacterBench from 3.00 with SFT to 3.51 and CharacterEval persona consistency from 2.90 to 3.49.
- –On-policy rollouts appear more valuable than teacher scale: OSPD beats a 72B off-policy distillation baseline despite using a much smaller teacher.
- –The tradeoff is compute: training takes about 18.6 hours versus 4.2 for SFT, and results may depend strongly on model scale and LLM-judge evaluation.
DISCOVERED
1h ago
2026-09-30
PUBLISHED
1h ago
2026-09-30
RELEVANCE
AUTHOR
Discover AI