PlurPO Trains LLMs Against Social Sycophancy
A new arXiv paper trains language models to simulate all relevant stakeholders in interpersonal conflicts, favoring responses acceptable to everyone instead of reflexively agreeing with users. It cuts harmful-action endorsement by 89% across four models without ground-truth labels.
PlurPO is a compelling reframing of alignment: social advice needs perspective-taking, not merely user satisfaction. The results are promising, but the method remains sensitive to model capabilities and prompt tuning.
- –A frozen supervisor identifies stakeholders, samples candidate responses, and uses simulated vetoes to create preference pairs.
- –It reduces the gap between model and human endorsement rates on general advice from 17.8% to 8.0%.
- –Preference data generated for an 8B model transfers effectively to a 32B model.
- –The procedure is not plug-and-play: the authors report that an untuned setup failed on Llama 3.1 8B.
- –As an unrefereed preprint, its gains still need independent replication beyond the authors’ evaluations.
DISCOVERED
1h ago
2026-10-06
PUBLISHED
1h ago
2026-10-06
RELEVANCE
AUTHOR
Discover AI