UniEvo-VL Learns From Image Errors
UniEvo-VL is an on-policy self-distillation framework that lets a multimodal model critique its own images, then transfer those corrections into future generations. Built on Qwen-image-2512, it raises GenEval from 0.747 to 0.808 and GenEval2 Soft-TIFA from 32.97 to 35.53.
UniEvo-VL is a promising step toward recursive model improvement, but its results show that self-critique is only as reliable as the model’s judging ability.
- –One model acts as both teacher and student: the teacher sees corrective feedback while the student receives only the original prompt.
- –Training matches denoising distributions along the model’s own sampling paths, avoiding corrected-image targets and traditional reward optimization.
- –Stronger external critics produce higher ceilings, suggesting evaluator quality remains a core bottleneck.
- –Improvements are uneven across tasks, particularly in text rendering, so benchmark gains do not guarantee universal quality improvements.
- –The approach could make open image models more capable without requiring a permanently larger teacher model.
DISCOVERED
1h ago
2026-10-02
PUBLISHED
1h ago
2026-10-02
RELEVANCE
AUTHOR
mark_k