DMAD Cuts Visual Generation to Few Steps
DMAD reframes distribution matching as adversarial classification, eliminating DMD’s separately fitted auxiliary score model. The research reports strong one-step image, four-step image, video, and audio-video generation results at substantially lower inference cost.
DMAD’s strongest contribution is a cleaner distillation objective that turns classifier log-density ratios into a direct training signal for fast visual generators.
- –Reports 1.04 FID for one-step ImageNet-64 generation and 14.47 FID for four-step SDXL on COCO-10K
- –Reaches 85.15 on VBench with four-step Wan2.1-T2V-14B generation
- –Four-step MiniMax-H3 students achieve 79.1% preference over DMD2 and 84.6% over rCM for joint audio-video generation
- –The speed-quality tradeoff remains workload-dependent: one-step image models lose local detail, while competing methods retain advantages on some semantic video metrics
- –Open code and models make DMAD immediately relevant for researchers deploying low-latency image and video generation
DISCOVERED
54m ago
2026-10-04
PUBLISHED
1h ago
2026-10-04
RELEVANCE
AUTHOR
AI Search