FastH3 V1 drops 14x-faster MiniMax video
FastVideo has open-sourced FastH3 Preview v1, a four-step DMD2-distilled MiniMax H3 model for text-to-video-and-audio with 90% sparse attention. The open weights and LoRA target up to 14× faster Blackwell inference, with 15-second 768p clips generated in under 13 seconds across eight B200 GPUs. [Announcement](https://haoailab.com/blogs/fasth3-preview/) [Model card](https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA)
FastH3 is a meaningful inference breakthrough, but its headline speed is as much a systems achievement as a model improvement.
- –Four DiT calls replace the base workflow’s longer sampling process, while 90% sparse VSA attention cuts the cost of video-scale computation.
- –The recommended path targets four B200 GPUs, CUDA 13, and FastVideo’s specialized VSA-H3 kernel; this is not a drop-in generic LoRA.
- –Data-free training, full checkpoints, LoRA adapters, dense references, and synthetic-data variants give developers useful customization and comparison paths.
- –The preview currently covers text-to-audio-video; image-reference and full reference-to-video workflows, improved motion quality, and eight-step checkpoints remain on the roadmap.
- –The MiniMax H3 Community License and substantial memory requirements should be evaluated before commercial deployment.
DISCOVERED
1h ago
2026-08-30
PUBLISHED
1h ago
2026-08-30
RELEVANCE
AUTHOR
AI Search