DeepSeek V4 Vision Model Goes Open Source
DeepSeek has released downloadable weights for its 305-billion-parameter sparse multimodal model under the MIT license. The experimental checkpoint adds vision and a one-million-token context window, but self-hosting requires serious multi-GPU infrastructure and its benchmark claims remain unverified independently.
This is a meaningful licensing win for developers, but the model’s hardware footprint makes it more strategically open than practically local.
- –MIT licensing permits commercial deployment, fine-tuning, modification, and redistribution with minimal friction.
- –The 305B FP8 checkpoint is a multi-GPU deployment project, not a consumer-GPU download; SGLang and Transformers support help, but serving costs remain substantial.
- –DeepSeek reports strong multimodal-agent results, including 36.5 on ApexBench and 35.0 on ZeroBench, but testing used its own harness and maximum reasoning settings.
- –Comparisons on ApexBench and Agents’ Last Exam are imperfect because the text-only baseline ignored multimodal inputs.
- –Developers should evaluate screenshot, chart, document, and UI-agent workflows directly before choosing it over hosted vision APIs or smaller open models.
DISCOVERED
1d ago
2026-09-02
PUBLISHED
1d ago
2026-09-02
RELEVANCE
AUTHOR
AI Revolution