North Micro Vision Drops 2.4B Vision Model
Cohere Labs released North Micro Vision, a 2.4B-parameter open-weight vision-language model built for document understanding, visual Q&A, charts, and structured extraction. Native-resolution image support and integrations with MLX, Axolotl, and NVIDIA AutoModel make it practical to run and fine-tune across local and GPU environments.
The real story is distribution: Cohere is pairing a compact vision model with an unusually broad fine-tuning stack from day one.
- –Native-resolution inputs preserve detail in documents, charts, and product imagery better than aggressively resized pipelines.
- –The 2.4B footprint makes local and edge deployment more plausible than larger enterprise VLMs.
- –MLX support opens a path to fast Apple Silicon experimentation, while Axolotl lowers the barrier to custom training.
- –NVIDIA AutoModel adds a reproducible LoRA/FSDP2 route for teams fine-tuning on GPUs.
- –Developers should still benchmark OCR accuracy, latency, and memory use on their own document mix before treating it as a production replacement.
DISCOVERED
1h ago
2026-08-12
PUBLISHED
1h ago
2026-08-12
RELEVANCE
AUTHOR
cohere