NAMAA releases Cohere Speech Tashkeel 2B
Cohere Speech Tashkeel 2B, developed by NAMAA, is an open-source Arabic speech recognition model that transcribes spoken Arabic into fully diacritized text. The model supports complete vocalization, including harakāt, tanwīn, sukūn, shadda, and grammatical case endings, and has already received community 4-bit quantizations, Ruby bindings, and containerized apps.
Precise diacritization in speech recognition addresses a longstanding gap in Arabic NLP, where unvocalized transcriptions often lose critical grammatical and semantic context.
• Comprehensive diacritization: Transcribing case endings and harakāt directly from audio greatly improves downstream translation, text-to-speech, and linguistic analysis.
• Robust community ecosystem: Immediate availability of GGUF/ONNX 4-bit quantizations and multi-platform Ruby bindings makes deploying the 2B model feasible on edge devices and consumer hardware.
• Focus on under-resourced languages: Demonstrates building open AI tooling tailored to specific linguistic complexities beyond English.
DISCOVERED
3h ago
2026-07-22
PUBLISHED
3h ago
2026-07-22
RELEVANCE
AUTHOR
cohere