Tencent's Youtu-Parsing-Omni Unifies Multimodal Parsing
Tencent released Youtu-Parsing-Omni, a 5B open-weight model that converts documents, images, charts, geometry figures, audio, and video into structured JSON. Its model card reports 96.96 on OmniDocBench v1.6 and 75.08 on OmniParsingBench, making it the strongest open-weight model in the latter benchmark and competitive with Gemini 3 Pro. [Hugging Face model card](https://huggingface.co/tencent/Youtu-Parsing-Omni)
Youtu-Parsing-Omni’s real breakthrough is workflow consolidation: developers can use one parser and one schema instead of stitching together OCR, chart extraction, ASR, and video-analysis services. The Gemini comparison is impressive but narrower than the headline suggests.
- –Beats Gemini 3 Pro on OmniDocBench v1.6, 96.96 to 92.91
- –Trails Gemini on OmniParsingBench overall, 75.08 to 77.44, while winning on chart and audio parsing
- –Produces layout boxes, text, formulas, tables, Mermaid diagrams, timestamps, ASR, acoustic events, and camera-motion metadata
- –Includes vLLM serving code and inference examples for self-hosted deployment
- –Its custom license excludes use within the European Union, limiting adoption despite the open weights
DISCOVERED
1h ago
2026-10-11
PUBLISHED
1h ago
2026-10-11
RELEVANCE
AUTHOR
0x0SojalSec
