Xberg Unifies Document Parsing, OCR, RAG
Xberg is an open-source, Rust-core document-intelligence engine that extracts text, tables, metadata, code structure, transcripts, and embeddings across 100 formats. Its 1.0 release succeeds Kreuzberg with native PDF parsing, multiple OCR backends, MCP support, and polyglot bindings.
Xberg’s appeal is consolidation: it turns fragmented ingestion pipelines into one local-first extraction layer, though real-world accuracy still needs validation against specialized parsers.
- –Supports PDFs, Office files, images, archives, audio, video, URLs, and source trees
- –Combines OCR, layout detection, table reconstruction, transcription, code intelligence, and RAG chunking
- –Rust core enables native performance, streaming, caching, WASM, CLI, REST, and MCP deployment
- –MIT licensing and bindings for 15 languages make it practical for teams building portable ingestion infrastructure
- –Its successor relationship to Kreuzberg means existing users should evaluate migration and API compatibility carefully
DISCOVERED
1h ago
2026-08-14
PUBLISHED
1h ago
2026-08-14
RELEVANCE
AUTHOR
techNmak