Perplexity Open-Sources Contextual Embedding Model
Perplexity released a 9B contextual embedding model that represents each chunk using the full document context, helping RAG systems retrieve both answers and supporting evidence. The preview is publicly available on Hugging Face, with API access planned.
Perplexity is targeting RAG’s most persistent weakness: isolated chunks that retrieve the answer but lose the context needed to interpret it.
- –A context-compression teacher turns token-level relevance into softer chunk-level training signals.
- –The model produces one vector per chunk without adding inference-time reranking or storage overhead.
- –It reports 45.5% Answer@10 and 40.6% Evidence Recall@10 on turbopuffer’s context-bench, ahead of Voyage Context 4.
- –1024-dimensional int8 vectors reduce storage to roughly 1 KB each, making the approach attractive for large-scale vector databases.
- –Turbopuffer’s object-storage economics provide a fitting infrastructure layer for this push toward cheaper, document-aware retrieval.
DISCOVERED
1h ago
2026-10-01
PUBLISHED
1h ago
2026-10-01
RELEVANCE
AUTHOR
stretchcloud