EmbeddingGemma 2 Launches On-Device Multimodal Search
Google DeepMind’s open 740-million-parameter model maps text, code, images, video, and audio into one shared embedding space. Its modular architecture targets private, offline search, classification, and multimodal RAG on phones and laptops.
EmbeddingGemma 2 makes multimodal retrieval practical at the edge, where privacy, latency, and memory matter more than peak cloud benchmark scores.
- –Shared embeddings remove the need to chain separate captioning, speech-to-text, and text-embedding models.
- –Modular encoders scale from 270M text/code parameters to 740M for full multimodal use.
- –Quantized deployments require roughly 191MB for text-only and 567MB for the full model on a Pixel 11 Pro.
- –Matryoshka embeddings can reduce vector-storage requirements while preserving much of the retrieval quality.
- –Code retrieval improves substantially over the first EmbeddingGemma, strengthening local code search and coding-agent pipelines.
DISCOVERED
1h ago
2026-10-07
PUBLISHED
1h ago
2026-10-07
RELEVANCE
AUTHOR
WorldofAI