Google Launches EmbeddingGemma 2 for Advanced Multimodal Search
Google has introduced EmbeddingGemma 2, a compact open model designed to enhance search and retrieval augmented generation (RAG) applications. Released under the Apache 2.0 license, this model offers superior multimodal performance, handling text, code, images, video, and audio within a unified 768-dimensional space. Based on Gemma 4, EmbeddingGemma 2 features a modular architecture, allowing developers to load only the necessary encoders, scaling from 270M parameters for text and code to 740M for full multimodal capabilities.
The model excels in native multimodal retrieval, superior code understanding, and modular memory footprint. It employs Matryoshka Representation Learning (MRL) to dynamically truncate dimensions, reducing storage requirements while retaining embedding quality. For instance, truncating to 256 dimensions retains most of the original quality for text and code and about 95% for image, video, and speech retrieval.
EmbeddingGemma 2 replaces chained models with modular encoders, ensuring all inputs are processed through a shared backbone. The model includes specialized encoders for text and code, vision, and audio, with the flexibility to load only the required modalities. This modular approach minimizes memory usage and enhances efficiency in on-device RAG pipelines.
Developers can integrate EmbeddingGemma 2 using the sentence-transformers library (v6.1.0 or later). The model supports various configurations, from text-only to full multimodal, and allows for task-specific prompts to steer representations. It also facilitates cross-modal search and interleaved inputs, enabling seamless comparison of embeddings from different modalities within the same dimensional space.