Google Releases EmbeddingGemma 2 for Advanced On-Device Search
Google has introduced EmbeddingGemma 2, a new open-source multimodal embedding model optimized for on-device use. The model, built on the Gemma 4 architecture, integrates text, images, audio, and video into a single embedding space and is available under an Apache 2.0 license.
The latest version boasts 740 million parameters and can be configured for text-only applications with as few as 270 million parameters. Optional vision and audio encoders add 170 million and 300 million parameters, respectively. On a Google Pixel 11 Pro, the text-only version consumes around 191MB of active RAM, while the full multimodal model requires 567MB with quantization.
EmbeddingGemma 2 features an 8,000-token context window, four times larger than its predecessor, allowing it to process up to 5.5 minutes of audio, 29 images, or 58 video frames locally. The model achieved a score of 78.68 on the Massive Text Embedding Benchmark Code test, a significant improvement over the previous version's 68.76.
Developers can download EmbeddingGemma 2 from Hugging Face and Kaggle, with support for deployment through various frameworks like MediaPipe, LiteRT, and transformers. Google also announced upcoming availability on the Gemini Enterprise Agent Platform Model Garden.