Google Unveils EmbeddingGemma 2 for Advanced On-Device Multimodal Search
Google has introduced EmbeddingGemma 2, a new open-source multimodal embedding model optimized for on-device applications. The model, released on Tuesday, integrates text, images, audio, and video into a unified embedding space and is available under the Apache 2.0 license.
EmbeddingGemma 2 boasts 740 million parameters and is based on the Gemma 4 architecture. For text-only tasks, it requires just 270 million parameters, with optional vision and audio encoders adding 170 million and 300 million parameters respectively. On a Google Pixel 11 Pro, the model uses 191MB of active RAM for text-only operations and 567MB for the full multimodal version with quantization.
The model features an 8,000-token context window, four times larger than its predecessor, EmbeddingGemma, which has been downloaded over 20 million times. This expanded window allows for processing up to 5.5 minutes of audio, 29 images, or 58 video frames on local hardware. EmbeddingGemma 2 scored 78.68 on the Massive Text Embedding Benchmark Code test, a significant improvement from EmbeddingGemma's score of 68.76.
The model supports Matryoshka Representation Learning, enabling developers to reduce output vectors from 768 to as few as 128 dimensions, offering up to 6x storage reduction for local vector databases. It is available for download on Hugging Face and Kaggle, with deployment support through various frameworks including MediaPipe, LiteRT, and transformers. Google also announced that Gemini Enterprise Agent Platform Model Garden availability is coming soon.