Google unveils EmbeddingGemma 2 for advanced on-device search
Google has launched EmbeddingGemma 2, an open-source multimodal embedding model designed for on-device applications. Announced on Tuesday, the model integrates text, images, audio, and video into a unified embedding space and is available under an Apache 2.0 license. Built on the Gemma 4 architecture, EmbeddingGemma 2 contains 740 million parameters, with text-only workloads requiring as little as 270 million parameters. Optional vision and audio encoders add 170 million and 300 million parameters, respectively. On a Google Pixel 11 Pro, the text-only version uses about 191MB of active RAM, while the full multimodal version requires 567MB with quantization.
The model features an 8,000-token context window, four times larger than its predecessor, EmbeddingGemma, which has been downloaded over 20 million times. This expanded window allows processing of up to 5.5 minutes of audio, 29 images, or 58 video frames on local hardware. EmbeddingGemma 2 scored 78.68 on the Massive Text Embedding Benchmark Code test, a significant improvement from EmbeddingGemma’s score of 68.76.
Developers can reduce output vectors from 768 dimensions to as few as 128 dimensions using Matryoshka Representation Learning, providing up to 6x storage reduction for local vector databases. The model is available for download on Hugging Face and Kaggle, with support for deployment through various platforms including MediaPipe, LiteRT, transformers, and sentence-transformers. Google also noted that Gemini Enterprise Agent Platform Model Garden availability is coming soon.