Google introduces EmbeddingGemma 2 for on-device multimodal search
Google has launched EmbeddingGemma 2, an open-source multimodal embedding model designed for on-device applications. The model, released on Tuesday, maps text, images, audio, and video into a unified embedding space and is available under an Apache 2.0 license.
The model contains 740 million parameters and is built on the Gemma 4 architecture. It requires as little as 270 million parameters for text-only workloads, with optional vision and audio encoders adding 170 million and 300 million parameters respectively. On a Google Pixel 11 Pro, the model uses approximately 191MB of active RAM for text-only weights and 567MB for the full multimodal version with quantization.
EmbeddingGemma 2 features an 8,000-token context window, four times larger than its predecessor EmbeddingGemma, which launched last year and has been downloaded more than 20 million times. The expanded context window allows processing of up to 5.5 minutes of audio, 29 images, or 58 video frames on local hardware.
The model scored 78.68 on the Massive Text Embedding Benchmark Code test, a 9.92-point improvement from EmbeddingGemma’s score of 68.76. It uses Matryoshka Representation Learning, enabling developers to reduce output vectors from 768 dimensions to as few as 128 dimensions, providing up to 6x storage reduction for local vector databases.
The model is available for download on Hugging Face and Kaggle, with support for deployment through various platforms including MediaPipe, LiteRT, transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio. Google stated that Gemini Enterprise Agent Platform Model Garden availability is coming soon.