Google DeepMind Releases EmbeddingGemma 2 for On-Device Multimodal Embeddings
Google DeepMind has unveiled EmbeddingGemma 2, a lightweight, multimodal embedding model designed for on-device applications. The model expands on the capabilities of its predecessor, EmbeddingGemma, by unifying text, code, images, video, and audio within a single embedding space. Built on the Gemma 4 architecture and released under an Apache 2.0 license, EmbeddingGemma 2 features 740 million parameters, making it highly efficient for on-device inference.
The new model excels in various tasks, including finding video clips from voice memos and searching through audio recordings using text queries. It achieves top-tier performance in benchmarks like MTEB (Massive Text Embedding Benchmark) and MAEB (Massive Audio Embedding Benchmark), outperforming many larger models in text, vision, and audio tasks. Additionally, it offers modular design, storage efficiency through Matryoshka Representation Learning (MRL), and optimized on-device performance.
EmbeddingGemma 2 is particularly strong in code performance, with a 9.92-point improvement in MTEB Code benchmarks. It also supports an extended context window of 8K tokens, allowing it to process up to 5.5 minutes of audio, 29 images, or 58 video frames directly on local hardware. The model is designed to enable semantic search, routing, and retrieval fully on-device, ensuring data privacy and reducing pipeline latency.
Developers can access the model weights on Hugging Face and Kaggle, with deployment options including Google AI Edge MediaPipe and LiteRT. The model is also compatible with popular development tools like transformers, sentence-transformers, and Qdrant for storing embedding vectors. Google DeepMind has collaborated with partners to ensure immediate usability, and a developer guide and documentation are available for inference and fine-tuning.