Google DeepMind Unveils EmbeddingGemma 2 for Advanced On-Device AI Search
Google DeepMind has introduced EmbeddingGemma 2, a cutting-edge multimodal embedding model designed to streamline semantic search and media retrieval on edge devices. This open-weight model, with a compact 740M parameter footprint, natively maps text, images, video frames, and audio into a single vector space, reducing latency and memory overhead for developers building privacy-first applications. Notably, it operates with as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model on a Google Pixel 11 Pro, making it ideal for on-device decision engines.
The model eliminates the need for separate image captioning, speech-to-text, and text-embedding models, offering ultra-low-latency intent routing without requiring any training data or fine-tuning. EmbeddingGemma 2 is now available through Google AI Edge’s interactive showcases, including Instant Media Search and Video Moments Finder, which allow users to search local media using natural language or example images. These features are accessible via the Google AI Edge Gallery app on mobile and the newly launched Google AI Edge Foresight on Mac, which provides a context-aware meeting companion with enhanced note-taking and cross-modal retrieval capabilities.
For developers, EmbeddingGemma 2 will soon be available as a service on Android through ML Kit, featuring NPU acceleration for optimized performance across various devices. Additionally, MediaPipe Tasks will support EmbeddingGemma 2, simplifying the integration of local search features for cross-platform applications. This model represents a significant advancement in on-device AI, ensuring privacy and uninterrupted productivity without relying on cloud services.