Google Introduces EmbeddingGemma 2 for On-Device Multimodal Processing
Google has launched EmbeddingGemma 2, a new open multimodal embedding model that can process text, code, images, audio, and video all on a single device like a phone or a small board. The model is licensed under Apache 2.0 and has 740 million parameters, with the ability to map all content into a single 768-dimensional space. It requires about 191MB of memory for text tasks on a Pixel 11 Pro or up to 567MB for all five modalities. Google emphasizes that no data needs to leave the device, ensuring privacy and security.
EmbeddingGemma 2 is built on the architecture of Gemma 4, released earlier this year, and shares its text tokenizer and audio encoder. This allows the models to run together with reduced combined memory usage. The model supports over 100 languages, though Google notes that performance may vary across them. It also lacks safety tuning and output moderation, relying instead on mitigations applied during training.
Google highlights that the Gemma family has surpassed a billion downloads, with the initial EmbeddingGemma accounting for 20 million of those. The new model can handle 8,192 tokens across all modalities, a significant increase from the previous version. Vectors can be compressed from 768 dimensions to 128, reducing storage needs by up to sixfold. This capability aligns with Europe's push for local data processing on hardware owned by European companies, all under an open license.