Google Expands EmbeddingGemma to Support Images, Audio, and Video
Google has expanded its EmbeddingGemma model beyond text to include images, audio, and video, introducing EmbeddingGemma 2. This multimodal embedding model, small enough to run on smartphones, was announced today. The original model, released in September 2025, only handled text. The new version allows apps to perform tasks like matching a voice memo to a specific moment in a video, all without leaving the device.
The response to the first version exceeded expectations, with over 20 million downloads by developers. Built on the Gemma 4 architecture, EmbeddingGemma 2 has grown to 740 million parameters, more than twice the size of the original. The expansion primarily comes from enhancements in vision and audio encoders, though text-only apps can leave these components out.
The new model has shown significant improvements in benchmark tests, particularly in code retrieval, where it scored 78.68 on the Massive Text Embedding Benchmark. Google claims that EmbeddingGemma 2 outperforms some specialist models more than twice its size on image, video, and audio tasks. The model also uses a training technique called Matryoshka Representation Learning, which allows developers to reduce the size of embeddings by up to six times while maintaining about 95% of their quality at 256 numbers.
EmbeddingGemma 2 shares a text tokenizer and an audio encoder with Gemma 4, reducing memory usage in on-device retrieval-augmented generation setups. The model weights are now available from Hugging Face and Google’s Kaggle under an Apache 2.0 license, permitting commercial use. Google also plans to add the model to the Model Garden in the Gemini Enterprise Agent Platform soon.