Google Releases Compact EmbeddingGemma 2 for Multimodal Content
Google has introduced EmbeddingGemma 2, an open-source model designed to convert text, images, video, audio, and code into numerical vectors. This allows similar content to be easily found and compared. The model stands out for its compact size, featuring just 740 million parameters, yet Google claims it outperforms rival models that are up to twice its size on multimodal embedding benchmarks.
The new model achieves a score of 78.68 on the Massive Text Embedding Benchmark (Code), marking an improvement of nearly 10 points over its predecessor, which scored 68.76. This places EmbeddingGemma 2 on par with much larger models in terms of performance. Additionally, the model operates locally without requiring an API key, with each query taking approximately 20 to 70 milliseconds via WebGPU in the browser. It also requires only around 191 MB of RAM and can reduce local vector database storage by up to six times.
For text-only tasks, a lighter version with 270 million parameters is available. When paired with small open models like Gemma 4, EmbeddingGemma 2 enables offline Retrieval-Augmented Generation (RAG) applications, eliminating the need to send data to external servers. The model's weights, along with a developer guide and documentation, are accessible on Hugging Face and Kaggle.