This Tuesday (6), Google DeepMind launched the EmbeddingGemma 2, its first open natively multimodal embedding model aimed at local execution. With up to 740 million parameters, it brings together text, code, images, video and audio in the same vector space, allowing searches across different types of content directly on phones, computers and other devices.

In practice, the technology allows finding an image using a text description, locating a specific moment in a video from a query or relating files of different formats without relying on cloud processing. Google also demonstrated searches on locally stored media that work without an internet connection.

The model was developed to reduce the need to combine separate systems for image recognition, audio transcription and embedding generation. According to Google, EmbeddingGemma 2 offers competitive performance for its size and expands the capabilities of the previous generation for tasks involving code, vision, video and audio.

EmbeddingGemma 2 leads or remains competitive in different text, vision and audio benchmarks.
Table compares EmbeddingGemma 2's performance with other models on text, vision and audio benchmarks.

Architecture allows loading only the necessary resources

EmbeddingGemma 2 uses a modular architecture. The configuration designed for text and code has 270 million parameters; when vision is added, it reaches 440 million; with text and audio, 570 million; and the full multimodal version uses 740 million. All of them project content into the same vector space of 768 dimensions.

In code, the model recorded 78.7 points on MTEB Code. Google states that the performance represents a 14% improvement over the previous generation. As this task uses only text and code, the benchmark can be run with the 270 million parameter configuration, without loading the vision and audio modules.

EmbeddingGemma 2 reaches 78.7 points on the MTEB Code benchmark with 740 million parameters.
Chart compares performance and size of code embedding models, with EmbeddingGemma 2 scoring 78.7 points.

The modular architecture also reduces memory consumption. In measurements released by Google, the text-only configuration uses about 191 MB of active RAM, while the full multimodal model requires approximately 567 MB on a Pixel 11 Pro.

The system also allows reducing the size of the stored vectors. With 256 dimensions, the required space drops to one third of that used by the standard 768-dimension configuration, maintaining approximately 95% of the original quality in the image, video and speech retrieval evaluations conducted by Google.

Another application is local RAG, in which documents, images and other files can be retrieved on the device itself before being sent to a generative model. Since EmbeddingGemma 2 is based on Gemma 4 and shares architecture components with that family, the models can be combined with less resource overlap.

Google also demonstrated the model in instant photo and video search features and in a tool capable of finding specific moments within local recordings. The company intends to bring these capabilities to ML Kit for Android in the coming weeks, with support for NPU acceleration on compatible devices.

EmbeddingGemma 2 was made available under the Apache 2.0license, with weights accessible to developers and support for tools such as Hugging Face Transformers, sentence-transformers, MLX, Ollama and LiteRT.

More from Radar