Google DeepMind has open-sourced EmbeddingGemma 2, a 740 million-parameter model built to process text, code, images, video and audio in a single shared vector space for search and retrieval-augmented generation, or RAG.
The previous EmbeddingGemma generation handled text only. EmbeddingGemma 2 adds multimodal support and can be loaded in parts depending on the task. The text-and-code setup runs at 270 million parameters. Adding vision brings it to 440 million. Loading the full model takes it to 740 million parameters.
Memory footprint and context window
Google said it tested the model on a Pixel 11 Pro. After quantization, the text version used about 191MB of active memory at the low end, while the full version used about 567MB.
The context window has been expanded from 2K in the previous generation to 8K. In one pass, the model can handle about 5.5 minutes of audio, 29 images or 58 video frames.
Retrieval performance
The biggest improvement showed up in code retrieval. Its MTEB Code score rose from 68.76 in the previous generation to 78.68. Multilingual text performance was largely flat, moving from 61.15 to 61.36.
Vector compression and licensing
The model outputs 768-dimensional vectors by default, and also supports compression to 512, 256 or 128 dimensions. At the lowest setting, vector storage can be reduced to one-sixth of the default footprint.
Google also changed the license. The original EmbeddingGemma used Google's own Gemma terms, while EmbeddingGemma 2 is released under Apache 2.0. Google said Gemma 4, released this year, also uses Apache 2.0, and its more recent open models are more accommodating for commercial use and downstream development.

