Google open-sources EmbeddingGemma 2 with 740 million parameters for multimodal search

Google open-sources EmbeddingGemma 2 with 740 million parameters for multimodal search

N
News Editor
2026-10-07 02:46:12
Google DeepMind has released EmbeddingGemma 2 as an open-source model, expanding the Gemma embedding line from text-only processing to a multimodal system that can handle text, code, images, video and audio in a shared vector space for search and retrieval-augmented generation. The model is modular rather than fixed at one size: it runs at 270 million parameters for text and code, 440 million with vision, and 740 million when all components are loaded. Google said tests on a Pixel 11 Pro showed the quantized text version using about 191MB of active memory, while the full version used about 567MB. Context length has also increased from 2K in the previous generation to 8K, allowing a single pass over roughly 5.5 minutes of audio, 29 images or 58 video frames. Performance gains were strongest in code retrieval, with the MTEB Code score rising from 68.76 to 78.68, while multilingual text moved from 61.15 to 61.36. Google also shifted the license from the original Gemma terms to Apache 2.0.

Google DeepMind has open-sourced EmbeddingGemma 2, a 740 million-parameter model built to process text, code, images, video and audio in a single shared vector space for search and retrieval-augmented generation, or RAG.

The previous EmbeddingGemma generation handled text only. EmbeddingGemma 2 adds multimodal support and can be loaded in parts depending on the task. The text-and-code setup runs at 270 million parameters. Adding vision brings it to 440 million. Loading the full model takes it to 740 million parameters.

Memory footprint and context window

Google said it tested the model on a Pixel 11 Pro. After quantization, the text version used about 191MB of active memory at the low end, while the full version used about 567MB.

The context window has been expanded from 2K in the previous generation to 8K. In one pass, the model can handle about 5.5 minutes of audio, 29 images or 58 video frames.

Retrieval performance

The biggest improvement showed up in code retrieval. Its MTEB Code score rose from 68.76 in the previous generation to 78.68. Multilingual text performance was largely flat, moving from 61.15 to 61.36.

Vector compression and licensing

The model outputs 768-dimensional vectors by default, and also supports compression to 512, 256 or 128 dimensions. At the lowest setting, vector storage can be reduced to one-sixth of the default footprint.

Google also changed the license. The original EmbeddingGemma used Google's own Gemma terms, while EmbeddingGemma 2 is released under Apache 2.0. Google said Gemma 4, released this year, also uses Apache 2.0, and its more recent open models are more accommodating for commercial use and downstream development.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.