Google on Oct. 6 open-sourced EmbeddingGemma 2, a 740 million-parameter model that brings text, code, image, video and audio retrieval into a single embedding system. Google said the model can run offline on phones, laptops and in browsers, with the fully quantized multimodal version using about 567MB of memory on a Pixel 11 Pro.

The company described it as Google’s first native multimodal open-source embedding model. CEO Sundar Pichai also promoted the release. A little more than an hour after launch, Hugging Face’s Victor M had already run it in a browser demo: after typing “birds singing,” the interface reordered results to surface both bird photos and bird-call audio clips. The query took 22 milliseconds and did not rely on a server or any API call.
A shared embedding space for text, images, audio and video
Embedding models usually sit behind AI search and retrieval-augmented generation systems. Their job is to convert content into vectors so a system can compare semantic distance and retrieve the closest matches.
The previous EmbeddingGemma generation handled text only. Searching images or audio often meant adding separate models or converting those inputs into text first. EmbeddingGemma 2 moves text, images, video and audio into the same 768-dimensional space. In practice, that means a cat photo, the word “cat,” and the sound of a cat can be linked by meaning rather than by file type.

That changes how retrieval works. A user can search images or videos with a sentence, use a voice clip to find a matching segment inside a video, or retrieve a full item that contains text, images and video together.
Google’s model card included a product-page example: a trail-running shoe listing with written descriptions, two detail images and a video showing slip-resistance testing on wet rock. The model can encode that bundle into one vector and match it against a query such as “waterproof shoes for trail running.”
8K context window and support for more than 100 languages
Google said EmbeddingGemma 2 also expands how much content can be processed at once. The model has an 8K context window, four times larger than the previous generation.
For single-modality inputs, it can take in about 5.5 minutes of audio, 29 images or 58 video frames. Google said it supports more than 100 languages.
In Google’s published comparisons against several models of similar size, the biggest gap showed up in image and video retrieval.
- Image retrieval score: 57.3 for EmbeddingGemma 2 versus 31.6 for Jina v5 Omni-Nano.
- Video retrieval score: 50.7 versus 31.2.
- Multilingual text performance was roughly in line with the previous generation.
Based on those figures, EmbeddingGemma 2 led Jina v5 Omni-Nano by 25.7 points in image retrieval and 19.5 points in video retrieval, even though Google said the Jina model has about 30% more parameters.
567MB for the full multimodal setup, with modular loading
Google broke the 740 million parameters into three parts: 270 million for text, 170 million for the vision encoder and 300 million for the audio encoder. Those components can be loaded as needed.
Text-only retrieval requires the 270 million-parameter text portion. Adding image search brings the total to 440 million parameters. On a Pixel 11 Pro, Google said the quantized text-only model used as little as about 191MB of memory, while the full multimodal model used about 567MB.
The model also supports a compressed search index that stores each item’s meaning with fewer numbers, cutting storage use by two-thirds. In Google’s multilingual text retrieval test, the score dropped by less than 1 point under that setup.
Google showed local-first apps and developers quickly picked it up
Google presented several application prototypes built around EmbeddingGemma 2.
- Instant Media Search in AI Edge Gallery lets users search photos in a gallery by meaning.
- Video Moments Finder can locate matching moments inside a video from a natural-language prompt.
- Foresight pairs the model with Gemma 4 to search files and meeting notes offline on a Mac.
Developer tooling is already moving around the release. llama.cpp supports the model, and Unsloth has published a quantized version for local deployment.

According to testing published by the Mac app Nativ, an 8-bit quantized build on an M5 Max reached a cosine similarity of 0.9997 against the original vectors and processed text embeddings at 817 items per second.
Developer tests pointed to gains in cross-language and typo-heavy search
Turkish developer Avenox shared a test that looked closer to everyday note retrieval. Most of his notes were written in English, but he usually asks questions in Turkish. With keyword matching, relevant notes could be missed when the wording did not line up.
He said Claude spent an hour building a test setup. He selected 40 notes, rephrased them, then queried in Turkish to see whether the correct note would appear in the top 18 candidates. In the results he published, keyword matching hit 15%, while EmbeddingGemma 2 reached 97%.
He then tested 30 real messages containing typos. EmbeddingGemma 2 found twice as many relevant notes as keyword matching, with each query taking about 50 milliseconds and running entirely on-device.
Developer Nick Lo pushed the model onto a Nano development board. After a text prompt, the model took about 15 milliseconds to turn the sentence into a search vector. Once the matching image was found, an ESP32-S3 microcontroller drew it line by line.
Code retrieval also improved
Beyond multimodal search, Google said code retrieval improved sharply. On the MTEB Code benchmark, EmbeddingGemma 2 scored 78.68, up from 68.76 for the previous generation, a gain of nearly 10 points. Google said that result leads among models of similar size.
That matters for coding agents. Tools such as Claude Code and Codex typically need to search a codebase for relevant files and snippets before they can act. With the 270 million-parameter text-only version, developers can build a local index for a code repository and query it in plain language, keeping both indexing and retrieval on-device.

From cloud APIs to on-device retrieval
Google had already released Gemini Embedding 2 in March as a cloud-based multimodal model accessed through an API and billed by usage. More than half a year later, it has moved multimodal retrieval into an open-source model small enough to run on a phone.
That gives photos, recordings and videos a local retrieval option. One setup Google highlighted pairs EmbeddingGemma 2 with Gemma 4 for offline, privacy-first RAG: one model retrieves the material, the other answers from it, and both steps stay on the device.
In many earlier workflows, AI search across photos, videos and recordings meant sending those files to the cloud. Google is now putting that capability into a package measured in a few hundred megabytes, making local search on personal devices much more practical.

