Google open-sources EmbeddingGemma 2, a 567MB multimodal retrieval model built to run offline on phones

Google open-sources EmbeddingGemma 2, a 567MB multimodal retrieval model built to run offline on phones

N
News Editor
2026-10-08 02:10:21
Google on Oct. 6 released EmbeddingGemma 2, its first native multimodal open-source embedding model, bringing text, code, image, video and audio retrieval into one system. The model has 740 million parameters and is designed to run locally across phones, laptops and even browsers, with the fully quantized multimodal version using about 567MB of memory on a Pixel 11 Pro, according to Google. The release expands Google’s earlier text-only embedding work by mapping multiple content types into a shared 768-dimensional space. That allows a text query such as “birds singing” to retrieve both bird photos and bird-call recordings, and it also enables voice-to-video matching and mixed-content retrieval across product pages, media libraries and local files. Google said the model supports an 8K context window, more than 100 languages, and larger single-input capacities including about 5.5 minutes of audio, 29 images or 58 video frames. Google’s published benchmarks showed EmbeddingGemma 2 ahead of Jina v5 Omni-Nano in image and video retrieval, while code retrieval on MTEB Code improved from 68.76 to 78.68 versus the previous generation. The company also pointed to local-first use cases, including offline RAG with Gemma 4, on-device media search, and codebase indexing without sending files to the cloud.

Google on Oct. 6 open-sourced EmbeddingGemma 2, a 740 million-parameter model that brings text, code, image, video and audio retrieval into a single embedding system. Google said the model can run offline on phones, laptops and in browsers, with the fully quantized multimodal version using about 567MB of memory on a Pixel 11 Pro.

Google open-sources EmbeddingGemma 2, a 567MB multimodal retrieval model built to run offline on phones 2

The company described it as Google’s first native multimodal open-source embedding model. CEO Sundar Pichai also promoted the release. A little more than an hour after launch, Hugging Face’s Victor M had already run it in a browser demo: after typing “birds singing,” the interface reordered results to surface both bird photos and bird-call audio clips. The query took 22 milliseconds and did not rely on a server or any API call.

A shared embedding space for text, images, audio and video

Embedding models usually sit behind AI search and retrieval-augmented generation systems. Their job is to convert content into vectors so a system can compare semantic distance and retrieve the closest matches.

The previous EmbeddingGemma generation handled text only. Searching images or audio often meant adding separate models or converting those inputs into text first. EmbeddingGemma 2 moves text, images, video and audio into the same 768-dimensional space. In practice, that means a cat photo, the word “cat,” and the sound of a cat can be linked by meaning rather than by file type.

Google open-sources EmbeddingGemma 2, a 567MB multimodal retrieval model built to run offline on phones 3

That changes how retrieval works. A user can search images or videos with a sentence, use a voice clip to find a matching segment inside a video, or retrieve a full item that contains text, images and video together.

Google’s model card included a product-page example: a trail-running shoe listing with written descriptions, two detail images and a video showing slip-resistance testing on wet rock. The model can encode that bundle into one vector and match it against a query such as “waterproof shoes for trail running.”

8K context window and support for more than 100 languages

Google said EmbeddingGemma 2 also expands how much content can be processed at once. The model has an 8K context window, four times larger than the previous generation.

For single-modality inputs, it can take in about 5.5 minutes of audio, 29 images or 58 video frames. Google said it supports more than 100 languages.

Google open-sources EmbeddingGemma 2, a 567MB multimodal retrieval model built to run offline on phones 4

In Google’s published comparisons against several models of similar size, the biggest gap showed up in image and video retrieval.

  • Image retrieval score: 57.3 for EmbeddingGemma 2 versus 31.6 for Jina v5 Omni-Nano.
  • Video retrieval score: 50.7 versus 31.2.
  • Multilingual text performance was roughly in line with the previous generation.

Based on those figures, EmbeddingGemma 2 led Jina v5 Omni-Nano by 25.7 points in image retrieval and 19.5 points in video retrieval, even though Google said the Jina model has about 30% more parameters.

567MB for the full multimodal setup, with modular loading

Google broke the 740 million parameters into three parts: 270 million for text, 170 million for the vision encoder and 300 million for the audio encoder. Those components can be loaded as needed.

Google open-sources EmbeddingGemma 2, a 567MB multimodal retrieval model built to run offline on phones 5

Text-only retrieval requires the 270 million-parameter text portion. Adding image search brings the total to 440 million parameters. On a Pixel 11 Pro, Google said the quantized text-only model used as little as about 191MB of memory, while the full multimodal model used about 567MB.

The model also supports a compressed search index that stores each item’s meaning with fewer numbers, cutting storage use by two-thirds. In Google’s multilingual text retrieval test, the score dropped by less than 1 point under that setup.

Google showed local-first apps and developers quickly picked it up

Google presented several application prototypes built around EmbeddingGemma 2.

  • Instant Media Search in AI Edge Gallery lets users search photos in a gallery by meaning.
  • Video Moments Finder can locate matching moments inside a video from a natural-language prompt.
  • Foresight pairs the model with Gemma 4 to search files and meeting notes offline on a Mac.

Developer tooling is already moving around the release. llama.cpp supports the model, and Unsloth has published a quantized version for local deployment.

Google open-sources EmbeddingGemma 2, a 567MB multimodal retrieval model built to run offline on phones 6

According to testing published by the Mac app Nativ, an 8-bit quantized build on an M5 Max reached a cosine similarity of 0.9997 against the original vectors and processed text embeddings at 817 items per second.

Developer tests pointed to gains in cross-language and typo-heavy search

Turkish developer Avenox shared a test that looked closer to everyday note retrieval. Most of his notes were written in English, but he usually asks questions in Turkish. With keyword matching, relevant notes could be missed when the wording did not line up.

He said Claude spent an hour building a test setup. He selected 40 notes, rephrased them, then queried in Turkish to see whether the correct note would appear in the top 18 candidates. In the results he published, keyword matching hit 15%, while EmbeddingGemma 2 reached 97%.

Google open-sources EmbeddingGemma 2, a 567MB multimodal retrieval model built to run offline on phones 7

He then tested 30 real messages containing typos. EmbeddingGemma 2 found twice as many relevant notes as keyword matching, with each query taking about 50 milliseconds and running entirely on-device.

Developer Nick Lo pushed the model onto a Nano development board. After a text prompt, the model took about 15 milliseconds to turn the sentence into a search vector. Once the matching image was found, an ESP32-S3 microcontroller drew it line by line.

Code retrieval also improved

Beyond multimodal search, Google said code retrieval improved sharply. On the MTEB Code benchmark, EmbeddingGemma 2 scored 78.68, up from 68.76 for the previous generation, a gain of nearly 10 points. Google said that result leads among models of similar size.

That matters for coding agents. Tools such as Claude Code and Codex typically need to search a codebase for relevant files and snippets before they can act. With the 270 million-parameter text-only version, developers can build a local index for a code repository and query it in plain language, keeping both indexing and retrieval on-device.

Google open-sources EmbeddingGemma 2, a 567MB multimodal retrieval model built to run offline on phones 8

From cloud APIs to on-device retrieval

Google had already released Gemini Embedding 2 in March as a cloud-based multimodal model accessed through an API and billed by usage. More than half a year later, it has moved multimodal retrieval into an open-source model small enough to run on a phone.

That gives photos, recordings and videos a local retrieval option. One setup Google highlighted pairs EmbeddingGemma 2 with Gemma 4 for offline, privacy-first RAG: one model retrieves the material, the other answers from it, and both steps stay on the device.

In many earlier workflows, AI search across photos, videos and recordings meant sending those files to the cloud. Google is now putting that capability into a package measured in a few hundred megabytes, making local search on personal devices much more practical.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.