← All articles

EmbeddingGemma 2 Makes On-Device Multimodal Search Work

2026-10-07 — Michael Leung

Part 6 of 6 · Local AI on an 8GB GPU

Useful already? ☕ Buy me a coffee and keep these articles free and ad-free.

EmbeddingGemma 2 Makes On-Device Multimodal Search Work

I've spent a lot of time working with local RAG pipelines, and until now, adding multimodal features always meant making a tough choice. You could either make your app bulky by combining separate vision, audio, and text models, or just send everything to a paid cloud API.

But on 6 October 2026, Google released EmbeddingGemma 2, which removes that compromise completely.

EmbeddingGemma 2 is built on the Gemma 4 architecture. It's an open model that maps text, code, images, video, and audio into one unified 768-dimensional vector space. It does all this natively with only 740 million parameters, understands 100+ languages, and has an 8K-token context window (4x larger than the first EmbeddingGemma).

Here's why this is important if you're building local apps:

Cross-Modal Search That Really Works

Since every type of data maps to the same vector space, you can search across them easily. For example, you could find a folder of video clips with a voice memo, or locate a specific moment in a long podcast by typing a text query. There's no need to transcribe audio or add captions to video first. The model understands the raw inputs. In one request it can take up to 5.5 minutes of audio, 29 images or 58 video frames (sampled at one frame per second by default), or a mix of them.

Cross-modal search: a voice memo finds the matching video frame in one shared vector space

Fully Local and Modular

The full multimodal model uses about 567MB of RAM (Google's figure on a Pixel 11 Pro), so it runs easily on a laptop or even a phone. If you don't need audio or video, you can skip loading those parts. The modular design lets you scale down to just 270 million parameters (about 191MB of RAM) for text and code retrieval only.

What you loadParametersWhat it can embed
Text backbone270MText and code (~191MB active RAM on a Pixel 11 Pro)
+ Vision encoder440MText, code, images, video frames
+ Audio encoder570MText, code, audio
Full model740MEverything (~567MB on a Pixel 11 Pro)

On code it's also a clear step up: Google reports MTEB Code rising from 68.76 to 78.68 compared with EmbeddingGemma 1. It already works with Sentence Transformers (v6.1.0+), Transformers, vLLM, SGLang, MLX, Ollama, LM Studio, LiteRT and Unsloth. If you're weighing local against cloud for the rest of your stack, see my local LLM vs cloud API guide.

EmbeddingGemma 2

Shrink Your Vector Database by 6x

EmbeddingGemma 2 uses Matryoshka Representation Learning (MRL). Simply put, you can reduce the embeddings from 768 dimensions to 512, 256 or even 128. At 128 dimensions that is a 6x storage saving: Google's example is 1 million vectors shrinking from about 1.5GB to about 250MB. At 256 dimensions retrieval quality hardly changes. At 128, text and code still hold up well, but video and audio give up noticeably more, so test it on your own data first.

EmbeddingGemma 2: quality kept when embeddings are truncatedShare of the 768-dimension benchmark score kept at 512, 256 and 128 dimensions. At 256 dimensions every modality keeps at least 95%. At 128 dimensions text keeps 96%, image 91%, code 91% and video 77%.Quality kept vs. the full 768-d embedding70%80%90%100%768d512d256d128dEmbedding size (dimensions)Text (MTEB English), 768d: 100.0% kept (score 68.46)Text (MTEB English), 512d: 99.9% kept (score 68.41)Text (MTEB English), 256d: 99.0% kept (score 67.78)Text (MTEB English), 128d: 95.9% kept (score 65.68)Code (MTEB Code), 768d: 100.0% kept (score 78.68)Code (MTEB Code), 512d: 98.2% kept (score 77.24)Code (MTEB Code), 256d: 96.8% kept (score 76.18)Code (MTEB Code), 128d: 90.8% kept (score 71.41)Image (MIEB Lite), 768d: 100.0% kept (score 64.64)Image (MIEB Lite), 512d: 99.5% kept (score 64.32)Image (MIEB Lite), 256d: 97.7% kept (score 63.13)Image (MIEB Lite), 128d: 91.4% kept (score 59.06)Video (MMEB v2), 768d: 100.0% kept (score 59.01)Video (MMEB v2), 512d: 98.9% kept (score 58.38)Video (MMEB v2), 256d: 95.3% kept (score 56.24)Video (MMEB v2), 128d: 77.4% kept (score 45.65)● Text 96%● Image 91%● Code 91%● Video 77%
Share of the 768-d score kept after truncation. Source: official EmbeddingGemma 2 model card (Google, Hugging Face).
Benchmark768d512d256d128d
MTEB English v268.4668.41 (100%)67.78 (99%)65.68 (96%)
MTEB Multilingual v261.3661.17 (100%)60.41 (98%)57.89 (94%)
MTEB Code v178.6877.24 (98%)76.18 (97%)71.41 (91%)
MIEB Lite (image)64.6464.32 (100%)63.13 (98%)59.06 (91%)
MMEB v2 (video)59.0158.38 (99%)56.24 (95%)45.65 (77%)
MSEB (audio retrieval)69.5469.18 (99%)66.76 (96%)56.71 (82%)
Matryoshka Representation Learning: smaller embeddings nested inside the full 768-dimension vector

Zero-Shot Decision Engine

Because it runs locally, you skip the network latency entirely. Google's AI Edge team demonstrated the model as a real-time decision engine in a chess game, evaluating 500 options per turn in less than 100 milliseconds. That speed unlocks instant intent routing directly on user hardware, without requiring any fine-tuning. Google also lists image embeddings at 37.3 ms each (26.9 images per second) on a MacBook M5 Pro GPU, and has added two demos to the Google AI Edge Gallery app: Instant Media Search and Video Moments Finder. ML Kit support for Android apps is coming in the next few weeks. If decision models interest you, I covered a similar idea in Ollama JEV-style decision models.

My Take

I'm usually skeptical of corporate "best-in-class" claims, but fitting text, code, images, video, and audio into one shared vector space under 1GB of RAM is a big change for local development. The model uses open weights under an Apache 2.0 license, so you don't have to worry about API keys or token limits. Together with the image model I run on my own 8GB card (Qwen Image 2.1 on an 8GB GPU), a fully local media pipeline is now realistic.

Download EmbeddingGemma 2 on Hugging Face

Sources

Checked 7 October 2026, Google sources only: Google Blog announcement; EmbeddingGemma 2: The Developer Guide; Google AI Edge with EmbeddingGemma 2; Gemma docs; official model card (google/embeddinggemma-2). RAM figures are Google's Pixel 11 Pro measurements, not my own. Research and chart prepared with AI assistance; images generated locally with Qwen Image 2.1.

Series: Local AI on an 8GB GPU

Running real AI models on an ordinary 8GB graphics card, and when the cloud is the better deal.

  1. The Death of Local LLMs? How Xiaomi's MiMo V2.6 Just Flipped the AI Economics in Australia
  2. Unlocking Local AI: Running Qwen-Image-2.1 on an 8GB GPU
  3. Local AI Gets a Big Boost: How Ollama's New Jev-Style Decision Models Change Everything
  4. Run Qwen3.8-Flash-Next on an 8 GB RTX 4060 Ti (Here's How)
  5. The Great LLM Debate 2026: Local Hardware vs. Cloud APIs (An Australian Dev's Guide)
  6. EmbeddingGemma 2 Makes On-Device Multimodal Search Work
Start the series from part 1 →

← All articles