
EmbeddingGemma 2, explained: 740M parameters that turn images, recordings and video into vectors your agent can search locally
Google's open EmbeddingGemma 2 maps text, code, images, audio and video into one 768-d space with 740M params (270M text-only). We unpack modular loading, MRL truncation and on-device RAG.
An agent that can only answer questions is half an agent. The other half is finding the material the answer depends on, and that material is rarely a neat paragraph of text. It is a screenshot buried in a chat thread, a forty-minute meeting recording, a twelve-second clip where the arm finally seats the connector. On October 6, 2026 Google released EmbeddingGemma 2, an open multimodal embedding model built on the Gemma 4 architecture and licensed under Apache 2.0: about 740 million parameters that map text, code, images, audio and video into one shared 768-dimensional vector space, sized to run on phones and laptops rather than in a datacenter.
It is worth being precise about what this release is not. EmbeddingGemma 2 is not a chat assistant, and it is not a search application you install and point at your hard drive. It is the layer underneath those products: the component that turns material into vectors a search system or an agent can compare. Everything below follows from that positioning.
1. You do not have to turn everything into text first
The common recipe for multimodal retrieval is a relay: caption every image, transcribe every recording, then hand the resulting text to a text embedding model. That pipeline works, and plenty of production systems run it. But every relay leg is another model to host, another latency hop, and another place where meaning leaks out of the signal. A caption says "a dog in a park"; the query was "the moment the dog catches the frisbee", and the moment lives in the frames, not in the caption.
EmbeddingGemma 2 offers the other route: encode each modality directly, so the captioning model and the speech-to-text model drop out of the critical path. Search for "waves breaking on the shore" and the candidates can be photographs or audio clips, ranked against the same query vector. This is cross-modal retrieval in the literal sense, and it comes with the honest caveat that semantic proximity is not a guarantee: some queries will still miss, and the miss rate is exactly what you measure on your own corpus before you trust the system.
The architectural point is not "every modality outputs 768 numbers". Any projection head can do that. The point is that the representations are aligned: cosine similarity between a query vector and a clip vector means something across modalities, which is what makes one index serve every file type.
One concept needs untangling here, because the marketing shorthand blurs it. Google lists PDFs, slides and charts among the visual documents the model handles. That does not mean every integration path accepts a .pdf file. In practice you render pages to images or extract their text, then encode what the interface actually takes. Skipping the captioning model does not skip file parsing, image resizing or video frame sampling; it skips the generative middle step, nothing more.
2. 740M is the full model; text-only workloads need 270M
The model is modular, and the modularity is the reason a 740M multimodal embedder can live on a phone. The text module carries about 270M parameters, the vision encoder about 170M, the audio encoder about 300M. You load the encoders your workload touches:
| Workload | Components loaded | Parameters |
|---|---|---|
| Text and code retrieval | Text module | 270M |
| Text, image and video-frame retrieval | Text module + vision encoder | 440M |
| Text and audio retrieval | Text module + audio encoder | 570M |
| Full multimodal retrieval | All three components | 740M |
These configurations come from a single checkpoint and share one vector space, which unlocks a useful asymmetry: encode the query with the cheap text-only configuration and match it against an index built with the full multimodal configuration. A search box should not pay for the audio encoder just because the corpus contains recordings.
Google publishes concrete on-device numbers for the quantized build: on a Pixel 11 Pro, active RAM drops to roughly 191MB for the text-only weights and about 567MB for the full multimodal model. Read those figures for what they are. They describe active weights on one device with one quantization recipe. The media pipeline, the inference runtime, the vector index and any generative model running alongside all bill separately. A "567MB model" inside a 4GB budget is not the same claim as a 567MB application.
Context is a shared budget, not per-modality allowances. The model takes 8,192 tokens, four times the first EmbeddingGemma, and at Google's default settings that fits roughly 29 images, 58 video frames or about 5.5 minutes of audio. An interleaved input divides the same budget among its parts, so an hour-long recording or a feature-length video still gets indexed in chunks. Nothing about MRL or modularity changes that arithmetic.
3. The index shrinks too: Matryoshka vectors
Model weights are a one-time download; the vector index is what grows forever. EmbeddingGemma 2 is trained with Matryoshka Representation Learning, which forces the leading dimensions of the vector to carry meaning on their own, so a 768-dimensional embedding can be truncated to 512, 256 or 128 dimensions after the fact. Truncation requires renormalization, and the query side and the index side must agree on the dimension, or you are comparing vectors from two different spaces.
| Dimensions | Raw float32 storage for 1M vectors | Reduction vs 768-d |
|---|---|---|
| 768 | 3.07 GB | 1.0x |
| 512 | 2.05 GB | 1.5x |
| 256 | 1.02 GB | 3.0x |
| 128 | 0.51 GB | 6.0x |
The arithmetic is items times dimensions times four bytes, and it covers the vectors alone: no index structure, no metadata, no original media. Google's "up to 6x storage reduction" refers to exactly this line item, which for a local vector database is often the dominant one, but it is not a claim that your whole application shrinks sixfold.
Compression is not free, and the model card says so plainly: at 128 dimensions multimodal quality degrades noticeably, and the short vectors are a better fit for text-first workloads. The sane selection procedure is unglamorous. Build a 768-dimensional baseline, measure recall on your own queries, then re-run at 512 and 256 and keep whichever dimension survives your recall floor. Image, video and audio retrieval in particular should not be traded away for a storage number.
4. Code retrieval is the headline gain; benchmark scores are not accuracy
| Benchmark | EmbeddingGemma (v1) | EmbeddingGemma 2 | Delta |
|---|---|---|---|
| MTEB Code | 68.76 | 78.68 | +9.92 |
| MTEB Multilingual | 61.15 | 61.36 | +0.21 |
Read those two rows together and the story is specific rather than generic: multilingual text quality holds roughly steady while code retrieval jumps by almost ten points. This is not an "everything improved" press release. The gain that matters is code, plus the new multimodal coverage, and the code gain is the one with immediate product consequences: local codebase indexing, semantic code search, and retrieval for coding agents.
For a deployment-shaped decision, reframe the question away from leaderboards: at the same latency and memory budget, does this model retrieve what your current stack misses? Can a natural-language description of a feature find the implementation? Can a question about a result find the chart on page nine of the paper? Can a recording be localized to a segment short enough to hand to a generative model? Those three questions are a better test set than any public benchmark, and they are cheap to assemble from material you already own.
5. For agents, the prize is a larger retrievable context
Google ships reference applications that show the intended shape. In the AI Edge Gallery, Instant Media Search ranks your local photos and clips against a text or image query, and Video Moments Finder locates the moment inside a video from a text or audio query. On macOS, AI Edge Foresight pairs EmbeddingGemma 2 with Gemma 4 for local meeting assistance and file retrieval. Because the embedder is built on Gemma 4 and shares its text tokenizer and audio encoder, the two models run as one pipeline with a lower combined footprint than two unrelated models would.
From those examples it is a short step to a research-assistant design: index paper pages, experiment screenshots, meeting clips and code files as separate collections, then let one query recall across all of them. Ask "what evidence do we have for the retrieval approach we discussed last week" and the system can return a code diff, a meeting segment and a results chart in one ranked list, instead of whatever someone happened to write down in a summary document.
flowchart LR Q["query: text or code"] --> TB["text module 270M"] I["images / video frames"] --> VE["vision encoder 170M"] A["audio"] --> AE["audio encoder 300M"] TB --> BB["shared backbone"] VE --> BB AE --> BB BB --> V["768-d vector, MRL-truncatable to 512/256/128"] V --> IDX["local vector index"] Q --> QV["query vector, text-only config"] QV -->|"cosine similarity, top-k"| IDX IDX --> HIT["retrieved chunks with source, page, timestamp"] HIT --> G4["Gemma 4 reads the evidence and answers"]
The diagram also marks the boundary of what embedding buys you. Retrieval and comprehension remain two steps: an embedding model outputs vectors, and vectors cannot read evidence or organize an answer. Google's own local RAG recipe keeps the division of labor explicit, with EmbeddingGemma 2 retrieving and Gemma 4 reasoning. The engineering consequences are concrete: keep the original files, keep permissions, page numbers and timestamps attached to every chunk, and hand the actual content to a model that can perceive the modality in question. Vector similarity is not established evidence, and retrieving a video is not understanding it.
6. A minimal example: one query against an image and a sound
The developer guide documents a Sentence Transformers path, and the multimodal loaders require version 6.1.0 or newer. The smallest honest example puts one image and one audio file next to the script, encodes a text query, and scores both media against it:
pip install -U "sentence-transformers[image,audio,video]>=6.1.0" transformers
from pathlib import Path
import torch
from sentence_transformers import SentenceTransformer
# Point these at your own local files.
media = {"image": Path("beach.jpg"), "audio": Path("waves.wav")}
for path in media.values():
if not path.is_file():
raise FileNotFoundError(f"missing media file: {path.resolve()}")
use_cuda = torch.cuda.is_available()
device = "cuda" if use_cuda else "cpu"
dtype = (
torch.bfloat16
if use_cuda and torch.cuda.is_bf16_supported()
else torch.float32
)
model = SentenceTransformer(
"google/embeddinggemma-2",
device=device,
model_kwargs={"torch_dtype": dtype},
)
query = model.encode(
"waves breaking on the shore",
prompt_name="SearchQuery",
normalize_embeddings=True,
)
for modality, path in media.items():
candidate = model.encode(
{modality: str(path)},
normalize_embeddings=True,
)
score = model.similarity(query, candidate).item()
print(f"{path.name}\tsimilarity: {score:.4f}")
Nothing in this snippet captions the image or transcribes the audio; the comparison happens entirely in vector space, and the printed number is a semantic similarity, not a probability of being correct. Turning it into a search system means persisting candidate vectors, building an index and ranking over it, which is where the MRL dimension choice from section 3 starts to matter.
Two deployment details are easy to miss and expensive to rediscover. First, the official guidance is bfloat16 or float32: switching to float16 can produce NaNs or silent quality loss. Second, this snippet is a desktop path with full-precision weights. The 191MB and 567MB figures from section 2 describe the quantized on-device build and say nothing about this process's resident memory.
What to take away
EmbeddingGemma 2 does not make agents more articulate. It makes the non-text half of your material findable, which for most real assistants is the harder half. A small open model, one aligned space for five modalities, and an on-device memory profile together add up to a retrieval path you can run without sending the corpus anywhere. The remaining work is the work that was always yours: decide what relevance means for your queries, measure recall before you truncate dimensions, and keep the evidence chain intact from vector back to file, page and timestamp.
Sources: the original Chinese analysis by ChallengeHub on WeChat (mp.weixin.qq.com/s/nGMVHB7po_tkitfc4RKWUw); Google's announcement, "EmbeddingGemma 2: an open, lightweight multimodal embedding model" (blog.google, October 6, 2026); model weights and model card on Hugging Face (litert-community/embeddinggemma-2-740m-litert-lm). Figures are Google's official diagrams as reproduced in the original article.
Source:ChallengeHub (WeChat)https://mp.weixin.qq.com/s/nGMVHB7po_tkitfc4RKWUw

