Open Source
Open Source

EmbeddingGemma 2: The Open Embedding Model With Fine Print

Google released the open embedding model EmbeddingGemma 2 on October 6, 2026 (HF repo google/embeddinggemma-2; API-verified: ungated, license tag apache-2.0, 21,148 downloads). Gemma 4 architecture: a 270M text/code backbone plus 170M vision and 300M audio encoders = 740M total, loading in four tiers (270M/440M/570M/740M) from one checkpoint; text, code, images, audio and video share one 768-dimensional space, with 8,192-token context (4x the first gen). On-device ledger (official): about 191MB active RAM for text-only and 567MB full multimodal on a quantized Pixel 11 Pro; weights are a 1.49GB bfloat16 safetensors. Benchmarks are all Google self-reported (no independent eval, no paper, MTEB submission unaccepted): MTEB Code 68.76 to 78.68 is the only meaningful gain; multilingual text moved just +0.21 and English slightly declined - the story is "same text quality, new modalities, much smaller deployment," not "better everywhere." Matryoshka trimming takes 768 dims down to 128 for up to 6x storage savings. License nuance, honestly: Apache 2.0 declared in metadata and ungated (v1 was gated under the Gemma license - a real opening up), but the repo root has no LICENSE file and the model card retains a Gemma Prohibited Use Policy line. Ecosystem: sentence-transformers 6.1.0+, transformers 5.19.0, Ollama, llama.cpp, MLX, Qdrant; the first-gen model passed 20 million downloads.

Published October 9, 202610 min read
<!-- embeddinggemma-2-open-source-resource | open-source | EmbeddingGemma 2: The Open Embedding Model With Fine Print -->

An open embedding model is the piece of infrastructure most RAG teams quietly depend on but rarely discuss: the component that turns your text, and increasingly your images and audio, into vectors that a search or retrieval system can compare. On 2026-10-06, Google released EmbeddingGemma 2, a lightweight multimodal embedding model that pushes that idea further than most open options on the table. The weights live at google/embeddinggemma-2 on Hugging Face -- you can also copy the address directly: huggingface.co/google/embeddinggemma-2.

The pitch is attractive. One checkpoint that loads at four different sizes, a single 768-dimensional vector space shared across text, image, and audio, an 8,192-token context window, and a footprint small enough that Google says the text-only quantized variant runs in about 191MB of active memory on a Pixel 11 Pro. The license metadata says Apache 2.0 and the repo is ungated -- a real change from the first EmbeddingGemma, which required accepting terms behind a gate.

But this is a resource article, not an ad, so the fine print comes early. Every benchmark number attached to this release is Google self-reported. There is no independent evaluation yet, no paper, and Google's own MTEB submission has not been accepted. If you are choosing an embedding model, that matters more than any headline score, because swapping an embedding model later means re-embedding your entire corpus -- the decision is expensive to undo, which is exactly why the embedding model upgrade SOP on this site treats it as a one-shot decision.

This article is based on the official blog.google announcement and the Hugging Face repository metadata as verified on 2026-10-09. Numbers are attributed to their source throughout.

1. What EmbeddingGemma 2 actually is

EmbeddingGemma 2 is built on the Gemma 4 architecture and is a multimodal embedding model, not a chat model and not a generation model. Its job is to map inputs into vectors. The full stack is 740M parameters: a 270M text and code backbone, a 170M vision encoder, and a 300M audio encoder.

The genuinely interesting design choice is that all of this ships as one checkpoint with four loading tiers:

  • 270M -- text and code only
  • 440M -- adds the vision encoder
  • 570M -- adds the audio encoder
  • 740M -- full multimodal

You download one set of weights (1.49GB in bfloat16 safetensors format) and load whichever tier your use case needs. If you only embed text, you never pay the memory cost of the vision and audio encoders.

On top of that sits a unified 768-dimensional vector space. "Unified" is the key word: a text query and an image document land in the same space, which means text-to-image search and text-to-audio or video search work without a separate cross-modal bridge model. Embed text, embed screenshots, embed meeting recordings, search all of it with the same query vector. Most open embedding models still handle one modality per model; this is the part that is actually new for a model this small.

The context window is 8,192 tokens, four times the first EmbeddingGemma's 2,048. For RAG work that means longer chunks without splitting, which reduces the retrieval errors that come from cutting a document mid-argument. The model also covers 100+ languages, per Google's own materials.

2. The license: Apache 2.0 with an asterisk

This is the section you should read carefully, because the release notes compress it into one word.

The Hugging Face repository metadata, verified via API on 2026-10-09, declares the license as apache-2.0 and the repo is ungated -- no login wall, no acceptance click, 21,148 downloads already logged at verification time. Compare that with the first EmbeddingGemma, which was gated and shipped under the custom Gemma license. On the open-source spectrum, this is a genuine loosening.

Now the asterisk, in three parts:

  1. No LICENSE file at the repo root. The 15-file listing contains no LICENSE file. Apache 2.0 is asserted in the repo metadata and linked from the model card, but the conventional file that lawyers and compliance teams look for is absent. If your organization requires a LICENSE file before adopting an artifact, this is a real blocker you will need to resolve or document.
  2. A Gemma Prohibited Use Policy line survives. The model card still carries a line stating that use must comply with the Gemma Prohibited Use Policy. That is a usage restriction layered on top of the permissive license. For nearly all ordinary applications -- search, RAG, classification -- it changes nothing. For edge-case deployments, read the policy yourself.
  3. Metadata and model card, not a signed document. The Apache 2.0 declaration lives in repo metadata and a model card link. That is normal for Hugging Face, but "normal" is not the same as "clean."

The honest summary: this is materially more open than version 1, and the metadata says Apache 2.0, but it is not a friction-free Apache project. For most readers the practical answer is "usable, with a note in your compliance doc." For commercial products with legal review, check the exact wording before you commit, because the RAG embedding model comparison on this site shows alternatives with fully unambiguous licensing, like BGE-M3 under MIT.

3. The benchmarks: read them with the caveat attached

Here are the numbers, and then the caveat, which is not optional context but the whole story.

Google reports that EmbeddingGemma 2 scores 78.68 on MTEB Code, up from 68.76 for the previous generation -- a gain of 9.92 points, and the only improvement Google's own materials describe as significant. On multilingual text tasks, the model improves over version 1 by just 0.21 points. On English, the score slightly declines. Those are all Google self-reported figures.

The caveat: there is no independent benchmark yet. No third party has tested this model publicly, there is no paper, and the model's MTEB submission has not been accepted, so it does not appear on the public leaderboard where cross-model comparisons usually happen. Every number in this section is Google's own account of its own model, measured against its own previous release.

How should you read that? Three calibrations:

  • The code gain is the claim worth testing. A +9.92 jump on MTEB Code is a big number, and code retrieval is a workload where embedding quality differences are easy to feel. But it is exactly the kind of claim that independent evaluation exists to check. If code search is your use case, this model is worth a pilot -- on your data, not on Google's.
  • The multilingual and English flatness is honest. Google did not hide that English went slightly down and multilingual barely moved. That transparency is a good sign about the process, and it also tells you the headline here is not "better embeddings everywhere." It is "same text quality, new modalities, much smaller deployment."
  • Your corpus is the benchmark that counts. Because changing embedding models later means a full re-embed of everything -- the same point the embedding model upgrade SOP opens with -- the only numbers that should drive your decision are ones you generate on your own documents. And if your chunks were sized for a tighter context window, the chunking strategy guide is worth a re-read before you commit to the new dimensions.

One more framing number from Google's side: the first EmbeddingGemma has accumulated more than 20 million downloads. Version 2 launches into an audience that already exists, which speeds up ecosystem validation -- but popularity is not accuracy, and the 21,148 downloads on the new repo at verification time say nothing about quality either way.

4. Footprint: the on-device story

The second pillar of the pitch is size, and here the numbers are concrete even if they carry the Google-claim label.

Google says the quantized text-only tier runs in about 191MB of active memory on a Pixel 11 Pro, and the full multimodal stack in about 567MB. Those are on-device figures for a flagship phone, which tells you the design target: embedding models that run where the data is, without a server round trip.

Pair that with MRL -- Matryoshka Representation Learning, the nesting-doll trick where a 768-dimensional vector can be truncated to shorter lengths while keeping most of its retrieval quality. With EmbeddingGemma 2, you can cut from 768 dimensions down to 128, and Google claims vector storage savings of up to 6x. For a corpus of millions of chunks, storage is not a rounding error; it is a recurring bill. Being able to store 128-dim vectors and still get usable recall changes the economics of self-hosted retrieval.

Put together: a 270M-parameter text tier, 191MB of memory, 128-dim vectors. This is a model built to be the default embedding layer on small hardware -- phones, edge boxes, laptops -- as much as on servers.

5. Ecosystem: day-one support is real

A model is only as usable as its tooling, and this is where the release is strongest on verifiable ground facts.

Day-one support includes sentence-transformers 6.1.0+ and transformers 5.19.0. The model also runs through MLX (Apple Silicon), Ollama, llama.cpp, vLLM, and LiteRT/MediaPipe for mobile deployment. Qdrant has a storage partnership for the vector side, and Unsloth publishes a fine-tuning guide if you want to adapt the model to your domain.

Practical notes from the facts worth flagging for implementation:

  • If you use sentence-transformers, the task prompt prefixes are a hard requirement, not a suggestion -- queries follow a task: {description} | query: {content} format and documents a title: {title} | text: {content} format, with seven task types each carrying their own prefix. Skipping them silently degrades accuracy.
  • Ollama support means you can be querying this model locally within minutes of reading this article, which is the fastest way to form your own opinion about its retrieval quality.

6. Who should actually pick this up

After all the calibration, here is the decision map.

Good fit: teams building on-device or edge retrieval, where a 191MB text-only footprint and MRL-truncated vectors solve real constraints. Teams that need one model to embed text, images, and audio into a shared space without gluing together multiple models. Teams doing code search who want to test the +9.92 claim on their own repositories. Self-hosters who want an ungated download and permissive metadata without a terms-acceptance flow.

Questionable fit: teams whose workload is English text retrieval above all -- the self-reported English score slightly declined, and with no independent benchmark you are taking the trade on faith. Teams that need a LICENSE file for compliance and cannot proceed without one at the repo root. Teams that just want the best-scoring API and zero ops, where a managed embedding endpoint still wins on convenience.

The honest bottom line: EmbeddingGemma 2's real contribution is architectural -- four sizes from one checkpoint, one vector space across modalities, a small footprint with MRL headroom -- plus a genuinely loosened license compared with version 1. Its benchmark story is entirely self-reported and unproven externally. That combination makes it a model to pilot now on your own data, not one to crown on the strength of a leaderboard it has not yet appeared on.

FAQ

Q1: Is EmbeddingGemma 2 really open source?

A1: The repository metadata declares Apache 2.0 and the repo is ungated, which is a real loosening compared with version 1's gated access and Gemma license. But there are caveats: no LICENSE file exists at the repo root, and the model card still carries a Gemma Prohibited Use Policy line. Practically it is open enough for most uses, but compliance teams should read the exact wording rather than trusting the one-word license tag.

Q2: Can EmbeddingGemma 2 search images and audio with a text query?

A2: Yes, per Google's design: text, image, and audio all map into one unified 768-dimensional vector space, so a text query can retrieve images, audio, or video directly. Note that you need to load the larger tiers (440M with vision, 570M with audio, or the full 740M) to get those modalities; the 270M tier is text and code only.

Q3: How much better is it than the first EmbeddingGemma?

A3: By Google's self-reported numbers: MTEB Code jumps from 68.76 to 78.68 (+9.92, the only significant gain), multilingual text improves by only 0.21, and English slightly declines. There is no independent benchmark yet -- no paper and an unaccepted MTEB submission -- so treat these as vendor claims and validate on your own data before committing, especially since switching embedding models later requires a full re-embed.

Q4: How small is it, really, and what does MRL do?

A4: Google says the quantized text-only tier needs about 191MB of active memory on a Pixel 11 Pro, with the full multimodal stack at about 567MB. MRL lets you truncate the 768-dimensional vectors down to 128 dimensions, cutting vector storage by up to 6x while retaining most retrieval usefulness -- the combination targets phones and edge devices, not just servers.

Q5: How do I start using it today?

A5: Download the weights from the Hugging Face repo (1.49GB bfloat16 safetensors), then use sentence-transformers 6.1.0+ or transformers 5.19.0, both supporting the model from release day. Ollama, llama.cpp, MLX, and vLLM also work depending on your platform. Remember the task prompt prefixes are mandatory for best accuracy, and start by benchmarking on your own retrieval data before any migration.

This article is AI-assisted and human-edited. Last updated: 2026-10-09

FAQ

Is EmbeddingGemma 2 really open source?
The repository metadata declares Apache 2.0 and the repo is ungated, which is a real loosening compared with version 1's gated access and Gemma license. But there are caveats: no LICENSE file exists at the repo root, and the model card still carries a Gemma Prohibited Use Policy line. Practically it is open enough for most uses, but compliance teams should read the exact wording rather than trusting the one-word license tag.
Can EmbeddingGemma 2 search images and audio with a text query?
Yes, per Google's design: text, image, and audio all map into one unified 768-dimensional vector space, so a text query can retrieve images, audio, or video directly. Note that you need to load the larger tiers (440M with vision, 570M with audio, or the full 740M) to get those modalities; the 270M tier is text and code only.
How much better is it than the first EmbeddingGemma?
By Google's self-reported numbers: MTEB Code jumps from 68.76 to 78.68 (+9.92, the only significant gain), multilingual text improves by only 0.21, and English slightly declines. There is no independent benchmark yet -- no paper and an unaccepted MTEB submission -- so treat these as vendor claims and validate on your own data before committing, especially since switching embedding models later requires a full re-embed.
How small is it, really, and what does MRL do?
Google says the quantized text-only tier needs about 191MB of active memory on a Pixel 11 Pro, with the full multimodal stack at about 567MB. MRL lets you truncate the 768-dimensional vectors down to 128 dimensions, cutting vector storage by up to 6x while retaining most retrieval usefulness -- the combination targets phones and edge devices, not just servers.
How do I start using it today?
Download the weights from the Hugging Face repo (1.49GB bfloat16 safetensors), then use sentence-transformers 6.1.0+ or transformers 5.19.0, both supporting the model from release day. Ollama, llama.cpp, MLX, and vLLM also work depending on your platform. Remember the task prompt prefixes are mandatory for best accuracy, and start by benchmarking on your own retrieval data before any migration.

Related

Open Source

Qwen-Image 2.1 Open Weights Are Back, But Apache Is Not

Alibaba's Qwen team open-sourced Qwen-Image-2.1 on September 20, 2026: text-to-image, image editing and native RGBA transparency in one 7B checkpoint - a 32-layer single-stream DiT (mixed-granularity attention plus prefix KV reuse) + a Qwen3-VL 8B text encoder (which also encodes reference images) + a 64-channel RGBA VAE. About 33GB in BF16, compressible to roughly 14GB via Comfy-Org quantized repacks. Capabilities: native 2048x2048 (seven ratios up to 2752x1536), 40-step inference, up to 10 reference images per call, circle/mask local edits, direct transparent-PNG output (official prompt phrasing provided), plus two 9B prompt rewriters. Reported throughput: vLLM-Omni at 1024 squared, 40 steps, BF16 peaks at 34GB and 3.3-4.5s per image (single source); Alibaba says a 3090 runs it with offload. Day-zero ecosystem: Diffusers, ComfyUI, vLLM-Omni, SGLang, LightX2V; 14 fine-tunes and 18 quantizations within the first day. The core of this piece: every prior Qwen-Image shipped Apache 2.0 - 2.1 switches to the Qwen Research License (non-commercial only; commercial use needs a separate agreement; fine-tunes inherit the restriction and must display "Built with Qwen"). The community "license trap" pushback and the "License renders this model useless" thread are reported as-is. Qwen-Image-Bench 60.28 (seventh) is vendor-designed and vendor-run. Lineage section: 3.0 is the API-only fork, 2.1 is the downloadable line - the weights came back, the Apache license did not.

Oct 8, 202610 min read
Open Source

Bilibili's 35B translation MoE: 150 languages, only 3B active

Bilibili's Index LLM team open-sourced the Index-Translate family on September 30, 2026 (weights on Hugging Face and ModelScope, Apache-2.0): 2B/9B/35B-A3B-preview text models built on Qwen3.5, targeting 150 languages (self-reported), with the 35B MoE activating only about 3B parameters per token and a 262,144-token context. Three branches split the work: Index-Echo produces dubbed speech that preserves the source speaker's voice (Chinese to/from English, Spanish, Japanese); Index-Homura controls syllable counts (the 9B lands within +-10% of the target on 81.92% of SandGlass cases - built for dubbing timing); Index-NativeLong (aka Index-Nailong) keeps long documents consistent (a 32K-token fantasy text keeps a name/royal-title pun intact, where chunk-by-chunk translation drifts). Constrained translation via instTrans: hard constraints enforce glossaries (e.g. carbon fiber) plus JSON/CSV/code/placeholder structure; soft constraints cover tone, domain disambiguation (plant to factory), cross-sentence consistency and LaTeX preservation. Local deployment is one line with the official GGUF (Q4_K_M, 21.71GB) via llama.cpp serve, plus a browser extension and a video dubbing pipeline. Honesty note: every score is self-reported with no independent replication yet (FLORES COMET-22 0.8794, WMT26 Judge 76.76, 2.4% off-target on low-resource pairs), and in the same official table DeepSeek-V4.1-Flash scores higher on WMT26 Judge at 83.55 - not first place on raw quality.

Oct 7, 202610 min read
Open Source

NVIDIA Locks Down AI Agents: Open Runtime, Silicon Watchdog

A deep dive into NVIDIA's Open Agent Safety Platform (official materials plus multi-source reporting, October 5, 2026 basis). Core thesis: agent safety cannot rest on the model behaving itself - when the agent is itself the software trying to escape its guardrails, application-layer protections fail, and agent drift (gradual divergence from operator intent via vague prompts, missing tools or policy conflicts) triggers no alarm. Two layers: OpenShell (Apache 2.0 open source, a kernel-level sandboxed secure runtime whose policies cover files, processes, credentials, tools, network and databases; intent-alignment checks before execution plus continuous drift monitoring; runs on Vera CPUs, extensible to Arm and Intel) and NVIDIA Sentry (an out-of-band watchdog on BlueField-4 DPUs that verifies agent identity via DOCA, produces attested telemetry, enforces granular policy, and quarantines an escaping agent within milliseconds; a reference design, not open hardware). Three layers stay separate: application, runtime, infrastructure. Over 100 organizations are on board (Anthropic integrating it with Claude Managed Agents, Salesforce for Slack, SpaceXAI for Cursor agents) under the Linux Foundation's Open Secure AI Alliance. The criticism is reported faithfully: Gartner notes OpenAI, Amazon and Google are absent; IDC estimates it addresses under 25% of enterprise agentic security problems; Control Risks notes it only governs known agents on infrastructure you own (shadow agents, SaaS-embedded ones and attacker-delivered ones are out of reach); lock-in risk discussed. Together with open-weights Kimi K3 it marks the two ends of the agent-safety map.

Oct 5, 202610 min read