Your RAG system's retrieval ceiling is set long before the language model writes a single word. It is set by the embedding model: the component that compresses documents and queries into vectors a machine can compare. Pick a weak one and no prompt engineering rescues you. Pick one that mismatches your deployment shape and you discover, months later, that switching means rebuilding your entire index from scratch.
This review lines up the four embedding models that dominate real-world RAG conversations in October 2026: Google DeepMind's EmbeddingGemma 2, Alibaba Qwen's Qwen3-Embedding, BAAI's BGE-M3, and OpenAI's text-embedding-3-large. Google's gemini-embedding-001 joins as a reference row, because it frames the central question of this model generation: what does the open-source route trade away against a frontier API model, and what does it gain back?
Ground rules first, as with every comparison on this site: we ran no proprietary stress tests, and we make no capability arbitration. Every benchmark figure below is cited with the table it comes from, self-reported numbers stay inside their own labeled context, and conclusions are organized by use case rather than by a crowned winner.
Why most embedding model comparisons mislead
The most common failure in this genre is also the least visible: lifting scores out of context and stacking them into one ranking. The MTEB ecosystem spans multiple boards, including MTEB English, MTEB Multilingual, and MTEB v2, each with a different task mix, language coverage, and evaluation conditions. A number published on one board is not commensurable with a number on another, even when both are casually called MTEB.
This review therefore enforces one citation rule without exception: every vendor's figures are quoted with the name of the table they were reported on, no cross-table ranking appears anywhere in this article, and provenance is always labeled. A vendor's own reporting says so, third-party relay is marked as relayed, and unpublished figures stay unpublished. EmbeddingGemma 2's code score and Qwen3-Embedding's multilingual score never meet in a sentence implying one beat the other. There is no third category in honest tech writing.
The four contenders at a glance
| Dimension | EmbeddingGemma 2 | Qwen3-Embedding | BGE-M3 | OpenAI text-embedding-3-large |
|---|---|---|---|---|
| Vendor | Google DeepMind | Alibaba Qwen | BAAI | OpenAI |
| Parameters | 270M to 740M, four load tiers | 0.6B / 4B / 8B, three sizes | About 568M | Undisclosed |
| Context window | 8,192 tokens | 32,768 tokens | 8,192 tokens | 8,191 tokens |
| Output dimensions | 768 (MRL down to 128) | 1024 / 2560 / 4096 (MRL) | 1024 | 3072 (MRL down to 256) |
| License (HF metadata verified 2026-10-09) | apache-2.0, ungated | apache-2.0, ungated | mit, ungated | Closed API |
| Price | Free, self-hosted | Free, self-hosted | Free, self-hosted | $0.13 per million tokens (official pricing page snapshot (2026-10-09)) |
| Headline specialty | On-device footprint, four modalities in one vector space | 32K context, task instruction prefixes | One model, three retrieval modes: dense, sparse, multi-vector | Zero ops, MRL support |
A reference row completes the picture: gemini-embedding-001 is API-only, outputs 3072 dimensions from a 2,048-token context, and, per Google's own reporting, holds top placement on three MTEB boards: multilingual, English, and code. If a closed, frontier-class model tops the boards its own maker reports on, the open-source contenders are not competing to dethrone it there; they are competing on what those boards never measure: deployment cost, data boundary, modality coverage, and license freedom.
Multilingual coverage, where published: EmbeddingGemma 2 lists 100+ languages, Qwen3-Embedding 100+ with reported strength in Chinese, BGE-M3 100+. OpenAI's total is unpublished, and we will not invent one.
EmbeddingGemma 2: the on-device multimodal play
Weights live at huggingface.co/google/embeddinggemma-2. A direct check of the Hugging Face API on 2026-10-09 confirmed the essentials: license apache-2.0, ungated, 21,148 downloads since the 2026-10-06 release.
Built on the Gemma 4 architecture, EmbeddingGemma 2 stacks a 270M text-and-code backbone with a 170M vision encoder and a 300M audio encoder, for 740M parameters at full load. The same checkpoint loads at four tiers: 270M text-only, 440M with vision, 570M with audio, and 740M full multimodal. Every tier projects into one unified 768-dimensional vector space, so text-to-image and text-to-audio search share one index, with an 8,192-token context that is four times the first generation's 2,048.
The deployment numbers, per Google's figures: quantized, text-only embedding runs in roughly 191MB of active memory on a Pixel 11 Pro, and full multimodal in roughly 567MB. MRL lets you truncate the 768-dimension output down to 128, cutting vector storage by as much as six times. The full checkpoint ships as 1.49GB of bfloat16 safetensors.
The license deserves a careful read. Apache 2.0 is declared in the repository metadata and linked from the model card, but the repo root contains no standalone LICENSE file, and the model card retains a line requiring compliance with the Gemma Prohibited Use Policy. Against the first generation's gated download and custom Gemma license, this is a genuine opening-up, just not the unconditional kind the bare string apache-2.0 might suggest.
Benchmarks: all figures are Google's own reporting, with no independent evaluations, no paper, and no accepted MTEB submission at review time. On Google's own MTEB Code table, EmbeddingGemma 2 moved from 68.76 to 78.68, a gain of 9.92 and the only clearly significant improvement. Multilingual text improved by just 0.21 over the first generation, and English registered a slight decline. This model is not a leaderboard play; it is a deployment play. Task prompt prefixes are a hard requirement per the model card, with queries formatted as task: {task description} | query: {content} and documents as title: {title} | text: {content}. Skipping them silently costs precision.
The ecosystem arrived fast: sentence-transformers 6.1.0+ and transformers 5.19.0 supported it on release day, with MLX, Ollama, llama.cpp, vLLM, and LiteRT/MediaPipe alongside, a Qdrant storage cooperation, and an Unsloth fine-tuning guide, while the first generation crossed 20 million cumulative downloads.
Qwen3-Embedding: three sizes, 32K context
Code and weights live at github.com/QwenLM/Qwen3.
Qwen3-Embedding ships in three sizes: 0.6B, 4B, and 8B parameters, with output dimensions of 1024, 2560, and 4096 respectively, each supporting MRL truncation. The standout spec is the context window: 32,768 tokens, four times the 8K-class windows of EmbeddingGemma 2, BGE-M3, and OpenAI's model. For retrieval over long documents, that difference changes your chunking strategy, not just your accuracy.
License status was verified against Hugging Face repository metadata on 2026-10-09: apache-2.0 and ungated, with the 8B repository showing roughly 2.446 million downloads. That is a commercially clean, self-hostable license at every size tier, which matters for teams that need legal simplicity across a model family.
Benchmark discipline applies with full force here. The Qwen3-Embedding-8B figure of about 70.6 on MTEB multilingual is self-reported and reached us through third-party relay; we cite it with exactly that provenance: one table, one label, no comparisons. What we can say without a scoreboard: 100+ language coverage with reported strength in Chinese, instruction-aware task prefixes, and a three-size scaling path make it the natural candidate when your corpus is long, multilingual, or Chinese-heavy and you want to right-size compute.
BGE-M3: one model, three retrieval modes
Code and weights live at github.com/BAAI/FlagEmbedding.
BGE-M3 runs about 568M parameters, outputs 1024 dimensions from an 8,192-token context, and was verified on 2026-10-09 as mit-licensed and ungated, with roughly 33.69 million downloads on Hugging Face: the most permissive license in this table and the largest verified download count. MIT means no usage-policy rider for your legal team to read.
Its defining feature is retrieval versatility: a single model returns dense, sparse, and multi-vector outputs. That is one checkpoint covering three retrieval paradigms, semantic dense search, lexical-style sparse matching, and late-interaction-style multi-vector scoring, which means hybrid retrieval pipelines without bolting a second model onto your stack.
Popularity, marked as relayed rather than measured: per LangChain and LlamaIndex telemetry as relayed by presenc, BGE-M3 ranks first in third-party production deployment popularity, ahead of Qwen3-Embedding in second, OpenAI's text-embedding-3-large in third, and Voyage in fourth. That is a popularity ranking, not a capability verdict, and per this review's citation rule we attach no cross-table benchmark score to it here.
OpenAI text-embedding-3-large: the zero-ops option
OpenAI's flagship embedding model outputs 3072 dimensions with MRL truncation down to 256, reads an 8,191-token context, and keeps its parameter count undisclosed, so we report it as undisclosed rather than guessing. Pricing is $0.13 per million tokens (official pricing page snapshot (2026-10-09)).
The value proposition is operational, not architectural: no GPUs to provision, no inference stack to maintain, no version pinning across your fleet, and MRL support to trim storage downstream. For teams whose core product is not retrieval infrastructure, that is frequently the deciding factor: the embedding layer stops being anyone's job.
The trade is equally plain. No weights to download and no self-hosting path, so every embedded document crosses your data boundary; per-token cost scales with corpus size and re-embedding frequency; and the total language count has not been published. Consistent with this review's citation rule, we attach no benchmark score to this model here: its self-reported figures come from its own tables, and folding them into one ranking would be exactly the cross-table mistake this article refuses to make.
Choosing by use case, not leaderboard
If the four contenders ranked cleanly on a single axis, this section would be one line long. They do not, so here is the map instead.
Run on-device, at the edge, or across modalities. EmbeddingGemma 2 is the only contender in this table designed for it: a 191MB quantized text-only footprint per Google's figures, four modalities sharing one 768-dimension space, MRL down to 128 dimensions for up to six-fold storage savings, and a four-tier checkpoint that lets you ship text-only first and grow into vision and audio without changing vector spaces. Accept the trade: its self-reported gains concentrate in code, while multilingual moved little and English slipped slightly.
Long documents, Chinese-heavy or broadly multilingual corpora, license-clean self-hosting. Qwen3-Embedding: a 32K context that reshapes chunking strategy, three sizes to right-size latency and memory, apache-2.0 verified at every tier, and 100+ languages with reported strength in Chinese.
Hybrid retrieval and a proven production track record. BGE-M3: dense, sparse, and multi-vector from one checkpoint, an MIT license with no riders, and, per relayed third-party telemetry, the largest production install base among these models. When you want the least-surprising choice, this is the least-surprising choice.
Managed, zero-ops, pay-as-you-go. OpenAI text-embedding-3-large: $0.13 per million tokens (official pricing page snapshot (2026-10-09)), 3072 dimensions truncatable to 256, and an embedding layer that stops being anyone's job. Accept the trade: no weights, and a data boundary crossing on every embedded document.
Two closing notes outlast any individual model. gemini-embedding-001's top placement on three MTEB boards per Google's own reporting is a reminder that the open models here are not chasing that trophy. Switching embedding models means a full re-embed of your entire corpus, so the decision is expensive to revisit. And if retrieval quality is lagging today, relayed production experience suggests adding a reranker first, a minutes-level change compared to a full re-embed.
FAQ
Q1: Can I directly compare the benchmark scores mentioned in this article? No. The figures cited here come from different tables: Google's own MTEB Code table for EmbeddingGemma 2, a self-reported MTEB multilingual figure relayed through third parties for Qwen3-Embedding-8B, and Google's own reporting for gemini-embedding-001. Different MTEB boards use different task mixes and conditions, so the numbers are not commensurable, and this article presents no cross-table ranking.
Q2: Is Qwen3-Embedding really Apache 2.0 licensed? Yes. Verified against Hugging Face repository metadata on 2026-10-09: apache-2.0 and ungated, with the 8B repository showing roughly 2.446 million downloads. Code and weights live at github.com/QwenLM/Qwen3.
Q3: What does it cost to switch embedding models after launch? A full re-embed of your entire corpus: re-encoding every document, rebuilding the vector index, and re-validating retrieval quality end to end. That is why this review organizes conclusions by use case instead of by winner.
Q4: Do I always need the maximum output dimension? Not when the model supports MRL. EmbeddingGemma 2 truncates from 768 to 128 dimensions with up to six-fold storage savings, and OpenAI's text-embedding-3-large truncates from 3072 to 256. Truncation trades a little fidelity for meaningful storage and memory savings, often the right exchange at scale.
Q5: My retrieval quality is weak. Should I swap the embedding model first? Probably not first. Relayed production experience suggests adding a reranker before replacing the embedding model, a minutes-level change compared to a full re-embed. If evaluation still shows the ceiling is the embedder, plan the migration deliberately.
Where to go next
If EmbeddingGemma 2's on-device story interests you, the EmbeddingGemma 2 open-source deep dive covers the repo, license fine print, and first-run experience. Planning a migration? The embedding model upgrade SOP walks through dual-run comparison, gradual traffic shifting, and consistency evaluation before a full re-embed. And once a new embedder is live, the RAG evaluation SOP is how you prove the upgrade actually helped.