Hardcore Reviews
Hardcore Reviews

Embedding Model Comparison: 4 RAG Contenders, One Honest Table

A four-way RAG embedding model comparison (pricing snapshots October 9, 2026): EmbeddingGemma 2, Qwen3-Embedding (0.6B/4B/8B), BGE-M3 and OpenAI text-embedding-3-large, with gemini-embedding-001 as a reference row. All licenses API-verified on Hugging Face (2026-10-09): google/embeddinggemma-2 and Qwen3-Embedding-8B are apache-2.0 and ungated; BAAI/bge-m3 is mit (33.7M downloads). Engineering ledger aligned item by item: context 8192 (EG2/BGE-M3/OpenAI) vs 32768 (Qwen3); dimensions 768 (MRL to 128) / 1024-4096 / 1024 / 3072 (MRL to 256); three self-host free vs OpenAI at $0.13 per million tokens. Discipline first: each vendor's benchmark numbers cited from their own tables only (EG2's MTEB Code +9.92 and gemini-embedding-001's three board leads are official self-reports; Qwen3-8B's ~70.6 multilingual is third-party relayed) - no cross-table ranking; production popularity (relayed from LangChain/LlamaIndex telemetry): BGE-M3 first, Qwen3 second, OpenAI third. Verdicts by scenario, no arbitration: on-device multimodal picks EG2, long documents and Chinese pick Qwen3, hybrid retrieval and mature ecosystem pick BGE-M3, zero-ops picks OpenAI; and since switching means a full re-embed, decide once, decide carefully.

Published October 9, 202610 min read
<!-- embedding-model-rag-comparison-review | review | Embedding Model Comparison: 4 RAG Contenders, One Honest Table -->

Your RAG system's retrieval ceiling is set long before the language model writes a single word. It is set by the embedding model: the component that compresses documents and queries into vectors a machine can compare. Pick a weak one and no prompt engineering rescues you. Pick one that mismatches your deployment shape and you discover, months later, that switching means rebuilding your entire index from scratch.

This review lines up the four embedding models that dominate real-world RAG conversations in October 2026: Google DeepMind's EmbeddingGemma 2, Alibaba Qwen's Qwen3-Embedding, BAAI's BGE-M3, and OpenAI's text-embedding-3-large. Google's gemini-embedding-001 joins as a reference row, because it frames the central question of this model generation: what does the open-source route trade away against a frontier API model, and what does it gain back?

Ground rules first, as with every comparison on this site: we ran no proprietary stress tests, and we make no capability arbitration. Every benchmark figure below is cited with the table it comes from, self-reported numbers stay inside their own labeled context, and conclusions are organized by use case rather than by a crowned winner.

Why most embedding model comparisons mislead

The most common failure in this genre is also the least visible: lifting scores out of context and stacking them into one ranking. The MTEB ecosystem spans multiple boards, including MTEB English, MTEB Multilingual, and MTEB v2, each with a different task mix, language coverage, and evaluation conditions. A number published on one board is not commensurable with a number on another, even when both are casually called MTEB.

This review therefore enforces one citation rule without exception: every vendor's figures are quoted with the name of the table they were reported on, no cross-table ranking appears anywhere in this article, and provenance is always labeled. A vendor's own reporting says so, third-party relay is marked as relayed, and unpublished figures stay unpublished. EmbeddingGemma 2's code score and Qwen3-Embedding's multilingual score never meet in a sentence implying one beat the other. There is no third category in honest tech writing.

The four contenders at a glance

DimensionEmbeddingGemma 2Qwen3-EmbeddingBGE-M3OpenAI text-embedding-3-large
VendorGoogle DeepMindAlibaba QwenBAAIOpenAI
Parameters270M to 740M, four load tiers0.6B / 4B / 8B, three sizesAbout 568MUndisclosed
Context window8,192 tokens32,768 tokens8,192 tokens8,191 tokens
Output dimensions768 (MRL down to 128)1024 / 2560 / 4096 (MRL)10243072 (MRL down to 256)
License (HF metadata verified 2026-10-09)apache-2.0, ungatedapache-2.0, ungatedmit, ungatedClosed API
PriceFree, self-hostedFree, self-hostedFree, self-hosted$0.13 per million tokens (official pricing page snapshot (2026-10-09))
Headline specialtyOn-device footprint, four modalities in one vector space32K context, task instruction prefixesOne model, three retrieval modes: dense, sparse, multi-vectorZero ops, MRL support

A reference row completes the picture: gemini-embedding-001 is API-only, outputs 3072 dimensions from a 2,048-token context, and, per Google's own reporting, holds top placement on three MTEB boards: multilingual, English, and code. If a closed, frontier-class model tops the boards its own maker reports on, the open-source contenders are not competing to dethrone it there; they are competing on what those boards never measure: deployment cost, data boundary, modality coverage, and license freedom.

Multilingual coverage, where published: EmbeddingGemma 2 lists 100+ languages, Qwen3-Embedding 100+ with reported strength in Chinese, BGE-M3 100+. OpenAI's total is unpublished, and we will not invent one.

EmbeddingGemma 2: the on-device multimodal play

Weights live at huggingface.co/google/embeddinggemma-2. A direct check of the Hugging Face API on 2026-10-09 confirmed the essentials: license apache-2.0, ungated, 21,148 downloads since the 2026-10-06 release.

Built on the Gemma 4 architecture, EmbeddingGemma 2 stacks a 270M text-and-code backbone with a 170M vision encoder and a 300M audio encoder, for 740M parameters at full load. The same checkpoint loads at four tiers: 270M text-only, 440M with vision, 570M with audio, and 740M full multimodal. Every tier projects into one unified 768-dimensional vector space, so text-to-image and text-to-audio search share one index, with an 8,192-token context that is four times the first generation's 2,048.

The deployment numbers, per Google's figures: quantized, text-only embedding runs in roughly 191MB of active memory on a Pixel 11 Pro, and full multimodal in roughly 567MB. MRL lets you truncate the 768-dimension output down to 128, cutting vector storage by as much as six times. The full checkpoint ships as 1.49GB of bfloat16 safetensors.

The license deserves a careful read. Apache 2.0 is declared in the repository metadata and linked from the model card, but the repo root contains no standalone LICENSE file, and the model card retains a line requiring compliance with the Gemma Prohibited Use Policy. Against the first generation's gated download and custom Gemma license, this is a genuine opening-up, just not the unconditional kind the bare string apache-2.0 might suggest.

Benchmarks: all figures are Google's own reporting, with no independent evaluations, no paper, and no accepted MTEB submission at review time. On Google's own MTEB Code table, EmbeddingGemma 2 moved from 68.76 to 78.68, a gain of 9.92 and the only clearly significant improvement. Multilingual text improved by just 0.21 over the first generation, and English registered a slight decline. This model is not a leaderboard play; it is a deployment play. Task prompt prefixes are a hard requirement per the model card, with queries formatted as task: {task description} | query: {content} and documents as title: {title} | text: {content}. Skipping them silently costs precision.

The ecosystem arrived fast: sentence-transformers 6.1.0+ and transformers 5.19.0 supported it on release day, with MLX, Ollama, llama.cpp, vLLM, and LiteRT/MediaPipe alongside, a Qdrant storage cooperation, and an Unsloth fine-tuning guide, while the first generation crossed 20 million cumulative downloads.

Qwen3-Embedding: three sizes, 32K context

Code and weights live at github.com/QwenLM/Qwen3.

Qwen3-Embedding ships in three sizes: 0.6B, 4B, and 8B parameters, with output dimensions of 1024, 2560, and 4096 respectively, each supporting MRL truncation. The standout spec is the context window: 32,768 tokens, four times the 8K-class windows of EmbeddingGemma 2, BGE-M3, and OpenAI's model. For retrieval over long documents, that difference changes your chunking strategy, not just your accuracy.

License status was verified against Hugging Face repository metadata on 2026-10-09: apache-2.0 and ungated, with the 8B repository showing roughly 2.446 million downloads. That is a commercially clean, self-hostable license at every size tier, which matters for teams that need legal simplicity across a model family.

Benchmark discipline applies with full force here. The Qwen3-Embedding-8B figure of about 70.6 on MTEB multilingual is self-reported and reached us through third-party relay; we cite it with exactly that provenance: one table, one label, no comparisons. What we can say without a scoreboard: 100+ language coverage with reported strength in Chinese, instruction-aware task prefixes, and a three-size scaling path make it the natural candidate when your corpus is long, multilingual, or Chinese-heavy and you want to right-size compute.

BGE-M3: one model, three retrieval modes

Code and weights live at github.com/BAAI/FlagEmbedding.

BGE-M3 runs about 568M parameters, outputs 1024 dimensions from an 8,192-token context, and was verified on 2026-10-09 as mit-licensed and ungated, with roughly 33.69 million downloads on Hugging Face: the most permissive license in this table and the largest verified download count. MIT means no usage-policy rider for your legal team to read.

Its defining feature is retrieval versatility: a single model returns dense, sparse, and multi-vector outputs. That is one checkpoint covering three retrieval paradigms, semantic dense search, lexical-style sparse matching, and late-interaction-style multi-vector scoring, which means hybrid retrieval pipelines without bolting a second model onto your stack.

Popularity, marked as relayed rather than measured: per LangChain and LlamaIndex telemetry as relayed by presenc, BGE-M3 ranks first in third-party production deployment popularity, ahead of Qwen3-Embedding in second, OpenAI's text-embedding-3-large in third, and Voyage in fourth. That is a popularity ranking, not a capability verdict, and per this review's citation rule we attach no cross-table benchmark score to it here.

OpenAI text-embedding-3-large: the zero-ops option

OpenAI's flagship embedding model outputs 3072 dimensions with MRL truncation down to 256, reads an 8,191-token context, and keeps its parameter count undisclosed, so we report it as undisclosed rather than guessing. Pricing is $0.13 per million tokens (official pricing page snapshot (2026-10-09)).

The value proposition is operational, not architectural: no GPUs to provision, no inference stack to maintain, no version pinning across your fleet, and MRL support to trim storage downstream. For teams whose core product is not retrieval infrastructure, that is frequently the deciding factor: the embedding layer stops being anyone's job.

The trade is equally plain. No weights to download and no self-hosting path, so every embedded document crosses your data boundary; per-token cost scales with corpus size and re-embedding frequency; and the total language count has not been published. Consistent with this review's citation rule, we attach no benchmark score to this model here: its self-reported figures come from its own tables, and folding them into one ranking would be exactly the cross-table mistake this article refuses to make.

Choosing by use case, not leaderboard

If the four contenders ranked cleanly on a single axis, this section would be one line long. They do not, so here is the map instead.

Run on-device, at the edge, or across modalities. EmbeddingGemma 2 is the only contender in this table designed for it: a 191MB quantized text-only footprint per Google's figures, four modalities sharing one 768-dimension space, MRL down to 128 dimensions for up to six-fold storage savings, and a four-tier checkpoint that lets you ship text-only first and grow into vision and audio without changing vector spaces. Accept the trade: its self-reported gains concentrate in code, while multilingual moved little and English slipped slightly.

Long documents, Chinese-heavy or broadly multilingual corpora, license-clean self-hosting. Qwen3-Embedding: a 32K context that reshapes chunking strategy, three sizes to right-size latency and memory, apache-2.0 verified at every tier, and 100+ languages with reported strength in Chinese.

Hybrid retrieval and a proven production track record. BGE-M3: dense, sparse, and multi-vector from one checkpoint, an MIT license with no riders, and, per relayed third-party telemetry, the largest production install base among these models. When you want the least-surprising choice, this is the least-surprising choice.

Managed, zero-ops, pay-as-you-go. OpenAI text-embedding-3-large: $0.13 per million tokens (official pricing page snapshot (2026-10-09)), 3072 dimensions truncatable to 256, and an embedding layer that stops being anyone's job. Accept the trade: no weights, and a data boundary crossing on every embedded document.

Two closing notes outlast any individual model. gemini-embedding-001's top placement on three MTEB boards per Google's own reporting is a reminder that the open models here are not chasing that trophy. Switching embedding models means a full re-embed of your entire corpus, so the decision is expensive to revisit. And if retrieval quality is lagging today, relayed production experience suggests adding a reranker first, a minutes-level change compared to a full re-embed.

FAQ

Q1: Can I directly compare the benchmark scores mentioned in this article? No. The figures cited here come from different tables: Google's own MTEB Code table for EmbeddingGemma 2, a self-reported MTEB multilingual figure relayed through third parties for Qwen3-Embedding-8B, and Google's own reporting for gemini-embedding-001. Different MTEB boards use different task mixes and conditions, so the numbers are not commensurable, and this article presents no cross-table ranking.

Q2: Is Qwen3-Embedding really Apache 2.0 licensed? Yes. Verified against Hugging Face repository metadata on 2026-10-09: apache-2.0 and ungated, with the 8B repository showing roughly 2.446 million downloads. Code and weights live at github.com/QwenLM/Qwen3.

Q3: What does it cost to switch embedding models after launch? A full re-embed of your entire corpus: re-encoding every document, rebuilding the vector index, and re-validating retrieval quality end to end. That is why this review organizes conclusions by use case instead of by winner.

Q4: Do I always need the maximum output dimension? Not when the model supports MRL. EmbeddingGemma 2 truncates from 768 to 128 dimensions with up to six-fold storage savings, and OpenAI's text-embedding-3-large truncates from 3072 to 256. Truncation trades a little fidelity for meaningful storage and memory savings, often the right exchange at scale.

Q5: My retrieval quality is weak. Should I swap the embedding model first? Probably not first. Relayed production experience suggests adding a reranker before replacing the embedding model, a minutes-level change compared to a full re-embed. If evaluation still shows the ceiling is the embedder, plan the migration deliberately.

Where to go next

If EmbeddingGemma 2's on-device story interests you, the EmbeddingGemma 2 open-source deep dive covers the repo, license fine print, and first-run experience. Planning a migration? The embedding model upgrade SOP walks through dual-run comparison, gradual traffic shifting, and consistency evaluation before a full re-embed. And once a new embedder is live, the RAG evaluation SOP is how you prove the upgrade actually helped.

This article is AI-assisted and human-edited. Last updated: 2026-10-09

FAQ

Can I directly compare the benchmark scores mentioned in this article?
No. The figures cited here come from different tables: Google's own MTEB Code table for EmbeddingGemma 2, a self-reported MTEB multilingual figure relayed through third parties for Qwen3-Embedding-8B, and Google's own reporting for gemini-embedding-001. Different MTEB boards use different task mixes and conditions, so the numbers are not commensurable, and this article presents no cross-table ranking.
Is Qwen3-Embedding really Apache 2.0 licensed?
Yes. Verified against Hugging Face repository metadata on 2026-10-09: apache-2.0 and ungated, with the 8B repository showing roughly 2.446 million downloads. Code and weights live at github.com/QwenLM/Qwen3.
What does it cost to switch embedding models after launch?
A full re-embed of your entire corpus: re-encoding every document, rebuilding the vector index, and re-validating retrieval quality end to end. That is why this review organizes conclusions by use case instead of by winner.
Do I always need the maximum output dimension?
Not when the model supports MRL. EmbeddingGemma 2 truncates from 768 to 128 dimensions with up to six-fold storage savings, and OpenAI's text-embedding-3-large truncates from 3072 to 256. Truncation trades a little fidelity for meaningful storage and memory savings, often the right exchange at scale.
My retrieval quality is weak. Should I swap the embedding model first?
Probably not first. Relayed production experience suggests adding a reranker before replacing the embedding model, a minutes-level change compared to a full re-embed. If evaluation still shows the ceiling is the embedder, plan the migration deliberately.

Related

Hardcore Reviews

Open Video Model Local Deployment: Five-Axis Comparison Review

A four-way open-source video model comparison on the five axes that matter for local deployment: license, VRAM, audio, duration/resolution, and regional restrictions - not generation quality, no cross-model benchmarks, no ranking. LTX-2.5 (22B, distilled/FP8 from 12GB, 24kHz stereo in-pass, LTX Community license with a $10M threshold, gated=auto); Wan 2.2 (TI2V-5B, 8GB with offload, Apache 2.0 and the most permissive, short clips 480p-720p, no native audio); HunyuanVideo 1.5 (8.3B, from 14GB, ~75s per clip on a 4090 step-distilled build, Tencent community license not valid in EU/UK/South Korea with 100M-MAU renegotiation, no native audio); MiniMax H3 (33B dense, 32kHz stereo with dialogue in 11 languages, 15s/768p default while 2K needs an unreleased regenerator, community license excludes EU/UK/South Korea/USA with written authorization required above $20M, VRAM unpublished so not invented). Sources labeled per vendor (bestfreewebresources/pinggy compilations plus official model cards); fal's H3 Max #1 on i2v-with-audio is third-party. Verdicts by scenario, no arbitration: most permissive license picks Wan 2.2, audiovisual-in-one picks LTX-2.5, a 24GB card outside restricted regions picks HunyuanVideo 1.5, reference audio and multilingual dialogue pick H3 (within licensed regions).

Oct 10, 20269 min read
Field SOP

Swap Your Embedding Model Without Wrecking RAG: The SOP

An embedding model upgrade SOP (basis October 9, 2026; pairs with the site's open-source and comparison pieces): switching embedding models means a full re-embed, and this SOP covers the decision and the drill in one pass. Step 0 audits your current setup: model and dims, context window, MRL usage, deployment form, prompt-prefix conventions - plus the production reminder (relayed) that a reranker often fixes retrieval misses in minutes before you re-embed anything. Five selection criteria: parameters/context/dimensions with MRL/license (HF API verified: EG2 apache-2.0 ungated but no LICENSE file, Qwen3 apache-2.0, BGE-M3 mit)/deployment form. EmbeddingGemma 2's task prompt prefixes are mandatory - queries as task: {...} | query: {...}, documents as title: {...} | text: {...}, seven task types each with a prefix; skip them and accuracy silently drops; encode(truncate_dim=512/256/128) trims vectors for up to 6x storage savings. Five execution steps: dual-run old vs new on sentence-transformers>=6.1.0 (same-prompt spot checks), canary cutover (framework only, no invented percentages), full re-embed with reconciliation and retirement, then an evaluation loop with Ragas/DeepEval linking to the site's RAG evaluation piece. Plain-text repos: github.com/UKPLab/sentence-transformers, huggingface.co/google/embeddinggemma-2.

Oct 9, 202610 min read
Hardcore Reviews

Small Model Price War: Haiku 5.5 Ties Luna Only Below 100K

A small-model price-war comparison (October 8, 2026 price snapshots; five columns, four rows): Claude Haiku 5.5, GPT-6 Luna, Sonnet 5.5, and Qwen3.8-27B as the open self-hosting row. Pricing: Haiku 5.5 at $0.10/$0.50 per million (under-100K prompts), Luna $0.10/$0.50 (up to 272K), Sonnet 5.5 at $2/$10 list with cache reads halved to $0.10, and Qwen3.8-27B free weights (Apache-2.0) on your own GPU. Core finding: Haiku 5.5 matches Luna only in the short-prompt tier - above 100K tokens its $0.50 input faces Luna's $0.20, so long-document batch bills invert the picture. Benchmarks cite each model's own official numbers, labeled and never mixed across suites: Haiku 5.5 Terminal-Bench 4.0 at 39.2% and OSWorld 72.4%; no fresh Luna numbers shipped with the release; Sonnet 5.5 at 70.6% on TB4; Qwen3.8-27B has no cross-suite table. Effort settings compared (Haiku 5.5 is the first Haiku-class with the dial; Luna runs none-to-max), contexts (1M for both hosted models; Qwen 262K extensible to 1M). Verdicts by use case - high-volume subagents, batch classification, main-line coding, private deployment - with the site's standing no-stress-test-no-arbitration stance. Four same-topic deep dives linked inline.

Oct 8, 202610 min read