Open Source
Open Source

From 2% to 47%: JetBrains' RL Bet on Mellum2.1

JetBrains shipped Mellum2.1 on October 8, 2026 (HF verified via hf-mirror: ungated, Apache 2.0, early-stage 198 downloads/77 likes). Repo: huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking. Architecture is identical to June's Mellum2 (12B MoE / 2.5B active / 131K context); what shipped is a training recipe - RL moved from a finishing stage to the main pipeline across millions of sandboxed runs in thousands of environments, with new math/competition/science/tool-use/software-engineering tasks and filtered sources. JetBrains self-reported benchmarks (Pi harness v0.73.1, thinking mode, no third-party rerun): SWE-bench Verified jumped 2.0 to 47.0, but Qwen3.5-9B sits at 50.0 - Qwen wins real-repo long tasks while Mellum wins LiveCodeBench v6 (82.0) and BFCL v4 (62.3); cite each vendor's own table, most figures exist only as charts. Speed: "almost twice Qwen3.5-9B throughput under heavy load on H200" is an official claim; the 1.6x single-request gain rides on MTP speculative decoding and the MTP head was not released with the weights (coming soon) - you cannot buy that speedup today; GGUF (llama.cpp/Ollama/LM Studio) is also coming soon. Deployment: vLLM supported today (reasoning-parser qwen3), BF16 weights around the 24GB class (estimate), no hosted API. Positioning: agent density on a fixed GPU budget with code that never leaves your servers - not a blanket Qwen beater.

Published October 10, 202610 min read
<!-- mellum-2-1-open-source-coding-resource | open-source | From 2% to 47%: JetBrains' RL Bet on Mellum2.1 -->

Same 12B architecture, same team, only the training method changed: SWE-bench Verified jumped from 2.0 to 47.0. That is not marketing spin for a new model — it is the self-reported scorecard for Mellum2.1, which JetBrains put on Hugging Face on October 8, 2026. The distance between 2% and 47% is not more parameters. It is one calculated gamble at the recipe level: promoting reinforcement learning from a finishing touch at the end of the pipeline to the main body of training itself.

For anyone tracking open-source coding models, this release deserves a close read: it demonstrates that in the agent era, the training method alone can be the biggest lever. One caveat before anything else — the 47 comes from JetBrains' own harness, and no third party has reproduced it yet.

The Repo and Release Context

The repo first. The full model name is Mellum2.1-12B-A2.5B-Thinking, hosted on Hugging Face: JetBrains/Mellum2.1-12B-A2.5B-Thinking.

Repository address: huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking

A few release-moment facts, verified via the HF API on 2026-10-10 through hf-mirror: the license is Apache 2.0, the repo is ungated, and all weights can be downloaded directly without logging into any account. Early numbers show 198 downloads and 77 likes — the state of a repo days after launch. Treat those as initial-release figures, not adoption data.

The ungated detail is worth emphasizing on its own. Last batch we covered LTX-2.5, whose repo is gated=auto, meaning you submit information and wait for approval before downloading. Mellum2.1 is fully open: open the page, click download, get the weights. For enterprise teams that want to validate quickly, that removes an entire process step.

On context: Mellum2.1 is a recipe upgrade to Mellum2, released in June 2026. The architecture is identical, the license is identical, and the only change is the training method. JetBrains announced it on the official blog the same day. There is no hosted API — if you want to use it, you pull the weights and deploy yourself, which is why this article spends real space on deployment boundaries.

Architecture Unchanged, the Recipe Is the Product

The spec: 12B total parameters in a MoE architecture, 2.5B active parameters, 131,072-token context. Identical to Mellum2. In other words, JetBrains did not add parameters, expand context, or change sparsity. What it changed lives in the training room.

The core change is the role of RL. In a conventional pipeline, RL comes last: pretraining and supervised fine-tuning first, then a short RL phase to align or sharpen a narrow set of capabilities. Mellum2.1 makes RL the main body of training, backed by an engineering scale of "thousands of environments and millions of sandbox runs" — the model repeatedly executes real tasks inside isolated sandboxes and learns from feedback.

The task mix expanded too. Mellum2 was mostly about code completion and software engineering; Mellum2.1 adds math, competitive programming, science, and tool-use tasks. On the data side, the approach is source-data filtering. In JetBrains' own words: open-source data often carries broken tests and unverifiable answers. Translated: RL reward signals must be verifiable. If a task's "correct answer" cannot be judged objectively, putting it in the training set only pollutes the model. So JetBrains filtered aggressively, keeping only tasks with clear decision criteria.

The result of that recipe is the number at the top: same architecture, same license, SWE-bench Verified from 2.0 to 47.0. A 23x improvement from a pure training-method change — that is what the "RL bet" in the title means. The bet paid off, at least inside JetBrains' own test environment.

Reading the Self-Test Scores Correctly

JetBrains ran its evaluation on Pi harness v0.73.1 with thinking mode enabled, comparing against Qwen3.5-9B in the same environment. Declare the caveat first: this is the vendor's self-test, no third party has rerun it, and the scores are mostly presented as charts, so some raw numbers cannot be verified item by item. Everything below is quoted per the official framing. Read it with that in mind.

BenchmarkMellum2.1Qwen3.5-9BWinner
SWE-bench Verified47.050.0Qwen
SWE-bench Pro28.038.0Qwen
Terminal-Bench 2.117.421.7Qwen
LiveCodeBench v682.075.4Mellum
BFCL v462.358.5Mellum
GPQA Diamond64.677.8Qwen

This table is easy to misread, so here is the discipline.

First, cite each vendor's table separately. Mellum2.1's 47.0 and Qwen3.5-9B's 50.0 come from the same JetBrains self-test, so the comparison is internally consistent. But if you see other Qwen3.5-9B scores elsewhere, those come from Qwen's own evaluations under different conditions, and the two sets cannot be mixed into one ranking. This article uses only the JetBrains table.

Second, look at the structure, not a total. Qwen3.5-9B leads on all three long-horizon, real-repository benchmarks (both SWE-bench variants and Terminal-Bench), while Mellum2.1 flips the script on competitive programming (LiveCodeBench) and tool calling (BFCL). Rough summary: Qwen wins on real-repo long tasks; Mellum wins on short tasks, competition-style coding, and tool calls. This is not universal dominance — it is two training orientations diverging across task shapes.

Third, do not read Mellum2's 2.0 as a verdict on the old model. Mellum2 was a completion-oriented model, and agent benchmarks like SWE-bench were never its home turf. The 2.0-to-47.0 jump shows the effect of the recipe change, not that "the old model was garbage."

Fourth, the 13-point deficit on GPQA Diamond (science QA) says the RL recipe targeted verifiable engineering tasks. Scientific reasoning is not the selling point, and using this model as a quiz machine is a mismatch.

Speed Claims and the Unreleased MTP Head

JetBrains' speed claim: on H200, throughput under heavy load approaches twice that of Qwen3.5-9B. Two qualifiers: it is the vendor's own framing with the concrete value shown only in a chart, and "heavy-load throughput" means high-concurrency scenarios, not single-request latency.

There is a separate claim for single requests: up to 1.6x via MTP (Multi-Token Prediction) speculative decoding. But here is the crucial fact — the MTP head was not released with the weights. JetBrains marks it "coming soon." In other words, the weights you pull from Hugging Face today contain no MTP head, and the 1.6x figure is currently unbuyable. It becomes reproducible only when the head ships.

This matters for deployment decisions. If you see tutorials claiming "Mellum2.1 is 1.6x faster," check whether they implemented MTP themselves or are simply waiting for the head — most likely they have repackaged a marketing number as a default capability. The same caution applies to GGUF: support for llama.cpp, Ollama, and LM Studio is also coming soon. There are no GGUF quantizations to pull today, and any "one-click Mellum2.1 in Ollama" bundle is ahead of reality.

The only working path right now is vLLM: officially supported, with reasoning-parser set to qwen3, and optional Hermes tool-calling templates.

Deployment Boundaries: Between 12GB and 24GB

JetBrains has not published hardware requirements. The verifiable estimate on record: BF16 weights are roughly in the 24GB range. From that number, the boundaries:

  • 24GB cards (4090/3090): full BF16 is feasible to start, but the KV cache for 131K context eats into headroom. Long-context scenarios want more VRAM or multiple cards.
  • 12GB cards: cannot hold the full weights. Wait for GGUF quantizations and re-evaluate.
  • 8GB cards: no path today.

Because the MoE architecture activates only 2.5B parameters, compute per token is modest; the bottleneck is memory capacity and bandwidth. That is the foundation of JetBrains' pitch — high agent density on a fixed GPU budget. A single 24GB card runs a dedicated agent model with throughput near twice the competitor's (per the official framing).

One more reminder: there is no hosted API. Everything is self-hosted, and for teams without GPU operations capability that barrier is real. The full vLLM walkthrough, including the reasoning-parser configuration that thinking models require, is in our deployment SOP (mellum-2-1-local-deploy-sop).

Where It Fits: Who Should Use It and Who Should Not

Placed into the local-model ecosystem as of batch 45, Mellum2.1's position is actually quite clear.

Good fits:

  • Running an agent farm on a fixed GPU budget. If the core need is "as many concurrent agent tasks as possible on one card," Mellum2.1's throughput-first design is aimed exactly there. A model that grew up on code completion is also naturally at home inside IDE agent workflows.
  • Data that never leaves your domain. Apache 2.0 license, pure self-hosting, code and task data never touching a third-party server. For many enterprises this is a hard requirement.
  • Workflows dense with short tasks and tool calls. The LiveCodeBench and BFCL leads show where its strengths concentrate.

Poor fits or caution:

  • If your main battlefield is long-horizon work in real repositories. Qwen3.5-9B is stronger on SWE-bench Pro and Terminal-Bench, and that signal is right there in the vendor's own table. Do not pick the weaker side just because it is newer.
  • 8GB GPU users. With GGUF not yet out, we are in a quantization gap. Qwen3.5-9B or another small model is your lane.
  • Multimodal input. Mellum is a pure code model with no image input.
  • Teams that do not want to operate GPUs. No hosted API and vLLM as the only path means the operational cost is yours.

The full four-way comparison — Mellum2.1, Qwen3.5-9B, DeepSeek-V4.1-Flash, and GLM-5.3-Flash across five axes including license, context, hardware floor, and runtime support — is in this batch's review (local-coding-model-comparison-review). One contrast worth flagging there: a "Flash" suffix does not mean small. GLM-5.3-Flash is a 321B node-class model, while the 12B Mellum is the only one of the four that runs on a consumer card.

If you want to slot it into an IDE workflow alongside tools like Claude Code or Cursor, see our earlier coding agent comparison (ai-coding-agent-comparison-review), and for choosing between inference tools, the local LLM tooling review (local-llm-deployment-comparison-review).

A final note on provenance: the model card ships the license text and the evaluation charts, but the raw prompts, seeds and serving configuration behind those charts are not published, so treat every figure in the table above as direction of travel rather than a league table.

FAQ

Q1: What is the difference between Mellum2.1 and Mellum2?

Architecture, parameter count, context length, and license are all identical. The only difference is the training recipe: Mellum2.1 makes RL the main body of training with thousands of environments and millions of sandbox runs, and adds math, competition, science, and tool-use tasks. The effect is SWE-bench Verified going from 2.0 to 47.0 (JetBrains' self-test; no third-party reproduction yet).

Q2: Is Mellum2.1 really better than Qwen3.5-9B?

No. Per the JetBrains self-test, Qwen3.5-9B wins on real-repo long tasks (SWE-bench Verified 47.0 vs 50.0, SWE-bench Pro 28.0 vs 38.0, Terminal-Bench 17.4 vs 21.7), while Mellum2.1 leads on competitive programming (LiveCodeBench 82.0 vs 75.4) and tool calling (BFCL v4 62.3 vs 58.5). Choose by your task shape; there is no universal superiority.

Q3: Can I get the claimed 1.6x speedup today?

No. The 1.6x relies on MTP speculative decoding, but the MTP head was not released with the weights and is marked coming soon. The weights you download today cannot reproduce that speedup. The "nearly 2x" heavy-load throughput figure is likewise vendor-reported with no published value. Only baseline vLLM performance is verifiable now.

Q4: Which runtimes support Mellum2.1? Can I run it in Ollama?

Only vLLM is officially supported today, with reasoning-parser set to qwen3. GGUF (llama.cpp/Ollama/LM Studio) is coming soon, not shipped. Community bundles claiming one-click Ollama support are ahead of reality; Ollama users need to wait for the GGUF release.

Q5: Do I need to log in or apply to download Mellum2.1?

No. The repo is ungated under the Apache 2.0 license — open the Hugging Face page and download all weights directly (verified via hf-mirror on 2026-10-10). Early figures of 198 downloads and 77 likes reflect the initial release. There is no hosted API; usage requires self-hosting.

Closing

The 2%-to-47% story is told, but the real test belongs to the community: JetBrains' self-reported numbers await third-party harness reproduction, and the release timing of the MTP head and GGUF weights will directly shape adoption. We will follow up with 12GB-card feasibility tests once quantizations land.

If you have already pulled the weights and run them, share your harness scores and tokens/s in the comments — especially side-by-side numbers against Qwen3.5-9B in the same environment, which is exactly where dual-run comparisons create the most value. And if you hit snags during deployment, the deployment SOP above has a comments section too.

This article is AI-assisted and human-edited. Last updated: 2026-10-10

FAQ

What is the difference between Mellum2.1 and Mellum2?
Architecture, parameter count, context length, and license are all identical. The only difference is the training recipe: Mellum2.1 makes RL the main body of training with thousands of environments and millions of sandbox runs, and adds math, competition, science, and tool-use tasks. The effect is SWE-bench Verified going from 2.0 to 47.0 (JetBrains' self-test; no third-party reproduction yet).
Is Mellum2.1 really better than Qwen3.5-9B?
No. Per the JetBrains self-test, Qwen3.5-9B wins on real-repo long tasks (SWE-bench Verified 47.0 vs 50.0, SWE-bench Pro 28.0 vs 38.0, Terminal-Bench 17.4 vs 21.7), while Mellum2.1 leads on competitive programming (LiveCodeBench 82.0 vs 75.4) and tool calling (BFCL v4 62.3 vs 58.5). Choose by your task shape; there is no universal superiority.
Can I get the claimed 1.6x speedup today?
No. The 1.6x relies on MTP speculative decoding, but the MTP head was not released with the weights and is marked coming soon. The weights you download today cannot reproduce that speedup. The "nearly 2x" heavy-load throughput figure is likewise vendor-reported with no published value. Only baseline vLLM performance is verifiable now.
Which runtimes support Mellum2.1? Can I run it in Ollama?
Only vLLM is officially supported today, with reasoning-parser set to qwen3. GGUF (llama.cpp/Ollama/LM Studio) is coming soon, not shipped. Community bundles claiming one-click Ollama support are ahead of reality; Ollama users need to wait for the GGUF release.
Do I need to log in or apply to download Mellum2.1?
No. The repo is ungated under the Apache 2.0 license — open the Hugging Face page and download all weights directly (verified via hf-mirror on 2026-10-10). Early figures of 198 downloads and 77 likes reflect the initial release. There is no hosted API; usage requires self-hosting.

Related

Open Source

Bilibili's 35B translation MoE: 150 languages, only 3B active

Bilibili's Index LLM team open-sourced the Index-Translate family on September 30, 2026 (weights on Hugging Face and ModelScope, Apache-2.0): 2B/9B/35B-A3B-preview text models built on Qwen3.5, targeting 150 languages (self-reported), with the 35B MoE activating only about 3B parameters per token and a 262,144-token context. Three branches split the work: Index-Echo produces dubbed speech that preserves the source speaker's voice (Chinese to/from English, Spanish, Japanese); Index-Homura controls syllable counts (the 9B lands within +-10% of the target on 81.92% of SandGlass cases - built for dubbing timing); Index-NativeLong (aka Index-Nailong) keeps long documents consistent (a 32K-token fantasy text keeps a name/royal-title pun intact, where chunk-by-chunk translation drifts). Constrained translation via instTrans: hard constraints enforce glossaries (e.g. carbon fiber) plus JSON/CSV/code/placeholder structure; soft constraints cover tone, domain disambiguation (plant to factory), cross-sentence consistency and LaTeX preservation. Local deployment is one line with the official GGUF (Q4_K_M, 21.71GB) via llama.cpp serve, plus a browser extension and a video dubbing pipeline. Honesty note: every score is self-reported with no independent replication yet (FLORES COMET-22 0.8794, WMT26 Judge 76.76, 2.4% off-target on low-resource pairs), and in the same official table DeepSeek-V4.1-Flash scores higher on WMT26 Judge at 83.55 - not first place on raw quality.

Oct 7, 202610 min read
Open Source

Xiaomi Open-Sourced the RL Stack That Trained MiMo-V2.6

A fact-check of XiaomiMiMo/verl (2026-09-27 GitHub API snapshot: 465 stars, 49 forks, Python, Apache-2.0, created 2026-09-21, pushed 2026-09-26). It is the Agentic RL training code behind MiMo-V2.6, forked from verl-project/verl (23,650 stars, HybridFlow) and built on verl 0.9.0.dev with reproduction code for five RL environment suites: Code software engineering (executable tests), Cyber vulnerability reproduction (rule checks), General knowledge work (rubric review), Visual web development (visual scoring), and Music symbolic composition (rule checks), each with launch scripts. The MiMo-V2.6-RL-oss training dataset on Hugging Face and the technical-report PDF are open too; same-day companions uni-agent (14 stars) and mimoagent (30 stars, MIT) complete the bundle. Divergence from upstream cannot be verified, so the article sticks to README-verifiable facts; treat the repository LICENSE file as the final license authority.

Sep 27, 20269 min read
Open Source

Alibaba Open-Sources Qwen3.8-Flash-Next, a Qwen4 Preview

On 2026-08-26 Alibaba open-sourced Qwen3.8-Flash-Next on Hugging Face and ModelScope: a 125B MoE model with 6B activated per token, the first open-weight preview of the Qwen4 architecture. Native context is 262K, extensible to 1M via YaRN; API pricing is \$0.16/\$0.47 per million tokens (about one-twelfth of flagship Qwen3.8-Max). Benchmarks: DeepSWE 58.7, SWE-bench Pro 62.5, CoWorkBench 73.9, AndroidWorld 84.5, MathVision 95.7, and rank 7 on the open Agent Arena. It ships under the qwen-community-1.0 license (not Apache 2.0), permitting commercial use and self-hosting, but verify the terms against the model page before commercial use.

Sep 4, 202610 min read