Same 12B architecture, same team, only the training method changed: SWE-bench Verified jumped from 2.0 to 47.0. That is not marketing spin for a new model — it is the self-reported scorecard for Mellum2.1, which JetBrains put on Hugging Face on October 8, 2026. The distance between 2% and 47% is not more parameters. It is one calculated gamble at the recipe level: promoting reinforcement learning from a finishing touch at the end of the pipeline to the main body of training itself.
For anyone tracking open-source coding models, this release deserves a close read: it demonstrates that in the agent era, the training method alone can be the biggest lever. One caveat before anything else — the 47 comes from JetBrains' own harness, and no third party has reproduced it yet.
The Repo and Release Context
The repo first. The full model name is Mellum2.1-12B-A2.5B-Thinking, hosted on Hugging Face: JetBrains/Mellum2.1-12B-A2.5B-Thinking.
Repository address: huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking
A few release-moment facts, verified via the HF API on 2026-10-10 through hf-mirror: the license is Apache 2.0, the repo is ungated, and all weights can be downloaded directly without logging into any account. Early numbers show 198 downloads and 77 likes — the state of a repo days after launch. Treat those as initial-release figures, not adoption data.
The ungated detail is worth emphasizing on its own. Last batch we covered LTX-2.5, whose repo is gated=auto, meaning you submit information and wait for approval before downloading. Mellum2.1 is fully open: open the page, click download, get the weights. For enterprise teams that want to validate quickly, that removes an entire process step.
On context: Mellum2.1 is a recipe upgrade to Mellum2, released in June 2026. The architecture is identical, the license is identical, and the only change is the training method. JetBrains announced it on the official blog the same day. There is no hosted API — if you want to use it, you pull the weights and deploy yourself, which is why this article spends real space on deployment boundaries.
Architecture Unchanged, the Recipe Is the Product
The spec: 12B total parameters in a MoE architecture, 2.5B active parameters, 131,072-token context. Identical to Mellum2. In other words, JetBrains did not add parameters, expand context, or change sparsity. What it changed lives in the training room.
The core change is the role of RL. In a conventional pipeline, RL comes last: pretraining and supervised fine-tuning first, then a short RL phase to align or sharpen a narrow set of capabilities. Mellum2.1 makes RL the main body of training, backed by an engineering scale of "thousands of environments and millions of sandbox runs" — the model repeatedly executes real tasks inside isolated sandboxes and learns from feedback.
The task mix expanded too. Mellum2 was mostly about code completion and software engineering; Mellum2.1 adds math, competitive programming, science, and tool-use tasks. On the data side, the approach is source-data filtering. In JetBrains' own words: open-source data often carries broken tests and unverifiable answers. Translated: RL reward signals must be verifiable. If a task's "correct answer" cannot be judged objectively, putting it in the training set only pollutes the model. So JetBrains filtered aggressively, keeping only tasks with clear decision criteria.
The result of that recipe is the number at the top: same architecture, same license, SWE-bench Verified from 2.0 to 47.0. A 23x improvement from a pure training-method change — that is what the "RL bet" in the title means. The bet paid off, at least inside JetBrains' own test environment.
Reading the Self-Test Scores Correctly
JetBrains ran its evaluation on Pi harness v0.73.1 with thinking mode enabled, comparing against Qwen3.5-9B in the same environment. Declare the caveat first: this is the vendor's self-test, no third party has rerun it, and the scores are mostly presented as charts, so some raw numbers cannot be verified item by item. Everything below is quoted per the official framing. Read it with that in mind.
| Benchmark | Mellum2.1 | Qwen3.5-9B | Winner |
|---|---|---|---|
| SWE-bench Verified | 47.0 | 50.0 | Qwen |
| SWE-bench Pro | 28.0 | 38.0 | Qwen |
| Terminal-Bench 2.1 | 17.4 | 21.7 | Qwen |
| LiveCodeBench v6 | 82.0 | 75.4 | Mellum |
| BFCL v4 | 62.3 | 58.5 | Mellum |
| GPQA Diamond | 64.6 | 77.8 | Qwen |
This table is easy to misread, so here is the discipline.
First, cite each vendor's table separately. Mellum2.1's 47.0 and Qwen3.5-9B's 50.0 come from the same JetBrains self-test, so the comparison is internally consistent. But if you see other Qwen3.5-9B scores elsewhere, those come from Qwen's own evaluations under different conditions, and the two sets cannot be mixed into one ranking. This article uses only the JetBrains table.
Second, look at the structure, not a total. Qwen3.5-9B leads on all three long-horizon, real-repository benchmarks (both SWE-bench variants and Terminal-Bench), while Mellum2.1 flips the script on competitive programming (LiveCodeBench) and tool calling (BFCL). Rough summary: Qwen wins on real-repo long tasks; Mellum wins on short tasks, competition-style coding, and tool calls. This is not universal dominance — it is two training orientations diverging across task shapes.
Third, do not read Mellum2's 2.0 as a verdict on the old model. Mellum2 was a completion-oriented model, and agent benchmarks like SWE-bench were never its home turf. The 2.0-to-47.0 jump shows the effect of the recipe change, not that "the old model was garbage."
Fourth, the 13-point deficit on GPQA Diamond (science QA) says the RL recipe targeted verifiable engineering tasks. Scientific reasoning is not the selling point, and using this model as a quiz machine is a mismatch.
Speed Claims and the Unreleased MTP Head
JetBrains' speed claim: on H200, throughput under heavy load approaches twice that of Qwen3.5-9B. Two qualifiers: it is the vendor's own framing with the concrete value shown only in a chart, and "heavy-load throughput" means high-concurrency scenarios, not single-request latency.
There is a separate claim for single requests: up to 1.6x via MTP (Multi-Token Prediction) speculative decoding. But here is the crucial fact — the MTP head was not released with the weights. JetBrains marks it "coming soon." In other words, the weights you pull from Hugging Face today contain no MTP head, and the 1.6x figure is currently unbuyable. It becomes reproducible only when the head ships.
This matters for deployment decisions. If you see tutorials claiming "Mellum2.1 is 1.6x faster," check whether they implemented MTP themselves or are simply waiting for the head — most likely they have repackaged a marketing number as a default capability. The same caution applies to GGUF: support for llama.cpp, Ollama, and LM Studio is also coming soon. There are no GGUF quantizations to pull today, and any "one-click Mellum2.1 in Ollama" bundle is ahead of reality.
The only working path right now is vLLM: officially supported, with reasoning-parser set to qwen3, and optional Hermes tool-calling templates.
Deployment Boundaries: Between 12GB and 24GB
JetBrains has not published hardware requirements. The verifiable estimate on record: BF16 weights are roughly in the 24GB range. From that number, the boundaries:
- 24GB cards (4090/3090): full BF16 is feasible to start, but the KV cache for 131K context eats into headroom. Long-context scenarios want more VRAM or multiple cards.
- 12GB cards: cannot hold the full weights. Wait for GGUF quantizations and re-evaluate.
- 8GB cards: no path today.
Because the MoE architecture activates only 2.5B parameters, compute per token is modest; the bottleneck is memory capacity and bandwidth. That is the foundation of JetBrains' pitch — high agent density on a fixed GPU budget. A single 24GB card runs a dedicated agent model with throughput near twice the competitor's (per the official framing).
One more reminder: there is no hosted API. Everything is self-hosted, and for teams without GPU operations capability that barrier is real. The full vLLM walkthrough, including the reasoning-parser configuration that thinking models require, is in our deployment SOP (mellum-2-1-local-deploy-sop).
Where It Fits: Who Should Use It and Who Should Not
Placed into the local-model ecosystem as of batch 45, Mellum2.1's position is actually quite clear.
Good fits:
- Running an agent farm on a fixed GPU budget. If the core need is "as many concurrent agent tasks as possible on one card," Mellum2.1's throughput-first design is aimed exactly there. A model that grew up on code completion is also naturally at home inside IDE agent workflows.
- Data that never leaves your domain. Apache 2.0 license, pure self-hosting, code and task data never touching a third-party server. For many enterprises this is a hard requirement.
- Workflows dense with short tasks and tool calls. The LiveCodeBench and BFCL leads show where its strengths concentrate.
Poor fits or caution:
- If your main battlefield is long-horizon work in real repositories. Qwen3.5-9B is stronger on SWE-bench Pro and Terminal-Bench, and that signal is right there in the vendor's own table. Do not pick the weaker side just because it is newer.
- 8GB GPU users. With GGUF not yet out, we are in a quantization gap. Qwen3.5-9B or another small model is your lane.
- Multimodal input. Mellum is a pure code model with no image input.
- Teams that do not want to operate GPUs. No hosted API and vLLM as the only path means the operational cost is yours.
The full four-way comparison — Mellum2.1, Qwen3.5-9B, DeepSeek-V4.1-Flash, and GLM-5.3-Flash across five axes including license, context, hardware floor, and runtime support — is in this batch's review (local-coding-model-comparison-review). One contrast worth flagging there: a "Flash" suffix does not mean small. GLM-5.3-Flash is a 321B node-class model, while the 12B Mellum is the only one of the four that runs on a consumer card.
If you want to slot it into an IDE workflow alongside tools like Claude Code or Cursor, see our earlier coding agent comparison (ai-coding-agent-comparison-review), and for choosing between inference tools, the local LLM tooling review (local-llm-deployment-comparison-review).
A final note on provenance: the model card ships the license text and the evaluation charts, but the raw prompts, seeds and serving configuration behind those charts are not published, so treat every figure in the table above as direction of travel rather than a league table.
FAQ
Q1: What is the difference between Mellum2.1 and Mellum2?
Architecture, parameter count, context length, and license are all identical. The only difference is the training recipe: Mellum2.1 makes RL the main body of training with thousands of environments and millions of sandbox runs, and adds math, competition, science, and tool-use tasks. The effect is SWE-bench Verified going from 2.0 to 47.0 (JetBrains' self-test; no third-party reproduction yet).
Q2: Is Mellum2.1 really better than Qwen3.5-9B?
No. Per the JetBrains self-test, Qwen3.5-9B wins on real-repo long tasks (SWE-bench Verified 47.0 vs 50.0, SWE-bench Pro 28.0 vs 38.0, Terminal-Bench 17.4 vs 21.7), while Mellum2.1 leads on competitive programming (LiveCodeBench 82.0 vs 75.4) and tool calling (BFCL v4 62.3 vs 58.5). Choose by your task shape; there is no universal superiority.
Q3: Can I get the claimed 1.6x speedup today?
No. The 1.6x relies on MTP speculative decoding, but the MTP head was not released with the weights and is marked coming soon. The weights you download today cannot reproduce that speedup. The "nearly 2x" heavy-load throughput figure is likewise vendor-reported with no published value. Only baseline vLLM performance is verifiable now.
Q4: Which runtimes support Mellum2.1? Can I run it in Ollama?
Only vLLM is officially supported today, with reasoning-parser set to qwen3. GGUF (llama.cpp/Ollama/LM Studio) is coming soon, not shipped. Community bundles claiming one-click Ollama support are ahead of reality; Ollama users need to wait for the GGUF release.
Q5: Do I need to log in or apply to download Mellum2.1?
No. The repo is ungated under the Apache 2.0 license — open the Hugging Face page and download all weights directly (verified via hf-mirror on 2026-10-10). Early figures of 198 downloads and 77 likes reflect the initial release. There is no hosted API; usage requires self-hosting.
Closing
The 2%-to-47% story is told, but the real test belongs to the community: JetBrains' self-reported numbers await third-party harness reproduction, and the release timing of the MTP head and GGUF weights will directly shape adoption. We will follow up with 12GB-card feasibility tests once quantizations land.
If you have already pulled the weights and run them, share your harness scores and tokens/s in the comments — especially side-by-side numbers against Qwen3.5-9B in the same environment, which is exactly where dual-run comparisons create the most value. And if you hit snags during deployment, the deployment SOP above has a comments section too.