If you want to run a purpose-built coding agent model on your own hardware, you have probably hit the same dilemma: cloud APIs are convenient, but your code and data leave your domain; small local models run fine, but they cannot handle real work. Mellum2.1-12B-A2.5B-Thinking, published by JetBrains on October 8, 2026, aims squarely at that gap: a 12B MoE with 2.5B active parameters that a consumer 24GB card can run. The catch is that it is a thinking model, the officially supported path is vLLM only, and the widely shared tutorials promising "1.6x speedup" or "one-command Ollama pull" are not reproducible today.
This SOP gives you a path that works right now: serving the model with vLLM, configuring the reasoning parser and tool call template, verifying 131K context, and defusing the three big traps (unreleased MTP head, no GGUF yet, thinking-output parsing). All commands follow the official model card. Performance conclusions should come from your own benchmarks on your own repositories.
Run this prerequisite checklist first. If any item fails, switch plans:
- NVIDIA GPU with 24GB or more VRAM (RTX 4090/3090 class) — an estimate, since JetBrains has published no official hardware numbers;
- Linux or WSL2 with working CUDA drivers, able to install a current vLLM;
- Acceptance of pure self-hosting — there is no official hosted API for Mellum2.1;
- Intended use of coding and agent tasks — do not treat it as a general chat model.
If your card has only 8GB to 12GB, bookmark this and go straight to the Qwen3.5-9B local deployment SOP, which has a mature quantization ecosystem.
What you are actually deploying
Mellum2.1 is a training-recipe upgrade over Mellum2. The architecture is unchanged: 12B total MoE parameters, 2.5B active, 131,072 context, Apache 2.0 license. What changed is the training method — RL moved from a finishing step to the core of training, across thousands of environments and millions of sandbox runs, adding math, competition, science, tool use, and software engineering tasks.
JetBrains' self-reported numbers (Pi harness v0.73.1, thinking mode): SWE-bench Verified 47.0 versus Qwen3.5-9B's 50.0; LiveCodeBench v6 at 82.0, ahead of Qwen's 75.4; BFCL v4 at 62.3, also above Qwen's 58.5. One-line reading: Mellum leads on short tasks, competition-style problems, and tool calling, while Qwen scores higher on long tasks in real repositories. This is not a blanket improvement — the pitch is high-throughput agent density on a fixed GPU budget, with data staying in your domain.
The repository is at huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking. As verified at writing time, it is ungated and Apache 2.0, pullable immediately. In plain text for easy copying: huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking. For background and a fuller read of the benchmarks, see the Mellum2.1 open-source analysis.
The vLLM main path in four steps
Step 1: Set up the environment
# Use a dedicated virtual environment
pip install -U vllm
# Confirm the version and GPU visibility
vllm --version
nvidia-smiExpected result: the vLLM version prints normally and nvidia-smi shows your GPU and driver. vLLM supports Mellum2.1 out of the box; use a current release, no patches needed.
Step 2: Start the server
vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking \
--reasoning-parser qwen3 \
--max-model-len 131072Expected result: after the weights download, the server starts and prints a listening address (port 8000 by default). The first run pulls roughly 24GB of BF16 weights (an estimate), so leave disk headroom. The --reasoning-parser qwen3 flag follows the official model card — this is a thinking model, and without the parser the reasoning text bleeds straight into the answer, which breaks agent integration. If VRAM is tight, start with --max-model-len 32768, confirm it works, then raise it step by step.
Step 3: Verify basic inference
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "JetBrains/Mellum2.1-12B-A2.5B-Thinking",
"messages": [{"role": "user", "content": "Write a Python HTTP request function with retries"}]
}'Expected result: the returned JSON places the reasoning in a reasoning_content field and a clean answer in content. If content contains think tags or large blocks of reasoning text, the parser is not active — go back to Step 2 and check the flag.
Step 4: Verify tool calls and long context
For tool calling, the model card offers an optional Hermes tool call template. Add --tool-call-parser hermes to the launch flags and send a request with a tools field:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "JetBrains/Mellum2.1-12B-A2.5B-Thinking",
"messages": [{"role": "user", "content": "Check the size of /var/log"}],
"tools": [{
"type": "function",
"function": {
"name": "run_shell",
"description": "Execute a shell command",
"parameters": {
"type": "object",
"properties": {"cmd": {"type": "string"}},
"required": ["cmd"]
}
}
}]
}'Expected result: the response contains a structured tool_calls field rather than plain text. Agent frameworks dispatch work through this field; if this step fails, everything downstream fails with it.
Long-context check: paste in a long code file and ask it to recall a specific function signature. Expected result: with 131,072 context configured, the long input is accepted and recall is accurate. Note that KV cache VRAM grows linearly with context, and a full 131K configuration costs several extra gigabytes — this sets up the memory math below.
Three traps, defused in advance
Trap one: the MTP head was not released with the weights, so the 1.6x speedup is not buyable today. The advertised single-request 1.6x gain relies on MTP speculative decoding, but the MTP head is marked coming soon and was not shipped with the main weights. Any tutorial teaching you speculative decoding flags right now is not reproducible. Run the base model stably today; add MTP when the head actually lands.
Trap two: no GGUF yet — Ollama and LM Studio users should not wait for a bundle. Support for llama.cpp, Ollama, and LM Studio is all still coming soon. "Announced support" and "pullable today" are different things, and any third party claiming to offer a Mellum2.1 GGUF bundle deserves suspicion — the model is two days old, and community bundles are neither trustworthy nor controlled. For running other models with Ollama, see the Ollama local deployment SOP, which also covers how to choose between inference tools.
Trap three: thinking-model output parsing. This is the most common failure point in agent integration. Mellum2.1 is a thinking model, and its raw output mixes reasoning with the answer. Downstream code without a reasoning parser will feed the "reasoning stream" into the agent loop as the final answer — at best verbose, at worst malformed tool calls and runaway loops. The --reasoning-parser qwen3 flag in vLLM (official model card) exists precisely for this. Do not omit it.
Why is vLLM the only route today? Because JetBrains shipped support specifications for vLLM only — the reasoning parser and the tool call template are both defined in vLLM's parameter system. The broader ecosystem (llama.cpp, Ollama, LM Studio) needs the GGUF weights first, then its own template and parser validation from scratch. In other words, every non-vLLM path is currently "waiting on upstream." Watch the official repository for updates instead of hunting for unofficial workarounds in community channels.
A note on the VRAM floor: a 12GB card cannot run the full BF16 weights; roughly 24GB is the estimated class, since JetBrains has published no official hardware numbers. For 12GB and below, the only realistic paths are waiting for GGUF quantization or switching to Qwen3.5-9B's quantization ecosystem.
Hardware and the VRAM ledger
A starting reference table based on estimates (no official hardware numbers exist; derived from roughly 24GB of BF16 weights — not measured, verify on your own card):
| Configuration | Verdict | Advice |
|---|---|---|
| Single 24GB card (4090/3090) | Viable floor | Start at 32K max-model-len, raise stepwise |
| 48GB (dual cards or A6000 class) | Comfortable | Full 131K context and concurrency become possible |
| 80GB (H100/A100) | Generous | Concurrency and longer KV headroom |
| 12GB or less | Cannot run full weights | Wait for GGUF, or use a 9B-class model |
Do the math carefully: the roughly 24GB of BF16 weights is only the floor. KV cache is a floating cost that grows with context length and concurrency. A single request at full 131K context is not light for a 24GB card, so start at 32K and watch nvidia-smi as you raise the limit. Comfortable full-131K KV headroom really begins at 48GB and up.
A word on budgets: Mellum2.1 is pitched as "high-throughput agent density on a fixed GPU budget," meaning the sparse 2.5B-active architecture lets the same card and the same power draw serve more concurrent requests. That advantage only materializes under concurrency, though — for a single serial request, the compute saved by sparse activation may not translate into noticeable speed. If your workload is the occasional manual question, a quantized Qwen3.5-9B is probably the simpler tenant; if you are running a resident agent farm or batch-processing tickets, Mellum's economics start to make sense.
For how this model compares on deployment thresholds across four local coding models, see the local coding model comparison review — that piece compares deployment barriers on five axes, while this one gets you running.
A two-model reconciliation with Qwen3.5-9B
JetBrains' benchmark table reflects only its own harness, and most scores appear only as chart values with no raw numbers to audit. For a production decision, reconcile it yourself:
- Fix one harness and one task set; point it at local endpoints for Mellum2.1 and Qwen3.5-9B in turn;
- Record two columns: tokens/s at fixed input/output, and tickets resolved on real tasks;
- Only compare tokens/s across matching precision tiers — Mellum runs full BF16, and if Qwen is quantized, label that clearly;
- Tickets resolved tracks your real returns better than any benchmark.
For directional guidance, the official reading holds: Mellum may lead on short tasks, competition problems, and tool calling, while Qwen3.5-9B scores higher on long tasks in real repositories — but let your own task distribution decide. JetBrains also claims near-2x throughput versus Qwen3.5-9B under heavy load on H200; that is official, chart-only, and should not be transplanted into a 24GB card budget. For the API-versus-hosted tradeoff, the comparison review covers that side as well.
Two failure modes deserve their own paragraph because they generate the most confused bug reports. First, out-of-memory at model load: the BF16 checkpoint needs roughly 24GB before any context, so if your card has less, vLLM will die during weight loading, not mid-generation - the fix is a bigger card or waiting for official quantized builds, not squeezing --gpu-memory-utilization, which only throttles KV cache after the weights already fit. Second, stalled downloads on the first run: the checkpoint is large enough that proxy timeouts look like a hung process; let the download finish once, keep the HF cache directory intact, and every later start becomes local. If you are migrating an existing vLLM server, stop it before loading Mellum2.1 on the same node - two 24GB-class models will not coexist on consumer hardware, and the OOM message will point at whichever model loaded second.
FAQ
Q1: Can a 16GB card run Mellum2.1? No. The full BF16 weights run around 24GB (estimated class), which a 16GB card cannot hold. The realistic paths are waiting for GGUF quantization or switching to Qwen3.5-9B, which has a mature quantization ecosystem.
Q2: Is --reasoning-parser qwen3 mandatory? Strongly recommended — it follows the official model card. Without it, reasoning text bleeds into the answers, and agent integration breaks with malformed formats and runaway loops.
Q3: Why can't I reproduce the advertised 1.6x speedup? That number relies on MTP speculative decoding, and the MTP head was not released with the weights (coming soon). Nobody can reproduce it today. Stabilize the base model first; add MTP when the head ships.
Q4: When will Ollama support this model? GGUF support is officially marked coming soon, with no date. "Announced" does not mean usable today; distrust third-party bundles and watch the huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking repository page for updates.
Q5: Can a 24GB card run the full 131K context? The weights alone consume roughly 24GB, and full-131K KV cache adds on top, so a 24GB card likely cannot open it fully. Start at 32K and raise stepwise; 48GB and up is where full context becomes comfortable.
The only deployment route for Mellum2.1 today is vLLM, and the three traps are explicit: no MTP head, no GGUF, and a mandatory reasoning parser for thinking output. A 24GB card is a viable floor (estimated), 131K context needs KV headroom, and performance claims should never be taken from charts — run both models under one harness and let your own repositories decide.
Three upcoming milestones are worth a calendar reminder: the MTP head release (time to retest single-request speedup), the GGUF release (the entry signal for small-VRAM cards and Ollama users), and the first independent harness replications (a second data source beyond JetBrains' self-reported table). Any one of these landing may change the deployment advice above, so treat the official repository as the source of truth.
If you have it running, drop your card model and tokens/s in the comments to help the next reader; if you are stuck, paste the error and let's debug it together.