Field SOP
Field SOP

Mellum2.1 Deploy SOP: Run a 12B Coding Agent on 24GB

A Mellum2.1 local deployment SOP (as of October 10, 2026; pairs with the open-source and comparison pieces). The vLLM main route in five steps: serve JetBrains/Mellum2.1-12B-A2.5B-Thinking, set reasoning-parser qwen3 (official model card; without it thinking text bleeds into answers), basic inference verification, tool calling (optional Hermes template) plus 131K-context validation, then an agent smoke test - each step with its expected result. Three traps first: the MTP head was not released with the weights, so the advertised 1.6x speedup is not reproducible today; GGUF is not out, so "announced" Ollama/LM Studio support is not usable today - distrust integration packs; and this is a thinking model, so a reasoning parser is mandatory. Hardware ledger: BF16 around the 24GB class (estimate; no official number), 16GB cards cannot hold it, 24GB (4090/3090) is the working floor and full-131K context adds KV cache on top. Two high-frequency failure modes get their own paragraphs: OOM dies at weight loading, not mid-generation (--gpu-memory-utilization only throttles KV cache, it cannot fit oversized weights), and first-run downloads of the large checkpoint can look hung - finish it once, keep the HF cache directory, and every later start is local. Dual-run reconciliation pairs with the comparison piece: race Qwen3.5-9B on the same harness for tokens/s and tickets solved; performance conclusions come from your own repo, not vendor charts. Repo: huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking.

Published October 10, 20269 min read
<!-- mellum-2-1-local-deploy-sop | sop | Mellum2.1 Deploy SOP: Run a 12B Coding Agent on 24GB -->

If you want to run a purpose-built coding agent model on your own hardware, you have probably hit the same dilemma: cloud APIs are convenient, but your code and data leave your domain; small local models run fine, but they cannot handle real work. Mellum2.1-12B-A2.5B-Thinking, published by JetBrains on October 8, 2026, aims squarely at that gap: a 12B MoE with 2.5B active parameters that a consumer 24GB card can run. The catch is that it is a thinking model, the officially supported path is vLLM only, and the widely shared tutorials promising "1.6x speedup" or "one-command Ollama pull" are not reproducible today.

This SOP gives you a path that works right now: serving the model with vLLM, configuring the reasoning parser and tool call template, verifying 131K context, and defusing the three big traps (unreleased MTP head, no GGUF yet, thinking-output parsing). All commands follow the official model card. Performance conclusions should come from your own benchmarks on your own repositories.

Run this prerequisite checklist first. If any item fails, switch plans:

  1. NVIDIA GPU with 24GB or more VRAM (RTX 4090/3090 class) — an estimate, since JetBrains has published no official hardware numbers;
  2. Linux or WSL2 with working CUDA drivers, able to install a current vLLM;
  3. Acceptance of pure self-hosting — there is no official hosted API for Mellum2.1;
  4. Intended use of coding and agent tasks — do not treat it as a general chat model.

If your card has only 8GB to 12GB, bookmark this and go straight to the Qwen3.5-9B local deployment SOP, which has a mature quantization ecosystem.


What you are actually deploying

Mellum2.1 is a training-recipe upgrade over Mellum2. The architecture is unchanged: 12B total MoE parameters, 2.5B active, 131,072 context, Apache 2.0 license. What changed is the training method — RL moved from a finishing step to the core of training, across thousands of environments and millions of sandbox runs, adding math, competition, science, tool use, and software engineering tasks.

JetBrains' self-reported numbers (Pi harness v0.73.1, thinking mode): SWE-bench Verified 47.0 versus Qwen3.5-9B's 50.0; LiveCodeBench v6 at 82.0, ahead of Qwen's 75.4; BFCL v4 at 62.3, also above Qwen's 58.5. One-line reading: Mellum leads on short tasks, competition-style problems, and tool calling, while Qwen scores higher on long tasks in real repositories. This is not a blanket improvement — the pitch is high-throughput agent density on a fixed GPU budget, with data staying in your domain.

The repository is at huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking. As verified at writing time, it is ungated and Apache 2.0, pullable immediately. In plain text for easy copying: huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking. For background and a fuller read of the benchmarks, see the Mellum2.1 open-source analysis.


The vLLM main path in four steps

Step 1: Set up the environment

bash
# Use a dedicated virtual environment
pip install -U vllm

# Confirm the version and GPU visibility
vllm --version
nvidia-smi

Expected result: the vLLM version prints normally and nvidia-smi shows your GPU and driver. vLLM supports Mellum2.1 out of the box; use a current release, no patches needed.

Step 2: Start the server

bash
vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking \
  --reasoning-parser qwen3 \
  --max-model-len 131072

Expected result: after the weights download, the server starts and prints a listening address (port 8000 by default). The first run pulls roughly 24GB of BF16 weights (an estimate), so leave disk headroom. The --reasoning-parser qwen3 flag follows the official model card — this is a thinking model, and without the parser the reasoning text bleeds straight into the answer, which breaks agent integration. If VRAM is tight, start with --max-model-len 32768, confirm it works, then raise it step by step.

Step 3: Verify basic inference

bash
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "JetBrains/Mellum2.1-12B-A2.5B-Thinking",
    "messages": [{"role": "user", "content": "Write a Python HTTP request function with retries"}]
  }'

Expected result: the returned JSON places the reasoning in a reasoning_content field and a clean answer in content. If content contains think tags or large blocks of reasoning text, the parser is not active — go back to Step 2 and check the flag.

Step 4: Verify tool calls and long context

For tool calling, the model card offers an optional Hermes tool call template. Add --tool-call-parser hermes to the launch flags and send a request with a tools field:

bash
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "JetBrains/Mellum2.1-12B-A2.5B-Thinking",
    "messages": [{"role": "user", "content": "Check the size of /var/log"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "run_shell",
        "description": "Execute a shell command",
        "parameters": {
          "type": "object",
          "properties": {"cmd": {"type": "string"}},
          "required": ["cmd"]
        }
      }
    }]
  }'

Expected result: the response contains a structured tool_calls field rather than plain text. Agent frameworks dispatch work through this field; if this step fails, everything downstream fails with it.

Long-context check: paste in a long code file and ask it to recall a specific function signature. Expected result: with 131,072 context configured, the long input is accepted and recall is accurate. Note that KV cache VRAM grows linearly with context, and a full 131K configuration costs several extra gigabytes — this sets up the memory math below.


Three traps, defused in advance

Trap one: the MTP head was not released with the weights, so the 1.6x speedup is not buyable today. The advertised single-request 1.6x gain relies on MTP speculative decoding, but the MTP head is marked coming soon and was not shipped with the main weights. Any tutorial teaching you speculative decoding flags right now is not reproducible. Run the base model stably today; add MTP when the head actually lands.

Trap two: no GGUF yet — Ollama and LM Studio users should not wait for a bundle. Support for llama.cpp, Ollama, and LM Studio is all still coming soon. "Announced support" and "pullable today" are different things, and any third party claiming to offer a Mellum2.1 GGUF bundle deserves suspicion — the model is two days old, and community bundles are neither trustworthy nor controlled. For running other models with Ollama, see the Ollama local deployment SOP, which also covers how to choose between inference tools.

Trap three: thinking-model output parsing. This is the most common failure point in agent integration. Mellum2.1 is a thinking model, and its raw output mixes reasoning with the answer. Downstream code without a reasoning parser will feed the "reasoning stream" into the agent loop as the final answer — at best verbose, at worst malformed tool calls and runaway loops. The --reasoning-parser qwen3 flag in vLLM (official model card) exists precisely for this. Do not omit it.

Why is vLLM the only route today? Because JetBrains shipped support specifications for vLLM only — the reasoning parser and the tool call template are both defined in vLLM's parameter system. The broader ecosystem (llama.cpp, Ollama, LM Studio) needs the GGUF weights first, then its own template and parser validation from scratch. In other words, every non-vLLM path is currently "waiting on upstream." Watch the official repository for updates instead of hunting for unofficial workarounds in community channels.

A note on the VRAM floor: a 12GB card cannot run the full BF16 weights; roughly 24GB is the estimated class, since JetBrains has published no official hardware numbers. For 12GB and below, the only realistic paths are waiting for GGUF quantization or switching to Qwen3.5-9B's quantization ecosystem.


Hardware and the VRAM ledger

A starting reference table based on estimates (no official hardware numbers exist; derived from roughly 24GB of BF16 weights — not measured, verify on your own card):

ConfigurationVerdictAdvice
Single 24GB card (4090/3090)Viable floorStart at 32K max-model-len, raise stepwise
48GB (dual cards or A6000 class)ComfortableFull 131K context and concurrency become possible
80GB (H100/A100)GenerousConcurrency and longer KV headroom
12GB or lessCannot run full weightsWait for GGUF, or use a 9B-class model

Do the math carefully: the roughly 24GB of BF16 weights is only the floor. KV cache is a floating cost that grows with context length and concurrency. A single request at full 131K context is not light for a 24GB card, so start at 32K and watch nvidia-smi as you raise the limit. Comfortable full-131K KV headroom really begins at 48GB and up.

A word on budgets: Mellum2.1 is pitched as "high-throughput agent density on a fixed GPU budget," meaning the sparse 2.5B-active architecture lets the same card and the same power draw serve more concurrent requests. That advantage only materializes under concurrency, though — for a single serial request, the compute saved by sparse activation may not translate into noticeable speed. If your workload is the occasional manual question, a quantized Qwen3.5-9B is probably the simpler tenant; if you are running a resident agent farm or batch-processing tickets, Mellum's economics start to make sense.

For how this model compares on deployment thresholds across four local coding models, see the local coding model comparison review — that piece compares deployment barriers on five axes, while this one gets you running.


A two-model reconciliation with Qwen3.5-9B

JetBrains' benchmark table reflects only its own harness, and most scores appear only as chart values with no raw numbers to audit. For a production decision, reconcile it yourself:

  1. Fix one harness and one task set; point it at local endpoints for Mellum2.1 and Qwen3.5-9B in turn;
  2. Record two columns: tokens/s at fixed input/output, and tickets resolved on real tasks;
  3. Only compare tokens/s across matching precision tiers — Mellum runs full BF16, and if Qwen is quantized, label that clearly;
  4. Tickets resolved tracks your real returns better than any benchmark.

For directional guidance, the official reading holds: Mellum may lead on short tasks, competition problems, and tool calling, while Qwen3.5-9B scores higher on long tasks in real repositories — but let your own task distribution decide. JetBrains also claims near-2x throughput versus Qwen3.5-9B under heavy load on H200; that is official, chart-only, and should not be transplanted into a 24GB card budget. For the API-versus-hosted tradeoff, the comparison review covers that side as well.


Two failure modes deserve their own paragraph because they generate the most confused bug reports. First, out-of-memory at model load: the BF16 checkpoint needs roughly 24GB before any context, so if your card has less, vLLM will die during weight loading, not mid-generation - the fix is a bigger card or waiting for official quantized builds, not squeezing --gpu-memory-utilization, which only throttles KV cache after the weights already fit. Second, stalled downloads on the first run: the checkpoint is large enough that proxy timeouts look like a hung process; let the download finish once, keep the HF cache directory intact, and every later start becomes local. If you are migrating an existing vLLM server, stop it before loading Mellum2.1 on the same node - two 24GB-class models will not coexist on consumer hardware, and the OOM message will point at whichever model loaded second.

FAQ

Q1: Can a 16GB card run Mellum2.1? No. The full BF16 weights run around 24GB (estimated class), which a 16GB card cannot hold. The realistic paths are waiting for GGUF quantization or switching to Qwen3.5-9B, which has a mature quantization ecosystem.

Q2: Is --reasoning-parser qwen3 mandatory? Strongly recommended — it follows the official model card. Without it, reasoning text bleeds into the answers, and agent integration breaks with malformed formats and runaway loops.

Q3: Why can't I reproduce the advertised 1.6x speedup? That number relies on MTP speculative decoding, and the MTP head was not released with the weights (coming soon). Nobody can reproduce it today. Stabilize the base model first; add MTP when the head ships.

Q4: When will Ollama support this model? GGUF support is officially marked coming soon, with no date. "Announced" does not mean usable today; distrust third-party bundles and watch the huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking repository page for updates.

Q5: Can a 24GB card run the full 131K context? The weights alone consume roughly 24GB, and full-131K KV cache adds on top, so a 24GB card likely cannot open it fully. Start at 32K and raise stepwise; 48GB and up is where full context becomes comfortable.


The only deployment route for Mellum2.1 today is vLLM, and the three traps are explicit: no MTP head, no GGUF, and a mandatory reasoning parser for thinking output. A 24GB card is a viable floor (estimated), 131K context needs KV headroom, and performance claims should never be taken from charts — run both models under one harness and let your own repositories decide.

Three upcoming milestones are worth a calendar reminder: the MTP head release (time to retest single-request speedup), the GGUF release (the entry signal for small-VRAM cards and Ollama users), and the first independent harness replications (a second data source beyond JetBrains' self-reported table). Any one of these landing may change the deployment advice above, so treat the official repository as the source of truth.

If you have it running, drop your card model and tokens/s in the comments to help the next reader; if you are stuck, paste the error and let's debug it together.

This article is AI-assisted and human-edited. Last updated: 2026-10-10

FAQ

Can a 16GB card run Mellum2.1?
No. The full BF16 weights run around 24GB (estimated class), which a 16GB card cannot hold. The realistic paths are waiting for GGUF quantization or switching to Qwen3.5-9B, which has a mature quantization ecosystem.
Is --reasoning-parser qwen3 mandatory?
Strongly recommended — it follows the official model card. Without it, reasoning text bleeds into the answers, and agent integration breaks with malformed formats and runaway loops.
Why can't I reproduce the advertised 1.6x speedup?
That number relies on MTP speculative decoding, and the MTP head was not released with the weights (coming soon). Nobody can reproduce it today. Stabilize the base model first; add MTP when the head ships.
When will Ollama support this model?
GGUF support is officially marked coming soon, with no date. "Announced" does not mean usable today; distrust third-party bundles and watch the huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking repository page for updates.
Can a 24GB card run the full 131K context?
The weights alone consume roughly 24GB, and full-131K KV cache adds on top, so a 24GB card likely cannot open it fully. Start at 32K and raise stepwise; 48GB and up is where full context becomes comfortable.

Related

Field SOP

ComfyUI Video Generation SOP: Run LTX-2.5 on 12GB VRAM

A ComfyUI local video generation SOP, using LTX-2.5 to get a 12GB card through the full pipeline. Model files and directories (int8_convrot builds, not BF16): the transformer goes to models/diffusion_models/, gemma4-12b-with-proj to models/text_encoders/, the video and audio VAEs (both bf16) to models/vae/, and the 2x latent upscaler to models/latent_upscale_models/. Pitfalls first: the audio VAE is a separate file - skip it and you render a silent video; the prompt enhancer is an extra ~5GB model adding 1-2 minutes per generation and is off by default; HF gated=auto means you must log in and accept terms to download (even to read the model card). Hardware floor: 12GB VRAM with distilled/int8, CUDA 12.7+, PyTorch ~2.7, Python 3.12+ (innfactory's compilation); on 8GB cards, switch to Wan 2.2 TI2V-5B with offload instead. Three workflow types: two-stage T2V/I2V (generate, then 2x latent upsample refinement), single-stage (fast, for previews), and A2V (audio-driven video). VRAM strategy: quantized/distilled first, ComfyUI native offload, short low-res generation then upscale. Repos: github.com/Lightricks/ComfyUI-LTXVideo, huggingface.co/Lightricks/LTX-2.5.

Oct 10, 20269 min read
Field SOP

MiMo-V2.6 Integration SOP: Desktop, API and Local Weights

A hands-on SOP for wiring MiMo-V2.6 into your workflow along three routes of rising effort: (1) the MiMo Desktop client (install, sign in or plug in your own API key, switch MiMo-V2.6-Pro/Flash in the model list, UltraSpeed mode, screenshot-feedback iteration); (2) the MiMo open-platform API (create an app for a key, pass model name, messages, tools and multimodal inputs; for agent tasks, wire environment logs, test pass rates, screenshots and verifier feedback into an execute-check-correct loop; validate price, latency and success rate on a small traffic slice before production); (3) local weights (download from the HF collection collections/XiaomiMiMo/mimo-v26; parameter counts are unpublished, so hardware floors defer to the model cards; for RL reproduction, go through the five environment scripts in XiaomiMiMo/verl). Includes a pitfall table and a pre-launch checklist; API pricing defers to the platform documentation.

Sep 27, 202610 min read
Field SOP

LLaDA-Image Local Deploy SOP: Setup, Inference, Production

A five-step SOP for running Ant's open-source 6B image model LLaDA-Image: (1) environment setup with dependencies and mirror-accelerated downloads; (2) choosing among four weight variants (Base 50-step / Turbo 4-step, each in BF16 or FP8, with ModelScope for China); (3) generating the first image with minimal Base and Turbo commands; (4) advanced work - reference-image editing, text rendering, ComfyUI integration, and degradation strategies when VRAM runs short; (5) productionizing with batch queues, concurrency sizing, cost monitoring, result storage and graceful failure modes. Includes 6 pitfalls and a 10-item launch checklist, with every command copied verbatim from the official README; note the repo license is null, so confirm rights before commercial use.

Sep 9, 202611 min read