The name "Flash" has become the biggest source of confusion in local deployment. GLM-5.3-Flash carries a lightweight-sounding suffix, so many assume an 8GB consumer GPU can pull it. Reality: its FP8 checkpoint runs about 306GiB, larger than the combined 282GB of two H200 cards, and the practical floor is an 8-GPU Hopper node. Meanwhile, the two models that actually fit consumer hardware carry no "Flash" in their names at all. This review strips away the naming filter and slots the four contenders -- JetBrains Mellum2.1-12B, Qwen3.5-9B, DeepSeek-V4.1-Flash, and GLM-5.3-Flash -- into their rightful places by deployment cost.
Why now? Because self-hosting a coding model has, for the first time, four serious open-weight options at four different scales, spanning the full hardware ladder from a single 8GB consumer card to a small machine room. Picking wrong means OOM after downloading 500GB, or a wasted server budget.
What We Compare, and What We Don't
This review compares deployment requirements only, not capability. The reason is simple: each vendor's self-reported benchmarks use different harnesses, environments, and task sets, and merging scores across tables is the most common abuse of self-reported data. JetBrains' official comparison against Qwen3.5-9B holds only within that table; we cite it axis by axis and never synthesize a single "who wins" leaderboard.
We fix five axes: parameter scale, open-source license, context length, hardware floor, and runtime support. The choice is deliberate -- parameters determine quantization headroom, the license determines commercial use, context length determines whether a whole repository fits, the hardware floor determines cost, and runtime support determines whether the model actually runs today. Every number carries its source attribution: official where official, third-party estimates labeled as such, unpublished is written as unpublished. License and download counts are our own verification, with the date noted.
The Five-Axis Table
| Axis | Mellum2.1-12B-A2.5B-Thinking | Qwen3.5-9B | DeepSeek-V4.1-Flash | GLM-5.3-Flash |
|---|---|---|---|---|
| Parameters | 12B MoE / 2.5B active | 9.7B dense (9.65B, official figure) | No unified published figure; checkpoint about 510GB (third-party compilation) | 321.3B MoE / 18B active (288 experts) |
| License | Apache 2.0, ungated | Apache 2.0, ungated | MIT, ungated | MIT, ungated |
| Context | 131,072 | 262,144 (256K) | Not published | 1,048,576 (1M; Z.ai self-reports 300K+ managed context) |
| Hardware floor | BF16 roughly 24GB class (third-party estimate), runs on a single card via vLLM | Q4 about 6GB, sweet spot for 8GB cards; FP16 about 18GB | Node-class: even two H200 cards struggle with V4-Flash; V4.1-Flash checkpoint about 510GB | Node-class: FP8 about 306GiB, exceeding two H200 cards at 282GB, no two-card entry point, practical floor an 8-GPU Hopper node; Q4 about 194.6GB |
| Runtime support | vLLM from day one; GGUF/Ollama/LM Studio not yet available; MTP head not shipped | Ollama/GGUF/vLLM ecosystem fully ready | vLLM, node-class deployment | vLLM; third-party hosted inference (Baseten/DeepInfra/Fireworks/Novita/Together, roughly $0.15/$0.50 per million tokens) |
The license row was verified by us on 2026-10-10 through the HF API via hf-mirror, repository by repository: all four are ungated, with licenses effective immediately for commercial use. Download counts, same-day figures: Mellum2.1, live for two days, has 198 downloads and 77 likes; Qwen3.5-9B has 8.386 million; DeepSeek-V4.1-Flash has 1.316 million; GLM-5.3-Flash has 6.39 million. A fair reminder: Mellum's 48-hour download count should not be compared directly against models that have accumulated for a year.
The four repository addresses, each in link and plain-text form:
- huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking
- huggingface.co/Qwen/Qwen3.5-9B
- huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- huggingface.co/zai-org/GLM-5.3-Flash
Plain text again for easy copying: huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking, huggingface.co/Qwen/Qwen3.5-9B, huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash, huggingface.co/zai-org/GLM-5.3-Flash.
Quick Takes: Who Should Use Which
Mellum2.1-12B-A2.5B-Thinking: the only agent-specialized model a single card can run
JetBrains published it on Hugging Face on 2026-10-08 as a training-recipe upgrade over Mellum2: reinforcement learning moved from a finishing step to the core of training, with thousands of environments and millions of sandbox runs, plus new coverage of math, competition, science, tool use, and software engineering. The architecture is unchanged: 12B MoE total parameters, 2.5B active, 131K context, code-only positioning (no image input). JetBrains said it plainly that open training data often carries broken tests and unverifiable answers, so they filtered at the source; the payoff is a broad lift in thinking-mode capability rather than a single-benchmark spike.
JetBrains' self-test figures (Pi harness v0.73.1, thinking mode, not independently reproduced) contain two contrasts: SWE-bench Verified 47.0 versus Qwen3.5-9B's 50.0, trailing on real-repository long tasks; SWE-bench Pro 28.0 versus 38.0, Terminal-Bench 2.1 17.4 versus 21.7, and GPQA Diamond 64.6 versus 77.8, all Qwen-favored. On the other side, LiveCodeBench v6 82.0 versus 75.4 and BFCL v4 62.3 versus 58.5 favor Mellum on short competition-style tasks and tool calling. That is where the in-table reading stops -- the real selling point is high-throughput agent density on a fixed GPU budget, with data never leaving your domain. JetBrains also cites a speed figure: on H200 under heavy load, throughput approaches twice that of Qwen3.5-9B, though the exact numbers appear only in charts, reported here as-is.
Three pitfalls to know up front. First, the MTP head was not shipped with the weights, so the official 1.6x single-request speedup cannot be reproduced today. Second, GGUF is still "coming soon," so Ollama and LM Studio users cannot pull it yet. Third, the BF16 weights run roughly 24GB class by third-party estimate, so a 12GB card cannot hold the full model. It is also a thinking model: when serving through vLLM, configure the qwen3 reasoning parser or the chain-of-thought will leak into output. Full deployment steps are in our Mellum2.1 local deploy SOP, and the ecosystem background is covered in our Mellum2.1 open-source resource roundup.
Who should use it: teams with a fixed 24GB-class GPU, batch agent workloads, and strict data-sovereignty requirements. A solo owner of one RTX 4090 can still pull it for fun, but keep expectations calibrated -- it is agent-specialized, and asking it to write marketing copy will disappoint.
Qwen3.5-9B: king of the 8GB card
The only true consumer-GPU sweet spot among the four. Official figures: 9.7B dense, 262K context, with Q4 quantization fitting in about 6GB -- an 8GB card still has room for KV cache and system overhead after that. A 262K context means a mid-sized repository fits whole for analysis, unthinkable two years ago. Ollama, GGUF, and vLLM support are all ready today; On capability, stick to in-table facts: in JetBrains' comparison it scores higher on real-repository tasks (SWE-bench Verified 50.0 versus 47.0, SWE-bench Pro 38.0 versus 28.0), and it accepts image input, making it the more general-purpose of the two. Everyday repository work -- fixing bugs, reviewing diffs, adjusting UI from screenshots -- will likely feel more natural here; running automated agents at scale on fixed GPUs is Mellum's territory.
Our earlier Qwen3.8 Flash local deployment SOP covers the same family; the Qwen local path has stayed consistent.
Who should use it: individual developers with an 8GB to 24GB card who want to write code, close tickets, and experiment at zero marginal cost. No suspense -- it is this one.
DeepSeek-V4.1-Flash: node-class, with unpublished parameters
Discipline first: DeepSeek has not published a unified parameter figure for V4.1-Flash, so we do not invent one. This is the only model in the review whose parameter axis stays blank -- we will not back into a parameter count from checkpoint size, a parlor trick common in the wild.
Hard facts that can be confirmed: third-party compilations put the checkpoint at about 510GB. For perspective, two H200 cards total 282GB of memory -- not enough even for V4-Flash -- so V4.1-Flash is node-only. Context length is also unpublished. The MIT license, ungated, is verified, and 1.316 million downloads show a real ecosystem. Runtime support is vLLM node-class deployment.
A pragmatic test: if DeepSeek API spending dominates your bill and data residency is a hard requirement, node-class self-hosting becomes worth discussing; otherwise owning a node rarely beats the official API.
Who should use it: institutions deep in the DeepSeek ecosystem that want weight autonomy and already budget for node-class inference hardware. 8GB-card users can skip straight past it.
GLM-5.3-Flash: the soul of this review -- "Flash" does not mean small
The starkest contrast of the four. 321.3B MoE total parameters, 18B active, 288 experts; the "Flash" refers to activation efficiency, not size. The FP8 weights run about 306GiB, exceeding the 282GB of two H200 cards, and there is no two-card entry point -- the practical floor is an 8-GPU Hopper node. Even Q4 quantization lands at about 194.6GB, far beyond a mainstream four-card workstation.
Why build it this way? Because its design target is different: the only 1M context in this group (Z.ai self-reports 300K+ managed-context capability) plus image input. To ingest an entire large monorepo or hundreds of documents in one request, parameter count and expert count cannot be squeezed down; "Flash" saves per-forward-pass compute -- 18B active means inference cost approaches an 18B model, provided you can first fit the weights somewhere.
If building a node is out of reach, third-party hosted inference (Baseten, DeepInfra, Fireworks, Novita, Together) runs roughly $0.15/$0.50 per million tokens as a middle path. The license is MIT, ungated, with 6.39 million downloads.
Who should use it: institutional scenarios -- very long document and codebase analysis, multimodal input, 1M context as a hard requirement, and a node budget in place. Individual developers can admire it from afar; this is not your battlefield.
The VRAM Ladder, At a Glance
- 8GB card: Qwen3.5-9B Q4 at about 6GB, the only sweet spot. Do not even consider the other three.
- 24GB card (RTX 4090/3090 class): Mellum2.1 BF16 at roughly 24GB class (estimate) is feasible, with extra memory leaving KV headroom for 131K context; Qwen3.5-9B FP16 at about 18GB runs comfortably with room for the 256K context. This tier is the main arena for individuals and small teams.
- Single machine, 2-4 cards: the awkward zone. GLM-5.3-Flash has no two-card entry, and DeepSeek-V4.1-Flash's 510GB checkpoint dwarfs even two H200 cards. Stay on Qwen or Mellum here; forcing multi-node inference will bleed the gains into network overhead.
- 8-GPU Hopper node: the home ground of GLM-5.3-Flash FP8 at about 306GiB and DeepSeek-V4.1-Flash at about 510GB, served through vLLM. At this tier the comparison shifts from "can it run" to per-token cost and concurrency density.
One sentence for the ladder: the loudest "Flash" lives on the most expensive floor, and the ones you can actually afford are quiet mid-size cups. One footnote: quantization quality loss does not appear in this table -- GLM-5.3-Flash's Q4 at about 194.6GB is a volume figure, not a quality guarantee, so run your own eval pass before committing.
FAQ
Q1: Which of the four models is the strongest? We decline to rank across tables. Each vendor's self-reported benchmarks use different harnesses and task sets, and merging scores across tables is data abuse. This review picks by deployment floor and scenario: 8GB card, Qwen3.5-9B; fixed GPU for agents, Mellum2.1; institutional 1M context, GLM-5.3-Flash; DeepSeek ecosystem, V4.1-Flash.
Q2: GLM-5.3-Flash has "Flash" in the name -- is it a small model an 8GB card can run? No. Flash refers to activation efficiency, not size: 321.3B total parameters, FP8 weights about 306GiB, exceeding two H200 cards combined, with a practical floor of an 8-GPU Hopper node; even Q4 lands at about 194.6GB. The only 8GB-card sweet spot among the four is Qwen3.5-9B (Q4 about 6GB, official figures).
Q3: Are Mellum2.1's 1.6x speedup and Ollama support available today? No, neither. The MTP head was not shipped with the weights, so the official 1.6x single-request speedup cannot be reproduced yet; GGUF is not released, and Ollama/LM Studio support is coming soon. Only vLLM works today, with the qwen3 reasoning parser configured or the chain-of-thought will leak into output.
Q4: How many parameters does DeepSeek-V4.1-Flash actually have? DeepSeek has not published a unified figure, and we do not invent one. Third-party compilations put the checkpoint at about 510GB, indicating a node-class model -- two H200 cards (282GB) cannot hold it, and deployment requires an 8-GPU node with vLLM. The MIT license, ungated, was verified on 2026-10-10.
Q5: Are these licenses really safe for commercial use? Yes. We verified through the HF API via hf-mirror on 2026-10-10: Mellum2.1 and Qwen3.5-9B are Apache 2.0, DeepSeek-V4.1-Flash and GLM-5.3-Flash are MIT, all four repositories ungated with no additional restrictions. Same-day download counts were 198, 8.386 million, 1.316 million, and 6.39 million respectively.
Closing
Closing: the conclusion is plain -- the goal is not the strongest model but the one that fits inside your VRAM. 8GB cards take Qwen3.5-9B without hesitation; 24GB cards running agent density should look at Mellum2.1; GLM-5.3-Flash and DeepSeek-V4.1-Flash wait for a node budget. Four options, no hierarchy -- just each in its rightful place. What is your GPU setup, and which one are you landing on? Tell us in the comments. If the tooling side is still undecided, our local LLM deployment tools review covers Ollama versus LM Studio -- that one compares tools, this one compares models, and they read well together.