Hardcore Reviews
Hardcore Reviews

Local Coding Models: From 8GB GPU to 8-GPU Node

A four-way local-deployment comparison of open coding models on five axes (parameters, license, context, hardware floor, runtime support) - not a quality ranking, no cross-model benchmarks: Mellum2.1-12B (12B MoE/2.5B active, Apache 2.0, 131K, ~24GB BF16 estimated, vLLM today with GGUF not yet out); Qwen3.5-9B (9.7B dense, Apache 2.0, 262K, ~6GB Q4 as the 8GB-card sweet spot, full Ollama/GGUF ecosystem today); DeepSeek-V4.1-Flash (MIT, checkpoint around 510GB per third-party compilation, parameters unpublished so not invented, node-class); GLM-5.3-Flash (321.3B MoE/18B active, MIT, 1M context, FP8 ~306GiB exceeding two H200s at 282GB, no two-GPU entry point, realistic floor an 8-GPU Hopper node, third-party hosting at $0.15/$0.50 per million). All four licenses HF-API verified via hf-mirror on 2026-10-10, all ungated (downloads 198/8.386M/1.316M/6.39M, dated). The soul of the piece: "Flash" does not mean small - GLM-5.3-Flash is a node-class beast while the 12B Mellum is the only consumer-card-runnable agent-specialist of the four. JetBrains self-reported scores cited from its own table only; verdicts by scenario: 8GB-card coding picks Qwen3.5-9B, fixed-GPU agent farms pick Mellum2.1, institutional 1M-context picks GLM-5.3-Flash, DeepSeek ecosystems pick V4.1-Flash (both node budget).

Published October 10, 20269 min read
<!-- local-coding-model-comparison-review | review | Local Coding Models: From 8GB GPU to 8-GPU Node -->

The name "Flash" has become the biggest source of confusion in local deployment. GLM-5.3-Flash carries a lightweight-sounding suffix, so many assume an 8GB consumer GPU can pull it. Reality: its FP8 checkpoint runs about 306GiB, larger than the combined 282GB of two H200 cards, and the practical floor is an 8-GPU Hopper node. Meanwhile, the two models that actually fit consumer hardware carry no "Flash" in their names at all. This review strips away the naming filter and slots the four contenders -- JetBrains Mellum2.1-12B, Qwen3.5-9B, DeepSeek-V4.1-Flash, and GLM-5.3-Flash -- into their rightful places by deployment cost.

Why now? Because self-hosting a coding model has, for the first time, four serious open-weight options at four different scales, spanning the full hardware ladder from a single 8GB consumer card to a small machine room. Picking wrong means OOM after downloading 500GB, or a wasted server budget.

What We Compare, and What We Don't

This review compares deployment requirements only, not capability. The reason is simple: each vendor's self-reported benchmarks use different harnesses, environments, and task sets, and merging scores across tables is the most common abuse of self-reported data. JetBrains' official comparison against Qwen3.5-9B holds only within that table; we cite it axis by axis and never synthesize a single "who wins" leaderboard.

We fix five axes: parameter scale, open-source license, context length, hardware floor, and runtime support. The choice is deliberate -- parameters determine quantization headroom, the license determines commercial use, context length determines whether a whole repository fits, the hardware floor determines cost, and runtime support determines whether the model actually runs today. Every number carries its source attribution: official where official, third-party estimates labeled as such, unpublished is written as unpublished. License and download counts are our own verification, with the date noted.

The Five-Axis Table

AxisMellum2.1-12B-A2.5B-ThinkingQwen3.5-9BDeepSeek-V4.1-FlashGLM-5.3-Flash
Parameters12B MoE / 2.5B active9.7B dense (9.65B, official figure)No unified published figure; checkpoint about 510GB (third-party compilation)321.3B MoE / 18B active (288 experts)
LicenseApache 2.0, ungatedApache 2.0, ungatedMIT, ungatedMIT, ungated
Context131,072262,144 (256K)Not published1,048,576 (1M; Z.ai self-reports 300K+ managed context)
Hardware floorBF16 roughly 24GB class (third-party estimate), runs on a single card via vLLMQ4 about 6GB, sweet spot for 8GB cards; FP16 about 18GBNode-class: even two H200 cards struggle with V4-Flash; V4.1-Flash checkpoint about 510GBNode-class: FP8 about 306GiB, exceeding two H200 cards at 282GB, no two-card entry point, practical floor an 8-GPU Hopper node; Q4 about 194.6GB
Runtime supportvLLM from day one; GGUF/Ollama/LM Studio not yet available; MTP head not shippedOllama/GGUF/vLLM ecosystem fully readyvLLM, node-class deploymentvLLM; third-party hosted inference (Baseten/DeepInfra/Fireworks/Novita/Together, roughly $0.15/$0.50 per million tokens)

The license row was verified by us on 2026-10-10 through the HF API via hf-mirror, repository by repository: all four are ungated, with licenses effective immediately for commercial use. Download counts, same-day figures: Mellum2.1, live for two days, has 198 downloads and 77 likes; Qwen3.5-9B has 8.386 million; DeepSeek-V4.1-Flash has 1.316 million; GLM-5.3-Flash has 6.39 million. A fair reminder: Mellum's 48-hour download count should not be compared directly against models that have accumulated for a year.

The four repository addresses, each in link and plain-text form:

Plain text again for easy copying: huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking, huggingface.co/Qwen/Qwen3.5-9B, huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash, huggingface.co/zai-org/GLM-5.3-Flash.

Quick Takes: Who Should Use Which

Mellum2.1-12B-A2.5B-Thinking: the only agent-specialized model a single card can run

JetBrains published it on Hugging Face on 2026-10-08 as a training-recipe upgrade over Mellum2: reinforcement learning moved from a finishing step to the core of training, with thousands of environments and millions of sandbox runs, plus new coverage of math, competition, science, tool use, and software engineering. The architecture is unchanged: 12B MoE total parameters, 2.5B active, 131K context, code-only positioning (no image input). JetBrains said it plainly that open training data often carries broken tests and unverifiable answers, so they filtered at the source; the payoff is a broad lift in thinking-mode capability rather than a single-benchmark spike.

JetBrains' self-test figures (Pi harness v0.73.1, thinking mode, not independently reproduced) contain two contrasts: SWE-bench Verified 47.0 versus Qwen3.5-9B's 50.0, trailing on real-repository long tasks; SWE-bench Pro 28.0 versus 38.0, Terminal-Bench 2.1 17.4 versus 21.7, and GPQA Diamond 64.6 versus 77.8, all Qwen-favored. On the other side, LiveCodeBench v6 82.0 versus 75.4 and BFCL v4 62.3 versus 58.5 favor Mellum on short competition-style tasks and tool calling. That is where the in-table reading stops -- the real selling point is high-throughput agent density on a fixed GPU budget, with data never leaving your domain. JetBrains also cites a speed figure: on H200 under heavy load, throughput approaches twice that of Qwen3.5-9B, though the exact numbers appear only in charts, reported here as-is.

Three pitfalls to know up front. First, the MTP head was not shipped with the weights, so the official 1.6x single-request speedup cannot be reproduced today. Second, GGUF is still "coming soon," so Ollama and LM Studio users cannot pull it yet. Third, the BF16 weights run roughly 24GB class by third-party estimate, so a 12GB card cannot hold the full model. It is also a thinking model: when serving through vLLM, configure the qwen3 reasoning parser or the chain-of-thought will leak into output. Full deployment steps are in our Mellum2.1 local deploy SOP, and the ecosystem background is covered in our Mellum2.1 open-source resource roundup.

Who should use it: teams with a fixed 24GB-class GPU, batch agent workloads, and strict data-sovereignty requirements. A solo owner of one RTX 4090 can still pull it for fun, but keep expectations calibrated -- it is agent-specialized, and asking it to write marketing copy will disappoint.

Qwen3.5-9B: king of the 8GB card

The only true consumer-GPU sweet spot among the four. Official figures: 9.7B dense, 262K context, with Q4 quantization fitting in about 6GB -- an 8GB card still has room for KV cache and system overhead after that. A 262K context means a mid-sized repository fits whole for analysis, unthinkable two years ago. Ollama, GGUF, and vLLM support are all ready today; On capability, stick to in-table facts: in JetBrains' comparison it scores higher on real-repository tasks (SWE-bench Verified 50.0 versus 47.0, SWE-bench Pro 38.0 versus 28.0), and it accepts image input, making it the more general-purpose of the two. Everyday repository work -- fixing bugs, reviewing diffs, adjusting UI from screenshots -- will likely feel more natural here; running automated agents at scale on fixed GPUs is Mellum's territory.

Our earlier Qwen3.8 Flash local deployment SOP covers the same family; the Qwen local path has stayed consistent.

Who should use it: individual developers with an 8GB to 24GB card who want to write code, close tickets, and experiment at zero marginal cost. No suspense -- it is this one.

DeepSeek-V4.1-Flash: node-class, with unpublished parameters

Discipline first: DeepSeek has not published a unified parameter figure for V4.1-Flash, so we do not invent one. This is the only model in the review whose parameter axis stays blank -- we will not back into a parameter count from checkpoint size, a parlor trick common in the wild.

Hard facts that can be confirmed: third-party compilations put the checkpoint at about 510GB. For perspective, two H200 cards total 282GB of memory -- not enough even for V4-Flash -- so V4.1-Flash is node-only. Context length is also unpublished. The MIT license, ungated, is verified, and 1.316 million downloads show a real ecosystem. Runtime support is vLLM node-class deployment.

A pragmatic test: if DeepSeek API spending dominates your bill and data residency is a hard requirement, node-class self-hosting becomes worth discussing; otherwise owning a node rarely beats the official API.

Who should use it: institutions deep in the DeepSeek ecosystem that want weight autonomy and already budget for node-class inference hardware. 8GB-card users can skip straight past it.

GLM-5.3-Flash: the soul of this review -- "Flash" does not mean small

The starkest contrast of the four. 321.3B MoE total parameters, 18B active, 288 experts; the "Flash" refers to activation efficiency, not size. The FP8 weights run about 306GiB, exceeding the 282GB of two H200 cards, and there is no two-card entry point -- the practical floor is an 8-GPU Hopper node. Even Q4 quantization lands at about 194.6GB, far beyond a mainstream four-card workstation.

Why build it this way? Because its design target is different: the only 1M context in this group (Z.ai self-reports 300K+ managed-context capability) plus image input. To ingest an entire large monorepo or hundreds of documents in one request, parameter count and expert count cannot be squeezed down; "Flash" saves per-forward-pass compute -- 18B active means inference cost approaches an 18B model, provided you can first fit the weights somewhere.

If building a node is out of reach, third-party hosted inference (Baseten, DeepInfra, Fireworks, Novita, Together) runs roughly $0.15/$0.50 per million tokens as a middle path. The license is MIT, ungated, with 6.39 million downloads.

Who should use it: institutional scenarios -- very long document and codebase analysis, multimodal input, 1M context as a hard requirement, and a node budget in place. Individual developers can admire it from afar; this is not your battlefield.

The VRAM Ladder, At a Glance

  • 8GB card: Qwen3.5-9B Q4 at about 6GB, the only sweet spot. Do not even consider the other three.
  • 24GB card (RTX 4090/3090 class): Mellum2.1 BF16 at roughly 24GB class (estimate) is feasible, with extra memory leaving KV headroom for 131K context; Qwen3.5-9B FP16 at about 18GB runs comfortably with room for the 256K context. This tier is the main arena for individuals and small teams.
  • Single machine, 2-4 cards: the awkward zone. GLM-5.3-Flash has no two-card entry, and DeepSeek-V4.1-Flash's 510GB checkpoint dwarfs even two H200 cards. Stay on Qwen or Mellum here; forcing multi-node inference will bleed the gains into network overhead.
  • 8-GPU Hopper node: the home ground of GLM-5.3-Flash FP8 at about 306GiB and DeepSeek-V4.1-Flash at about 510GB, served through vLLM. At this tier the comparison shifts from "can it run" to per-token cost and concurrency density.

One sentence for the ladder: the loudest "Flash" lives on the most expensive floor, and the ones you can actually afford are quiet mid-size cups. One footnote: quantization quality loss does not appear in this table -- GLM-5.3-Flash's Q4 at about 194.6GB is a volume figure, not a quality guarantee, so run your own eval pass before committing.

FAQ

Q1: Which of the four models is the strongest? We decline to rank across tables. Each vendor's self-reported benchmarks use different harnesses and task sets, and merging scores across tables is data abuse. This review picks by deployment floor and scenario: 8GB card, Qwen3.5-9B; fixed GPU for agents, Mellum2.1; institutional 1M context, GLM-5.3-Flash; DeepSeek ecosystem, V4.1-Flash.

Q2: GLM-5.3-Flash has "Flash" in the name -- is it a small model an 8GB card can run? No. Flash refers to activation efficiency, not size: 321.3B total parameters, FP8 weights about 306GiB, exceeding two H200 cards combined, with a practical floor of an 8-GPU Hopper node; even Q4 lands at about 194.6GB. The only 8GB-card sweet spot among the four is Qwen3.5-9B (Q4 about 6GB, official figures).

Q3: Are Mellum2.1's 1.6x speedup and Ollama support available today? No, neither. The MTP head was not shipped with the weights, so the official 1.6x single-request speedup cannot be reproduced yet; GGUF is not released, and Ollama/LM Studio support is coming soon. Only vLLM works today, with the qwen3 reasoning parser configured or the chain-of-thought will leak into output.

Q4: How many parameters does DeepSeek-V4.1-Flash actually have? DeepSeek has not published a unified figure, and we do not invent one. Third-party compilations put the checkpoint at about 510GB, indicating a node-class model -- two H200 cards (282GB) cannot hold it, and deployment requires an 8-GPU node with vLLM. The MIT license, ungated, was verified on 2026-10-10.

Q5: Are these licenses really safe for commercial use? Yes. We verified through the HF API via hf-mirror on 2026-10-10: Mellum2.1 and Qwen3.5-9B are Apache 2.0, DeepSeek-V4.1-Flash and GLM-5.3-Flash are MIT, all four repositories ungated with no additional restrictions. Same-day download counts were 198, 8.386 million, 1.316 million, and 6.39 million respectively.

Closing

Closing: the conclusion is plain -- the goal is not the strongest model but the one that fits inside your VRAM. 8GB cards take Qwen3.5-9B without hesitation; 24GB cards running agent density should look at Mellum2.1; GLM-5.3-Flash and DeepSeek-V4.1-Flash wait for a node budget. Four options, no hierarchy -- just each in its rightful place. What is your GPU setup, and which one are you landing on? Tell us in the comments. If the tooling side is still undecided, our local LLM deployment tools review covers Ollama versus LM Studio -- that one compares tools, this one compares models, and they read well together.

This article is AI-assisted and human-edited. Last updated: 2026-10-10

FAQ

Which of the four models is the strongest?
We decline to rank across tables. Each vendor's self-reported benchmarks use different harnesses and task sets, and merging scores across tables is data abuse. This review picks by deployment floor and scenario: 8GB card, Qwen3.5-9B; fixed GPU for agents, Mellum2.1; institutional 1M context, GLM-5.3-Flash; DeepSeek ecosystem, V4.1-Flash.
GLM-5.3-Flash has "Flash" in the name -- is it a small model an 8GB card can run?
No. Flash refers to activation efficiency, not size: 321.3B total parameters, FP8 weights about 306GiB, exceeding two H200 cards combined, with a practical floor of an 8-GPU Hopper node; even Q4 lands at about 194.6GB. The only 8GB-card sweet spot among the four is Qwen3.5-9B (Q4 about 6GB, official figures).
Are Mellum2.1's 1.6x speedup and Ollama support available today?
No, neither. The MTP head was not shipped with the weights, so the official 1.6x single-request speedup cannot be reproduced yet; GGUF is not released, and Ollama/LM Studio support is coming soon. Only vLLM works today, with the qwen3 reasoning parser configured or the chain-of-thought will leak into output.
How many parameters does DeepSeek-V4.1-Flash actually have?
DeepSeek has not published a unified figure, and we do not invent one. Third-party compilations put the checkpoint at about 510GB, indicating a node-class model -- two H200 cards (282GB) cannot hold it, and deployment requires an 8-GPU node with vLLM. The MIT license, ungated, was verified on 2026-10-10.
Are these licenses really safe for commercial use?
Yes. We verified through the HF API via hf-mirror on 2026-10-10: Mellum2.1 and Qwen3.5-9B are Apache 2.0, DeepSeek-V4.1-Flash and GLM-5.3-Flash are MIT, all four repositories ungated with no additional restrictions. Same-day download counts were 198, 8.386 million, 1.316 million, and 6.39 million respectively.

Related

Hardcore Reviews

Open Video Model Local Deployment: Five-Axis Comparison Review

A four-way open-source video model comparison on the five axes that matter for local deployment: license, VRAM, audio, duration/resolution, and regional restrictions - not generation quality, no cross-model benchmarks, no ranking. LTX-2.5 (22B, distilled/FP8 from 12GB, 24kHz stereo in-pass, LTX Community license with a $10M threshold, gated=auto); Wan 2.2 (TI2V-5B, 8GB with offload, Apache 2.0 and the most permissive, short clips 480p-720p, no native audio); HunyuanVideo 1.5 (8.3B, from 14GB, ~75s per clip on a 4090 step-distilled build, Tencent community license not valid in EU/UK/South Korea with 100M-MAU renegotiation, no native audio); MiniMax H3 (33B dense, 32kHz stereo with dialogue in 11 languages, 15s/768p default while 2K needs an unreleased regenerator, community license excludes EU/UK/South Korea/USA with written authorization required above $20M, VRAM unpublished so not invented). Sources labeled per vendor (bestfreewebresources/pinggy compilations plus official model cards); fal's H3 Max #1 on i2v-with-audio is third-party. Verdicts by scenario, no arbitration: most permissive license picks Wan 2.2, audiovisual-in-one picks LTX-2.5, a 24GB card outside restricted regions picks HunyuanVideo 1.5, reference audio and multilingual dialogue pick H3 (within licensed regions).

Oct 10, 20269 min read
Hardcore Reviews

Half Price vs One-Fifth: The Flagship Alternative Shake-Up

A value-focused comparison of flagship-adjacent coding models as of October 2026: Claude Sonnet 5.5 (official basis via machine-intelligence press: 30%+ faster, Terminal-Bench 4.0 up from 10.3% to 70.6%, priced at half of Opus 5.5) vs GPT-6.1 Sol (official basis: near-Astra intelligence at one-fifth the API price, with a paid Ultrafast speed tier) vs Claude Opus 5.5 (the reference point, $4/$20 per million tokens) vs DeepSeek V4.1 (the open-weights reference). Discipline: harnesses differ so cross-model scores cannot be compared - this piece only uses relative-to-own-flagship ratios, does not run its own evals, and does not compute Sonnet 5.5's unpublished dollar pricing. Four scenario verdicts: budget-conscious daily coding, flagship-ceiling complex work, and compliance-driven private deployment each have a winner. Complements the site's 2026-08 comprehensive and flagship-reasoning comparisons.

Oct 1, 202610 min read
Hardcore Reviews

Embedding Model Comparison: 4 RAG Contenders, One Honest Table

A four-way RAG embedding model comparison (pricing snapshots October 9, 2026): EmbeddingGemma 2, Qwen3-Embedding (0.6B/4B/8B), BGE-M3 and OpenAI text-embedding-3-large, with gemini-embedding-001 as a reference row. All licenses API-verified on Hugging Face (2026-10-09): google/embeddinggemma-2 and Qwen3-Embedding-8B are apache-2.0 and ungated; BAAI/bge-m3 is mit (33.7M downloads). Engineering ledger aligned item by item: context 8192 (EG2/BGE-M3/OpenAI) vs 32768 (Qwen3); dimensions 768 (MRL to 128) / 1024-4096 / 1024 / 3072 (MRL to 256); three self-host free vs OpenAI at $0.13 per million tokens. Discipline first: each vendor's benchmark numbers cited from their own tables only (EG2's MTEB Code +9.92 and gemini-embedding-001's three board leads are official self-reports; Qwen3-8B's ~70.6 multilingual is third-party relayed) - no cross-table ranking; production popularity (relayed from LangChain/LlamaIndex telemetry): BGE-M3 first, Qwen3 second, OpenAI third. Verdicts by scenario, no arbitration: on-device multimodal picks EG2, long documents and Chinese pick Qwen3, hybrid retrieval and mature ecosystem pick BGE-M3, zero-ops picks OpenAI; and since switching means a full re-embed, decide once, decide carefully.

Oct 9, 202610 min read