Open Source
Open Source

Bilibili's 35B translation MoE: 150 languages, only 3B active

Bilibili's Index LLM team open-sourced the Index-Translate family on September 30, 2026 (weights on Hugging Face and ModelScope, Apache-2.0): 2B/9B/35B-A3B-preview text models built on Qwen3.5, targeting 150 languages (self-reported), with the 35B MoE activating only about 3B parameters per token and a 262,144-token context. Three branches split the work: Index-Echo produces dubbed speech that preserves the source speaker's voice (Chinese to/from English, Spanish, Japanese); Index-Homura controls syllable counts (the 9B lands within +-10% of the target on 81.92% of SandGlass cases - built for dubbing timing); Index-NativeLong (aka Index-Nailong) keeps long documents consistent (a 32K-token fantasy text keeps a name/royal-title pun intact, where chunk-by-chunk translation drifts). Constrained translation via instTrans: hard constraints enforce glossaries (e.g. carbon fiber) plus JSON/CSV/code/placeholder structure; soft constraints cover tone, domain disambiguation (plant to factory), cross-sentence consistency and LaTeX preservation. Local deployment is one line with the official GGUF (Q4_K_M, 21.71GB) via llama.cpp serve, plus a browser extension and a video dubbing pipeline. Honesty note: every score is self-reported with no independent replication yet (FLORES COMET-22 0.8794, WMT26 Judge 76.76, 2.4% off-target on low-resource pairs), and in the same official table DeepSeek-V4.1-Flash scores higher on WMT26 Judge at 83.55 - not first place on raw quality.

Published October 7, 202610 min read
<!-- bilibili-index-translate-resource | open-source | Bilibili's 35B translation MoE: 150 languages, only 3B active -->

On September 30, 2026, Bilibili's Index LLM team released something surprising for a video company: a family of open-weight models built specifically for translation. Published under Apache-2.0, the family covers 2B, 9B and 35B-A3B-preview text models, trained to work across what the team counts as 150 languages. The flagship is a mixture-of-experts model with 35 billion total parameters that activates only about 3 billion per token, and it holds up to 262,144 tokens in context.

For anyone who has watched Chinese fan-sub groups do their unpaid, frame-accurate magic on Japanese anime, or seen bilingual danmaku comments light up under a clip from Seoul, the origin story makes instant sense. Bilibili sits in one of the most multilingual corners of the Chinese internet, and it has now turned that community need into weights you can download, quantize and run on your own hardware.

This article walks through what was released, why a video platform cares about translation, how the three specialized branches differ, what the benchmarks honestly say, and how to run the flagship locally with a single llama.cpp command.

Why a video platform builds translation models

The easy assumption is a cheaper internal API bill. The reality is more interesting. Chinese otaku culture has been shaped by fan translation for two decades: volunteer groups time, translate and typeset subtitles for anime, variety shows and game trailers, often within hours of release. Quality can be superb, but the economics are fragile and long-tail coverage spotty. Meanwhile the platform's content ecosystem is inherently cross-border: Japanese gameplay footage with Chinese commentary, Korean dance covers with English reactions, and a comment layer that mixes Chinese, English, Japanese and Korean in a single thread.

Machine translation inside this ecosystem is not a general chat problem. A subtitle line must land within a syllable budget if it will be spoken aloud. A game announcement keeps its JSON structure and hashtags intact, or the downstream renderer breaks. A fantasy novel translated paragraph by paragraph quietly forgets how a title was rendered ten chapters ago. These are the failure modes general-purpose assistants hit, and precisely the problems the Index team scoped its models around.

There is also a strategic point: a translation stack Bilibili owns rather than rents compounds over time, and a permissive license buys goodwill with the exact developer community that builds fan tools and subtitle pipelines.

The family: three sizes on a Qwen3.5 base

The lineup consists of 2B, 9B and 35B-A3B-preview text models, built on Alibaba's Qwen3.5 as the foundation architecture. The stated coverage of 150 languages is an official self-reported figure, not something an independent benchmark has audited. The weights live on Hugging Face under the IndexTeam organization and on ModelScope, with code, a technical report and an online demo on GitHub at bilibili/Index-Translate. The license tag on the model card reads Apache-2.0, so commercial use, redistribution and fine-tuning are all on the table.

The flagship's name, 35B-A3B-preview, describes its economics: a sparse mixture-of-experts model with 35 billion parameters in total, of which roughly 3 billion activate for any given token. Inference cost tracks active parameters, so you get the capacity of a mid-30s-billion model while paying compute closer to a 3B dense model. For a platform translating millions of subtitle lines a day, that gap is a real budget line, not a rounding error.

The other headline number is context: the 35B model accepts 262,144 tokens, far beyond what a subtitle file needs, but exactly right for long documents where earlier context must inform the next paragraph.

Three branches for three jobs that generic models fail

The most distinctive part of this release is not the base models but the three specialized branches built on top of them.

Index-Echo handles speech. It is a speech-to-text-to-speech pipeline that transcribes audio, translates the transcript and speaks the result back, preserving the original speaker's voice timbre. Supported directions are Chinese to and from English, Spanish and Japanese. In practice this is the dubbing branch: take a talky video, keep the original voice's character, and let viewers hear it in another language.

Index-Homura tackles a constraint most people never consider until they try dubbing: syllable counts. A translated line spoken over original footage must fit roughly the same duration as the source, and syllable count is the best proxy for duration. Homura is trained to control output length at the syllable level. On the SandGlass test, the 9B variant lands within plus or minus 10 percent of the target syllable count 81.92 percent of the time, per the team's own evaluation; in subtitle-alignment workflows, that number decides whether a human retimes every third line or almost none.

Index-NativeLong, also known by the ID Index-Nailong, addresses long-document consistency. The team's demonstration runs a 32,000-token fantasy text through the model in which character names and royal-title puns recur across many passages, and the translations keep those terms consistent from first token to last. The contrast with a naive chunk-by-chunk approach is instructive: a 9B model translating the same text in independent blocks visibly drifts partway through. This is the branch that justifies the 262K context window, because consistency over long spans is easier when the model can actually see the long span.

Constrained translation: when "translate this" is not enough

Beyond the branches, the core models are tuned for what the team calls constrained translation, evaluated under the instTrans benchmark with a self-reported IFscore of 0.8336. The constraints split into two tiers.

Hard constraints must not be violated. Glossary enforcement forces a fixed rendering for listed terms: if your documentation says a source term must always become a specific target phrase, the model obeys even when a freer translation would read more naturally. Structure preservation keeps machine-readable formats intact: JSON, CSV, code and placeholder tokens pass through translation without being mangled. The official project page shows the 9B model translating a JSON game-maintenance announcement into Korean while keeping the structure and hashtags untouched, exactly the job where a general assistant tends to helpfully reformat your payload into prose.

Soft constraints are preferences the model should respect without being absolute. Tone and register can be specified, domain ambiguity can be steered (the classic example is "plant," which should become a factory in an industrial text and a potted one in a gardening blog), cross-sentence consistency can be requested, and LaTeX fragments can be preserved. One project-page example keeps gamer slang intact: "When the FFXIV itch hits, just go play." survives as slang rather than being flattened into formal English.

The practical read: this family treats translation as an engineering interface with requirements, not a free-text conversation. That is what separates a product-scoped model from a general assistant wearing a system prompt.

Benchmarks: strong numbers, self-reported, and not the whole story

Every number in this section is officially self-reported, with no independent replication yet, so treat them as the team's own homework rather than a graded exam.

On FLORES with COMET-22 as the metric, the flagship scores 0.8794. On WMT26 Judge it posts 76.76. The low-resource slice is more interesting: FLORES COMET-22 of 0.8168 on low-resource pairs, and an off-target rate of 2.4 percent, meaning the share of outputs that come back in the wrong language entirely. The 9B variant posts a low-resource instruction-following score of 0.7725 and an off-target rate of 3.47 percent, which the team notes is the lowest among the models compared in their table. Off-target rate is an unglamorous but important production metric: one response in the wrong language is a user-facing bug.

Now the honest part. In the same official table, DeepSeek-V4.1-Flash scores higher on WMT26 Judge, at 83.55 against Index-Translate's 76.76. Translation quality on that axis is not first place, and pretending otherwise would do readers a disservice. We covered that release in our DeepSeek-V4.1-Flash open-source write-up; if your primary workload is raw translation quality on well-resourced language pairs, that table row matters.

The counterweight is scope. Index-Translate packages translation with syllable control, voice-preserving dubbing, long-document consistency and format constraints in one Apache-2.0 bundle. Whether that bundle beats a higher raw score depends on whether your pipeline needs those constraints. If it does, rebuilding them on top of a general model is real engineering work, and that is what you skip by picking a specialized family.

Running it locally: one line and 21.71GB

The team released official GGUF quantizations, so llama.cpp works out of the box with no conversion steps. The recommended quantization is Q4_K_M at 21.71GB. For the quality-conscious, the team validated it by measuring KL divergence against the F16 weights on an A100, a more rigorous sanity check than most community quantizations receive.

The one-line command to serve the recommended quantization:

bash
llama serve -hf IndexTeam/Index-Translate-35B-A3B-preview-GGUF:Q4_K_M

That command pulls the weights from Hugging Face and starts a local server. Note that 21.71GB implies a 24GB-class GPU or a machine with enough system memory, though the 9B and 2B variants offer lighter options if the flagship does not fit.

The official inference recipe is deliberately minimal. The prompt template is a single Chinese-language instruction telling the model to translate the given text into the target language and output the translation directly, without explanation. Two settings matter: disable thinking mode (enable_thinking set to false, since the Qwen3.5 lineage has a thinking mode that only slows translation down), and set temperature to 0 for deterministic output. Translation is one workload where you almost never want creativity.

If you are comparing local deployment stories across open releases, our DeepSeek-V4-Flash-Vision-Exp resource guide covers another open-weight model you can self-host, and the DeepSeek-V4.1-Flash integration SOP walks through API integration patterns that apply to most OpenAI-compatible endpoints, including locally served ones.

From browser pages to dubbed videos

Two pieces of tooling ship alongside the models. The first is a browser extension that runs a local model to translate web pages on the fly; your browsing text never leaves the machine. The second is a video dubbing pipeline that chains the translation models with Index-Echo, closing the loop from a source video to a dubbed output with the speaker's voice preserved.

The team has also published its roadmap: an official stable release of 35B-A3B-preview, an open-sourced evaluation benchmark, more languages for Index-Echo, and larger models. None of these carry dates, so treat them as intentions rather than promises.

The open-source read

The interesting signal in this release is not any single benchmark number. It is that a major content platform decided translation was core infrastructure worth owning, specializing and giving away. Apache-2.0 puts it in the same camp as the most permissive open releases of the year, in contrast to the custom-license compromises some much larger models have shipped with. For fan communities, indie localization studios and anyone building multilingual tools, a specialized translation family with 3B-class inference costs and hard-constraint obedience is a genuinely new option on the shelf.

It also fits a broader pattern: open-weight releases are increasingly scoped rather than general, with focused families for translation, vision and speech, each priced and licensed to be adopted rather than admired. Our million-token output comparison looked at this economics story from the output-capacity angle, and the lesson is the same: raw scale matters less than what the architecture costs you per token. Whether Index-Translate becomes the default substrate for the next generation of fan-sub tooling is a question for the community, but the raw materials are on the table, and the price of admission is one llama.cpp command.

FAQ

Q1: Can I use Index-Translate commercially?

A1: Yes. The models are released under the Apache-2.0 license, confirmed by the license tag on the official Hugging Face model card. Commercial use, redistribution and fine-tuning are permitted under the license terms. As always, read the actual license text and any usage policies attached to the repository before shipping a product.

Q2: What hardware do I need to run the 35B model locally?

A2: The official GGUF quantization at Q4_K_M is 21.71GB, so a 24GB-class GPU or a machine with sufficient system memory works. If that is too heavy, the same family ships 9B and 2B variants for lighter hardware. The team validated the quantization against F16 weights via KL divergence on an A100, so the quality loss at Q4 is documented rather than guessed.

Q3: Is Index-Translate better than DeepSeek-V4.1-Flash at translation?

A3: On the WMT26 Judge metric in the official table, no: DeepSeek-V4.1-Flash scores 83.55 against Index-Translate's 76.76, and both numbers are self-reported. Index-Translate's pitch is the bundle: syllable control, voice-preserving dubbing, 32K-token consistency and hard constraints like glossaries and JSON preservation, which general high-scoring models do not ship as built-in capabilities. Pick based on whether your pipeline needs the constraints.

Q4: Does the model handle audio and video directly?

A4: The text models do not. Speech is Index-Echo's job: a speech-to-text-to-speech pipeline that transcribes the audio, translates the transcript and generates speech in the original speaker's timbre, currently covering Chinese to and from English, Spanish and Japanese. The released dubbing pipeline chains Echo with the translation models for complete video workflows.

Q5: What is the off-target rate, and why does it matter?

A5: It is the percentage of outputs returned in the wrong language entirely. The 35B model reports 2.4 percent on low-resource pairs, and the 9B variant reports 3.47 percent, the lowest among compared models in the official table, though all figures are self-reported. It matters because in production, one answer in the wrong language is a user-facing bug regardless of translation quality.

Search Keywords

  • how to run index-translate locally
  • open source translation model vs commercial api
  • index-translate gguf download

This article was drafted with AI assistance and reviewed by a human editor.

This article is AI-assisted and human-edited. Last updated: 2026-10-07

FAQ

Can I use Index-Translate commercially?
Yes. The models are released under the Apache-2.0 license, confirmed by the license tag on the official Hugging Face model card. Commercial use, redistribution and fine-tuning are permitted under the license terms. As always, read the actual license text and any usage policies attached to the repository before shipping a product.
What hardware do I need to run the 35B model locally?
The official GGUF quantization at Q4_K_M is 21.71GB, so a 24GB-class GPU or a machine with sufficient system memory works. If that is too heavy, the same family ships 9B and 2B variants for lighter hardware. The team validated the quantization against F16 weights via KL divergence on an A100, so the quality loss at Q4 is documented rather than guessed.
Is Index-Translate better than DeepSeek-V4.1-Flash at translation?
On the WMT26 Judge metric in the official table, no: DeepSeek-V4.1-Flash scores 83.55 against Index-Translate's 76.76, and both numbers are self-reported. Index-Translate's pitch is the bundle: syllable control, voice-preserving dubbing, 32K-token consistency and hard constraints like glossaries and JSON preservation, which general high-scoring models do not ship as built-in capabilities. Pick based on whether your pipeline needs the constraints.
Does the model handle audio and video directly?
The text models do not. Speech is Index-Echo's job: a speech-to-text-to-speech pipeline that transcribes the audio, translates the transcript and generates speech in the original speaker's timbre, currently covering Chinese to and from English, Spanish and Japanese. The released dubbing pipeline chains Echo with the translation models for complete video workflows.
What is the off-target rate, and why does it matter?
It is the percentage of outputs returned in the wrong language entirely. The 35B model reports 2.4 percent on low-resource pairs, and the 9B variant reports 3.47 percent, the lowest among compared models in the official table, though all figures are self-reported. It matters because in production, one answer in the wrong language is a user-facing bug regardless of translation quality.

Related

Open Source

NVIDIA Locks Down AI Agents: Open Runtime, Silicon Watchdog

A deep dive into NVIDIA's Open Agent Safety Platform (official materials plus multi-source reporting, October 5, 2026 basis). Core thesis: agent safety cannot rest on the model behaving itself - when the agent is itself the software trying to escape its guardrails, application-layer protections fail, and agent drift (gradual divergence from operator intent via vague prompts, missing tools or policy conflicts) triggers no alarm. Two layers: OpenShell (Apache 2.0 open source, a kernel-level sandboxed secure runtime whose policies cover files, processes, credentials, tools, network and databases; intent-alignment checks before execution plus continuous drift monitoring; runs on Vera CPUs, extensible to Arm and Intel) and NVIDIA Sentry (an out-of-band watchdog on BlueField-4 DPUs that verifies agent identity via DOCA, produces attested telemetry, enforces granular policy, and quarantines an escaping agent within milliseconds; a reference design, not open hardware). Three layers stay separate: application, runtime, infrastructure. Over 100 organizations are on board (Anthropic integrating it with Claude Managed Agents, Salesforce for Slack, SpaceXAI for Cursor agents) under the Linux Foundation's Open Secure AI Alliance. The criticism is reported faithfully: Gartner notes OpenAI, Amazon and Google are absent; IDC estimates it addresses under 25% of enterprise agentic security problems; Control Risks notes it only governs known agents on infrastructure you own (shadow agents, SaaS-embedded ones and attacker-delivered ones are out of reach); lock-in risk discussed. Together with open-weights Kimi K3 it marks the two ends of the agent-safety map.

Oct 5, 202610 min read
Open Source

LTX-2 Open-Sourced: One 22B Model Generates Video and Sound

A deep dive into LTX-2 as open weights (GitHub verified 2026-10-02): Lightricks/LTX-2 at 9,569 stars, Python, last push October 2; per the official README, the first DiT-based audio-video foundation model that generates picture and synchronized sound in one pass. LTX-2.5 composition: a 22B distilled transformer (bf16) plus a custom Gemma 4 12B text encoder (not interchangeable with Google stock), video and audio VAEs, spatial and temporal upscalers - roughly 66 GiB in total; DistilledPipeline runs on just 8 preset sigmas (8-step stage 1 + 4-step stage 2) for the fastest path, FP8 quantization and CPU/disk offload cut memory, and 4K means 3840x2176 (the README explicitly says not 2160). Twelve pipelines include DFR production quality, DubIt re-dubbing with lip-sync preserved, Retake partial regeneration, and SDR-to-HDR (BT.2020/HLG plus ACEScct EXR); ltx-trainer covers LoRA, full fine-tuning and IC-LoRA, with an official ComfyUI plugin. License verification: the GitHub badge reads NOASSERTION because LTX-2 ships a custom LTX Community License (applying to LTX-2.5 since August 11, 2026), not an OSI-approved open-source license - free for personal non-commercial use, but entities with annual revenue of USD 10 million or more must purchase a Commercial Use Agreement for any commercial use; derivatives include distillation, and redistribution must carry the full agreement. The self-hosting case: data stays in-house, marginal cost of batch generation approaches electricity, and LoRA styles remain your own asset.

Oct 3, 202610 min read
Open Source

Open-source Dots: 2,400 stars in three days

A teardown of feder-cr/dots (GitHub API snapshot 2026-10-02): 2,416 stars, 421 forks, MIT-licensed Python, created 2026-09-29 - an open-source isotope that hit escape velocity the day after OpenAI's closed Dots debut. Its core thesis: the model is swappable with one flag, the browser is what the website actually sees. A real Firefox engine patched in C++ decides the fingerprint inside the engine rather than as a page-inspectable JavaScript coat; one seed equals one consistent identity (screen, fonts, GPU, timezone and language agree, reproducibly); nothing for a page to find (no WebDriver flag, no DevTools protocol, no automation globals); the pointer travels before it clicks and keys are pressed one at a time so events arrive trusted; --profile-dir keeps logins across runs; with --proxy the timezone and language follow the exit. Models come from OpenRouter with a one-flag swap, and the companion repo invisible_playwright_mcp exposes the same browser as an MCP server for Claude Code, Codex and Gemini CLI. The README states plainly it is not affiliated with OpenAI. Includes a compliance boundary note: respect target-site terms and local law; no fraud, ticket-scalping or bulk sign-ups.

Oct 2, 20269 min read