On September 30, 2026, Bilibili's Index LLM team released something surprising for a video company: a family of open-weight models built specifically for translation. Published under Apache-2.0, the family covers 2B, 9B and 35B-A3B-preview text models, trained to work across what the team counts as 150 languages. The flagship is a mixture-of-experts model with 35 billion total parameters that activates only about 3 billion per token, and it holds up to 262,144 tokens in context.
For anyone who has watched Chinese fan-sub groups do their unpaid, frame-accurate magic on Japanese anime, or seen bilingual danmaku comments light up under a clip from Seoul, the origin story makes instant sense. Bilibili sits in one of the most multilingual corners of the Chinese internet, and it has now turned that community need into weights you can download, quantize and run on your own hardware.
This article walks through what was released, why a video platform cares about translation, how the three specialized branches differ, what the benchmarks honestly say, and how to run the flagship locally with a single llama.cpp command.
Why a video platform builds translation models
The easy assumption is a cheaper internal API bill. The reality is more interesting. Chinese otaku culture has been shaped by fan translation for two decades: volunteer groups time, translate and typeset subtitles for anime, variety shows and game trailers, often within hours of release. Quality can be superb, but the economics are fragile and long-tail coverage spotty. Meanwhile the platform's content ecosystem is inherently cross-border: Japanese gameplay footage with Chinese commentary, Korean dance covers with English reactions, and a comment layer that mixes Chinese, English, Japanese and Korean in a single thread.
Machine translation inside this ecosystem is not a general chat problem. A subtitle line must land within a syllable budget if it will be spoken aloud. A game announcement keeps its JSON structure and hashtags intact, or the downstream renderer breaks. A fantasy novel translated paragraph by paragraph quietly forgets how a title was rendered ten chapters ago. These are the failure modes general-purpose assistants hit, and precisely the problems the Index team scoped its models around.
There is also a strategic point: a translation stack Bilibili owns rather than rents compounds over time, and a permissive license buys goodwill with the exact developer community that builds fan tools and subtitle pipelines.
The family: three sizes on a Qwen3.5 base
The lineup consists of 2B, 9B and 35B-A3B-preview text models, built on Alibaba's Qwen3.5 as the foundation architecture. The stated coverage of 150 languages is an official self-reported figure, not something an independent benchmark has audited. The weights live on Hugging Face under the IndexTeam organization and on ModelScope, with code, a technical report and an online demo on GitHub at bilibili/Index-Translate. The license tag on the model card reads Apache-2.0, so commercial use, redistribution and fine-tuning are all on the table.
The flagship's name, 35B-A3B-preview, describes its economics: a sparse mixture-of-experts model with 35 billion parameters in total, of which roughly 3 billion activate for any given token. Inference cost tracks active parameters, so you get the capacity of a mid-30s-billion model while paying compute closer to a 3B dense model. For a platform translating millions of subtitle lines a day, that gap is a real budget line, not a rounding error.
The other headline number is context: the 35B model accepts 262,144 tokens, far beyond what a subtitle file needs, but exactly right for long documents where earlier context must inform the next paragraph.
Three branches for three jobs that generic models fail
The most distinctive part of this release is not the base models but the three specialized branches built on top of them.
Index-Echo handles speech. It is a speech-to-text-to-speech pipeline that transcribes audio, translates the transcript and speaks the result back, preserving the original speaker's voice timbre. Supported directions are Chinese to and from English, Spanish and Japanese. In practice this is the dubbing branch: take a talky video, keep the original voice's character, and let viewers hear it in another language.
Index-Homura tackles a constraint most people never consider until they try dubbing: syllable counts. A translated line spoken over original footage must fit roughly the same duration as the source, and syllable count is the best proxy for duration. Homura is trained to control output length at the syllable level. On the SandGlass test, the 9B variant lands within plus or minus 10 percent of the target syllable count 81.92 percent of the time, per the team's own evaluation; in subtitle-alignment workflows, that number decides whether a human retimes every third line or almost none.
Index-NativeLong, also known by the ID Index-Nailong, addresses long-document consistency. The team's demonstration runs a 32,000-token fantasy text through the model in which character names and royal-title puns recur across many passages, and the translations keep those terms consistent from first token to last. The contrast with a naive chunk-by-chunk approach is instructive: a 9B model translating the same text in independent blocks visibly drifts partway through. This is the branch that justifies the 262K context window, because consistency over long spans is easier when the model can actually see the long span.
Constrained translation: when "translate this" is not enough
Beyond the branches, the core models are tuned for what the team calls constrained translation, evaluated under the instTrans benchmark with a self-reported IFscore of 0.8336. The constraints split into two tiers.
Hard constraints must not be violated. Glossary enforcement forces a fixed rendering for listed terms: if your documentation says a source term must always become a specific target phrase, the model obeys even when a freer translation would read more naturally. Structure preservation keeps machine-readable formats intact: JSON, CSV, code and placeholder tokens pass through translation without being mangled. The official project page shows the 9B model translating a JSON game-maintenance announcement into Korean while keeping the structure and hashtags untouched, exactly the job where a general assistant tends to helpfully reformat your payload into prose.
Soft constraints are preferences the model should respect without being absolute. Tone and register can be specified, domain ambiguity can be steered (the classic example is "plant," which should become a factory in an industrial text and a potted one in a gardening blog), cross-sentence consistency can be requested, and LaTeX fragments can be preserved. One project-page example keeps gamer slang intact: "When the FFXIV itch hits, just go play." survives as slang rather than being flattened into formal English.
The practical read: this family treats translation as an engineering interface with requirements, not a free-text conversation. That is what separates a product-scoped model from a general assistant wearing a system prompt.
Benchmarks: strong numbers, self-reported, and not the whole story
Every number in this section is officially self-reported, with no independent replication yet, so treat them as the team's own homework rather than a graded exam.
On FLORES with COMET-22 as the metric, the flagship scores 0.8794. On WMT26 Judge it posts 76.76. The low-resource slice is more interesting: FLORES COMET-22 of 0.8168 on low-resource pairs, and an off-target rate of 2.4 percent, meaning the share of outputs that come back in the wrong language entirely. The 9B variant posts a low-resource instruction-following score of 0.7725 and an off-target rate of 3.47 percent, which the team notes is the lowest among the models compared in their table. Off-target rate is an unglamorous but important production metric: one response in the wrong language is a user-facing bug.
Now the honest part. In the same official table, DeepSeek-V4.1-Flash scores higher on WMT26 Judge, at 83.55 against Index-Translate's 76.76. Translation quality on that axis is not first place, and pretending otherwise would do readers a disservice. We covered that release in our DeepSeek-V4.1-Flash open-source write-up; if your primary workload is raw translation quality on well-resourced language pairs, that table row matters.
The counterweight is scope. Index-Translate packages translation with syllable control, voice-preserving dubbing, long-document consistency and format constraints in one Apache-2.0 bundle. Whether that bundle beats a higher raw score depends on whether your pipeline needs those constraints. If it does, rebuilding them on top of a general model is real engineering work, and that is what you skip by picking a specialized family.
Running it locally: one line and 21.71GB
The team released official GGUF quantizations, so llama.cpp works out of the box with no conversion steps. The recommended quantization is Q4_K_M at 21.71GB. For the quality-conscious, the team validated it by measuring KL divergence against the F16 weights on an A100, a more rigorous sanity check than most community quantizations receive.
The one-line command to serve the recommended quantization:
llama serve -hf IndexTeam/Index-Translate-35B-A3B-preview-GGUF:Q4_K_MThat command pulls the weights from Hugging Face and starts a local server. Note that 21.71GB implies a 24GB-class GPU or a machine with enough system memory, though the 9B and 2B variants offer lighter options if the flagship does not fit.
The official inference recipe is deliberately minimal. The prompt template is a single Chinese-language instruction telling the model to translate the given text into the target language and output the translation directly, without explanation. Two settings matter: disable thinking mode (enable_thinking set to false, since the Qwen3.5 lineage has a thinking mode that only slows translation down), and set temperature to 0 for deterministic output. Translation is one workload where you almost never want creativity.
If you are comparing local deployment stories across open releases, our DeepSeek-V4-Flash-Vision-Exp resource guide covers another open-weight model you can self-host, and the DeepSeek-V4.1-Flash integration SOP walks through API integration patterns that apply to most OpenAI-compatible endpoints, including locally served ones.
From browser pages to dubbed videos
Two pieces of tooling ship alongside the models. The first is a browser extension that runs a local model to translate web pages on the fly; your browsing text never leaves the machine. The second is a video dubbing pipeline that chains the translation models with Index-Echo, closing the loop from a source video to a dubbed output with the speaker's voice preserved.
The team has also published its roadmap: an official stable release of 35B-A3B-preview, an open-sourced evaluation benchmark, more languages for Index-Echo, and larger models. None of these carry dates, so treat them as intentions rather than promises.
The open-source read
The interesting signal in this release is not any single benchmark number. It is that a major content platform decided translation was core infrastructure worth owning, specializing and giving away. Apache-2.0 puts it in the same camp as the most permissive open releases of the year, in contrast to the custom-license compromises some much larger models have shipped with. For fan communities, indie localization studios and anyone building multilingual tools, a specialized translation family with 3B-class inference costs and hard-constraint obedience is a genuinely new option on the shelf.
It also fits a broader pattern: open-weight releases are increasingly scoped rather than general, with focused families for translation, vision and speech, each priced and licensed to be adopted rather than admired. Our million-token output comparison looked at this economics story from the output-capacity angle, and the lesson is the same: raw scale matters less than what the architecture costs you per token. Whether Index-Translate becomes the default substrate for the next generation of fan-sub tooling is a question for the community, but the raw materials are on the table, and the price of admission is one llama.cpp command.
FAQ
Q1: Can I use Index-Translate commercially?
A1: Yes. The models are released under the Apache-2.0 license, confirmed by the license tag on the official Hugging Face model card. Commercial use, redistribution and fine-tuning are permitted under the license terms. As always, read the actual license text and any usage policies attached to the repository before shipping a product.
Q2: What hardware do I need to run the 35B model locally?
A2: The official GGUF quantization at Q4_K_M is 21.71GB, so a 24GB-class GPU or a machine with sufficient system memory works. If that is too heavy, the same family ships 9B and 2B variants for lighter hardware. The team validated the quantization against F16 weights via KL divergence on an A100, so the quality loss at Q4 is documented rather than guessed.
Q3: Is Index-Translate better than DeepSeek-V4.1-Flash at translation?
A3: On the WMT26 Judge metric in the official table, no: DeepSeek-V4.1-Flash scores 83.55 against Index-Translate's 76.76, and both numbers are self-reported. Index-Translate's pitch is the bundle: syllable control, voice-preserving dubbing, 32K-token consistency and hard constraints like glossaries and JSON preservation, which general high-scoring models do not ship as built-in capabilities. Pick based on whether your pipeline needs the constraints.
Q4: Does the model handle audio and video directly?
A4: The text models do not. Speech is Index-Echo's job: a speech-to-text-to-speech pipeline that transcribes the audio, translates the transcript and generates speech in the original speaker's timbre, currently covering Chinese to and from English, Spanish and Japanese. The released dubbing pipeline chains Echo with the translation models for complete video workflows.
Q5: What is the off-target rate, and why does it matter?
A5: It is the percentage of outputs returned in the wrong language entirely. The 35B model reports 2.4 percent on low-resource pairs, and the 9B variant reports 3.47 percent, the lowest among compared models in the official table, though all figures are self-reported. It matters because in production, one answer in the wrong language is a user-facing bug regardless of translation quality.
Search Keywords
- how to run index-translate locally
- open source translation model vs commercial api
- index-translate gguf download
This article was drafted with AI assistance and reviewed by a human editor.