One Stack, Three Repositories, Listen Translate and Speak
At the end of September, NetEase Youdao pushed its Confucius Live series of real-time interaction models onto the open-source shelf. What makes this release unusual is completeness: three repositories map onto the three stages of a real-time voice pipeline, listening, translating, and speaking, and the LICENSE file of each repository is the verbatim full text of the standard Apache License 2.0. For commercial deployment, that places all three at the permissive end of the licensing spectrum.
The division of labor can be stated in one line:
Listen Confucius4-R2T2 true-streaming ASR
Translate Confucius4-T3PO 14B streaming text-to-text interpreter
Speak Confucius4-TTS zero-shot speech synthesisAccording to the official press release as relayed by machine intelligence media, R2T2 and T3PO simultaneously topped the Hugging Face Trending charts for ASR and for Translation. As of a GitHub API snapshot taken on 2026-09-29, the three repositories hold 680, 110, and 821 stars respectively. The loudest of the trio is not the ASR that listens but the TTS that speaks.
Anyone who has shipped a real-time voice product knows the three obstacles. The first is revision: many systems marketed as streaming are chunked pseudo-streaming, and later audio can overturn a sentence already displayed, so subtitles and downstream logic keep jittering. The second is the latency paradox of translation: wait for the period and the delay is unusable, do not wait and the interpreter risks the wrong reading. The third is voice customization: a fixed voice traditionally means studio-grade recording or handing data to a cloud service. The Confucius Live release puts a solution to all three on the table at once, which is what a one-stop real-time voice stack means in practice.
This is an open-source teardown. I first unpack the stable-prefix mechanism and append-only output of R2T2, where its real value for live subtitles and agent pipelines lives, then T3PO's Pareto-aware reinforcement learning that ties quality and latency into one optimization problem, and finally the engineering adaptation across the trio plus the line between this article and the site's other translation coverage.
R2T2: True Streaming Is a Discipline, Not a Latency Number
The repository is netease-youdao/Confucius4-R2T2, written in Python, with 680 stars and a last push dated 2026-09-24 in the same snapshot. The name expands to Real Real-Time Transcription. The doubled Real is not a typo; it insists on genuine streaming. Average latency runs 200 to 600 milliseconds, and the decoding chunk is configurable from 80 milliseconds up to 2 seconds.
The line between true and fake streaming is drawn not by the latency figure but by output discipline. R2T2's output is append-only: once text has been emitted it is permanently committed and never revised. That sentence sounds modest, yet its engineering weight is heavy. In live captioning, append-only means a subtitle freezes the moment it appears; viewers never watch already-read text get rewritten wholesale. In agent voice pipelines, recognition results must be fed immediately to a downstream LLM. If the text keeps changing, the downstream either rebuilds its context over and over or maintains a bespoke diff mechanism for revised transcripts. Append-only drives both classes of cost to zero.
Backing this discipline is the training-side stable-prefix machinery. R2T2 is built on Qwen3-ASR. Training uses stable-prefix data construction, forced time-alignment data, and token-level audio segmentation, combined with the LSP learning paradigm, short for Longest Stable Prefix; a technical report is promised. The intuition is straightforward: during training the model is repeatedly required to settle early on the text corresponding to audio it has already heard, rather than leaving everything to a later global optimization pass. Only a model trained that way can afford to promise append-only output at inference time.
The accuracy claims deserve a careful label. The README officially reports state-of-the-art latency and recognition quality among open-source models, competitive with closed systems. The evaluation charts include an English WER table and a Chinese CER table, benchmarked against Qwen3-ASR, X-ASR, WhisperRT, Nemotron, Voxtral, AssemblyAI, and commercial systems. Under the 160-millisecond chunk setting, the English table lists 2.13 WER on LS-clean and 4.88 on LS-other. The Chinese table mirrors the English one in structure; I cite no single Chinese number here. One warning is mandatory: the competitor scores inside those tables come from R2T2's own official evaluation setup. Test harnesses differ across projects, so comparing raw accuracy numbers across repositories is unreliable.
The engineering coverage is generous. Backends include vLLM for high-throughput serving, HuggingFace transformers for quick prototyping, and llama.cpp for lightweight local machines, plus a bundled WebSocket Server so a browser frontend can connect to the streaming service directly. Context and hotword prompts are supported, which lets vertical industries prop up proper nouns with a hotword list. Optimization focuses on Chinese and English, with multilingual support on top. One intranet box running vLLM, a laptop browser over WebSocket, and the guest list preloaded as hotwords: names stop coming out as wrong homophones in the captions.
T3PO: Pareto Reinforcement Learning Ties Quality to Latency
The second repository is netease-youdao/Confucius4-T3PO, with 110 stars and a last push dated 2026-09-21. The name expands to simulTaneous Translation via pareTo Policy Optimization. It is a 14-billion-parameter text-to-text simultaneous interpretation model. Note the qualifier: it is purely text in, text out, with no speech input. To interpret speech into text, the officially recommended combination is cascading it behind R2T2, which converts the audio stream into a text stream before T3PO takes over.
Training proceeds in three stages: segment-aligned data construction, streaming cold start, and Pareto-aware reinforcement learning. The first two stages teach the model to interpret while listening; the third tackles the trade-off between translating well and translating fast. The inherent tension of simultaneous interpretation is that reading a few more words before committing raises quality but inflates latency, while writing early cuts latency and raises the risk of heading down the wrong path. Traditional pipelines fix a latency target first and optimize quality within it. T3PO instead runs quality and latency as a joint Pareto optimization, letting the model produce multiple operating points along the trade-off curve, then exposing them as selectable latency tiers so each business picks the point that fits its scenario.
The inference protocol also rewards a close look. T3PO consumes the input stream as character- and word-level fine-grained chunks and dynamically decides whether to keep reading or to write an incremental piece of translation. Committed translation output is likewise append-only and never rewritten. An interleaved history protocol supports KV-cache reuse, which amortizes inference cost over long sessions. The model also retains the general instruction-following ability of its Qwen base, so instructions such as terminology constraints can be layered on top of translation.
The honest boundaries belong in print. For the untrained Japanese-to-Chinese direction, the README self-reports a degree of streaming generalization, while the team concedes that quality outside the Chinese-English pair has not been rigorously evaluated. In other words, Chinese-English interpretation carries both training and evaluation backing; other language pairs run at your own risk. On evaluation, the official charts compare the open-source InfiniSST and EAST against two commercial simultaneous interpretation systems, plotted along COMET and word-CW dimensions. I cite no specific scores here; readers who need them should consult the original charts in the repository. An online demo lives at t3po.youdao.com.
TTS: Why the Loudest Repository Is the One That Speaks
The third repository is netease-youdao/Confucius4-TTS, the most popular of the trio at 821 stars in the snapshot, with a last push dated 2026-09-03 and an accompanying paper at arXiv 2608.11650. It is an LLM-based TTS built from a speech encoder plus an LLM architecture, covering 14 languages including Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, and Vietnamese.
Its three headline capabilities aim straight at the third obstacle above. Zero-shot cloning works without a reference transcript, meaning it can replicate a target voice with no text transcript of that speaker. Cross-lingual timbre preservation keeps a cloned voice from drifting when Chinese text drives English output. Emotion transfer lets synthesized speech carry the sentiment of the reference audio. For developers building a fixed persona voice or multilingual broadcast, the combination is not hard to imagine being useful. No objective benchmark figures for the TTS are covered by the fact base of this article, so audition the synthesis yourself.
One logistics note: all three repositories publish weights through both Hugging Face and ModelScope, so users in China are not blocked at the first step.
Ecosystem Fit and How This Article Divides Labor with the Site
Assembled together, the community adaptation across the three repositories already covers the practical ground. On recognition, R2T2 offers three backend routes, vLLM, transformers, and llama.cpp, corresponding to high-concurrency serving, fast prototyping, and low-spec machines, plus its own WebSocket Server. On translation, T3PO ships an online demo and open weights. On synthesis, TTS rides the standard inference stack. The R2T2 and T3PO cascade is the officially recommended combination; add TTS and the synthesis stage stays local too. The entire real-time voice chain then runs on Apache-2.0 weights with no data leaving the intranet, a practical option for customer service, meetings, and live captioning.
The division of labor with existing coverage is worth drawing precisely. This article focuses on the open-source teardown of the Youdao trio. The Kyutai interpretation model Hibiki is a separate deep dive into a speech-to-speech route that supports only French to English. The Qwen live-translate hotspot tracks the Qwen camp's real-time translation news. The translation tools comparison reviews cloud translation services, not self-hosted weights. The speech-to-speech resource roundup indexes the S2S landscape. If you cannot decide between deployment forms, see the local versus cloud voice AI review. These pieces interlink and each covers one segment without repeating the others.
Cold Reflection: Snapshots, Labels, and Boundaries
First, stars are only a heat reference and carry an explicit timestamp: every star count in this article is a GitHub API snapshot from 2026-09-29. The gradient of 821 stars for TTS, 680 for ASR, and 110 for simultaneous interpretation reflects how community attention distributes: synthesis has the broadest audience while interpretation is the most vertical niche, a pattern Hibiki and comparable projects have faced.
Second, performance conclusions need their labels. The dual-SOTA claim for R2T2 is a README self-report with a self-built evaluation harness, and the competitor scores sit under that same setup; the same caveat applies to T3PO's comparison charts. Such tables are fine for judging rough magnitudes and out of bounds for precise rankings.
Third, the language boundaries should be acknowledged. R2T2 is optimized for Chinese and English with multilingual support. T3PO has rigorous evaluation only in the Chinese-English direction, and its Japanese-to-Chinese generalization is self-reported. TTS spans 14 languages, but variation in cloning quality across languages falls outside the fact base of this piece. Testing your own language pair before committing beats reading any table.
Fourth, the licensing is the cleanest part of this release. The LICENSE file of each repository is the verbatim full text of the standard Apache License 2.0. It is permissive, non-copyleft, and carries an explicit patent grant, so the compliance burden of commercial deployment is low. That is a substantive difference from many research-only open-source voice projects.
The value of the Confucius Live trio lies not in topping any single leaderboard but in open-sourcing the listen, translate, and speak segments of a real-time voice pipeline under one uniform license, with the key mechanisms, append-only output, the LSP stable prefix, and Pareto joint optimization, all given explicit technical accounts. For developers building their own real-time voice pipeline, this is a rare and complete starting point in 2026.
FAQ
Q1: Must the three repositories be used together, or can each stand alone?
A1: Each works entirely on its own. R2T2 is a standalone streaming recognition service with a bundled WebSocket Server; T3PO is a pure text-to-text model that can do streaming translation by itself; TTS is likewise an independent repository. Only speech-input interpretation requires cascading R2T2 with T3PO into an S2T pipeline.
Q2: Why does append-only output matter so much?
A2: Because it directly determines the architectural complexity of everything downstream. Live captions freeze the moment they appear instead of being rewritten wholesale, and agent pipelines can feed recognition results straight to an LLM without building revision and context-rebuilding machinery. Revision-style streaming saves effort on the model side and passes the cost to every consumer downstream.
Q3: Does T3PO accept speech input?
A3: No. T3PO is a pure text-to-text model whose input and output are both text streams. For speech-to-text interpretation, the official recommendation is to cascade it with R2T2: R2T2 turns the audio stream into a text stream, and T3PO takes over with streaming translation, forming an S2T pipeline.
Q4: What should commercial users know about the licenses?
A4: The LICENSE file of each repository is the verbatim full text of the standard Apache License 2.0, a mainstream permissive license with an explicit patent grant and a low barrier to commercial deployment; keeping the license and copyright notices suffices under convention. By contrast, some open-source voice projects use custom licenses whose terms must be reviewed before any commercial use, which makes this trio notably cleaner.
Q5: What are the sources for the star counts and performance numbers?
A5: All star counts are a GitHub API snapshot from 2026-09-29. R2T2's latency and WER figures come from the official README, and the competitor scores inside its evaluation tables share that same table's setup; T3PO's comparison charts are described by dimension only, without specific scores; the Chinese CER table is summarized structurally. Every number is labeled with its source, and results from different harnesses are never mixed.