If you have ever watched a live stream with auto-generated subtitles, you have probably seen this: half a sentence appears on screen, then gets wholesale replaced two seconds later, punctuation reshuffled and words swapped. That is not a captioner's mistake. It is the classic failure mode of pseudo-streaming speech recognition: slice the audio, transcribe each chunk as a batch, recompute every time new audio arrives, and let earlier text change at will. True streaming has a one-word answer: append-only. Once text is emitted, it is committed permanently and never revised. In late September, NetEase Youdao open-sourced Confucius4-R2T2, which topped the Hugging Face ASR Trending chart (official press release wording, as relayed by Synced), pushing true streaming back into the spotlight. This review puts five open-source options on one table, R2T2, faster-whisper, FunASR, moonshine, and Kyutai delayed-streams-modeling, and compares them along five axes: streaming authenticity, latency, deployment cost, Chinese and English support, and license risk.
Scope: What This Article Can and Cannot Compare
The most common way a comparison article fails is ranking accuracy numbers produced by different evaluation harnesses. This article refuses to do that. Three ground rules:
- No cross-comparison of accuracy numbers. Different projects use different test sets, audio segmentation, and post-processing. Scores from different harnesses are not comparable, and ranking by them is meaningless.
- Any R2T2 WER/CER figure cited is labeled. Specific scores quoted here are all under the "R2T2 official evaluation wording (README)", and the competitor scores in that table were measured under the same wording, not independently re-tested by a third party. For example, under the official 160ms chunk wording, English LibriSpeech-clean is 2.13 WER and test-other is 4.88. We did not run our own benchmarks.
- This article compares streaming recognition and transcription: dictation, live subtitles, and the voice front end of agents, the "turn speech into text" job. Translation and speech-to-speech interpreting are a different track. Our AI translation tools comparison covers cloud translation services, our Qwen3-8-LiveTranslate hotspot covers end-to-end live translation models, and our Hibiki resource page is dedicated to Kyutai's speech-to-speech model. The four articles interlink without overlapping.
All star counts are a 2026-09-29 GitHub API snapshot; licenses are verified against each repository's LICENSE file.
The Five Contenders: One Table for the Big Picture
| Option | Stars (2026-09-29 snapshot) | License | Streaming mode | Latency wording | Chinese / English | Deployment | Commercial risk |
|---|---|---|---|---|---|---|---|
| Confucius4-R2T2 (NetEase Youdao) | 680 | Apache-2.0 (LICENSE text verified) | True streaming, append-only, no revision | Officially reported average 200-600ms, chunk configurable from 80ms to 2s | Chinese and English optimized, multilingual | vLLM / transformers / llama.cpp, built-in WebSocket server | Low |
| faster-whisper (SYSTRAN) | 25,617 | MIT | Not true streaming; pseudo-streaming via slicing | No official streaming latency figure | Leverages Whisper's multilingual base | CTranslate2 inference engine | Low |
| FunASR (Alibaba DAMO) | 20,536 | MIT | Known for Paraformer-series segment transcription | No unified official streaming latency figure found | Deepest Chinese ecosystem | Python toolchain, ModelScope ecosystem | Low |
| moonshine (usefulsensors) | 11,154 | Custom (NOASSERTION) | On-device lightweight route | No unified official streaming latency figure found | No official wording verified | Targets resource-constrained devices | Must review custom terms |
| delayed-streams-modeling (Kyutai) | 3,032 | Apache-2.0 | True streaming (delayed streams modeling) | Designed via delayed streams; no official figure cited here | No official wording verified | Research-oriented integration | Low |
One-line profiles:
- Confucius4-R2T2: Youdao's newly open-sourced true-streaming ASR, built on Qwen3-ASR, with append-only output as the headline feature.
- faster-whisper: SYSTRAN's improved inference engine for Whisper (CTranslate2), the de facto community standard, though Whisper itself is segment-based rather than truly streaming.
- FunASR: Alibaba DAMO Academy's open-source project with the deepest Chinese ecosystem, still receiving updates as of 2026-09-28.
- moonshine: usefulsensors' on-device lightweight route.
- delayed-streams-modeling: Kyutai's true-streaming approach built on delayed streams modeling.
Axis One: True or Pseudo Streaming, Does It Dare to Commit
The dividing line between true and pseudo streaming is a single question: does already-emitted text ever change?
R2T2 makes append-only a training objective, not just an inference-time restraint. Its training uses stable-prefix data, forced time-alignment data, and token-level audio segmentation, combined with the LSP (Longest Stable Prefix) learning paradigm, teaching the model that words, once spoken, stick (a technical report is forthcoming). In practice: once a subtitle lands on screen it stays there, so downstream features like word-by-word caption animation, log archiving, and agent intent parsing never have to deal with the messy state of rewritten history.
faster-whisper is a different story. It optimizes Whisper's inference to a high standard via the CTranslate2 engine and is the default choice for many transcription workloads, but Whisper itself is segment-based and not truly streaming. Engineering teams commonly emulate streaming with rolling slices and overlap, which means recomputing every time a new slice arrives and letting earlier text change. For offline transcription that is irrelevant; for live subtitles it is exactly where the "jumping captions" come from.
Kyutai's delayed-streams-modeling takes a third path: delayed streams modeling, where the model learns during training how long to wait before emitting output, so its output is naturally true streaming. For FunASR and moonshine, this article does not rule on their streaming posture on the officials' behalf: FunASR is best known for its Paraformer-series segment transcription ecosystem, and moonshine targets on-device lightweight use. Whether either meets your no-revision requirement is best verified with your own audio.
Axis Two: Latency, Only Numbers With a Source
In the latency column, this article only writes numbers with a source: R2T2's officially reported average latency is 200-600ms, with decode chunks configurable from 80ms to 2s. It is the only one of the five that publishes both a clear range and a configurable window. The value is at both ends: around 200ms approaches the perceptual floor of text appearing as you speak, while pushing chunks toward 2s trades latency for more stable recognition quality. Products can pick their point on that curve.
For the other four, we found no unified official streaming latency figures, and we will not invent any. Beware of "millisecond-level" marketing copy: ask whether the figure is first-token or stable-output latency, measured on what hardware and what audio pipeline, before wiring anything into production.
Latency sensitivity also differs by scenario. Live subtitles are the most sensitive to stable-output latency because captions must track the speaker. Meeting minutes can tolerate a one-to-two-second window in exchange for lower error rates. An agent voice front end has the tightest budget of all: the entire pipeline from end of user speech to the agent's first response usually has less than a second to spend, and the ASR share must be budgeted in milliseconds. All three are "streaming", but they define "fast" differently. Decide which fast you need first.
Axis Three: Deployment, From vLLM to CPU
R2T2 offers the most complete backend selection: vLLM for high throughput, HuggingFace transformers for tinkering, llama.cpp to cover CPU and edge scenarios, plus a built-in WebSocket server with context and hotword support. For teams building a voice front end, this covers the path from development to production in one repo.
faster-whisper's edge is maturity: the CTranslate2 engine plus a vast body of community know-how means integrating it into an existing system is nearly frictionless. FunASR rides the ModelScope ecosystem with the deepest reserve of Chinese tooling and pretrained models, keeping onboarding costs low for teams in China. moonshine is designed for on-device lightweight use, aiming squarely at resource-constrained hardware. Kyutai's delayed-streams-modeling is research-oriented, suited to teams who want to reproduce the modeling paradigm rather than drop a production component into place.
For the broader methodology of running models locally, see our local LLM deployment comparison and Ollama deployment SOP. The logic transfers: fix your hardware budget first, then the inference backend, then pick the model.
Axis Four: Chinese and English, Trust Tuning Not Claims
R2T2's official wording is Chinese and English optimized with multilingual support, and its hotword mechanism helps with Chinese proper nouns. FunASR has the deepest Chinese ecosystem, with the Paraformer series accumulated over years in the Chinese community. faster-whisper leans on Whisper's multilingual base. For moonshine and Kyutai's delayed-streams-modeling, we found no official wording on language coverage and will not guess.
A practical test: check whether your target language is a first-class citizen in the project's README and evaluation tables. R2T2's README includes a full Chinese CER table and an English WER table (both under its official evaluation wording), which at least shows both languages made the starting lineup. If your main language gets one passing sentence in the docs, discount your expectations accordingly.
Axis Five: Licenses and Commercial Risk, moonshine Flagged Separately
The good news is that four of the five are commercially friendly. R2T2 and Kyutai's delayed-streams-modeling are Apache-2.0 (R2T2's LICENSE file was verified as the standard Apache License 2.0 text), and faster-whisper and FunASR are MIT. Closed-source commercial use, redistribution, and shipping products are routine legal work for all four.
moonshine must be flagged separately: its license is custom, with GitHub's license field marked NOASSERTION. It is not a standard recognized open-source license, and the fine print determines whether you can use it commercially, whether you must open-source your own code, and what attribution you owe. If your product wants it, step one is not writing code; it is running the license text past legal, clause by clause. Until that review is done, treating it as non-commercial is the safe default.
Recommendations: Find Your Row
- Production pipeline needing true streaming, Chinese and English, complete backends: R2T2. Append-only plus a WebSocket server plus vLLM throughput is essentially purpose-built for the voice-front-end job.
- Already on the Whisper ecosystem, just want faster offline transcription: faster-whisper. Do not expect it to become true streaming.
- Chinese-heavy scenarios, teams in China: FunASR, with the smoothest ecosystem and toolchain.
- On-device, offline first: moonshine points the right direction, but clear the custom license first.
- Researching streaming modeling itself: Kyutai's delayed-streams-modeling.
- Still deciding between cloud and local: read our cloud vs local voice AI review and settle the cost and compliance math before choosing.
A final dose of perspective: R2T2's 680 stars against faster-whisper's 25,617 is not a fair ecosystem fight. Choosing R2T2 is a bet that the true-streaming differentiator is worth real product money. If captions that never rewrite themselves are your pain point, the bet is justified; otherwise, maturity wins.
Zooming out: Whisper itself sits at 109,704 stars (MIT, 2026-09-29 snapshot) as the undisputed leader of open-source speech recognition, and faster-whisper is essentially an inference accelerator built around it; vosk-api, at 15,153 stars (Apache-2.0), represents the older offline lightweight route that avoids large-model architectures. Neither made our five-way table, but both underline a fact: streaming large-model ASR is still a young branch of this field, and what R2T2 and Kyutai are exploring may be the roadbed for the next generation of real-time voice products. Few stars today does not mean a weak direction; many stars does not mean true streaming.
FAQ
Q1: How can I tell true streaming from pseudo streaming at a glance?
A: Watch the already-emitted text. True streaming is append-only: committed words stay committed. Pseudo streaming recomputes per slice and rewrites earlier text. Feed thirty seconds of continuous audio and see whether the earlier words change; the answer will be obvious.
Q2: Can R2T2's 2.13 WER be compared directly with a faster-whisper score?
A: No. That figure is under the R2T2 official evaluation wording (README), and the competitor scores in that table were measured under the same wording. Harnesses differ across projects, so accuracy numbers cannot be cross-compared. To compare fairly, you must re-test everything on one test set with one segmentation scheme.
Q3: Can moonshine be used commercially?
A: It carries a custom license (GitHub marks it NOASSERTION), not a standard open-source license, so commercial use depends on the terms themselves. Review the license text clause by clause before any commercial use, and involve legal if needed. Until the review is done, treat it as non-commercial.
Q4: Can these run without a GPU?
A: There are options. R2T2 officially supports a llama.cpp backend, giving CPU and edge scenarios an official path, and moonshine is an on-device lightweight route by design. For faster-whisper and FunASR CPU support, check their documentation for the current version; this article does not rule on their behalf. One more planning note: true streaming loads inference continuously, unlike batch workloads with peaks, so size capacity for steady state.
Q5: First choice for Chinese live subtitles?
A: Choose between R2T2 and FunASR. If you need true streaming with no revision and production-ready backends, pick R2T2 (Chinese and English optimized with hotwords, per its official wording). If you value ecosystem and toolchain depth for Chinese, pick FunASR. The other three each have strengths, but for the specific job of Chinese live subtitles, we found no official wording strong enough to back a first pick.
References
- Confucius4-R2T2 repository and README (latency, evaluation wording, backends): https://github.com/netease-youdao/Confucius4-R2T2
- faster-whisper repository: https://github.com/SYSTRAN/faster-whisper
- FunASR repository: https://github.com/modelscope/FunASR
- moonshine repository (license NOASSERTION): https://github.com/usefulsensors/moonshine
- Kyutai delayed-streams-modeling repository: https://github.com/kyutai-labs/delayed-streams-modeling
- GitHub API (star snapshot 2026-09-29): https://api.github.com