Hardcore Reviews
Hardcore Reviews

Which Open-Source Streaming ASR Never Rewrites Your Words?

Five open-source streaming/local speech-recognition options compared: R2T2 (680 stars, Apache-2.0, true streaming with append-only output, 200-600ms), faster-whisper (25,617 stars, MIT, the community workhorse though Whisper itself fakes streaming via slicing), FunASR (20,536 stars, MIT, deepest Chinese ecosystem), moonshine (11,154 stars, custom license - review terms before commercial use, edge-first lightweight), and Kyutai delayed-streams-modeling (3,032 stars, Apache-2.0, the delayed-streams true-streaming route). Star counts are a 2026-09-29 GitHub API snapshot. The axes: true vs pseudo streaming (whether committed text gets revised), latency scale, deployment floor, Chinese/English support, and commercial risk; accuracy figures cannot be compared across harnesses, and R2T2-table numbers are quoted as official-basis only. The opening states the boundary against the site's simultaneous-translation review and cloud voice-service pieces.

Published September 29, 202610 min read
<!-- streaming-asr-local-comparison-review | review | Which Open-Source Streaming ASR Never Rewrites Your Words? -->

If you have ever watched a live stream with auto-generated subtitles, you have probably seen this: half a sentence appears on screen, then gets wholesale replaced two seconds later, punctuation reshuffled and words swapped. That is not a captioner's mistake. It is the classic failure mode of pseudo-streaming speech recognition: slice the audio, transcribe each chunk as a batch, recompute every time new audio arrives, and let earlier text change at will. True streaming has a one-word answer: append-only. Once text is emitted, it is committed permanently and never revised. In late September, NetEase Youdao open-sourced Confucius4-R2T2, which topped the Hugging Face ASR Trending chart (official press release wording, as relayed by Synced), pushing true streaming back into the spotlight. This review puts five open-source options on one table, R2T2, faster-whisper, FunASR, moonshine, and Kyutai delayed-streams-modeling, and compares them along five axes: streaming authenticity, latency, deployment cost, Chinese and English support, and license risk.

Scope: What This Article Can and Cannot Compare

The most common way a comparison article fails is ranking accuracy numbers produced by different evaluation harnesses. This article refuses to do that. Three ground rules:

  • No cross-comparison of accuracy numbers. Different projects use different test sets, audio segmentation, and post-processing. Scores from different harnesses are not comparable, and ranking by them is meaningless.
  • Any R2T2 WER/CER figure cited is labeled. Specific scores quoted here are all under the "R2T2 official evaluation wording (README)", and the competitor scores in that table were measured under the same wording, not independently re-tested by a third party. For example, under the official 160ms chunk wording, English LibriSpeech-clean is 2.13 WER and test-other is 4.88. We did not run our own benchmarks.
  • This article compares streaming recognition and transcription: dictation, live subtitles, and the voice front end of agents, the "turn speech into text" job. Translation and speech-to-speech interpreting are a different track. Our AI translation tools comparison covers cloud translation services, our Qwen3-8-LiveTranslate hotspot covers end-to-end live translation models, and our Hibiki resource page is dedicated to Kyutai's speech-to-speech model. The four articles interlink without overlapping.

All star counts are a 2026-09-29 GitHub API snapshot; licenses are verified against each repository's LICENSE file.

The Five Contenders: One Table for the Big Picture

OptionStars (2026-09-29 snapshot)LicenseStreaming modeLatency wordingChinese / EnglishDeploymentCommercial risk
Confucius4-R2T2 (NetEase Youdao)680Apache-2.0 (LICENSE text verified)True streaming, append-only, no revisionOfficially reported average 200-600ms, chunk configurable from 80ms to 2sChinese and English optimized, multilingualvLLM / transformers / llama.cpp, built-in WebSocket serverLow
faster-whisper (SYSTRAN)25,617MITNot true streaming; pseudo-streaming via slicingNo official streaming latency figureLeverages Whisper's multilingual baseCTranslate2 inference engineLow
FunASR (Alibaba DAMO)20,536MITKnown for Paraformer-series segment transcriptionNo unified official streaming latency figure foundDeepest Chinese ecosystemPython toolchain, ModelScope ecosystemLow
moonshine (usefulsensors)11,154Custom (NOASSERTION)On-device lightweight routeNo unified official streaming latency figure foundNo official wording verifiedTargets resource-constrained devicesMust review custom terms
delayed-streams-modeling (Kyutai)3,032Apache-2.0True streaming (delayed streams modeling)Designed via delayed streams; no official figure cited hereNo official wording verifiedResearch-oriented integrationLow

One-line profiles:

  • Confucius4-R2T2: Youdao's newly open-sourced true-streaming ASR, built on Qwen3-ASR, with append-only output as the headline feature.
  • faster-whisper: SYSTRAN's improved inference engine for Whisper (CTranslate2), the de facto community standard, though Whisper itself is segment-based rather than truly streaming.
  • FunASR: Alibaba DAMO Academy's open-source project with the deepest Chinese ecosystem, still receiving updates as of 2026-09-28.
  • moonshine: usefulsensors' on-device lightweight route.
  • delayed-streams-modeling: Kyutai's true-streaming approach built on delayed streams modeling.

Axis One: True or Pseudo Streaming, Does It Dare to Commit

The dividing line between true and pseudo streaming is a single question: does already-emitted text ever change?

R2T2 makes append-only a training objective, not just an inference-time restraint. Its training uses stable-prefix data, forced time-alignment data, and token-level audio segmentation, combined with the LSP (Longest Stable Prefix) learning paradigm, teaching the model that words, once spoken, stick (a technical report is forthcoming). In practice: once a subtitle lands on screen it stays there, so downstream features like word-by-word caption animation, log archiving, and agent intent parsing never have to deal with the messy state of rewritten history.

faster-whisper is a different story. It optimizes Whisper's inference to a high standard via the CTranslate2 engine and is the default choice for many transcription workloads, but Whisper itself is segment-based and not truly streaming. Engineering teams commonly emulate streaming with rolling slices and overlap, which means recomputing every time a new slice arrives and letting earlier text change. For offline transcription that is irrelevant; for live subtitles it is exactly where the "jumping captions" come from.

Kyutai's delayed-streams-modeling takes a third path: delayed streams modeling, where the model learns during training how long to wait before emitting output, so its output is naturally true streaming. For FunASR and moonshine, this article does not rule on their streaming posture on the officials' behalf: FunASR is best known for its Paraformer-series segment transcription ecosystem, and moonshine targets on-device lightweight use. Whether either meets your no-revision requirement is best verified with your own audio.

Axis Two: Latency, Only Numbers With a Source

In the latency column, this article only writes numbers with a source: R2T2's officially reported average latency is 200-600ms, with decode chunks configurable from 80ms to 2s. It is the only one of the five that publishes both a clear range and a configurable window. The value is at both ends: around 200ms approaches the perceptual floor of text appearing as you speak, while pushing chunks toward 2s trades latency for more stable recognition quality. Products can pick their point on that curve.

For the other four, we found no unified official streaming latency figures, and we will not invent any. Beware of "millisecond-level" marketing copy: ask whether the figure is first-token or stable-output latency, measured on what hardware and what audio pipeline, before wiring anything into production.

Latency sensitivity also differs by scenario. Live subtitles are the most sensitive to stable-output latency because captions must track the speaker. Meeting minutes can tolerate a one-to-two-second window in exchange for lower error rates. An agent voice front end has the tightest budget of all: the entire pipeline from end of user speech to the agent's first response usually has less than a second to spend, and the ASR share must be budgeted in milliseconds. All three are "streaming", but they define "fast" differently. Decide which fast you need first.

Axis Three: Deployment, From vLLM to CPU

R2T2 offers the most complete backend selection: vLLM for high throughput, HuggingFace transformers for tinkering, llama.cpp to cover CPU and edge scenarios, plus a built-in WebSocket server with context and hotword support. For teams building a voice front end, this covers the path from development to production in one repo.

faster-whisper's edge is maturity: the CTranslate2 engine plus a vast body of community know-how means integrating it into an existing system is nearly frictionless. FunASR rides the ModelScope ecosystem with the deepest reserve of Chinese tooling and pretrained models, keeping onboarding costs low for teams in China. moonshine is designed for on-device lightweight use, aiming squarely at resource-constrained hardware. Kyutai's delayed-streams-modeling is research-oriented, suited to teams who want to reproduce the modeling paradigm rather than drop a production component into place.

For the broader methodology of running models locally, see our local LLM deployment comparison and Ollama deployment SOP. The logic transfers: fix your hardware budget first, then the inference backend, then pick the model.

Axis Four: Chinese and English, Trust Tuning Not Claims

R2T2's official wording is Chinese and English optimized with multilingual support, and its hotword mechanism helps with Chinese proper nouns. FunASR has the deepest Chinese ecosystem, with the Paraformer series accumulated over years in the Chinese community. faster-whisper leans on Whisper's multilingual base. For moonshine and Kyutai's delayed-streams-modeling, we found no official wording on language coverage and will not guess.

A practical test: check whether your target language is a first-class citizen in the project's README and evaluation tables. R2T2's README includes a full Chinese CER table and an English WER table (both under its official evaluation wording), which at least shows both languages made the starting lineup. If your main language gets one passing sentence in the docs, discount your expectations accordingly.

Axis Five: Licenses and Commercial Risk, moonshine Flagged Separately

The good news is that four of the five are commercially friendly. R2T2 and Kyutai's delayed-streams-modeling are Apache-2.0 (R2T2's LICENSE file was verified as the standard Apache License 2.0 text), and faster-whisper and FunASR are MIT. Closed-source commercial use, redistribution, and shipping products are routine legal work for all four.

moonshine must be flagged separately: its license is custom, with GitHub's license field marked NOASSERTION. It is not a standard recognized open-source license, and the fine print determines whether you can use it commercially, whether you must open-source your own code, and what attribution you owe. If your product wants it, step one is not writing code; it is running the license text past legal, clause by clause. Until that review is done, treating it as non-commercial is the safe default.

Recommendations: Find Your Row

  • Production pipeline needing true streaming, Chinese and English, complete backends: R2T2. Append-only plus a WebSocket server plus vLLM throughput is essentially purpose-built for the voice-front-end job.
  • Already on the Whisper ecosystem, just want faster offline transcription: faster-whisper. Do not expect it to become true streaming.
  • Chinese-heavy scenarios, teams in China: FunASR, with the smoothest ecosystem and toolchain.
  • On-device, offline first: moonshine points the right direction, but clear the custom license first.
  • Researching streaming modeling itself: Kyutai's delayed-streams-modeling.
  • Still deciding between cloud and local: read our cloud vs local voice AI review and settle the cost and compliance math before choosing.

A final dose of perspective: R2T2's 680 stars against faster-whisper's 25,617 is not a fair ecosystem fight. Choosing R2T2 is a bet that the true-streaming differentiator is worth real product money. If captions that never rewrite themselves are your pain point, the bet is justified; otherwise, maturity wins.

Zooming out: Whisper itself sits at 109,704 stars (MIT, 2026-09-29 snapshot) as the undisputed leader of open-source speech recognition, and faster-whisper is essentially an inference accelerator built around it; vosk-api, at 15,153 stars (Apache-2.0), represents the older offline lightweight route that avoids large-model architectures. Neither made our five-way table, but both underline a fact: streaming large-model ASR is still a young branch of this field, and what R2T2 and Kyutai are exploring may be the roadbed for the next generation of real-time voice products. Few stars today does not mean a weak direction; many stars does not mean true streaming.

FAQ

Q1: How can I tell true streaming from pseudo streaming at a glance?

A: Watch the already-emitted text. True streaming is append-only: committed words stay committed. Pseudo streaming recomputes per slice and rewrites earlier text. Feed thirty seconds of continuous audio and see whether the earlier words change; the answer will be obvious.

Q2: Can R2T2's 2.13 WER be compared directly with a faster-whisper score?

A: No. That figure is under the R2T2 official evaluation wording (README), and the competitor scores in that table were measured under the same wording. Harnesses differ across projects, so accuracy numbers cannot be cross-compared. To compare fairly, you must re-test everything on one test set with one segmentation scheme.

Q3: Can moonshine be used commercially?

A: It carries a custom license (GitHub marks it NOASSERTION), not a standard open-source license, so commercial use depends on the terms themselves. Review the license text clause by clause before any commercial use, and involve legal if needed. Until the review is done, treat it as non-commercial.

Q4: Can these run without a GPU?

A: There are options. R2T2 officially supports a llama.cpp backend, giving CPU and edge scenarios an official path, and moonshine is an on-device lightweight route by design. For faster-whisper and FunASR CPU support, check their documentation for the current version; this article does not rule on their behalf. One more planning note: true streaming loads inference continuously, unlike batch workloads with peaks, so size capacity for steady state.

Q5: First choice for Chinese live subtitles?

A: Choose between R2T2 and FunASR. If you need true streaming with no revision and production-ready backends, pick R2T2 (Chinese and English optimized with hotwords, per its official wording). If you value ecosystem and toolchain depth for Chinese, pick FunASR. The other three each have strengths, but for the specific job of Chinese live subtitles, we found no official wording strong enough to back a first pick.

References

This article is AI-assisted and human-edited. Last updated: 2026-09-29

FAQ

How can I tell true streaming from pseudo streaming at a glance?
A: Watch the already-emitted text. True streaming is append-only: committed words stay committed. Pseudo streaming recomputes per slice and rewrites earlier text. Feed thirty seconds of continuous audio and see whether the earlier words change; the answer will be obvious.
Can R2T2's 2.13 WER be compared directly with a faster-whisper score?
A: No. That figure is under the R2T2 official evaluation wording (README), and the competitor scores in that table were measured under the same wording. Harnesses differ across projects, so accuracy numbers cannot be cross-compared. To compare fairly, you must re-test everything on one test set with one segmentation scheme.
Can moonshine be used commercially?
A: It carries a custom license (GitHub marks it NOASSERTION), not a standard open-source license, so commercial use depends on the terms themselves. Review the license text clause by clause before any commercial use, and involve legal if needed. Until the review is done, treat it as non-commercial.
Can these run without a GPU?
A: There are options. R2T2 officially supports a llama.cpp backend, giving CPU and edge scenarios an official path, and moonshine is an on-device lightweight route by design. For faster-whisper and FunASR CPU support, check their documentation for the current version; this article does not rule on their behalf. One more planning note: true streaming loads inference continuously, unlike batch workloads with peaks, so size capacity for steady state.
First choice for Chinese live subtitles?
A: Choose between R2T2 and FunASR. If you need true streaming with no revision and production-ready backends, pick R2T2 (Chinese and English optimized with hotwords, per its official wording). If you value ecosystem and toolchain depth for Chinese, pick FunASR. The other three each have strengths, but for the specific job of Chinese live subtitles, we found no official wording strong enough to back a first pick.

Related

Hardcore Reviews

Cloud vs Local Voice AI: A Cost and Control Showdown

This review ignores capability and runs the cost-and-control numbers on two routes for voice AI: the cloud real-time speech API versus local open-source tooling (explicitly scoped apart from our 8-26 image-model capability review, batch-22 image cost ledger, and batch-23 agent long-context cost ledger). It opens by arguing voice cost is harder to model than text or images: real-time behavior, concurrency, duration distribution, language/dialect coverage, and privacy compliance stack five dimensions at once. It then compares the two routes dimension by dimension - unit price and billing, latency and real-time, privacy/compliance, controllability/customization, language coverage - pitting cloud representative GPT-Live-1 (closed-source, metered, real-time out of the box) against local representative VoiceStudio (open-source, one-time compute, data stays local, engines swappable), with a five-dimension scorecard and a five-scenario selection table, and concludes by scale: individual, small team, bulk. Every unit price is symbolic (P_cloud / C_local) or marked "per official pricing page"; magnitude judgments are engineering estimates. Cold take: a vendor's "pay-as-you-go saves" only covers the one workload inside its chosen sweet spot; concurrency N, duration distribution T, and mandatory real-time decide the actual bill, so measure yourself.

Sep 13, 20269 min read
Hardcore Reviews

Picking an Open-Source Agentic RL Framework: Five Contenders

Five open-source Agentic RL post-training frameworks compared on positioning, ecosystem and engineering shape rather than benchmarks (self-reported harness figures cannot be compared): verl (23,650-star upstream HybridFlow with a 465-star Xiaomi fork carrying five production environment suites), TRL (19,398 stars, the official Hugging Face on-ramp), OpenRLHF (10,045 stars, Ray-based PPO/DAPO/REINFORCE++), AReaL (5,796 stars, Ant-lineage asynchronous agent RL), and NeMo-RL (2,033 stars, the NVIDIA enterprise ticket). Four scenario verdicts: pick TRL for LoRA-scale trials; put verl and OpenRLHF into your PoC for large-scale verifiable-reward RL; look at AReaL for long-horizon asynchronous agent training; pick NeMo-RL inside NVIDIA all-in stacks. All Apache-2.0 (LICENSE files are the final authority); star counts are a 2026-09-27 snapshot and self-reported performance figures are not trusted.

Sep 27, 202610 min read
Hardcore Reviews

Flagship Price Math: $2 Sol, $0.10 Luna, $4 Opus

Four flagship vendors compared on money only, not intelligence: GPT-6 Sol ($2/$10, 90% cache-read discount), GPT-6 Luna ($0.10/$0.50), Claude Opus 5.5 ($4/$20, exactly 2x Sol, cache read $0.20 tying Sol's), Grok 4.7 ($2/$6, no published cache discount), and DeepSeek V4.1 Flash's time-of-use pricing. At an 80% cache hit rate, the real bill: Opus is 2x on paper, about 1.9x in practice, with nearly the whole gap on output; Grok posts the lowest output rate but its cache savings cannot be booked. Benchmarks next to prices: on DeepSWE, Grok's 71.0% buys the most points per dollar, yet it scores only 38.0% on Terminal-Bench 4.0 — one model, two stories. Three red lines: currencies cannot be compared directly, unpublished cells stay unpublished, and different harnesses cannot settle conclusions.

Sep 27, 202610 min read