Hardcore Reviews
Hardcore Reviews

Realtime video models compared: who edits while you talk

This review skips image quality and compares form and ownership only: it lines up Vidu S2, Kling, Jimeng, Seedance, Sora 2 and HiDream-O1-Video across real-time interaction and editing ability, delivery form (web, API, open weights, commercial product), open versus closed ownership and fit-for-purpose scenarios. The core claim is that what you buy in a real-time video tool is not fidelity but workflow, whether you can edit while talking, how expensive revisions are, and who owns the artifact. It closes with a selection table by scenario, e-commerce try-on, virtual-host livestreaming, ad shorts and personal tinkering, and warns that real-time quality and cost lack a unified third-party benchmark, so do not let the real-time label set the pace.

Published September 17, 202610 min read
<!-- ai-video-model-realtime-comparison-review | review | Realtime video models compared: who edits while you talk -->

Why we split out the realtime and interactive dimension

For the past two years the main arena of video generation has been offline text to video. You drop in a prompt, the model runs for tens of seconds to a few minutes, and it returns a clip of a few seconds to a few tens of seconds. In that paradigm generation is one shot and one way. You cannot interrupt the clip while it is being made, and you cannot say swap the background to snow or change the lead actor's outfit. If a creator wants a change, they regenerate or pull the result into an editor and fix it by hand.

In 2026 a second route started to appear: realtime and interactive video generation. Its core is not sharper pictures but a changed workflow. The user and the model share one timeline, so you can edit while you talk and adjust while you watch. Shengshu's Vidu S2 pushed this route to the front: its S2-Avatar swaps reference images dynamically inside a conversation and keeps continuous motion and state, while S2-Editing turns style transfer, virtual try on, character replacement, and background replacement into realtime edits.

So this review deliberately refuses to compare who renders more realistic frames. Absolute quality is everyone's own story until there is a unified third party benchmark, and offline and realtime generation optimize for different goals anyway: realtime wants low latency and interruptibility, offline wants peak quality. Forcing both onto one quality leaderboard only produces misleading conclusions. We compare shape and ownership: who can interact in realtime, who can edit in realtime, what form it ships in, and whether it is open or closed.

One point often missed is that realtime is not a single shape. Some realtime is conversation driven virtual humans, some is realtime erase and redraw during generation, and some is generation plus editing in one loop. Different shapes fit different businesses. Below we put shape first, not vendor fame.

Back on the business side, realtime ability decides the reason to pay. Offline generation sells a finished clip, billed per clip or per second. Realtime generation sells time and interaction, billed per concurrent session or per online hour. These two billing models map to two customer types: one wants a usable ad, the other wants a virtual employee that can keep talking. Ranking the reviewed models on this axis beats ranking them on quality. Because the billing logic changes, small teams evaluating should not only watch single shot quality but also calculate the long online cost.

Overview comparison

ModelVendorDelivery formRealtime interactionRealtime editingOwnership
Vidu S2ShengshuWeb plus commercial product (S2-Avatar and S2-Editing)Yes, 720P live conversationYes, four edit typesClosed commercial
KlingKuaishouWeb plus API plus creator platformUnverifiedUnverifiedClosed commercial
JimengByteDance and JianyingWeb plus Jianying integrationUnverifiedUnverifiedClosed commercial
SeedanceByteDance and VolcanoAPI plus Doubao ecosystemUnverifiedUnverifiedClosed commercial
Sora 2OpenAIWeb plus APIUnverifiedUnverifiedClosed commercial
HiDream-O1-Video-1.0HiDreamWeb plus open weights historyUnverifiedUnverifiedPartially open

Note: except for Vidu S2, whose shape and realtime ability have explicit official information (released 2026-09-16), the specific realtime interaction and editing abilities of the other models are described qualitatively from public information. Where we could not confirm, we mark it unverified and we do not invent parameters. Resolution caps, clip length, and price are pending official release or pending measurement unless noted.

Dimension by dimension

DimensionVidu S2KlingJimengSeedanceSora 2HiDream
Realtime interaction720P live talk, state keptUnverifiedUnverifiedUnverifiedUnverifiedUnverified
Realtime editingStyle, try on, swap person, swap bgUnverifiedUnverifiedUnverifiedUnverifiedUnverified
ControllabilityTalk driven plus dynamic ref imageHigh expressiveness (known)Ecosystem integrationFast iterationPhysics consistency (known)Multimodal input
Delivery formWeb commercialWeb and APIWeb and JianyingAPIWeb and APIWeb and weights
EcosystemStandalone productKuaishouJianying and DouyinDoubao and VolcanoChatGPTRelatively standalone
Open weightsNoNoNoNoNoOpen weights history

Vendor by vendor

Vidu S2 (Shengshu)

Shengshu pulled realtime from a demo reel to a usable product. The selling point of S2-Avatar is not mere realtime but dynamic reference image swapping inside a conversation while continuous motion and state are kept, which means a virtual human can change looks mid talk instead of restarting each time. The four realtime edit types of S2-Editing cover the most painful scenes for ecommerce and live streaming. The caveat: 720P realtime is the current official line, and higher resolution and longer clips are pending measurement. There is no public code repository, so self hosting is off the table. Its risk lies in the stability and marginal cost of realtime editing, which only become clear after large scale commercial use. For ecommerce and live clients, its biggest selling point now is not flashy tech but compressing a try on and reshoot workflow that once needed a team into one chat window.

Kling (Kuaishou)

Kling sits in the top tier of domestic video generation, and versions 2.0 and 2.5 already perform well on expressiveness. Its strength is the Kuaishou ecosystem and creator platform, with a clear commercialization path. Whether realtime interaction and editing are already product capabilities is not yet explicit in public information, so we mark it unverified. Its forte looks more like high quality offline generation plus a strong ecosystem than Vidu's conversation style realtime editing. For short video teams, Kling's value is in output quality and distribution, not a realtime workflow. If it later adds realtime editing, its distribution scale at Kuaishou would immediately threaten Vidu's first mover spot, so it is worth watching.

Jimeng (ByteDance and Jianying)

Jimeng's difference is Jianying integration: creators close the loop from generation to editing inside one workflow. For short video creators this is stickier than a standalone generation model. Realtime ability is also unverified, but its ecosystem advantage (Douyin distribution plus Jianying editing) makes it fit people who create and publish. Its weakness is that realtime interaction is not yet publicly verified, and deep binding to the ByteDance ecosystem raises cross platform migration cost. For people already working in Jianying, this binding is convenience, and for newcomers it is a low barrier to start.

Seedance (ByteDance and Volcano)

Seedance is the video model under Doubao, with dense iteration in 2026 and an API plus Doubao ecosystem route. Its positioning leans toward being called as a capability rather than a standalone creator product. Realtime interaction is unverified, but ByteDance's engineering and distribution mean that once it works, scale up is fast. Its watch point is not how realtime it is today but whether it can later ship realtime as a low cost base service to mass applications. Once ByteDance pushes Seedance's realtime into Doubao and Jianying, the fight shifts from model quality to traffic entry.

Sora 2 (OpenAI)

Sora 2 is known for high fidelity and physics consistency, a candidate for the quality benchmark among closed routes. The question is whether its realtime interaction and editing are open, and in what form, and public information is limited, so we mark it unverified. Its value likely sits in peak quality rather than edit while you talk. Domestic users also face access and compliance cost, which makes it not necessarily the first pick for local realtime workflows. Its halo on brand and peak quality easily makes people miss that its realtime interaction is still a blank to fill.

HiDream-O1-Video-1.0 (HiDream)

HiDream is the only participant with an open weights history. It is a native multimodal video model, supports multimodal input, and generates five to twenty second 1080p in one click. Media reports cite it at fourth globally on the Artificial Analysis image to video leaderboard and eighth on Arena.ai (source is media citation, not our measurement). Its meaning: if weights open later, researchers and small teams get a self hostable realtime and interactive base. Specific realtime behavior is pending measurement, and open weights do not equal open realtime editing, so the two must be read separately. For universities and startups, its biggest value is a base you can run locally, modify, and retrain, not a finished product.

Selection advice

ScenarioPickReason
Ecommerce virtual try onVidu S2S2-Editing includes try on with a live swap loop
Virtual human live streamVidu S2S2-Avatar live talk with state kept
Fast ad shortsKling or JimengHigh quality generation plus distribution
Research visualizationHiDream (pending test)Open weights history, self hostable to verify
Personal hobbyJimeng or KlingWeb ready, low barrier

A sober sign off

Realtime video generation is a real need, but realtime is still more a marketing word than a unified standard. Resolution cap, clip length, and per unit cost are mostly pending official release or pending measurement except for Vidu S2's 720P realtime official line. The absence of a unified third party benchmark means any who is stronger conclusion should be discounted first.

Our call: from the second half of 2026 to 2027, what truly separates winners is not who shouts realtime first, but who turns realtime editing into a stable commercial workflow and pushes unit price low enough for small teams. Buying early for the word realtime, we suggest verifying in a small scope before scaling. Treat realtime as a tool, not a slogan, and you will not be burned when pricing lands. In the end, realtime is a means and business is the end. Whoever first lets customers earn money with realtime truly wins this round.

FAQ

Q1: How much more does realtime generation cost than offline generation?

A1: Qualitatively, realtime demands low latency and interruptible inference, so per unit compute cost is usually higher than offline batch processing. But the exact gap cannot be stated because vendor pricing is not published. Wait for official pricing and measurement before comparing.

Q2: Which one is open source and can be self hosted?

A2: Among those reviewed, Vidu S2, Kling, Jimeng, Seedance, and Sora 2 are all closed commercial products with no public code repository, so self hosting is basically impossible. HiDream has an open weights history and is the most likely self hostable base for researchers and small teams today, but whether it supports realtime interaction is pending measurement.

Q3: How should an individual creator choose?

A3: If you only want fast clips and easy publishing, pick Jimeng or Kling, web ready with smooth ecosystem. For virtual human or try on realtime play, prefer Vidu S2. If budget is tight and you want local tinkering, watch HiDream's weight release progress.

Q4: Will realtime editing replace editors?

A4: Not in the short term. Realtime editing is good at preview and fast iteration while you talk, but pacing, narrative, and taste still come from people. It is more like freeing editors from repetitive labor than replacing creative decisions.

Q5: Is it worth paying for realtime now?

A5: It depends on the scenario. Strong realtime needs like ecommerce try on and virtual human live stream can be tested in a small scope now. For general creation, wait for pricing and third party measurement before deciding, and do not let the realtime marketing word set your pace.

This article is AI-assisted and human-edited. Last updated: 2026-09-17

FAQ

How much more does realtime generation cost than offline generation?
Qualitatively, realtime demands low latency and interruptible inference, so per unit compute cost is usually higher than offline batch processing. But the exact gap cannot be stated because vendor pricing is not published. Wait for official pricing and measurement before comparing.
Which one is open source and can be self hosted?
Among those reviewed, Vidu S2, Kling, Jimeng, Seedance, and Sora 2 are all closed commercial products with no public code repository, so self hosting is basically impossible. HiDream has an open weights history and is the most likely self hostable base for researchers and small teams today, but whether it supports realtime interaction is pending measurement.
How should an individual creator choose?
If you only want fast clips and easy publishing, pick Jimeng or Kling, web ready with smooth ecosystem. For virtual human or try on realtime play, prefer Vidu S2. If budget is tight and you want local tinkering, watch HiDream's weight release progress.
Will realtime editing replace editors?
Not in the short term. Realtime editing is good at preview and fast iteration while you talk, but pacing, narrative, and taste still come from people. It is more like freeing editors from repetitive labor than replacing creative decisions.
Is it worth paying for realtime now?
It depends on the scenario. Strong realtime needs like ecommerce try on and virtual human live stream can be tested in a small scope now. For general creation, wait for pricing and third party measurement before deciding, and do not let the realtime marketing word set your pace.

Related

Hardcore Reviews

One compromised agent loses everything: a comparison of four credential and permission governance approaches

Credentials went from a config item to an attack surface, yet most teams' defenses are still stuck at "put the agent in a sandbox." This review splits cleanly from our sandbox-isolation comparison: the sandbox governs where code runs; credential governance governs how secrets are used, who approves actions, and whether they can leave. It contrasts four approaches — OpenClaw 2.0, OpenWorker, OpenHuman and traditional secret storage — across six lifecycle stages (store / use / approve / exfiltrate / audit / multi-agent): OpenClaw with masked requests plus an opt-in proxy allowlist; OpenWorker with hard floors, an autonomy ladder, a reviewer model and a circuit breaker, and never self-approving unattended; OpenHuman with Privacy Mode enforced in the Rust core and E2E-encrypted inter-agent comms. Secondhand data (SaaS Sentinel transcription, no primary source located) shows compromise probability 0.24 with one agent rising to 0.86 with seven — risk grows superlinearly with count, under the premise "any agent proposes, execute."

Sep 1, 202611 min read
Hardcore Reviews

Comparing 11 Models by Real Token Cost After the August 31 Repricing: Peak Hours, Cache Hits, and Tokenizer Effects

A model's list price wears at least three more layers. Time of day: DeepSeek moved to peak and off-peak pricing on August 17, charging peak rates on weekdays from 09:00-12:00 and 14:00-18:00, halving them off-peak, and applying off-peak rates all weekend, so the same model costs twice as much at 3pm as at 10pm. Caching: prefix cache hits are billed far below standard input, and the variable sits with your prompt structure rather than the vendor. Tokenization: Sonnet 5 changed tokenizers, so the same input now maps to 1.0x to 1.35x more tokens, and the multiplier floats with content type. This comparison fixes one unit throughout, blended rate equals input plus output divided by two, assuming equal token volumes, as a neutral starting point, then recalculates under three realistic load profiles across 11 models, covering list price, cached input, peak and off-peak, and post-tokenizer position. The finding is not which model is cheapest, it is that no model is cheapest, only cheapest for your particular load: any comparison that ignores input-output ratio, cache hit rate, and content type is comparing list prices, not costs. Chinese model prices come from a page-by-page check of official pricing pages on 2026-08-24, re-confirmed on 08-28; overseas prices from a 2026-08-31 roundup. Conflicts are flagged per line. No live benchmarking was performed.

Aug 31, 20269 min read