Why we split out the realtime and interactive dimension
For the past two years the main arena of video generation has been offline text to video. You drop in a prompt, the model runs for tens of seconds to a few minutes, and it returns a clip of a few seconds to a few tens of seconds. In that paradigm generation is one shot and one way. You cannot interrupt the clip while it is being made, and you cannot say swap the background to snow or change the lead actor's outfit. If a creator wants a change, they regenerate or pull the result into an editor and fix it by hand.
In 2026 a second route started to appear: realtime and interactive video generation. Its core is not sharper pictures but a changed workflow. The user and the model share one timeline, so you can edit while you talk and adjust while you watch. Shengshu's Vidu S2 pushed this route to the front: its S2-Avatar swaps reference images dynamically inside a conversation and keeps continuous motion and state, while S2-Editing turns style transfer, virtual try on, character replacement, and background replacement into realtime edits.
So this review deliberately refuses to compare who renders more realistic frames. Absolute quality is everyone's own story until there is a unified third party benchmark, and offline and realtime generation optimize for different goals anyway: realtime wants low latency and interruptibility, offline wants peak quality. Forcing both onto one quality leaderboard only produces misleading conclusions. We compare shape and ownership: who can interact in realtime, who can edit in realtime, what form it ships in, and whether it is open or closed.
One point often missed is that realtime is not a single shape. Some realtime is conversation driven virtual humans, some is realtime erase and redraw during generation, and some is generation plus editing in one loop. Different shapes fit different businesses. Below we put shape first, not vendor fame.
Back on the business side, realtime ability decides the reason to pay. Offline generation sells a finished clip, billed per clip or per second. Realtime generation sells time and interaction, billed per concurrent session or per online hour. These two billing models map to two customer types: one wants a usable ad, the other wants a virtual employee that can keep talking. Ranking the reviewed models on this axis beats ranking them on quality. Because the billing logic changes, small teams evaluating should not only watch single shot quality but also calculate the long online cost.
Overview comparison
| Model | Vendor | Delivery form | Realtime interaction | Realtime editing | Ownership |
|---|---|---|---|---|---|
| Vidu S2 | Shengshu | Web plus commercial product (S2-Avatar and S2-Editing) | Yes, 720P live conversation | Yes, four edit types | Closed commercial |
| Kling | Kuaishou | Web plus API plus creator platform | Unverified | Unverified | Closed commercial |
| Jimeng | ByteDance and Jianying | Web plus Jianying integration | Unverified | Unverified | Closed commercial |
| Seedance | ByteDance and Volcano | API plus Doubao ecosystem | Unverified | Unverified | Closed commercial |
| Sora 2 | OpenAI | Web plus API | Unverified | Unverified | Closed commercial |
| HiDream-O1-Video-1.0 | HiDream | Web plus open weights history | Unverified | Unverified | Partially open |
Note: except for Vidu S2, whose shape and realtime ability have explicit official information (released 2026-09-16), the specific realtime interaction and editing abilities of the other models are described qualitatively from public information. Where we could not confirm, we mark it unverified and we do not invent parameters. Resolution caps, clip length, and price are pending official release or pending measurement unless noted.
Dimension by dimension
| Dimension | Vidu S2 | Kling | Jimeng | Seedance | Sora 2 | HiDream |
|---|---|---|---|---|---|---|
| Realtime interaction | 720P live talk, state kept | Unverified | Unverified | Unverified | Unverified | Unverified |
| Realtime editing | Style, try on, swap person, swap bg | Unverified | Unverified | Unverified | Unverified | Unverified |
| Controllability | Talk driven plus dynamic ref image | High expressiveness (known) | Ecosystem integration | Fast iteration | Physics consistency (known) | Multimodal input |
| Delivery form | Web commercial | Web and API | Web and Jianying | API | Web and API | Web and weights |
| Ecosystem | Standalone product | Kuaishou | Jianying and Douyin | Doubao and Volcano | ChatGPT | Relatively standalone |
| Open weights | No | No | No | No | No | Open weights history |
Vendor by vendor
Vidu S2 (Shengshu)
Shengshu pulled realtime from a demo reel to a usable product. The selling point of S2-Avatar is not mere realtime but dynamic reference image swapping inside a conversation while continuous motion and state are kept, which means a virtual human can change looks mid talk instead of restarting each time. The four realtime edit types of S2-Editing cover the most painful scenes for ecommerce and live streaming. The caveat: 720P realtime is the current official line, and higher resolution and longer clips are pending measurement. There is no public code repository, so self hosting is off the table. Its risk lies in the stability and marginal cost of realtime editing, which only become clear after large scale commercial use. For ecommerce and live clients, its biggest selling point now is not flashy tech but compressing a try on and reshoot workflow that once needed a team into one chat window.
Kling (Kuaishou)
Kling sits in the top tier of domestic video generation, and versions 2.0 and 2.5 already perform well on expressiveness. Its strength is the Kuaishou ecosystem and creator platform, with a clear commercialization path. Whether realtime interaction and editing are already product capabilities is not yet explicit in public information, so we mark it unverified. Its forte looks more like high quality offline generation plus a strong ecosystem than Vidu's conversation style realtime editing. For short video teams, Kling's value is in output quality and distribution, not a realtime workflow. If it later adds realtime editing, its distribution scale at Kuaishou would immediately threaten Vidu's first mover spot, so it is worth watching.
Jimeng (ByteDance and Jianying)
Jimeng's difference is Jianying integration: creators close the loop from generation to editing inside one workflow. For short video creators this is stickier than a standalone generation model. Realtime ability is also unverified, but its ecosystem advantage (Douyin distribution plus Jianying editing) makes it fit people who create and publish. Its weakness is that realtime interaction is not yet publicly verified, and deep binding to the ByteDance ecosystem raises cross platform migration cost. For people already working in Jianying, this binding is convenience, and for newcomers it is a low barrier to start.
Seedance (ByteDance and Volcano)
Seedance is the video model under Doubao, with dense iteration in 2026 and an API plus Doubao ecosystem route. Its positioning leans toward being called as a capability rather than a standalone creator product. Realtime interaction is unverified, but ByteDance's engineering and distribution mean that once it works, scale up is fast. Its watch point is not how realtime it is today but whether it can later ship realtime as a low cost base service to mass applications. Once ByteDance pushes Seedance's realtime into Doubao and Jianying, the fight shifts from model quality to traffic entry.
Sora 2 (OpenAI)
Sora 2 is known for high fidelity and physics consistency, a candidate for the quality benchmark among closed routes. The question is whether its realtime interaction and editing are open, and in what form, and public information is limited, so we mark it unverified. Its value likely sits in peak quality rather than edit while you talk. Domestic users also face access and compliance cost, which makes it not necessarily the first pick for local realtime workflows. Its halo on brand and peak quality easily makes people miss that its realtime interaction is still a blank to fill.
HiDream-O1-Video-1.0 (HiDream)
HiDream is the only participant with an open weights history. It is a native multimodal video model, supports multimodal input, and generates five to twenty second 1080p in one click. Media reports cite it at fourth globally on the Artificial Analysis image to video leaderboard and eighth on Arena.ai (source is media citation, not our measurement). Its meaning: if weights open later, researchers and small teams get a self hostable realtime and interactive base. Specific realtime behavior is pending measurement, and open weights do not equal open realtime editing, so the two must be read separately. For universities and startups, its biggest value is a base you can run locally, modify, and retrain, not a finished product.
Selection advice
| Scenario | Pick | Reason |
|---|---|---|
| Ecommerce virtual try on | Vidu S2 | S2-Editing includes try on with a live swap loop |
| Virtual human live stream | Vidu S2 | S2-Avatar live talk with state kept |
| Fast ad shorts | Kling or Jimeng | High quality generation plus distribution |
| Research visualization | HiDream (pending test) | Open weights history, self hostable to verify |
| Personal hobby | Jimeng or Kling | Web ready, low barrier |
A sober sign off
Realtime video generation is a real need, but realtime is still more a marketing word than a unified standard. Resolution cap, clip length, and per unit cost are mostly pending official release or pending measurement except for Vidu S2's 720P realtime official line. The absence of a unified third party benchmark means any who is stronger conclusion should be discounted first.
Our call: from the second half of 2026 to 2027, what truly separates winners is not who shouts realtime first, but who turns realtime editing into a stable commercial workflow and pushes unit price low enough for small teams. Buying early for the word realtime, we suggest verifying in a small scope before scaling. Treat realtime as a tool, not a slogan, and you will not be burned when pricing lands. In the end, realtime is a means and business is the end. Whoever first lets customers earn money with realtime truly wins this round.
FAQ
Q1: How much more does realtime generation cost than offline generation?
A1: Qualitatively, realtime demands low latency and interruptible inference, so per unit compute cost is usually higher than offline batch processing. But the exact gap cannot be stated because vendor pricing is not published. Wait for official pricing and measurement before comparing.
Q2: Which one is open source and can be self hosted?
A2: Among those reviewed, Vidu S2, Kling, Jimeng, Seedance, and Sora 2 are all closed commercial products with no public code repository, so self hosting is basically impossible. HiDream has an open weights history and is the most likely self hostable base for researchers and small teams today, but whether it supports realtime interaction is pending measurement.
Q3: How should an individual creator choose?
A3: If you only want fast clips and easy publishing, pick Jimeng or Kling, web ready with smooth ecosystem. For virtual human or try on realtime play, prefer Vidu S2. If budget is tight and you want local tinkering, watch HiDream's weight release progress.
Q4: Will realtime editing replace editors?
A4: Not in the short term. Realtime editing is good at preview and fast iteration while you talk, but pacing, narrative, and taste still come from people. It is more like freeing editors from repetitive labor than replacing creative decisions.
Q5: Is it worth paying for realtime now?
A5: It depends on the scenario. Strong realtime needs like ecommerce try on and virtual human live stream can be tested in a small scope now. For general creation, wait for pricing and third party measurement before deciding, and do not let the realtime marketing word set your pace.