Frontline Hotspot
Frontline Hotspot

Wan3.0 Officially Launches: 30-Second Single Clips, PPT-to-Video, and API Pricing Starting at 0.3 RMB per Second

On August 24, Alibaba Cloud officially launched the Wan3.0 video generation model (moving beyond the early-August teaser): single-clip length doubles from Wan 2.7's 15 seconds to 30 seconds, and for the first time it accepts direct uploads of five office formats - doc, xls, ppt, pdf, and md. Reference-to-video takes files or public webpage links. API pricing is per second: 480P at 0.3 RMB, 720P at 0.6 RMB, 1080P at 1.2 RMB, with a 30% launch discount on Bailian and the Qwen AI platform from August 24 to September 23 - a 15-second 720P clip costs just 6.3 RMB. Eight entry points are open, with Meitu and JD Lingjing among early integrators; note it lacks Function Calling, web search, batch inference, and context caching - it is a generation engine, not an agent.

Published August 25, 20266 min read
<!-- alibaba-wan3-0-official-launch-hotspot | hotspot | Wan3.0 Officially Launches: 30-Second Single Clips, PPT-to-Video, and API Pricing Starting at 0.3 RMB per Second -->

When Alibaba announced Wan3.0 as its "new generation" on August 8, we wrote here that there is often a gap between announcement and availability. On August 24, that gap closed - Alibaba Cloud officially launched the Wan3.0 video generation model: pricing published, entry points open, API callable, and office documents can be thrown straight in as input. From generation-change teaser to shipped product, the keyword this time isn't "stronger" - it's "usable."

Scope note: this article is based on reporting by Guancha.cn and Sina Finance (Sanyan Tech) from 2026-08-24, plus Alibaba Cloud's developer community docs and the Bailian official documentation. Prices are snapshots; the official pages prevail.

1. The Biggest Upgrade: PPT Straight to Video

The most notable change in this launch isn't a model parameter - it's the input. Wan3.0 is the first in the family to accept direct uploads of five office formats: doc, xls, ppt, pdf, and md. Previously you had to copy-paste document contents into plain text, losing all layout, charts, and structure; now you drop in a PPT and the model reads it itself.

The surrounding capabilities are in place too: text-to-video, image-to-video (first frame / first-last frame), and reference-to-video are all live. Reference-to-video accepts either a file (PDF/PPT/DOCX) or a link (public webpage) as the source - the two can't be combined. Multi-shot storyboards can be described in natural language with timestamps, and Alibaba highlights production-grade character consistency and realistic audio. With this input pipeline, "one product deck -> one 30-second promo video" runs end to end without reformatting anything.

2. 30 Seconds: From "Shot Material" to "Complete Narrative"

Single-generation length doubles from 15 seconds on the previous Wan 2.7 to 30 seconds (any integer duration from 2-30 seconds). This is more than a doubled number: in short-video grammar, 15 seconds is "one shot plus half a transition," while 30 seconds covers a complete narrative - setup, development, and payoff in one clip, with no stitching of multiple 15-second fragments needed to keep characters and lighting consistent. Stitching is precisely where AI video burns the most human labor; doubling the single-clip length saves the post-production sewing cost.

3. Priced by the Second: A 720P Clip for 6.3 RMB at Launch

API pricing on Alibaba Cloud Bailian is cleanly tiered, billed per second:

ResolutionUnit price30s list price30% launch discount
480P0.3 RMB/s9 RMB6.3 RMB
720P0.6 RMB/s18 RMB12.6 RMB
1080P1.2 RMB/s36 RMB25.2 RMB

Limited offer: from August 24 to September 23, API calls on Bailian and the Qwen AI platform are 30% off. Run the numbers: a 15-second 720P clip costs 6.3 RMB during the launch window - already below what most freelance editors charge for subtitles alone. Multi-region deployment (Beijing, Singapore at a slight premium, Tokyo, Frankfurt) means overseas teams can call from nearby. There are eight ways to try it: Alibaba Cloud Bailian, the Wanxiang site, Wanjing Yike, the Qwen AI platform, Qwen Creation on PC, the Qwen app, Duiyou, and IF STUDIO. Early integrators include Tiaoyue Shijie, Deevid, Juhuo, JD Lingjing, Meitu, and Jingmeng.

4. The Boundary: It's a Generation Engine, Not an Agent

Per Alibaba Cloud's developer documentation, Wan3.0 does not support Function Calling, web search, batch inference, or context caching, and calls must be submitted asynchronously. That defines its position: a per-second-billed video generation engine, not an agent that decomposes tasks on its own. If you want it inside an automated pipeline, you still have to build the workflow orchestration in the layer above - which is exactly the problem our companion piece, the Wan3.0 hands-on SOP, tackles.

One-line closer: when AI video's input side plugs into office documents and its pricing drops to 0.3 RMB per second, it stops being a novelty toy and becomes a quote sheet for a productivity tool.


References

This article is based on public reporting and official documentation (as of 2026-08-24). Prices and promotions are snapshots; the official pages prevail. Not investment advice.

This article is AI-assisted and human-edited. Last updated: 2026-08-25

Related

Frontline Hotspot

Triple Strike: Alibaba Wan3.0, ChatGPT Free Text Chat, and Seedance 2.5 API All Land on the Same Day

Three announcements on 2026-08-08: Alibaba's next-gen video model Wan3.0, OpenAI opening free-text chat with ChatGPT for free users, and Volcengine bringing Seedance 2.5 API online. Spanning video generation, conversation, and capability access, together they push AI capability on two fronts at once: stronger and more accessible. A trend integration based on public announcements, not a hands-on benchmark.

Aug 8, 20264 min read
Frontline Hotspot

ChatGPT Images 2.5: Half the Latency, Real Consistency

OpenAI launched ChatGPT Images 2.5 on 2026-09-09: up to 50% lower latency than 2.0, better preservation of reference-photo subjects and multi-turn edit consistency; ChatGPT adds sketch mode, templates, image comments and prompt sharing; the API ships two models, Flare and Sunburst. This piece breaks down each upgrade, argues the real leap is latency plus consistency rather than raw image quality, reads the two-model split as capability tiering and pricing segmentation (analysis, not official wording), and weighs the long-term lock-in cost of closed APIs.

Sep 9, 20269 min read
Frontline Hotspot

Nvidia's $13B Hugging Face Deal: What It Means for Open Source

Reported 2026-09-04 (Cailianspress and others): NVIDIA announced the acquisition of Hugging Face for about \$13B — \$11.9B to investors and \$1B for employee equity retention — one of the largest deals in NVIDIA's history. Jensen Huang committed to keeping HF an open platform without forcing NVIDIA compute. This piece breaks down the deal structure, why a compute hegemon would buy the open-source ecosystem's front door, how much developers should trust the promise ("not forced" is not the same as "not default"), and the hosting-platform implications.

Sep 8, 20269 min read