Before You Touch Anything: Align on What Is Actually Available
This SOP answers one concrete question: how do you use Kling AI, and once the 4.0 generation of capabilities lands in your account, how do you work through them step by step? Before following any step, align on the timeline first - every claim in this article is current as of October 3, 2026, the date of publication.
According to Kling's official announcements and media reports: on September 28, 2026, Kling officially announced that Kling 4.0 would formally launch in October; on the same day, the lightweight Kling 4.0 Flash opened for small-scale trial to black-gold annual-card members. In other words, as of this writing: the full 4.0 has not launched yet; 4.0 Flash is only open to black-gold annual-card members in a limited rollout, so most accounts will not see any 4.0-series entry point yet; pricing, credit costs, and public beta feedback have not been published, and this article does not invent any of them. For the full spec sheet and background, see our Kling 4.0 announcement breakdown.
That shapes how this article is written: Step 1 and Step 2 are the general onboarding path you can verify right now; the 8,000-token prompt method in Step 3, the 15-item multimodal reference workflow in Step 4, and the extend-to-2-minutes continuation workflow in Step 5 are the usage framework for after 4.0 or 4.0 Flash becomes available to your account. The spec numbers come from official announcements; the actual entry points depend on what your account shows, and the official pages are always the source of truth.
Step 1: Pick Your Entry Point and Prepare Your Account
There are two entry points: the web at klingai.com and the Kling app. Accounts are shared between them and creation history syncs across, so use whichever you prefer. The steps:
- Open klingai.com or the Kling app and register or log in. Expected result: you land on the creation home page with the AI video entry points visible.
- Check your membership status. If you want to try 4.0 Flash, confirm on the membership page whether your plan is the black-gold annual card and whether the 4.0 Flash trial eligibility has been provisioned to your account. Expected result: the membership page clearly shows your current plan and benefits. If it shows no 4.0 trial eligibility, your account is outside the limited rollout for now - wait for the wider release after the full 4.0 launches in October.
- Browse the creation page and confirm which model versions are currently visible to you. Expected result: you know exactly which version you can use right now, rather than guessing.
One real constraint to note: Kling has not committed to a unified rollout schedule for feature entry points, and different accounts may see different features. If you cannot find a button at any step, wait for the next wave of the limited rollout or for the October full launch - do not hunt for unofficial workarounds in third-party tutorials.
Step 2: Get Your First Text-to-Video Through the Loop
Whichever version you have, the goal of the first video is not a great film - it is closing the minimal loop of "input, generate, review, download."
- On the creation page, select video generation and pick the newest model version available to you. Expected result: the input box is ready.
- Write a one-sentence prompt, for example: "An early-morning convenience store, warm light, a woman in a dark coat pushes the door open, camera tracking slowly sideways past the shelves." Expected result: the system accepts the input and shows a generate button.
- Keep the default duration and resolution, hit generate, and wait for the task to finish. Expected result: a playable video. A first generation commonly takes minutes; do not resubmit while waiting.
- Play it back, then download and save it. Expected result: the file is on your local disk, and the original prompt is stored in your own asset library.
The real value of this step is establishing your baseline: the visual quality, wait time, lip sync, and on-screen text behavior of your current version with the same prompt become the reference for judging whether the 4.0 upgrade is worth it. For the full methodology of general video generation, see our AI video generation SOP.
Step 3: Using the 8,000-Token Prompt Ceiling Correctly
One line in Kling 4.0's official spec is easy to underestimate: the prompt input limit has been raised to 8,000 tokens. This does not mean you should fill 8,000 tokens every time - it means you can feed a "screenplay-grade" structured description in one pass: multiple shots, multiple scenes, dialogue, sound design, and style constraints can all live in a single prompt instead of being split across generations and stitched later.
Per the official positioning, 4.0 upgrades narrative completeness, multilingual lip sync, text rendering, and multi-keyframe control - all of which depend on writing the description properly. A reusable skeleton for long prompts:
- Opening block: one sentence setting the tone - genre, style, framing mood. Expected result: the model locks onto the overall register first.
- Storyboard block: describe shot by shot in chronological order, each shot covering camera position and movement, subject and appearance, on-frame content, and emotion. Expected result: the model organizes the visuals along your shot order instead of improvising.
- Audio block: music, rhythm, dialogue, and language. Expected result: lip sync and language follow your settings; Chinese, English, Japanese, Korean, Spanish, Portuguese, German, French, Hindi and more, with dialects and accents, are in the officially announced supported range.
- Constraint block: elements that must appear and must be avoided. Expected result: fewer unwanted surprises during the gacha rounds.
For the concrete shot-description template, see Step 3 of our Wan3.0 document-to-video SOP: the five-element formula of camera plus movement plus subject plus frame plus emotion transfers across models. (Readers of the English edition: that article is in Chinese; the formula itself is language-independent.)
One more cross-model methodology worth importing, from LTX-2's official prompting guide (note: this is advice from the LTX-2 official documentation, not Kling's official docs - Kling's own prompt spec is defined by its official pages): treat the prompt as a cinematographer writing a shot list - start with the main action in the very first sentence; describe chronologically in one flowing paragraph; make actions, gestures, appearance, camera angles, environment, and lighting specific; around 200 words works best. The principles of "action first, chronological order, concrete camera language" apply equally to Kling's storyboard block. More on that open-source model in our LTX-2 open-source breakdown.
Step 4: How to Feed the 15 Multimodal References
This is the 4.0-generation capability most worth rehearsing in advance. Official spec: a single generation task supports up to 15 multimodal references, with per-category caps of 10 images, 5 videos, and 7 subjects, all usable combined in one task. In practice you can feed a character sheet, scene reference images, camera-move example videos, and the lead subject together, and have the model align all of them within one task.
State the boundary first: this capability belongs to the post-launch framework for 4.0 and 4.0 Flash - run it once the entry point appears in your account, and let the creation page define the exact interaction. When it does, practice in this order:
- Single-subject alignment. Feed just 1 subject image and keep one character consistent across shots. Expected result: you learn how faithfully the model reproduces your asset and build a feel for how much it honors what you feed it.
- Images plus subject. Feed 3 to 5 images plus 1 subject - images carry scene and style, the subject carries character consistency. Expected result: same prompt, new scene, same face.
- Add a video reference. Feed 1 example video for camera movement so the motion style aligns. Expected result: the generated camera work resembles your reference clip.
- Push toward the cap. Combine up to 7 subjects with a full set of images and videos for multi-character scenes or brand multi-asset composites. Expected result: this is the hardest mode. Define every asset's role in the prompt with small combinations first, then scale up - otherwise multiple subjects contaminate each other.
Key lesson: the more references you add, the more explicitly the prompt must state each item's role - which image is the lead, which is the scene, which video contributes only its camera move and not its content. Piling up assets without assigning roles is the most common way this feature goes wrong.
Step 5: The Continuation Workflow, from 30 Seconds to 2 Minutes
Official spec: a single generation runs up to 30 seconds, and multiple continuation rounds can extend a video to 2 minutes. This workflow also belongs to the post-launch framework, but it is worth designing for now because it changes your storyboard structure:
- Slice the content into 30-second blocks. Opening and setup in the first block, development in the second, climax and resolution in the third. Expected result: every block closes its own narrative loop while hooking the next one.
- After the first block generates, inspect the final frame. Expected result: the character pose and scene lighting at the end are a suitable starting point for the next block; if the character is mid-motion, the next block's opening description picks up from that pose.
- Continue with the continuation feature, stating the handoff point in each block's prompt. Expected result: subject, style, and lighting stay continuous across blocks.
- When all blocks are done, review duration and pacing as a whole. Expected result: one film up to 2 minutes with a complete narrative - not four unrelated fragments.
Worth knowing alongside this: the official creation-page upgrade points include a unified input box, grid and list layouts, timeline-based continuous creation, and a new canvas Agent experience. Timeline-based continuous creation pairs directly with the continuation workflow - once visible, use the timeline to manage multi-block projects.
FAQ
Q1: Can I use Kling 4.0 on klingai.com right now? A1: Probably not. As of October 3, 2026, the full 4.0 awaits its October launch; 4.0 Flash is open only to black-gold annual-card members in a limited rollout. Use whichever version your account shows; if no 4.0-series entry point appears, wait for the wider release. The official pages are the source of truth.
Q2: What is the relationship between 4.0 Flash and the full 4.0? A2: Per official announcements, 4.0 Flash is the lightweight version, opened for small-scale trial to black-gold annual-card members on September 28; the full 4.0 is planned for October. Kling has not published a complete list of capability differences between the two, and this article does not guess.
Q3: Does 15 references mean 22 at once? A3: No. The cap is 15. The 10 images, 5 videos, and 7 subjects are per-category ceilings; all references in one task together cannot exceed 15. These are published 4.0-series specs - once the entry point opens, the creation page's actual limits prevail.
Q4: Do I need to write prompts in English? A4: No. Per the official announcement, Kling supports Chinese, English, Japanese, Korean, Spanish, Portuguese, German, French, Hindi and more, with dialects and accents - you can write prompts directly in your language. Structure matters more than language: chronological order, action first, concrete camera language. Those three beat switching to English.
Q5: Is 30 seconds enough? What about content longer than a minute? A5: Once 4.0 or 4.0 Flash opens to you, use continuation: 30 seconds per generation, up to 2 minutes via multiple rounds, per the official spec. The sturdier habit is to storyboard in 30-second blocks regardless of total length - each block a self-contained scene, then joined by continuation. For cross-model long-form thinking, see the AI video generation SOP on this site.
Join the Discussion
What pitfalls have you hit in the 4.0 Flash limited trial, and which combination of the 15 references have you already made work? Leave your prompt structure and workflow in the comments. If you want to rehearse the multimodal-reference mindset on other platforms first, our 30-second club comparison breaks down the trade-offs among four current representative models - picking by need beats chasing novelty.
References
- Kling official announcements and media reports (Shanghai Securities News 2026-09-29, Cailian Press 2026-09-30, Geek Park coverage 2026-09-28): Kling 4.0 October launch, 4.0 Flash small-scale trial for black-gold annual-card members, 4K and 1080p 10-bit HDR output, 30 seconds per generation, extension to 2 minutes, 8,000-token prompt limit, 15 multimodal references (10 images, 5 videos, 7 subjects), multilingual lip sync, and the creation-page upgrade points
- LTX-2 official repository README, "Prompting for LTX-2" section (Lightricks/LTX-2, pulled 2026-10-02): the prompting method of chronological description within about 200 words, concrete camera language, and action-first structure (attributed in the text to the LTX-2 official guide, not Kling's official documentation)
This article is based on official announcements and media reports (as of 2026-10-03) and is not an official partnership or promotion; feature rollout pace, entry points, and capability boundaries follow Kling's official pages in real time. Pricing and credit rules have not been announced, and this article does not speculate.