Kling AI has unveiled Kling 4.0 Preview, a major update to its flagship video generation model that extends clip length, enhances reference handling, and adds native 4K resolution with synchronized multilingual audio. The preview is now rolling out, according to an announcement from Kuaishou's Kling AI lineup.
The update addresses a frequent limitation of earlier versions: short native runtimes. Kling 4.0 Preview now generates clips up to 30 seconds natively, with an experimental Long Video mode capable of producing continuous 120-second sequences at 1080p. Clips can also be extended to 60 seconds after generation, a feature designed to ease planning for narrative and advertising projects. This matters because longer, stable clips reduce the need for stitching multiple segments, streamlining production workflows for content creators and small teams.
Central to the release is Omni Reference, a system that allows a single generation to draw on up to 15 reference elements from a pool of 50 uploaded files—30 images, 10 video clips, and 10 audio clips. Creators can lock a character's face across shots, match a product's exact color from a photo, or carry a specific lighting setup between scenes without re-describing details in each prompt. This capability could significantly reduce manual prompt engineering and improve consistency in AI-generated video, a persistent challenge for the industry.
Multi-shot generation leverages the same reference system, enabling a single prompt to produce a sequence of shots with spatial continuity. A room remains consistent as the camera moves, and characters retain the same clothing and face across cuts. Historically, AI video tools often visibly reset scenes when the camera angle changed, breaking immersion. By maintaining continuity, Kling 4.0 Preview may raise the bar for narrative coherence in AI video, benefiting filmmakers, advertisers, and marketers who require reliable multi-shot sequences.
On the audio front, dialogue, ambient sound, and music are generated in the same pass as the picture, with lip movement synced to spoken lines across multiple languages and accents. Different lines can be assigned to different characters within the same scene, eliminating a manual dubbing step that previously required separate audio software. This integration could simplify post-production and make AI-generated videos more accessible to non-experts.
Native 4K output (3840×2160) rounds out the release, aimed at preserving fine texture, sharp edges, and depth of field through fast-motion shots without relying on post-generation upscaling. Creators evaluating AI video tools have increasingly flagged native resolution as a key test for whether a clip survives the move from preview to publish. The addition of native 4K could make Kling 4.0 Preview more viable for professional-grade content.
Kling 4.0 is accessible via Kling4.org, which provides direct access to generation modes as they roll out, including standard, Pro, and native 4K tiers. The site also offers a Kling 4.0 video prompt library for creators seeking practical starting points. As the preview becomes more broadly available, these tools may help creators, marketers, and small teams test prompt-to-clip workflows, persistent-character generation, and multi-shot sequencing. The implications for the AI video industry are notable: longer clips, better consistency, and integrated audio could accelerate adoption and set new expectations for what AI video tools can deliver.

