Skip to main content
Kling 3.0 is one of the most advanced multi-modal generation systems, now available in ComfyUI via Partner Nodes. This release includes Kling Video 3.0, Kling Video 3.0 Omni, Kling Image 3.0, and Kling Image 3.0 Omni, bringing video, image, audio-visual, and narrative generation capabilities directly into your node-based workflows.

What Kling 3.0 is good at

  • Multi-shot generation: Generate multiple shots with duration control in a single generation
  • Locked subject consistency: Maintain character and object identity across camera motion and scene evolution
  • Native audio with character awareness: Multilingual dialogue with precise lip sync and facial expression alignment
  • Native text rendering: Clear, structured on-screen text for ads, UI scenes, and branded content

Example outputs

Text-to-image generation with the Kling 3.0 model: Kling 3.0: Text to Image example

Available models

Learn more

For detailed examples, prompts, and video demonstrations of Kling 3.0’s capabilities, check out our blog post:

Kling 3.0 Models Are Now Available

Read the full announcement with video examples, prompt tips, and feature deep-dives.

Use it in ComfyUI

Kling 3.0 workflows

Run the text-to-image, video generation, and first-last-frame workflows in ComfyUI, locally or on Comfy Cloud