Skip to main content
The Grok Imagine Video 1.5 Partner Node enables high-quality video generation with native audio from a single image input. Powered by xAI’s latest Grok model, it produces realistic motion with synchronized audio and supports up to 1080p resolution. The node supports two model variants selected via its model parameter:
  • grok-imagine-video: the previous generation model, supports optional image input
  • grok-imagine-video-1.5: the latest model, always requires an input image and supports 1080p output
Both variants generate native audio: sound effects, ambience, and dialogue are synthesized in the same pass, with no separate audio pipeline needed. Video duration ranges from 1 to 15 seconds.

What Grok Imagine Video 1.5 is good at

  • Image-to-video generation: produces high-quality video from a single input image
  • Native audio: sound effects, ambience, and dialogue are synthesized in the same pass, with no separate audio pipeline needed
  • Realistic motion: motion stays synchronized with the generated audio
  • Up to 1080p output: the 1.5 model supports 1080p resolution
  • Flexible duration: video length from 1 to 15 seconds

Use it in ComfyUI

Grok Imagine Video 1.5 workflow

Run the image-to-video workflow in ComfyUI, locally or on Comfy Cloud