Skip to main content
The WanMoveTrackToVideo node prepares conditioning and latent space data for video generation, incorporating optional motion tracking information. It encodes a starting image sequence into a latent representation and can blend in positional data from object tracks to guide the motion in the generated video. The node outputs modified positive and negative conditioning along with an empty latent tensor ready for a video model.

Inputs

Note: The strength parameter only has an effect when tracks are provided and strength is greater than 0.0; track conditioning is applied only when start_image is also provided. If tracks are not provided or strength is 0.0, the track blending is skipped. When track blending is active, the positive conditioning receives the track-blended latent image, while the negative conditioning receives the unmodified latent image. If start_image is not provided, no latent image and mask conditioning is created; the positive and negative conditioning pass through unchanged (except that clip_vision_output is still added if provided), and the node outputs an empty latent. Note: When start_image is provided, the image sequence is resized to the target width and height and truncated to the first length frames. If the sequence is shorter than length, the remaining frames are filled with neutral gray frames (value 0.5) before VAE encoding. The resulting conditioning includes a concat_mask with value 0 at the temporal positions corresponding to the start image frames and 1 elsewhere.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): b02a1a359d349a0136d84ed77a510c46cb2c8b565650ed54d5fca6c87cd0ab1f