IMAGINE / make-imagine-video-15s

Imagine

Make a Grok Imagine video from a prompt or a still

Make a Grok Imagine video from a prompt or a still

Video 1.5 Fast is on grok.com/imagine, iOS, and Android. The API model id is grok-imagine-video-1.5. Clips run up to 15 seconds. Native 1080p is available for text-to-video and image-to-video.

Text-to-video

  1. Open grok.com/imagine.
  2. Sign in if Grok asks.
  3. Switch the session to video (the control sits with Image / Speed / Quality).
  4. Describe one shot: subject, camera move, and what happens in time. No starting image.
  5. Generate. Play it once. Change one variable (duration, camera, action) and generate again.

Text-to-video builds a first frame, then animates it. You send one request; Grok does not hand you that intermediate still.

Image-to-video

  1. Start from a still you like (Quality Mode still, or an upload).
  2. Describe the motion only: push-in, head turn, cloth, weather.
  3. Pick duration and resolution. 6-second 720p on Video 1.5 Fast is the speed xAI quotes (about 25 seconds to generate).
  4. Generate. If the subject melts, shorten the action and keep the still.

References (plan gate)

Image and voice references landed first in the US on SuperGrok Heavy and SuperGrok Plus, on grok.com/imagine and iOS, then rolled out. Up to seven reference images per generation. Each reference locks one thing: a face, a product, a location. Voice references hold a speaking voice across scenes.

If your account does not show reference slots yet, stay on text-to-video or image-to-video. Those two plus native 1080p are generally available on grok.com/imagine, iOS, and Android.

Pitfalls

  • Audio (ambience, effects, dialogue) is generated in the same pass. If you need silence, say so in the prompt.
  • App Store, x.ai/grok, and SuperGrok Plus bullets disagree on max length and resolution. Teach 15 seconds and 1080p from the Video 1.5 references post, then read the control you actually see after login.
  • Do not mix the consumer page with the API. The API takes image_url / reference_image_urls, duration, aspect_ratio, and resolution.