Turn Text Concepts and Static Images into Short Video Shots
grok-imagine-video:reverse is the video generation entry point for the xAI Grok Imagine series, supporting both video creation from text descriptions and dynamic shot design starting from static images. It is suitable for producing 1–15 second creative clips, product showcases, and storyboard previews. When creating, you can organize inputs around the subject, action, scene, and camera movement; after generation, obtain a video link and integrate it into editing or content publishing workflows.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API Features
Creation Method
Text-to-video, image-to-video
Duration for This Endpoint
1–15 seconds, with duration defaulting to 6 seconds
Text Input
prompt is required for text-to-video; image-to-video can include an additional action prompt
Image Input
Submit the input image link through image_url
API Endpoint
POST /grok/videos, explicitly specify grok-imagine-video:reverse
Task Management
Supports async and callback_url; use task_id to query tasks
Result Delivery
video_url video link and pending, succeeded, failed statuses
The durations and operating methods above apply to this API endpoint and do not treat the native specifications of other versions in the series as guarantees for this model.
Core Capabilities
Build Short Shots from Text
When no ready-made image is available, you can directly describe the subject, environment, action, and camera changes to generate a video. Prompts work well as a clear single shot—for example, how the subject moves through a space, how the camera approaches, and what atmosphere the lighting creates. Keeping the focus on one action makes it easier to compare different creative options.
Bring Static Images into Dynamic Storytelling
When you already have a product image, character image, or scene image, you can enter the image-to-video workflow through image_url and then describe the action you want to occur in text. The image serves as the visual starting point, while the prompt adds motion intent. This is suitable for exploring short segments such as a character turning, environmental changes, or camera push-ins, rather than redefining the entire image from text.
Integrate Generation into Task Workflows
Video generation can be submitted asynchronously: first obtain a task_id, then query the status or receive a completion callback. Results provide video_url, allowing applications to handle submission, waiting, previewing, and downloading separately. For content production systems, this task-based delivery is easier to organize than requiring users to remain on a waiting page throughout the process.
Use Cases
Product clips and showcase assets
Use static product images as input, add background atmosphere, motion direction, and camera distance, and generate candidate short videos for presentation. Start by trying one clear action, then pass suitable clips to the editing stage to add selling-point text, logos, and music. Suitable for teams that need to expand dynamic content from existing visual assets.
Script shot previsualization
Rewrite a shot from the script as descriptions of the subject, scene, action, and camera movement, then create a watchable draft through text-to-video. Deliverables are used to discuss pacing, composition, and action intent, without needing to prepare complete live-action footage early on; different descriptions can be generated separately to help creators select a more suitable direction of expression.
Illustration and character motion proposals
Use existing illustrations or character images as a starting point, specify subtle movements, viewpoint changes, or environmental motion, and create dynamic proposals. Suitable for exploring presentation formats for event visuals, character showcases, and social content. The details that need to be retained in the final video can be specified in the prompt and checked item by item when selecting clips for subject form and motion results.
How to Choose This Model
Choose this model for short shots; consider fast for longer clips
If you need to retain both text-to-video and image-to-video modes, and each clip is within the 1–15 second range, choose this model. grok-imagine-video-1.5-fast:reverse has a duration range of 6–30 seconds, making it more suitable for tasks requiring longer clips. They are not the same invocation ID; when switching, adjust the duration settings accordingly rather than merely replacing the name while keeping all parameters.
Differentiate related endpoints by creative starting point
grok-imagine-video:official also provides text-to-video and image-to-video modes and is another public invocation endpoint; grok-imagine-video-1.5:official supports image-to-video only and requires an image. When you only have a text concept, this model can be used directly; when you already have images and are ready to compare other versions, you can use the same assets and action descriptions to test each one separately.
Getting Started
Choose a text or image starting point
For text-to-video, provide a prompt; for image-to-video, use image_url to define the visual starting point and add motion instructions. Configure reference image capabilities according to this public variant guide.
Specify the complete public invocation ID
Specify model=grok-imagine-video:reverse for /grok/videos, starting with a small 6-second task; use durations of 1–15 seconds. Choose aspect_ratio and resolution based on the actual footage.
Retrieve generation task results
Use async or callback_url to track the task, and use task_id to query /grok/tasks; wait for succeeded before reading video_url, check the subject, motion, visual continuity, and ending, then move the returned video into editing as appropriate.
Trial Recommendation: Dynamic Continuation of the First Frame
Input and Goal
Use a photo of greenery by a window as the first frame, with the leaves gently swaying and the camera slowly pushing in, while preserving the flowerpot, window frame, and composition.
Acceptance Criteria and Next Steps
Provide image_url and clearly specify subtle motion, selecting 1–15 seconds; do not apply the 6–30 second range of fast:reverse.
Usage Boundaries
The duration per generation for this model is 1–15 seconds. Do not submit a 30-second task simply because a longer range appears in a shared interface. When a complete long-form video or multi-part narrative is needed, split the content into shots and edit them afterward; a single generation does not automatically complete the arrangement of an entire video.
Image-to-video is a creative method for generating dynamic visuals and should not be treated as a pixel-perfect animation tool. When there are strict requirements for product text, character facial features, logos, or structures, inspect key frames in the final video and arrange subtitles and brand elements that must be preserved precisely during post-production.
Do not apply the Beta fun, normal, or spicy modes, or the audio and high-resolution capabilities of other versions, to this model. Creative styles can be expressed through prompts; voice-over, music, and precise editing should be planned as separate production steps rather than relying on generation controls that are not explicitly supported.
Frequently Asked Questions
Can I use this model without an image?
Yes. This model supports text-to-video; simply provide a prompt to describe the content you want to generate. It is recommended to specify the subject, scene, action, and camera movement, organizing the information around a single shot first. When calling it, explicitly specify grok-imagine-video:reverse to avoid using other default models.
Is a prompt required for image-to-video?
When providing image_url, prompt can be left blank. However, if you want the subject to move in a specific direction, or want the camera to push in or rotate, adding an action description makes your intent easier to express. The image provides the visual starting point, while the text explains how you want the scene to change; they serve different roles.
Can this model generate a 30-second video?
This model supports durations of 1–15 seconds, with a default of 6 seconds, and is not suitable for directly submitting a 30-second task. For longer single-shot videos, consider grok-imagine-video-1.5-fast:reverse, which supports durations of 6–30 seconds; for multiple shots, you can also generate them separately and edit them together.
What should I do after a generation request returns task_id?
task_id is a task identifier, not a video URL. When using async, you can check progress through POST /grok/tasks; when callback_url is set, you can wait for the completion notification. After confirming that the result status is succeeded, read video_url; pending indicates that processing is still in progress.
Is :reverse an independent native video model?
:reverse is a suffix for the public call ID, used to distinguish endpoints, and should not be treated as an independent vendor model name. When using this endpoint, retain the complete ID. It shares the basic text-to-video and image-to-video creation methods with the :official endpoint, but this does not mean that all optional parameters are exactly the same.