All models

grok-imagine-video:reverse

xAIVideo
Get your API key
grok-imagine-video:reverse

Turn Text Concepts and Static Images into Short Video Shots

grok-imagine-video:reverse is the video generation entry point in the xAI Grok Imagine series, supporting video creation from text descriptions as well as dynamic shot design starting from static images. It is suitable for producing 1–15 second creative clips, product showcases, and storyboard previews. When creating, organize your input around the subject, action, scene, and camera movement; after generation, obtain a video link to integrate into editing or content publishing workflows.

xAIModel Brand
VideoModel Type
VideoTask Capability

Specifications and API Features

Clarify capacity, inputs and outputs, and invocation methods before choosing a model.

Creation Methods
Text-to-video, image-to-video
Duration for This Entry Point
1–15 seconds, with duration defaulting to 6 seconds
Text Input
prompt is required for text-to-video; image-to-video may include additional action prompts
Image Input
Submit the input image link through image_url
API Endpoint
POST /grok/videos, explicitly specify grok-imagine-video:reverse
Task Management
Supports async and callback_url; use task_id to query tasks
Result Delivery
video_url video link and pending, succeeded, failed statuses

The durations and operating methods above apply to this API entry point and do not treat the native specifications of other versions in the series as guarantees for this model.

Core Capabilities

Learn what grok-imagine-video:reverse can bring to your work.

Build Short Shots from Text

When no ready-made image is available, you can directly describe the subject, environment, action, and camera changes to generate a video. Prompts work best when written as a clear shot, such as where the subject moves, how the camera approaches, and what atmosphere the lighting conveys; focus on a single action to make it easier to compare different creative options.

Bring Static Images into Dynamic Narratives

When you already have product, character, or scene images, you can enter the image-to-video workflow through image_url, then use text to describe the action you want to happen. The image provides the visual starting point, while the prompt adds motion intent, making it suitable for exploring short clips such as a character turning, environmental changes, or camera push-ins, rather than redefining the entire scene from text.

Integrate Generation into Task Workflows

Video generation can use asynchronous submission: first obtain a task_id, then check the status or receive a completion callback. The result provides video_url, allowing applications to handle submission, waiting, previewing, and downloading separately. For content production systems, this task-based delivery is easier to organize than keeping users on a waiting page throughout the process.

Use Cases

Start with specific tasks to find where the model can be most effective.

Product Clips and Showcase Assets

Use a static product image as input, add background atmosphere, motion direction, and camera distance, and generate candidate short videos for presentation. Start by trying one clear action, then pass suitable clips to the editing stage to add selling-point text, branding, and music. This is well suited for teams that need to expand existing visual assets into dynamic content.

Script Shot Previsualization

Rewrite a shot from a script as a description of the subject, scene, action, and camera movement, then create a watchable draft through text-to-video generation. The deliverable can be used to discuss pacing, composition, and action intent, without requiring complete live-action footage in the early stage; different descriptions can be generated separately to help creators select a more suitable direction of expression.

Illustration and Character Motion Proposals

Use existing illustrations or character images as a starting point, specify subtle movements, viewpoint changes, or environmental motion, and create dynamic proposals. This is suitable for exploring formats for event visuals, character showcases, and social content. Details that need to be retained in the final clip can be specified clearly in the prompt and checked item by item when selecting clips for subject form and motion results.

How to choose this model

Choose based on task complexity, input materials, and expected results.

Choose this model for short shots; consider fast for longer clips

If you need to retain both text-to-video and image-to-video modes, and each clip is within the 1–15 second range, you can choose this model. grok-imagine-video-1.5-fast:reverse supports durations of 6–30 seconds, making it better suited for tasks requiring longer clips. The two do not use the same invocation ID; when switching, adjust the duration settings accordingly rather than simply replacing the name while retaining all parameters.

Differentiate related entry points by creative starting point

grok-imagine-video:official also provides text-to-video and image-to-video modes and is another public invocation entry point; grok-imagine-video-1.5:official supports image-to-video only and requires an image. When you only have a text concept, you can start directly with this model; when you already have an image and are ready to compare other versions, you can test each separately using the same material and action description.

Getting started

From a small-scale task to production integration.

01

Prepare the task and materials

Clarify the goal, required inputs, and output requirements, using real business examples as a starting point.

02

Try it in the API testing area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.

03

Integrate according to the API documentation

Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage boundaries

Before formal use, understand the output quality and capability scope.

  • This model supports a single duration of 1–15 seconds. Do not submit a 30-second task simply because a longer range appears in a shared interface. For complete long videos or multi-part narratives, split the content into shots and then edit them together; a single generation does not automatically complete the arrangement of an entire video.
  • Image-to-video is a creative method for generating moving visuals and should not be treated as a pixel-perfect animation tool. When product text, character facial features, logos, or structures have strict requirements, inspect key frames in the final output and add subtitles and brand elements that must be preserved precisely during post-production.
  • Do not apply the Beta fun, normal, or spicy modes, or the audio and high-resolution capabilities of other versions, to this model. Creative style can be expressed through prompts; voiceover, music, and precise editing should be planned as separate production steps rather than relying on generation controls that are not explicitly supported.

Frequently Asked Questions

Answers to common questions about using grok-imagine-video:reverse.

Can I use this model without an image?

Yes. This model supports text-to-video; simply provide a prompt to describe what you want to generate. It is recommended to specify the subject, scene, action, and camera movement, and organize the information around a single shot first. When calling it, explicitly enter grok-imagine-video:reverse to avoid using another default model.

Is a prompt required for image-to-video?

When providing image_url, prompt can be left blank. However, if you want the subject to move in a specific direction, or want the camera to push in or rotate, adding an action description makes it easier to express your intent. The image provides the visual starting point, while the text explains how you want the scene to change; they serve different purposes.

Can this model generate a 30-second video?

This model supports durations of 1–15 seconds, with a default of 6 seconds, and is not suitable for directly submitting a 30-second task. For longer single-shot videos, consider grok-imagine-video-1.5-fast:reverse, which supports durations of 6–30 seconds; for multiple shots, you can also generate them separately and edit them together.

What should I do after a generation request returns task_id?

task_id is a task identifier, not a video URL. When using async, you can check progress through POST /grok/tasks; when callback_url is set, you can wait for the completion notification. After confirming that the result status is succeeded, read video_url; pending means it is still being processed.

Is :reverse an independent native video model?

:reverse is a suffix for the public call ID used to distinguish entry points, and should not be treated as an independent vendor model name. When using this entry point, retain the full ID. It shares the basic text-to-video and image-to-video creation methods with the :official entry point, but this does not mean that all available parameters are exactly the same.

Model information · Updated: 2026-10-01. For calling parameters and billing rules, see the API and pricing sections.

Use grok-imagine-video:reverse for your next task

Start with a clear goal and judge from actual results whether it suits your work.