Video generation
$0.1960/video secondGrok Imagine Video 1.5 Image-to-Video
Generate video from an input image and text prompt with duration and resolution controls.
Loading tools…
Compare text, image and reference-guided video models before choosing a production mode.
Use verified tools with pay-per-use pricing. Provider access is included. Tools that work in your own apps ask you to connect that account.
Connect your first agent to receive $1 trial credit.
Video generation
$0.1960/video secondGenerate video from an input image and text prompt with duration and resolution controls.
Video generation
$0.1960/video secondGenerate video from one to seven reference images and a tagged text prompt with duration and aspect-ratio controls.
Video generation
$0.1960/video secondGenerate video from a text prompt with duration, resolution and aspect-ratio controls.
Video generation
$0.1120/video secondGenerate video from a text prompt with bounded duration, resolution and aspect-ratio controls.
Video generation
$0.0034/1k tokensGenerate video from text prompts with duration, resolution and aspect-ratio controls.
Video generation
$0.0150/1k tokensGenerate video with duration, resolution and native audio controls.
Video generation
$0.0150/1k tokensGenerate video with duration, resolution and native audio controls.
Video generation
$0.0150/1k tokensGenerate video with duration, resolution and native audio controls.
Choose the production route from the assets you already have: a written shot, a source image or visual references. Your agent can then use the narrow video capability that fits that starting point.
Use text-to-video when the shot starts from a written idea, image-to-video when one source frame should come alive, and reference-to-video when supplied visuals must guide identity, style or composition across the result.
Break a concept into short shots with one subject action and one camera intention each. Define aspect ratio, duration and continuity requirements before generating expensive final footage.
Check identity, wardrobe, objects, screen direction, lighting and transitions between shots. Give feedback on one observable problem at a time so the agent can revise the right part of the sequence.