# Text to video API for AI agents

Scrollport gives AI agents 4 tools to turn text into video, from ByteDance, MiniMax and xAI, through one connection. Prices are per 1k tokens or video second, depending on the tool, paid per use from one prepaid wallet with no subscription, provider accounts or API keys. Agents use them over MCP, the CLI or the HTTP API.

Updated 15 September 2026

## Create video from text with Scrollport

[Get started](https://scrollport.com/start)

Use verified tools with pay-per-use pricing. Provider access is included. Tools that work in your own apps ask you to connect that account.

Connect your first agent to receive $1 trial credit.

## Compare Text to video API tools

Prices are per 1k tokens or video second, depending on the tool, so compare what each tool returns and a typical run's cost before choosing.

| Tool | Provider | What it returns | Price |
| --- | --- | --- | --- |
| [Seedance 1.5 Pro](https://scrollport.com/tools/providers/bytedance/seedance-1-5-pro) | ByteDance | Generate video from text prompts with duration, resolution and aspect-ratio controls. | $0.0034/1k tokens |
| [Seedance 2.5 Text-to-Video](https://scrollport.com/tools/providers/bytedance/seedance-2-5-text) | ByteDance | Generate video with duration, resolution and native audio controls. | $0.0150/1k tokens |
| [MiniMax H3 Text-to-Video](https://scrollport.com/tools/providers/minimax/h3-text-to-video) | MiniMax | Generate video from a text prompt with bounded duration, resolution and aspect-ratio controls. | $0.1120/video second |
| [Grok Imagine Video 1.5 Text-to-Video](https://scrollport.com/tools/providers/xai/grok-imagine-video-1-5-text-to-video) | xAI | Generate video from a text prompt with duration, resolution and aspect-ratio controls. | $0.1960/video second |

## 4 available tools

Generate videos

$0.0034/1k tokens

### Seedance 1.5 Pro

Generate video from text prompts with duration, resolution and aspect-ratio controls.

ByteDance

Generate videos

$0.0150/1k tokens

### Seedance 2.5 Text-to-Video

Generate video with duration, resolution and native audio controls.

ByteDance

Generate videos

$0.1120/video second

### MiniMax H3 Text-to-Video

Generate video from a text prompt with bounded duration, resolution and aspect-ratio controls.

MiniMax

Generate videos

$0.1960/video second

### Grok Imagine Video 1.5 Text-to-Video

Generate video from a text prompt with duration, resolution and aspect-ratio controls.

xAI

## Questions about Text to video API

### How much do text to video API tools cost?

Prices are per 1k tokens or video second, depending on the tool. You pay per use from one prepaid wallet with no subscription. Your agent sees the exact price before each run, and connecting your first agent adds $1 of trial credit.

### What is the best text to video API for an AI agent?

It depends on what you need back and how much you will run it. Seedance 1.5 Pro (ByteDance, $0.0034/1k tokens): Generate video from text prompts with duration, resolution and aspect-ratio controls. Seedance 2.5 Text-to-Video (ByteDance, $0.0150/1k tokens): Generate video with duration, resolution and native audio controls. MiniMax H3 Text-to-Video (MiniMax, $0.1120/video second): Generate video from a text prompt with bounded duration, resolution and aspect-ratio controls. Your agent can inspect the closest match and see its exact price before it runs.

## Learn more about text-to-video generation

Describe one filmable shot with a clear subject, action, setting and camera intention. Your agent can turn that direction into a bounded first generation.

**Write a shot, not a whole story**

Specify what is visible during a short period: subject appearance, action, environment, framing, camera movement, lighting and mood. Split scene changes into separate shots.

**How detailed should a text-to-video prompt be?**

Include details that affect the visible result and omit instructions the camera cannot show. Prioritise subject consistency and physical action over long backstory or several competing style references.

**Approve a short draft before extending**

Review motion, anatomy, camera behaviour, unwanted objects and the ending frame. Refine the shot before requesting more duration or matching shots for a sequence.

**Example prompt**

This prompt keeps the generation to one observable shot.

Example prompt to give your agent

`Use Scrollport's Text-to-video tools to create a [DURATION] [ASPECT RATIO] shot for [USE]. Show [SUBJECT] [ACTION] in [SETTING]. Frame it as [SHOT SIZE], move the camera [CAMERA MOVEMENT], and use [LIGHTING AND MOOD]. Preserve [IMPORTANT DETAILS], avoid [EXCLUSIONS], and show me the price before the paid run.`

---

Canonical page: https://scrollport.com/tools/text-to-video-api
