Generation API
Generate videos
Create motion from text, images, and supported media references.
Read as Markdown ↗POST /api/generate/video
Requires generate scope and Idempotency-Key. Uses the same model, prompt, parameters, referenceAssetIds, and maxGems envelope as images. Returns a media task.
Parameters
| Parameter | Use |
|---|---|
mode | text-to-video or image-to-video, according to the model. |
sourceImage | Owned image asset ID used as a first frame for image-to-video. |
endFrame | Owned image asset ID used as a last frame where supported. |
duration | Supported video duration in seconds. |
ratio | Supported aspect ratio. |
resolution | Supported resolution. |
audio | Audio setting supported by the selected model. |
referenceVideo / referenceVideos | Owned video IDs where supported. |
referenceAudios | Owned audio IDs where supported. |
Top-level referenceAssetIds map to visual references, not automatically to the first frame. Use sourceImage for a required first frame. Do not pass remote URLs or provider routing parameters.
Example: text to video
Choose a discovered video model that supports text-to-video. Add only duration, ratio, and resolution values confirmed by its capabilities.
{
"model": "VIDEO_MODEL_ID_FROM_DISCOVERY",
"prompt": "A paper boat drifts slowly across a still pond",
"parameters": {"mode": "text-to-video"},
"maxGems": 100
}Example: image to video
{
"model": "VIDEO_MODEL_ID_FROM_DISCOVERY",
"prompt": "Slow camera push toward the subject",
"parameters": {
"mode": "image-to-video",
"sourceImage": "OWNED_IMAGE_ASSET_ID"
},
"maxGems": 100
}Upload input files with Upload assets. Respect capability-specific counts and durations. Poll the returned task with backoff; task reads do not trigger provider polling or generation processing.