Bytedance OmniHuman v1.5
/v1/runOmnihuman 1.5 is Bytedance's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.
Request body
Submit an async generation request. The model field selects the model; other fields are model-specific input parameters.
Model identifier. Set to bytedance/omnihuman/1.5.
Default: bytedance/omnihuman/1.5
The text prompt used to guide the video generation.
The URL of the image used to generate the video
The URL of the mask image to apply to the image. Only the person in the white area of the mask will speak.
The URL of the audio file to generate the video. Audio must be under 30s long for 1080p generation and under 60s long for 720p generation.
The resolution of the generated video. Defaults to 1080p. 720p generation is faster and higher in quality. 1080p generation is limited to 30s audio and 720p generation is limited to 60s audio.
Allowed values: 720p, 1080p
Default: 1080p
Generate a video at a faster rate with a slight quality trade-off.
Default: false
Response Schema
The submit endpoint returns an accepted generation task. Poll the result endpoint with the returned id for terminal outputs or errors.
Error message if the task failed. Empty on success.
Unique identifier for the generation task.
Model ID used for the prediction.
Array of generated content. Empty when status is not completed.
Status of the task: pending, running, completed, failed, or timeout.
Allowed values: pending, running, completed, failed, timeout
Model capabilities
Capabilities declared by the model registry.
Default: image-to-video
Execution mode declared by the model registry.
Default: async

