Skip to content

Audio (TTS/STT)

POST/v1/run

Run a text-to-speech or speech-to-text model. Asynchronous audio models may include webhook_url for a task callback.

Request body

ImageEditParams
stringmodelrequired

Use gpt-image-2 on this endpoint.

stringpromptrequired

Text description of the requested edit.

file | file[]imagerequired

One or more source image files. Repeat the multipart field for multiple images.

Optional<file>mask

Optional mask image for inpainting.

Optional<integer>n

Number of edited images to return when supported.

Optional<string>size

Output dimensions supported by the model.

Optional<string>quality

Output quality supported by the model.

Optional<string>background

Output background setting.

Optional<string>input_fidelity

How closely supported models should preserve source-image details.

Optional<string>output_format

Output encoding such as png, webp, or jpeg.

Optional<integer>output_compression

Output compression level from 0 to 100.

Optional<string>user

Provider-compatible end-user identifier.