Skip to content

ElevenLabs Speech to Text

POST/v1/run

Speech To Text by ElevenLabs - accurate speech-to-text transcription with AI. Convert audio and video to text with high accuracy, multilingual support, and speaker identification.

Request body

Submit an async generation request. The model field selects the model; other fields are model-specific input parameters.

stringmodelrequired

Model identifier. Set to elevenlabs/speech-to-text.

Default: elevenlabs/speech-to-text

Optional<string>audio

URL of the audio file to transcribe

Optional<boolean>diarize

Whether to annotate who is speaking

Default: true

Optional<string>language_code

Language code of the audio

Optional<boolean>tag_audio_events

Tag audio events like laughter, applause, etc.

Default: true

Response Schema

The submit endpoint returns an accepted generation task. Poll the result endpoint with the returned id for terminal outputs or errors.

Optional<string>error

Error message if the task failed. Empty on success.

stringidrequired

Unique identifier for the generation task.

Optional<string>model

Model ID used for the prediction.

Optional<array>outputs

Array of generated content. Empty when status is not completed.

stringstatusrequired

Status of the task: pending, running, completed, failed, or timeout.

Allowed values: pending, running, completed, failed, timeout

Model capabilities

array<string>capability_tagsrequired

Capabilities declared by the model registry.

Default: speech-to-text

stringexecution_moderequired

Execution mode declared by the model registry.

Default: async