Skip to content

Audio Generation ​

SandBase currently publishes API reference pages for 46 enabled audio generation models across 10 providers. Choose a provider in the left navigation, then open a model page for its exact API identifier, supported capabilities, and a working request.

Audio Generation models use the async SandBase generation protocol declared in each model registry file. Submit a request, receive a task id, then poll the result endpoint until the generation is completed, failed, or timed out.

Providers ​

OpenAI ​

  • Wizper (Whisper v3) — Wizper by OpenAI - accurate speech-to-text transcription with AI. Convert audio and video to text with high accuracy, multilingual support, and speaker identification.

ElevenLabs ​

  • Scribe V2 — Scribe V2 is ElevenLabs's speech recognition model. Transcribe audio content with industry-leading accuracy across multiple languages and accents.
  • Music — Music is ElevenLabs's AI audio generation model. Produce high-quality music tracks, sound effects, and audio landscapes from natural language prompts.
  • Text To Dialogue — Text To Dialogue by ElevenLabs - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
  • Sound Effects V2 — Sound Effects V2 by ElevenLabs - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
  • V3 — V3 by ElevenLabs - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
  • Turbo V2.5 — Turbo V2.5 by ElevenLabs - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
  • Multilingual V2 — Multilingual V2 is ElevenLabs's AI audio generation model. Produce high-quality music tracks, sound effects, and audio landscapes from natural language prompts.
  • ElevenLabs Audio Isolation — Audio Isolation by ElevenLabs - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
  • Voice Changer — Voice Changer by ElevenLabs - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Google ​

  • Gemini 3.1 Flash Tts — Gemini 3.1 Flash Tts by Google - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
  • Gemini TTS — Gemini TTS by Google - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
  • Lyria2 — Lyria 2 is Google's AI audio generation model. Produce high-quality music tracks, sound effects, and audio landscapes from natural language prompts.

MiniMax ​

  • MiniMax Music 2.5 — MiniMax Music 2.5 text-to-music model with lyrics support, instrumental generation, and configurable audio output.
  • MiniMax Music 2.6 — MiniMax Music 2.6 text-to-music model with enhanced quality, lyrics support, instrumental generation, and configurable audio output.
  • MiniMax Speech 2.8 Turbo — MiniMax Speech 2.8 Turbo text-to-speech model with fast voice synthesis, enhanced expressiveness, 40+ language support, and voice cloning.
  • Minimax Music — Music V2 by MiniMax - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
  • MiniMax Speech 2.6 Turbo — MiniMax Speech 2.6 Turbo text-to-speech model with fast voice synthesis, 40+ language support, voice cloning, and expressive emotion control.
  • MiniMax Speech 2.6 HD — MiniMax Speech 2.6 HD text-to-speech model with high-definition voice synthesis, 40+ language support, voice cloning, and expressive emotion control.
  • MiniMax (Hailuo AI) Music v1.5 — Music V1.5 by MiniMax - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
  • Minimax — Preview Speech 2.5 Hd is MiniMax's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.
  • MiniMax Speech 2.5 Turbo — Preview Speech 2.5 Turbo by MiniMax - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
  • MiniMax Voice Design — Voice Design by MiniMax - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
  • MiniMax Voice Cloning — Voice Clone is MiniMax's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.
  • MiniMax Speech-02 Turbo — Speech 02 Turbo is MiniMax's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.
  • …and 3 more models in the sidebar.

Mirelo ​

  • Mirelo SFX1.6 — Sfx1.6 by Mirelo - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.

stability-ai ​

  • Stable Audio 2.5 Audio to Audio — Stable Audio 2.5 Audio To Audio by stability-ai - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
  • Stable Audio 2.5 Text to Audio — Stable Audio 2.5 by stability-ai - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
  • Stable Audio 2.5 Inpaint — Stable Audio 2.5 Inpaint by stability-ai - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
  • Stable Audio Open — Stable Audio Open is stability-ai's AI audio generation model. Produce high-quality music tracks, sound effects, and audio landscapes from natural language prompts.

Alibaba ​

  • Qwen 3 TTS - Voice Design [1.7B] — Qwen 3 Tts Voice Design 1.7b by Alibaba - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.
  • Qwen 3 TTS - Text to Speech [1.7B] — Qwen 3 Tts 1.7b is Alibaba's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.
  • Qwen 3 TTS - Text to Speech [0.6B] — Qwen 3 Tts 0.6b is Alibaba's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.
  • Qwen 3 TTS - Clone Voice [1.7B] — Qwen 3 Tts Clone Voice 1.7b by Alibaba - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
  • Qwen 3 TTS - Clone Voice [0.6B] — Qwen 3 Tts Clone Voice 0.6b by Alibaba - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

ace ​

  • ACE-Step Audio Outpaint — Ace Step Audio Outpaint by ace - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
  • ACE-Step — Ace Step Audio Inpaint by ace - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
  • ACE-Step — Ace Step Audio To Audio by ace - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
  • ACE-Step — Ace Step Prompt To Audio by ace - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.
  • ACE-Step — Ace Step by ace - generate music, sound effects, and audio from text descriptions with AI. Create original compositions, ambient sounds, and audio content for any creative project.

KwaiVGI ​

  • Kling Video — Kling Video Video To Audio by KwaiVGI - advanced AI model for video-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.
  • Kling TTS — Kling Tts is KwaiVGI's text-to-speech AI model. Generate human-like voiceovers with expressive intonation, multilingual support, and customizable voice characteristics.

xAI ​

  • xAI Text to Speech — Grok Tts by xAI - convert text to natural-sounding speech with AI. Supports multiple voices, languages, emotions, and speaking styles for content creation and accessibility.

Capability coverage ​

audio-to-audio, speech-to-text, text-to-audio, text-to-music, text-to-speech, video-to-audio