Skip to content

ACE-Step

POST/v1/run

Ace Step Audio To Audio by ace - advanced AI model for audio-to-audio. Delivers high-quality results with fast inference, suitable for both creative and production workflows.

Request body

Submit an async generation request. The model field selects the model; other fields are model-specific input parameters.

stringmodelrequired

Model identifier. Set to ace/ace-step/audio-to-audio.

Default: ace/ace-step/audio-to-audio

Optional<string>audio

URL of the audio file to be outpainted.

Optional<string>lyrics

Lyrics to be sung in the audio. If not provided or if [inst] or [instrumental] is the content of this field, no lyrics will be sung. Use control structures like [verse], [chorus] and [bridge] to control the structure of the song.

Default:

Optional<number>guidance_scale

Guidance scale for the generation.

Range: 0 to 200

Default: 15

Optional<integer>seed

Random seed for reproducibility. If not provided, a random seed will be used.

Optional<number>minimum_guidance_scale

Minimum guidance scale for the generation after the decay.

Range: 0 to 200

Default: 3

Optional<number>tag_guidance_scale

Tag guidance scale for the generation.

Range: 0 to 10

Default: 5

Optional<number>lyric_guidance_scale

Lyric guidance scale for the generation.

Range: 0 to 10

Default: 1.5

Optional<string>edit_mode

Whether to edit the lyrics only or remix the audio.

Allowed values: lyrics, remix

Default: remix

Optional<number>guidance_interval

Guidance interval for the generation. 0.5 means only apply guidance in the middle steps (0.25 * infer_steps to 0.75 * infer_steps)

Range: 0 to 1

Default: 0.5

Optional<integer>original_seed

Original seed of the audio file.

Optional<string>scheduler

Scheduler to use for the generation process.

Allowed values: euler, heun

Default: euler

Optional<integer>granularity_scale

Granularity scale for the generation process. Higher values can reduce artifacts.

Range: -100 to 100

Default: 10

Optional<string>guidance_type

Type of CFG to use for the generation process.

Allowed values: cfg, apg, cfg_star

Default: apg

Optional<string>original_lyrics

Original lyrics of the audio file.

Default:

Optional<string>original_tags

Original tags of the audio file.

Optional<number>guidance_interval_decay

Guidance interval decay for the generation. Guidance scale will decay from guidance_scale to min_guidance_scale in the interval. 0.0 means no decay.

Range: 0 to 1

Default: 0

Optional<integer>number_of_steps

Number of steps to generate the audio.

Range: 3 to 60

Default: 27

Optional<string>tags

Comma-separated list of genre tags to control the style of the generated audio.

Response Schema

The submit endpoint returns an accepted generation task. Poll the result endpoint with the returned id for terminal outputs or errors.

Optional<string>error

Error message if the task failed. Empty on success.

stringidrequired

Unique identifier for the generation task.

Optional<string>model

Model ID used for the prediction.

Optional<array>outputs

Array of generated content. Empty when status is not completed.

stringstatusrequired

Status of the task: pending, running, completed, failed, or timeout.

Allowed values: pending, running, completed, failed, timeout

Model capabilities

array<string>capability_tagsrequired

Capabilities declared by the model registry.

Default: audio-to-audio

stringexecution_moderequired

Execution mode declared by the model registry.

Default: async