Skip to content

Arcee AI: Maestro Reasoning

POST/v1/chat/completions

Maestro Reasoning is Arcee's flagship analysis model: a 32 B‑parameter derivative of Qwen 2.5‑32 B tuned with DPO and chain‑of‑thought RL for step‑by‑step logic. Compared to the earlier 7 B preview, the production 32 B release widens the context window to 128 k tokens and doubles pass‑rate on MATH and GSM‑8K, while also lifting code completion accuracy. Its instruction style encourages structured "thought → answer" traces that can be parsed or hidden according to user preference. That transparency pairs well with audit‑focused industries like finance or healthcare where seeing the reasoning path matters. In Arcee Conductor, Maestro is automatically selected for complex, multi‑constraint queries that smaller SLMs bounce.

Request body

Parameters supported by this model. Values, defaults, and limits are read from the model registry.

stringmodelrequired

Model identifier. Set to arcee-ai/maestro-reasoning.

Default: arcee-ai/maestro-reasoning

array<object>messagesrequired

Conversation messages in system, user, or assistant order.

Optional<integer>max_tokens

Maximum number of tokens the model may generate in the response.

Range: −∞ to 32000

Optional<number>temperature

Sampling temperature. Lower values are more deterministic; higher values are more creative.

Range: 0 to 2

Default: 1

Optional<number>top_p

Nucleus sampling threshold. Use this or temperature, but usually not both.

Range: 0 to 1

Optional<boolean>stream

When true, returns incremental Server-Sent Events instead of one completed response.

Default: false

Optional<array<string>>stop

Sequences that stop generation when the model produces one of them.

Optional<number>presence_penalty

Penalizes tokens that already appeared, encouraging the model to introduce new topics.

Range: -2 to 2

Optional<number>frequency_penalty

Penalizes repeated tokens, reducing repetition in the generated response.

Range: -2 to 2

Optional<integer>top_k

Request parameter supported by this model.

Range: 0 to ∞

Response Schema

Fields returned by this model API response.

array<object>choicesrequired

Generated completion choices.

stringidrequired

Unique chat completion identifier.

stringmodelrequired

Model that generated the response.

Optional<object>usage

Token usage when available.

Model capabilities

array<string>capability_tagsrequired

Capabilities declared by the model registry.

Default: chat

integercontext_lengthrequired

Maximum context window accepted by this model.

Default: 131072 tokens

stringexecution_moderequired

Execution mode declared by the model registry.

Default: sync