Skip to content

First API call ​

This guide walks through the anatomy of a SandBase API request and response in detail. By the end, you'll understand every field in the request body, how to read the response, and how to use streaming.

Before you start ​

You need an active SandBase organization, an API key, and a model ID from Supported Models. Export the key so examples do not place secrets in source code:

bash
export SANDBASE_API_KEY="sk-YOUR_API_KEY"

Start with a non-streaming request. Once authentication and response handling work, add streaming and production retry behavior.

Request Anatomy ​

This walkthrough uses the OpenAI-compatible Chat Completions route:

bash
POST https://api.sandbase.ai/v1/chat/completions

Required Headers ​

HeaderValueDescription
AuthorizationBearer sk-YOUR_API_KEYYour SandBase API key
Content-Typeapplication/jsonRequest body format

Request Body ​

json
{
  "model": "deepseek/deepseek-v4-flash",
  "messages": [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "What is the capital of France?"}
  ],
  "temperature": 0.7,
  "max_tokens": 256
}

Body Parameters ​

ParameterTypeRequiredDescription
modelstringYesAn enabled model ID returned by GET /v1/models (for example, deepseek/deepseek-v4-flash)
messagesarrayYesConversation history as an array of message objects
temperaturenumberNoSampling temperature (0–2). Support and defaults can vary by model.
max_tokensintegerNoMaximum tokens to generate in the response
top_pnumberNoNucleus sampling parameter (0–1)
streambooleanNoWhether to stream the response. Default: false
stopstring or arrayNoStop sequences — generation stops when these are encountered
frequency_penaltynumberNoProvider-compatible repetition penalty when supported
presence_penaltynumberNoProvider-compatible presence penalty when supported

Message Roles ​

RolePurpose
systemSets the assistant's behavior and personality
userThe human's input
assistantPrevious assistant responses (for multi-turn conversations)
developer, toolAdditional OpenAI-compatible roles when supported by the selected model and route

Full Request Example ​

bash
curl https://api.sandbase.ai/v1/chat/completions \
  -H "Authorization: Bearer $SANDBASE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4-flash",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "What is the capital of France?"}
    ],
    "temperature": 0.7,
    "max_tokens": 256
  }'
python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["SANDBASE_API_KEY"],
    base_url="https://api.sandbase.ai/v1"
)

response = client.chat.completions.create(
    model="deepseek/deepseek-v4-flash",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What is the capital of France?"}
    ],
    temperature=0.7,
    max_tokens=256
)

print(response.choices[0].message.content)
javascript
import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: process.env.SANDBASE_API_KEY,
  baseURL: 'https://api.sandbase.ai/v1',
});

const response = await client.chat.completions.create({
  model: 'deepseek/deepseek-v4-flash',
  messages: [
    { role: 'system', content: 'You are a helpful assistant.' },
    { role: 'user', content: 'What is the capital of France?' },
  ],
  temperature: 0.7,
  max_tokens: 256,
});

console.log(response.choices[0].message.content);

Response Structure ​

Typical non-streaming response ​

json
{
  "id": "chatcmpl-abc123def456",
  "object": "chat.completion",
  "created": 1719000000,
  "model": "deepseek/deepseek-v4-flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The capital of France is Paris."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 24,
    "completion_tokens": 8,
    "total_tokens": 32
  }
}

Response Fields Explained ​

FieldDescription
idUnique identifier for this completion
objectOpenAI-compatible response object, normally "chat.completion" on this route
createdUnix timestamp of when the response was generated
modelThe model that generated the response
choicesArray of completion choices (typically one)
choices[].indexIndex of this choice in the array
choices[].message.roleRole returned for the generated message, normally "assistant"
choices[].message.contentThe generated text
choices[].finish_reasonWhy generation stopped (see below)
usage.prompt_tokensTokens in your input
usage.completion_tokensTokens in the generated output
usage.total_tokensSum of prompt + completion tokens

Finish reasons ​

Finish reasons are provider-dependent. Common OpenAI-compatible values include:

ValueTypical meaning
stopNatural end of response or hit a stop sequence
lengthHit max_tokens limit — response was truncated

Treat unknown values as valid provider output rather than rejecting the response.

Streaming Responses ​

For real-time output (like a chatbot typing), use streaming. The response arrives as Server-Sent Events (SSE):

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["SANDBASE_API_KEY"],
    base_url="https://api.sandbase.ai/v1"
)

stream = client.chat.completions.create(
    model="deepseek/deepseek-v4-flash",
    messages=[{"role": "user", "content": "Write a haiku about coding."}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
print()  # newline at the end
javascript
import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: process.env.SANDBASE_API_KEY,
  baseURL: 'https://api.sandbase.ai/v1',
});

const stream = await client.chat.completions.create({
  model: 'deepseek/deepseek-v4-flash',
  messages: [{ role: 'user', content: 'Write a haiku about coding.' }],
  stream: true,
});

for await (const chunk of stream) {
  const content = chunk.choices[0]?.delta?.content;
  if (content) process.stdout.write(content);
}
console.log();
bash
curl https://api.sandbase.ai/v1/chat/completions \
  -H "Authorization: Bearer $SANDBASE_API_KEY" \
  -H "Content-Type: application/json" \
  -N \
  -d '{
    "model": "deepseek/deepseek-v4-flash",
    "messages": [{"role": "user", "content": "Write a haiku about coding."}],
    "stream": true
  }'

Streaming SSE Format ​

A compatible stream commonly emits Server-Sent Events like these:

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1719000000,"model":"deepseek/deepseek-v4-flash","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1719000000,"model":"deepseek/deepseek-v4-flash","choices":[{"index":0,"delta":{"content":"The"},"finish_reason":null}]}

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1719000000,"model":"deepseek/deepseek-v4-flash","choices":[{"index":0,"delta":{"content":" capital"},"finish_reason":null}]}

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1719000000,"model":"deepseek/deepseek-v4-flash","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

Key points:

  • consume chunks in order and append any delta.content fragments
  • tolerate metadata-only chunks, empty deltas, and provider-compatible fields you do not recognize
  • stop when the stream closes or its terminal marker arrives; Chat Completions commonly uses data: [DONE]
  • do not assume every model emits the exact same first or final chunk shape

Using the Anthropic SDK ​

SandBase also exposes an Anthropic-compatible endpoint at POST /v1/messages. Use the Anthropic SDK by changing the base_url:

python
import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["SANDBASE_API_KEY"],
    base_url="https://api.sandbase.ai"
)

message = client.messages.create(
    model="anthropic/claude-sonnet-5",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "What is the capital of France?"}
    ]
)
print(message.content[0].text)
javascript
import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic({
  apiKey: process.env.SANDBASE_API_KEY,
  baseURL: 'https://api.sandbase.ai',
});

const message = await client.messages.create({
  model: 'anthropic/claude-sonnet-5',
  max_tokens: 1024,
  messages: [{ role: 'user', content: 'What is the capital of France?' }],
});
console.log(message.content[0].text);

Anthropic Response Structure ​

The Anthropic-compatible endpoint returns an Anthropic-style response. A typical text response looks like this:

json
{
  "id": "msg_abc123",
  "type": "message",
  "role": "assistant",
  "content": [
    {
      "type": "text",
      "text": "The capital of France is Paris."
    }
  ],
  "model": "anthropic/claude-sonnet-5",
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 14,
    "output_tokens": 8
  }
}
FieldDescription
idMessage ID (prefixed with msg_)
typeMessage object type, normally "message"
roleGenerated message role, normally "assistant"
contentArray of typed content blocks; do not assume every block is text
modelThe model that generated the response
stop_reason"end_turn" (natural stop), "max_tokens" (hit limit), or "stop_sequence"
usage.input_tokensTokens in your input
usage.output_tokensTokens in the generated output

Anthropic Streaming ​

Streaming with the Anthropic SDK works the same way — just pass stream=True:

python
import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["SANDBASE_API_KEY"],
    base_url="https://api.sandbase.ai"
)

with client.messages.stream(
    model="anthropic/claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Write a haiku about coding."}]
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
print()

Choosing a Model ​

When selecting a model, check its live model card for availability, pricing, context length, supported inputs, and capabilities. These values change independently, so do not rely on a copied model comparison.

  • Browse the current Supported Models.
  • Open the Model API Reference for model-specific request fields.
  • Query GET /v1/models when an integration needs current model metadata programmatically.

For models that share the same compatible interface, switching usually starts with the model parameter. Recheck the selected model's capabilities and schema before carrying over optional fields, tools, media inputs, or structured-output settings.

Next Steps ​

Before production, set request timeouts, retry only transient failures with bounded exponential backoff, log request IDs when available, and monitor usage without logging prompts or API keys.

  • Model API Reference — Model endpoints, parameters, and model-specific references
  • Streaming Guide — Advanced streaming patterns and error handling
  • Models — Model discovery, capabilities, and current pricing guidance
  • Errors — Handle failures and retries