Kling Video O3
kwaivgi/kling-video/o3Kling Video O3 is KwaiVGI's unified omni video generation model. Generate or transform videos from prompts, images, reference videos, and multi-shot text while choosing standard or pro mode per request.
- Base price
- $0.56USD / run
- Execution
- async
- Model type
- video
- Input fields
- 13
Try the model
Playground
PNG, JPEG, WebP, or GIF · 20 MiB maximum
PNG, JPEG, WebP, or GIF · 20 MiB maximum
PNG, JPEG, WebP, or GIF · 20 MiB maximum each
Your output will appear here
Complete the inputs, then click Run.
Specifications
Pricing
- Base price
- $0.56 / run
- Billing formula
- params.resolution == "4k" ? params.duration * 0.42 : (params.video ? (params.resolution == "pro" ? params.duration * 0.168 : params.duration * 0.126) : (params.resolution == "pro" ? (params.generate_audio == false ? params.duration * 0.112 : params.duration * 0.14) : (params.generate_audio == false ? params.duration * 0.084 : params.duration * 0.112)))
Context & modalities
- Input
- Schema-defined
- Output
- video
Capabilities
- Chat
- Not supported
- Vision
- Not supported
- Reasoning
- Not supported
- Structured output
- Not supported
- Function calling
- Not supported
- Audio input
- Not supported
Access
- Provider
- KwaiVGI
- Model ID
- kwaivgi/kling-video/o3
- Execution
- async
- API
- Unified Run API
- Endpoint
- /v1/run
API README
Kling Video O3
Kling Video O3 is a unified omni-video model that can create, transform, and continue video from combinations of text, images, and source footage. It is intended for workflows that need one model to reason across characters, scenes, motion references, camera direction, and sound rather than selecting a separate endpoint for every creative operation.
The model supports multi-shot storytelling and reference-led generation while maintaining characters and visual elements across a sequence. It can preserve or regenerate source audio, create synchronized sound, and use image or video orientation to guide identity and performance, making it suitable for narrative scenes, branded characters, and complex audiovisual edits.
Highlights
Unified multimodal video creation. Combines text, images, and video references within one generation and transformation model.
Reference consistency. Carries characters, objects, and visual identity across new actions, viewpoints, and scenes.
Multi-shot storytelling. Produces coherent sequences with multiple shots, camera changes, and structured scene progression.
Integrated audiovisual control. Generates synchronized sound or retains source audio while coordinating it with the new visual sequence.
Pricing
| Configuration | Mode | Price |
|---|---|---|
| 3s | Standard, no source video, no generated audio | $0.252 |
| 3s | Standard, no source video, generated audio | $0.336 |
| 3s | Pro, no source video, no generated audio | $0.336 |
| 3s | Pro, no source video, generated audio | $0.420 |
| 3s | Standard with source video | $0.378 |
| 3s | Pro with source video | $0.504 |
| 3s | 4K | $1.260 |
| 4s | Standard, no source video, no generated audio | $0.336 |
| 4s | Standard, no source video, generated audio | $0.448 |
| 4s | Pro, no source video, no generated audio | $0.448 |
| 4s | Pro, no source video, generated audio | $0.560 |
| 4s | Standard with source video | $0.504 |
| 4s | Pro with source video | $0.672 |
| 4s | 4K | $1.680 |
| 5s | Standard, no source video, no generated audio | $0.420 |
| 5s | Standard, no source video, generated audio | $0.560 |
| 5s | Pro, no source video, no generated audio | $0.560 |
| 5s | Pro, no source video, generated audio | $0.700 |
| 5s | Standard with source video | $0.630 |
| 5s | Pro with source video | $0.840 |
| 5s | 4K | $2.100 |
| 6s | Standard, no source video, no generated audio | $0.504 |
| 6s | Standard, no source video, generated audio | $0.672 |
| 6s | Pro, no source video, no generated audio | $0.672 |
| 6s | Pro, no source video, generated audio | $0.840 |
| 6s | Standard with source video | $0.756 |
| 6s | Pro with source video | $1.008 |
| 6s | 4K | $2.520 |
| 7s | Standard, no source video, no generated audio | $0.588 |
| 7s | Standard, no source video, generated audio | $0.784 |
| 7s | Pro, no source video, no generated audio | $0.784 |
| 7s | Pro, no source video, generated audio | $0.980 |
| 7s | Standard with source video | $0.882 |
| 7s | Pro with source video | $1.176 |
| 7s | 4K | $2.940 |
| 8s | Standard, no source video, no generated audio | $0.672 |
| 8s | Standard, no source video, generated audio | $0.896 |
| 8s | Pro, no source video, no generated audio | $0.896 |
| 8s | Pro, no source video, generated audio | $1.120 |
| 8s | Standard with source video | $1.008 |
| 8s | Pro with source video | $1.344 |
| 8s | 4K | $3.360 |
| 9s | Standard, no source video, no generated audio | $0.756 |
| 9s | Standard, no source video, generated audio | $1.008 |
| 9s | Pro, no source video, no generated audio | $1.008 |
| 9s | Pro, no source video, generated audio | $1.260 |
| 9s | Standard with source video | $1.134 |
| 9s | Pro with source video | $1.512 |
| 9s | 4K | $3.780 |
| 10s | Standard, no source video, no generated audio | $0.840 |
| 10s | Standard, no source video, generated audio | $1.120 |
| 10s | Pro, no source video, no generated audio | $1.120 |
| 10s | Pro, no source video, generated audio | $1.400 |
| 10s | Standard with source video | $1.260 |
| 10s | Pro with source video | $1.680 |
| 10s | 4K | $4.200 |
| 11s | Standard, no source video, no generated audio | $0.924 |
| 11s | Standard, no source video, generated audio | $1.232 |
| 11s | Pro, no source video, no generated audio | $1.232 |
| 11s | Pro, no source video, generated audio | $1.540 |
| 11s | Standard with source video | $1.386 |
| 11s | Pro with source video | $1.848 |
| 11s | 4K | $4.620 |
| 12s | Standard, no source video, no generated audio | $1.008 |
| 12s | Standard, no source video, generated audio | $1.344 |
| 12s | Pro, no source video, no generated audio | $1.344 |
| 12s | Pro, no source video, generated audio | $1.680 |
| 12s | Standard with source video | $1.512 |
| 12s | Pro with source video | $2.016 |
| 12s | 4K | $5.040 |
| 13s | Standard, no source video, no generated audio | $1.092 |
| 13s | Standard, no source video, generated audio | $1.456 |
| 13s | Pro, no source video, no generated audio | $1.456 |
| 13s | Pro, no source video, generated audio | $1.820 |
| 13s | Standard with source video | $1.638 |
| 13s | Pro with source video | $2.184 |
| 13s | 4K | $5.460 |
| 14s | Standard, no source video, no generated audio | $1.176 |
| 14s | Standard, no source video, generated audio | $1.568 |
| 14s | Pro, no source video, no generated audio | $1.568 |
| 14s | Pro, no source video, generated audio | $1.960 |
| 14s | Standard with source video | $1.764 |
| 14s | Pro with source video | $2.352 |
| 14s | 4K | $5.880 |
| 15s | Standard, no source video, no generated audio | $1.260 |
| 15s | Standard, no source video, generated audio | $1.680 |
| 15s | Pro, no source video, no generated audio | $1.680 |
| 15s | Pro, no source video, generated audio | $2.100 |
| 15s | Standard with source video | $1.890 |
| 15s | Pro with source video | $2.520 |
| 15s | 4K | $6.300 |
When to Use
| ✅ Good fit | ❌ Consider alternatives |
|---|---|
| The model's named workflow matches the source material and intended output | A different input modality or model route is required |
| A managed asynchronous result is suitable for the production pipeline | A synchronous, interactive editor is essential |
| The documented controls cover the required duration, framing, or format | The project needs controls outside this endpoint's schema |
| Creative iteration benefits from a repeatable request structure | Exact deterministic pixels, frames, geometry, or samples are mandatory |
| A finished downloadable media asset is the desired deliverable | Editable source layers or a native project file are required |
Prompt Guide
For image-conditioned generation, state the intended result first, then add the subject or source treatment, progression, style, and delivery constraints. Keep one creative variable per phrase, use the documented field names for controls, and change one setting at a time when comparing results.
{
"aspect_ratio": "16:9",
"duration": 5,
"image": "https://static.sandbase.ai/examples/mirrored/f77b827ca02a-TNErq9yD7ZxGRATjfAqnh_EIgJSN67.png",
"images": [
"https://example.com/reference.png"
],
"prompt": "Based on @Video1, make the character from @Image1 dance.",
"resolution": "standard",
"video": "https://static.sandbase.ai/examples/mirrored/f1ecbb471a16-hklvF__w53diz6Rve7f5__JuDW2xl0mr6sJ_Kjz3Vxe_vidoeook--1-_1.mp4"
}
Technical Specs
| Spec | Value |
|---|---|
| Model ID | kwaivgi/kling-video/o3 |
| Inputs | aspect_ratio, character_orientation, duration, end_image, generate_audio, image, images, keep_original_sound, multi_prompt, prompt, resolution, shot_type, video |
| Required inputs | prompt |
| Output fields | content_type, url |
| Execution | Async (submit, then poll for result) |
| Duration | 3 / 4 / 5 / 6 / 7 / 8 / 9 / 10 / 11 / 12 / 13 / 14 / 15 |
| Resolution | standard / pro / 4k |
| Aspect Ratio | 16:9 / 9:16 / 1:1 |
Related Models
kwaivgi/kling-video/o3/4k/image-to-video— Compare a nearby route in the same local model family.kwaivgi/kling-video/o3/4k/reference-to-video— Compare a nearby route in the same local model family.kwaivgi/kling-video/o3/4k/text-to-video— Compare a nearby route in the same local model family.
Start building
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
# The quoted heredoc keeps Unicode and shell metacharacters unchanged.
result=$(curl --fail-with-body --silent \
-X POST "https://api.sandbase.ai/v1/run" \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
-H "Content-Type: application/json" \
--data-binary @- <<'SANDBASE_JSON'
{
"model": "kwaivgi/kling-video/o3",
"image": "https://static.sandbase.ai/examples/mirrored/f77b827ca02a-TNErq9yD7ZxGRATjfAqnh_EIgJSN67.png",
"video": "https://static.sandbase.ai/examples/mirrored/f1ecbb471a16-hklvF__w53diz6Rve7f5__JuDW2xl0mr6sJ_Kjz3Vxe_vidoeook--1-_1.mp4",
"prompt": "Based on @Video1, make the character from @Image1 dance.",
"duration": 5,
"shot_type": "customize",
"resolution": "standard",
"aspect_ratio": "16:9",
"generate_audio": true,
"keep_original_sound": true,
"character_orientation": "image"
}
SANDBASE_JSON
)
run_id=$(printf '%s' "$result" | jq -r .id)
for attempt in $(seq 1 120); do
status=$(printf '%s' "$result" | jq -r .status)
case "$status" in completed|failed|timeout) break ;; esac
sleep 2
result=$(curl --fail-with-body --silent \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
"https://api.sandbase.ai/v1/run/$run_id")
done
status=$(printf '%s' "$result" | jq -r .status)
[ "$status" = completed ] || { echo "Generation ended: $status" >&2; exit 1; }
printf '%s\n' "$result"Choose your model
Compare models
| Model | Context | Input / 1M | Output / 1M | Released |
|---|---|---|---|---|
Kling Video O3This model KwaiVGI | — | — | — | Jul 6, 2026 |
KwaiVGI | — | — | — | Jul 6, 2026 |
KwaiVGI | — | — | — | Jul 3, 2026 |
KwaiVGI | — | — | — | Jul 3, 2026 |
KwaiVGI | — | — | — | Jul 3, 2026 |
KwaiVGI | — | — | — | Jul 3, 2026 |
