Kling Video O3 Pro

kwaivgi/kling-video/o3/pro/image-to-video

Kling Video O3 Pro is KwaiVGI's image-to-video AI model. Bring static images to life with fluid animation, consistent character motion, and professional-grade video output.

Base price
$0.70USD / run
Execution
async
Model type
video
Input fields
7

Try the model

Playground

Open playground
Input
Text prompt for video generation. Either prompt or multi_prompt must be provided, but not both.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of the start frame image.

PNG, JPEG, WebP, or GIF · 20 MiB maximum

URL of the end frame image (optional).
Video duration in seconds (3-15s). Allowed values: 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15.
Whether to generate native audio for the video.
List of prompts for multi-shot video generation.
The type of multi-shot video generation. 'intelligent' lets the model automatically determine shot structure. Allowed values: customize, intelligent.
OutputReady

Your output will appear here

Complete the inputs, then click Run.

Specifications

Pricing

Base price
$0.70 / run
Billing formula
params.generate_audio ? params.duration * 0.14 : params.duration * 0.112

Context & modalities

Input
Schema-defined
Output
video

Capabilities

Chat
Not supported
Vision
Not supported
Reasoning
Not supported
Structured output
Not supported
Function calling
Not supported
Audio input
Not supported

Access

Provider
KwaiVGI
Model ID
kwaivgi/kling-video/o3/pro/image-to-video
Execution
async
API
Unified Run API
Endpoint
/v1/run

API README

Kling Video O3 Pro

Kling Video O3 Pro is a image-led video-generation route in the Kling Video O3 family. It turns written scene direction and a starting image into a multi-second sequence with controlled subject action, camera movement, shot structure, and optional synchronized sound.

The Pro route is intended for creators who need dependable temporal continuity across a directed clip. Prompts can describe staging, motion, lens behavior, atmosphere, and audio cues; multi-prompt controls can divide the sequence into several shots while reference inputs keep important visual elements anchored.

Highlights

  • Image-anchored animation. Uses the supplied opening image to preserve composition and identity as motion develops.
  • Multi-shot direction. Supports structured prompt segments for sequences that need more than one planned shot.
  • Native audio option. Can generate sound together with the visible action when the route exposes audio generation.
  • Professional-detail output. Uses the Pro route when finer visual fidelity is more important than the lower-cost Standard tier.

Pricing

ConfigurationBilling unitPrice
3 secondsWithout generated audio$0.336
3 secondsWith generated audio$0.420
4 secondsWithout generated audio$0.448
4 secondsWith generated audio$0.560
5 secondsWithout generated audio$0.560
5 secondsWith generated audio$0.700
6 secondsWithout generated audio$0.672
6 secondsWith generated audio$0.840
7 secondsWithout generated audio$0.784
7 secondsWith generated audio$0.980
8 secondsWithout generated audio$0.896
8 secondsWith generated audio$1.120
9 secondsWithout generated audio$1.008
9 secondsWith generated audio$1.260
10 secondsWithout generated audio$1.120
10 secondsWith generated audio$1.400
11 secondsWithout generated audio$1.232
11 secondsWith generated audio$1.540
12 secondsWithout generated audio$1.344
12 secondsWith generated audio$1.680
13 secondsWithout generated audio$1.456
13 secondsWith generated audio$1.820
14 secondsWithout generated audio$1.568
14 secondsWith generated audio$1.960
15 secondsWithout generated audio$1.680
15 secondsWith generated audio$2.100

When to Use

✅ Good fit❌ Consider alternatives
The project needs creating a directed video sequenceThe goal is a different media task or endpoint
The available inputs match the required local schemaRequired source media or permissions are unavailable
The brief can specify subject, composition, style, and deliveryThe result must be deterministic at pixel or sample level
The supported formats and controls match final placementDelivery requires unsupported dimensions, codecs, or duration
An asynchronous generated result fits the workflowA live, frame-synchronous, or real-time response is mandatory

Prompt Guide

Write the request as a production brief: identify the main subject or source material, state the intended transformation, describe composition or timing, and finish with style, atmosphere, and delivery constraints. Keep preservation requirements separate from requested changes, and use only fields exposed by this route.

{
  "prompt": "A cinematic, precisely directed scene with clear subject action, camera movement, lighting, and atmosphere",
  "image": "https://example.com/reference.jpg",
  "duration": 3,
  "end_image": "https://example.com/reference.jpg",
  "shot_type": "customize"
}

Technical Specs

SpecValue
Model IDkwaivgi/kling-video/o3/pro/image-to-video
Input fieldsimage (string)<br>prompt (string)<br>duration (integer; 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15)<br>end_image (string)<br>shot_type (string; customize, intelligent)<br>multi_prompt (array)<br>generate_audio (boolean)
Required inputprompt, image
Output fieldsurl, content_type
ExecutionAsynchronous job

Related Models

  • kwaivgi/kling-video/o3/standard/reference-to-video
  • kwaivgi/kling-video/o3/standard/text-to-video
  • kwaivgi/kling-video/o3/pro/text-to-video

Start building

Send your first request

OpenAI-compatible endpoint with unified authentication and usage tracking.

Production API
Unified Run API endpoint
https://api.sandbase.ai/v1/run
Model ID
kwaivgi/kling-video/o3/pro/image-to-video
# The quoted heredoc keeps Unicode and shell metacharacters unchanged.
result=$(curl --fail-with-body --silent \
  -X POST "https://api.sandbase.ai/v1/run" \
  -H "Authorization: Bearer $SANDBASE_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @- <<'SANDBASE_JSON'
{
  "model": "kwaivgi/kling-video/o3/pro/image-to-video",
  "image": "https://static.sandbase.ai/examples/mirrored/0ff4703452ec-8ABMp4n9rh3kfD2Rq8fHd_start_frame.png",
  "prompt": "The character walks forward slowly, with the camera following from behind.",
  "duration": 10,
  "end_image": "https://static.sandbase.ai/examples/alibaba/qwen-image-edit-plus-lora-gallery/add-background/output_url_1.png",
  "shot_type": "customize",
  "generate_audio": false
}
SANDBASE_JSON
)
run_id=$(printf '%s' "$result" | jq -r .id)
for attempt in $(seq 1 120); do
  status=$(printf '%s' "$result" | jq -r .status)
  case "$status" in completed|failed|timeout) break ;; esac
  sleep 2
  result=$(curl --fail-with-body --silent \
    -H "Authorization: Bearer $SANDBASE_API_KEY" \
    "https://api.sandbase.ai/v1/run/$run_id")
done
status=$(printf '%s' "$result" | jq -r .status)
[ "$status" = completed ] || { echo "Generation ended: $status" >&2; exit 1; }
printf '%s\n' "$result"

Choose your model

Compare models

All language models
ModelContextInput / 1MOutput / 1MReleased
KwaiVGI———Feb 4, 2026
KwaiVGI———Jul 6, 2026
KwaiVGI———Jul 6, 2026
KwaiVGI———Jul 3, 2026
KwaiVGI———Jul 3, 2026
KwaiVGI———Jul 3, 2026