Kling Video V3 Standard
kwaivgi/kling-video/v3/standard/text-to-videoKling Video V3 Standard is KwaiVGI's text-to-video AI model. Turn written scripts and prompts into professional-quality video clips with realistic motion, lighting, and scene composition.
- Base price
- $0.63USD / run
- Execution
- async
- Model type
- video
- Input fields
- 6
Try the model
Playground
Your output will appear here
Complete the inputs, then click Run.
Specifications
Pricing
- Base price
- $0.63 / run
- Billing formula
- params.duration * 0.126
Context & modalities
- Input
- Schema-defined
- Output
- video
Capabilities
- Chat
- Not supported
- Vision
- Not supported
- Reasoning
- Not supported
- Structured output
- Not supported
- Function calling
- Not supported
- Audio input
- Not supported
Access
- Provider
- KwaiVGI
- Model ID
- kwaivgi/kling-video/v3/standard/text-to-video
- Execution
- async
- API
- Unified Run API
- Endpoint
- /v1/run
API README
Kling Video V3 Standard Text-to-Video
Kling Video V3 Standard generates video directly from a written scene. The prompt describes the subject, action, camera, light, atmosphere, and sound without requiring an opening image. Landscape, portrait, and square frame shapes are available.
The route produces clips from 3 to 15 seconds and returns a video URL. Native audio is enabled by default, with documented language handling in the local contract. Custom and intelligent shot structure are exposed as request settings.
Highlights
- Cinematic visual quality. The exact Standard text-to-video route is documented for cinematic visuals. This is the route's stated visual positioning.
- Fluid motion. Fluid motion is explicitly identified for Standard text-to-video. Movement is generated from the written scene direction.
- Improved visual quality. V3 Standard text-to-video documentation identifies improved visual quality. The statement is scoped to this exact version, tier, and route.
- Improved motion consistency. Motion consistency is also identified as improved for V3 Standard text-to-video. It describes the generated motion rather than a schema control.
Pricing
| Duration | Price |
|---|---|
| 3 seconds | $0.378 |
| 4 seconds | $0.504 |
| 5 seconds | $0.630 |
| 6 seconds | $0.756 |
| 7 seconds | $0.882 |
| 8 seconds | $1.008 |
| 9 seconds | $1.134 |
| 10 seconds | $1.260 |
| 11 seconds | $1.386 |
| 12 seconds | $1.512 |
| 13 seconds | $1.638 |
| 14 seconds | $1.764 |
| 15 seconds | $1.890 |
When to Use
| ✅ Good fit | ❌ Consider alternatives |
|---|---|
| Turn a written scene into a short video without an image input. | Preserve a specific opening composition; use image-to-video. |
| Describe action and camera movement in one continuous prompt. | Require a clip longer than 15 seconds in one request. |
| Deliver landscape, vertical, or square video. | Require an aspect ratio outside the documented set. |
| Include generated ambience, effects, or speech. | Require voice behavior outside the documented language handling. |
| Use the Standard tier for this text-led workflow. | Choose Pro when its exact tier capabilities are required. |
Prompt Guide
Write one continuous scene in temporal order. Start with the required prompt; add duration, aspect ratio, audio, or shot settings only when the request needs them.
Scene: [setting, time, atmosphere]
Subject: [appearance and action]
Camera: [framing and movement]
Lighting: [direction and change]
Motion: [subject and environment]
Audio: [ambience, effects, or speech when enabled]
{
"prompt": "Cinematic drone shot through ancient stone ruins at golden hour. The camera rises through archways and reveals a misty valley.",
"duration": 5,
"aspect_ratio": "16:9",
"generate_audio": true
}
Technical Specs
| Spec | Value |
|---|---|
| Required input | prompt |
| Prompt length | Up to 2,500 characters |
| Duration | 3–15 seconds; default 5 |
| Aspect ratios | 16:9, 9:16, 1:1 |
| Shot types | customize, intelligent |
| Multi-prompt field | Exposed locally, but prompt remains required by the request contract |
| Native audio | Optional; default enabled |
| Output | Video URL with optional file metadata |
Related
- Kling Video V3 Standard Image-to-Video — Animate a supplied opening image.
- Kling Video V3 Pro Text-to-Video — Compare the sibling tier.
Start building
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
# The quoted heredoc keeps Unicode and shell metacharacters unchanged.
result=$(curl --fail-with-body --silent \
-X POST "https://api.sandbase.ai/v1/run" \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
-H "Content-Type: application/json" \
--data-binary @- <<'SANDBASE_JSON'
{
"model": "kwaivgi/kling-video/v3/standard/text-to-video",
"prompt": "Cinematic drone shot flying through ancient stone ruins covered in moss and vines at golden hour. Camera starts low, rises through crumbling archways, revealing a vast misty valley beyond. Volumetric light rays pierce through gaps in the stone. Epic scale, photorealistic, 8K quality.",
"duration": 5,
"shot_type": "customize",
"generate_audio": true
}
SANDBASE_JSON
)
run_id=$(printf '%s' "$result" | jq -r .id)
for attempt in $(seq 1 120); do
status=$(printf '%s' "$result" | jq -r .status)
case "$status" in completed|failed|timeout) break ;; esac
sleep 2
result=$(curl --fail-with-body --silent \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
"https://api.sandbase.ai/v1/run/$run_id")
done
status=$(printf '%s' "$result" | jq -r .status)
[ "$status" = completed ] || { echo "Generation ended: $status" >&2; exit 1; }
printf '%s\n' "$result"Choose your model
Compare models
| Model | Context | Input / 1M | Output / 1M | Released |
|---|---|---|---|---|
Kling Video V3 StandardThis model KwaiVGI | — | — | — | Feb 4, 2026 |
KwaiVGI | — | — | — | Jul 6, 2026 |
KwaiVGI | — | — | — | Jul 6, 2026 |
KwaiVGI | — | — | — | Jul 3, 2026 |
KwaiVGI | — | — | — | Jul 3, 2026 |
KwaiVGI | — | — | — | Jul 3, 2026 |
