Seedance 2.0 Reference to Video
bytedance/seedance/2.0/reference-to-videoByteDance's most advanced reference-to-video model generating cinematic video guided by reference content, with native audio, multi-shot editing, and director-level camera control for professional-grade video creation.
- Base price
- $0.74USD / run
- Execution
- async
- Model type
- video
- Input fields
- 7
Try the model
Playground
PNG, JPEG, WebP, or GIF · 20 MiB maximum each
Your output will appear here
Complete the inputs, then click Run.
Specifications
Pricing
- Base price
- $0.74 / run
- Billing formula
- (usage.total_seconds ?? ((params.duration ?? 5) == 'auto' ? 5 : (params.duration ?? 5))) * ((params.resolution ?? '720p') == '480p' ? 0.065519287834 : (params.resolution ?? '720p') == '720p' ? 0.147418397626 : (params.resolution ?? '720p') == '1080p' ? 0.367744807122 : (params.resolution ?? '720p') == '4k' ? 1.800000000000 : (0 / 0)) + (params.has_input_videos ? usage.input_video_seconds : 0) * ((params.resolution ?? '720p') == '480p' ? 0.039881305638 : (params.resolution ?? '720p') == '720p' ? 0.089732937685 : (params.resolution ?? '720p') == '1080p' ? 0.223531157270 : (params.resolution ?? '720p') == '4k' ? 1.107692307692 : (0 / 0))
Context & modalities
- Input
- Schema-defined
- Output
- video
Capabilities
- Chat
- Not supported
- Vision
- Not supported
- Reasoning
- Not supported
- Structured output
- Not supported
- Function calling
- Not supported
- Audio input
- Not supported
Access
- Provider
- Bytedance
- Model ID
- bytedance/seedance/2.0/reference-to-video
- Execution
- async
- API
- Unified Run API
- Endpoint
- /v1/run
API README
Seedance 2.0 Reference to Video
ByteD's flagship multi-reference video generation model — combines images, videos, and audio references to produce cinematic output with precise style and motion control.
- Need text-only generation? Try Seedance 2.0 Text to Video.
- Need single-image driven generation? Try Seedance 2.0 Image to Video.
- Need faster & cheaper? Try Seedance 2.0 Fast.
Highlights
Multi-modal references — Combine reference images, videos, and audio in a single request for rich, guided generation.
Native audio generation — Automatically creates synchronized sound effects, ambient audio, and lip-synced speech. Toggle with generate_audio parameter.
Style transfer — Use reference images and videos to guide the visual style, motion patterns, and composition.
Flexible duration — Generate 4 to 15 second clips in a single request. Use -1 for model-intelligent duration selection.
Up to 4K — 4K output with support for 6 aspect ratios including cinematic 21:9.
Pricing
Non-4K rates use Volcengine's non-promotional list price converted at 1 USD = 6.74 CNY. All 4K output rates use the confirmed reference price.
| Resolution | Output price per second | Input-video price per second |
|---|---|---|
| 480p | $0.07 | $0.04 |
| 720p | $0.15 | $0.09 |
| 1080p | $0.37 | $0.22 |
| 4K | $1.80 | $1.11 |
Default: 5 output seconds at 720p = $0.74. Video references add the input-video rate using the server-probed trusted aggregate input duration. Only the exact Seedance 2.0 standard, Fast, and Mini reference-to-video modes and Seedance 2.5 reference-to-video mode, each with non-empty videos, explicitly use the isolated lenient-duration policy; the default media probe, Webhook validation, empty-video requests, and every other model or caller remain strict. Every legal domain in this isolated policy uses the system resolver, standard HTTP transport, and ProxyFromEnvironment without all-address public-DNS validation, IP pinning, or DNS-rebinding protection. Each redirect hop repeats the base URL and redirect validation. Invalid URL or host, redirect-policy, count-limit, per-item or aggregate size-limit, parser-resource-limit, and successfully parsed aggregate duration-limit failures fail closed. All other fixed failures—dns_unavailable, connect_error, network_timeout, tls_error, http_status, download_error, empty_body, mime_mismatch, magic_mismatch, unsupported_media, malformed_media, and duration_unavailable—use the 15-second family cap for the whole request once, not per video. Logs and errors expose only fixed stage/reason, video index/count, host hash, bytes, elapsed time, fallback marker, and policy version; complete URLs, queries, signatures, response bodies, and media content remain redacted. New requests freeze trusted-probe policy v4; existing v3, v2, and legacy predictions retain their frozen settlement semantics. This isolated application policy requires no new server, Kubernetes, proxy, egress, or environment configuration. Image and audio references do not add this video-input charge. Audio generation is included at no extra cost.
When to Use
| ✅ Good fit | ❌ Consider alternatives |
|---|---|
| Brand-consistent video production | Real-time video streaming |
| Style-guided content creation | Long-form video (>15s) |
| Video-to-video style transfer | Simple text prompts (use T2V) |
| Audio-synced video generation | Single image animation (use I2V) |
| Multi-reference scene composition | Quick prototyping without assets |
Input Requirements
- Reference images/videos/audio as publicly accessible URLs
- Supported image formats: PNG, JPEG, WebP (< 5MB)
- Supported video formats: MP4 (< 50MB)
- Supported audio formats: MP3, WAV (< 50MB)
- Audio cannot be used alone — must pair with image, video, or text
Technical Specs
| Spec | Value |
|---|---|
| Input | Text + reference images/videos/audio |
| Output | MP4 video with optional audio |
| Resolution | 480p / 720p / 1080p / 4K |
| Aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 |
| Duration | 4–15 seconds |
| Execution | Async (submit → poll for result) |
| Typical latency | 2–8 minutes |
Related
- Seedance 2.0 Text to Video — Generate video from text prompts
- Seedance 2.0 Image to Video — Single image to video
- Seedance 2.0 Fast Reference to Video — ~20% cheaper, lower latency
Start building
Send your first request
OpenAI-compatible endpoint with unified authentication and usage tracking.
# The quoted heredoc keeps Unicode and shell metacharacters unchanged.
result=$(curl --fail-with-body --silent \
-X POST "https://api.sandbase.ai/v1/run" \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
-H "Content-Type: application/json" \
--data-binary @- <<'SANDBASE_JSON'
{
"model": "bytedance/seedance/2.0/reference-to-video",
"images": [
"https://static.sandbase.ai/examples/bytedance/seedance/2.0/reference-to-video/input_image_safe_ocean_0.jpg"
],
"prompt": "Use Image1 as the visual reference and Video1 as the motion reference. Create a peaceful cinematic shoreline scene with rolling ocean waves at golden hour, soft foam, natural water movement, and no people.",
"videos": [
"https://static.sandbase.ai/examples/pixverse/sound-effects/input_video_0.mp4"
],
"duration": 5,
"resolution": "720p"
}
SANDBASE_JSON
)
run_id=$(printf '%s' "$result" | jq -r .id)
for attempt in $(seq 1 120); do
status=$(printf '%s' "$result" | jq -r .status)
case "$status" in completed|failed|timeout) break ;; esac
sleep 2
result=$(curl --fail-with-body --silent \
-H "Authorization: Bearer $SANDBASE_API_KEY" \
"https://api.sandbase.ai/v1/run/$run_id")
done
status=$(printf '%s' "$result" | jq -r .status)
[ "$status" = completed ] || { echo "Generation ended: $status" >&2; exit 1; }
printf '%s\n' "$result"Choose your model
Compare models
| Model | Context | Input / 1M | Output / 1M | Released |
|---|---|---|---|---|
Seedance 2.0 Reference to VideoThis model Bytedance | — | — | — | Apr 1, 2026 |
Bytedance | — | — | — | Jul 1, 2026 |
Bytedance | — | — | — | Jul 1, 2026 |
Bytedance | — | — | — | Jul 1, 2026 |
Bytedance | — | — | — | Jun 24, 2026 |
Bytedance | — | — | — | Jun 24, 2026 |
