Seedance 2.0 Reference to Video

bytedance/seedance/2.0/reference-to-video

ByteDance's most advanced reference-to-video model generating cinematic video guided by reference content, with native audio, multi-shot editing, and director-level camera control for professional-grade video creation.

Base price
$0.74USD / run
Execution
async
Model type
video
Input fields
7

Try the model

Playground

Open playground
Input
The text prompt to generate a video from.

PNG, JPEG, WebP, or GIF · 20 MiB maximum each

The URLs of reference images to guide video generation.
The URLs of reference videos to guide video generation. Up to 3 MP4 or MOV files, 2-15 seconds combined duration, under 50 MB total.
The URLs of reference audio files to guide video generation.
The duration of the generated video in seconds. Allowed values: 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15.
The aspect ratio of the generated video. Allowed values: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9.
The resolution of the video to generate. Allowed values: 480p, 720p, 1080p, 4k.
OutputReady

Your output will appear here

Complete the inputs, then click Run.

Specifications

Pricing

Base price
$0.74 / run
Billing formula
(usage.total_seconds ?? ((params.duration ?? 5) == 'auto' ? 5 : (params.duration ?? 5))) * ((params.resolution ?? '720p') == '480p' ? 0.065519287834 : (params.resolution ?? '720p') == '720p' ? 0.147418397626 : (params.resolution ?? '720p') == '1080p' ? 0.367744807122 : (params.resolution ?? '720p') == '4k' ? 1.800000000000 : (0 / 0)) + (params.has_input_videos ? usage.input_video_seconds : 0) * ((params.resolution ?? '720p') == '480p' ? 0.039881305638 : (params.resolution ?? '720p') == '720p' ? 0.089732937685 : (params.resolution ?? '720p') == '1080p' ? 0.223531157270 : (params.resolution ?? '720p') == '4k' ? 1.107692307692 : (0 / 0))

Context & modalities

Input
Schema-defined
Output
video

Capabilities

Chat
Not supported
Vision
Not supported
Reasoning
Not supported
Structured output
Not supported
Function calling
Not supported
Audio input
Not supported

Access

Provider
Bytedance
Model ID
bytedance/seedance/2.0/reference-to-video
Execution
async
API
Unified Run API
Endpoint
/v1/run

API README

Seedance 2.0 Reference to Video

ByteD's flagship multi-reference video generation model — combines images, videos, and audio references to produce cinematic output with precise style and motion control.

Highlights

Multi-modal references — Combine reference images, videos, and audio in a single request for rich, guided generation.

Native audio generation — Automatically creates synchronized sound effects, ambient audio, and lip-synced speech. Toggle with generate_audio parameter.

Style transfer — Use reference images and videos to guide the visual style, motion patterns, and composition.

Flexible duration — Generate 4 to 15 second clips in a single request. Use -1 for model-intelligent duration selection.

Up to 4K — 4K output with support for 6 aspect ratios including cinematic 21:9.

Pricing

Non-4K rates use Volcengine's non-promotional list price converted at 1 USD = 6.74 CNY. All 4K output rates use the confirmed reference price.

ResolutionOutput price per secondInput-video price per second
480p$0.07$0.04
720p$0.15$0.09
1080p$0.37$0.22
4K$1.80$1.11

Default: 5 output seconds at 720p = $0.74. Video references add the input-video rate using the server-probed trusted aggregate input duration. Only the exact Seedance 2.0 standard, Fast, and Mini reference-to-video modes and Seedance 2.5 reference-to-video mode, each with non-empty videos, explicitly use the isolated lenient-duration policy; the default media probe, Webhook validation, empty-video requests, and every other model or caller remain strict. Every legal domain in this isolated policy uses the system resolver, standard HTTP transport, and ProxyFromEnvironment without all-address public-DNS validation, IP pinning, or DNS-rebinding protection. Each redirect hop repeats the base URL and redirect validation. Invalid URL or host, redirect-policy, count-limit, per-item or aggregate size-limit, parser-resource-limit, and successfully parsed aggregate duration-limit failures fail closed. All other fixed failures—dns_unavailable, connect_error, network_timeout, tls_error, http_status, download_error, empty_body, mime_mismatch, magic_mismatch, unsupported_media, malformed_media, and duration_unavailable—use the 15-second family cap for the whole request once, not per video. Logs and errors expose only fixed stage/reason, video index/count, host hash, bytes, elapsed time, fallback marker, and policy version; complete URLs, queries, signatures, response bodies, and media content remain redacted. New requests freeze trusted-probe policy v4; existing v3, v2, and legacy predictions retain their frozen settlement semantics. This isolated application policy requires no new server, Kubernetes, proxy, egress, or environment configuration. Image and audio references do not add this video-input charge. Audio generation is included at no extra cost.

When to Use

✅ Good fit❌ Consider alternatives
Brand-consistent video productionReal-time video streaming
Style-guided content creationLong-form video (>15s)
Video-to-video style transferSimple text prompts (use T2V)
Audio-synced video generationSingle image animation (use I2V)
Multi-reference scene compositionQuick prototyping without assets

Input Requirements

  • Reference images/videos/audio as publicly accessible URLs
  • Supported image formats: PNG, JPEG, WebP (< 5MB)
  • Supported video formats: MP4 (< 50MB)
  • Supported audio formats: MP3, WAV (< 50MB)
  • Audio cannot be used alone — must pair with image, video, or text

Technical Specs

SpecValue
InputText + reference images/videos/audio
OutputMP4 video with optional audio
Resolution480p / 720p / 1080p / 4K
Aspect ratios16:9, 9:16, 1:1, 4:3, 3:4, 21:9
Duration4–15 seconds
ExecutionAsync (submit → poll for result)
Typical latency2–8 minutes

Related

Start building

Send your first request

OpenAI-compatible endpoint with unified authentication and usage tracking.

Production API
Unified Run API endpoint
https://api.sandbase.ai/v1/run
Model ID
bytedance/seedance/2.0/reference-to-video
# The quoted heredoc keeps Unicode and shell metacharacters unchanged.
result=$(curl --fail-with-body --silent \
  -X POST "https://api.sandbase.ai/v1/run" \
  -H "Authorization: Bearer $SANDBASE_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @- <<'SANDBASE_JSON'
{
  "model": "bytedance/seedance/2.0/reference-to-video",
  "images": [
    "https://static.sandbase.ai/examples/bytedance/seedance/2.0/reference-to-video/input_image_safe_ocean_0.jpg"
  ],
  "prompt": "Use Image1 as the visual reference and Video1 as the motion reference. Create a peaceful cinematic shoreline scene with rolling ocean waves at golden hour, soft foam, natural water movement, and no people.",
  "videos": [
    "https://static.sandbase.ai/examples/pixverse/sound-effects/input_video_0.mp4"
  ],
  "duration": 5,
  "resolution": "720p"
}
SANDBASE_JSON
)
run_id=$(printf '%s' "$result" | jq -r .id)
for attempt in $(seq 1 120); do
  status=$(printf '%s' "$result" | jq -r .status)
  case "$status" in completed|failed|timeout) break ;; esac
  sleep 2
  result=$(curl --fail-with-body --silent \
    -H "Authorization: Bearer $SANDBASE_API_KEY" \
    "https://api.sandbase.ai/v1/run/$run_id")
done
status=$(printf '%s' "$result" | jq -r .status)
[ "$status" = completed ] || { echo "Generation ended: $status" >&2; exit 1; }
printf '%s\n' "$result"

Choose your model

Compare models

All language models
ModelContextInput / 1MOutput / 1MReleased
Bytedance———Apr 1, 2026
Bytedance———Jul 1, 2026
Bytedance———Jul 1, 2026
Bytedance———Jul 1, 2026
Bytedance———Jun 24, 2026
Bytedance———Jun 24, 2026