seedance-v2

Hasilkan video pendek sinematik dengan ByteDance Seedance 2.0 Pro di RunComfy. Mendokumentasikan keunggulan Seedance 2.0 Pro (referensi multimodal — hingga 9 gambar, 3 video, 3 audio — audio tersinkronisasi dalam video dengan lip-sync alami, penyempurnaan

npx skills add https://github.com/runcomfy-com/skills --skill seedance-v2

Seedance 2.0 Pro — Pro Pack on RunComfy

runcomfy.com · Seedance 2.0 Pro · GitHub

ByteDance Seedance 2.0 Pro — multimodal cinematic video generator with native lip-synced audio — hosted on the RunComfy Model API.

npx skills add agentspace-so/runcomfy-skills --skill seedance-v2 -g

When to pick this model (vs siblings)

Seedance 2.0 Pro's distinct strength is multi-modal cinematic short-form: combine character images + scene videos + reference audio into one coherent shot. Pick it when fidelity to a reference identity / scene matters and you want native lip-sync.

You wantUse
Lip-synced spokesperson / dialogue adSeedance 2.0 Pro
Multi-modal references (image + video + audio)Seedance 2.0 Pro
Brand-consistent multi-language narrativeSeedance 2.0 Pro
Currently-#1 blind-vote video qualityHappyHorse 1.0
Audio-driven lip-sync from your own trackWan 2.7 (audio_url)
Motion editing on existing footageKling Video O1
Ultra-fast iterationLTX 2

If the user said "Seedance" / "Seedance 2" / "ByteDance video" explicitly, route here regardless.

Prerequisites

  1. RunComfy CLInpm i -g @runcomfy/cli
  2. RunComfy accountruncomfy login opens a browser device-code flow.
  3. CI / containers — set RUNCOMFY_TOKEN=<token> instead of runcomfy login.

Endpoints + input schema

bytedance/seedance-v2/pro

FieldTypeRequiredDefaultNotes
promptstringyesCN ≤ 500 chars OR EN ≤ 1000 words.
image_urlarrayno[]0–9 references (JPEG/PNG/WebP/BMP/TIFF/GIF).
video_urlarrayno[]0–3 clips (MP4/MOV), 2–15s each.
audio_urlarrayno[]0–3 audio refs (WAV/MP3), 2–15s, < 15MB each.
aspect_ratioenumnoadaptiveadaptive, 16:9, 9:16, 4:3, 3:4, 1:1, 21:9.
durationintno54–15 (whole seconds).
resolutionenumno720p480p or 720p.
generate_audioboolnotrueIn-pass synchronized speech / SFX / music.
seedintnoReproducibility.

How to invoke

Default (text only, 5s, 720p with audio):

runcomfy run bytedance/seedance-v2/pro \
  --input '{"prompt": "<user prompt>"}' \
  --output-dir <absolute/path>

Lip-synced ad with character reference (image-stable, text-evolves):

runcomfy run bytedance/seedance-v2/pro \
  --input '{
    "prompt": "Medium close-up. The woman explains today'\''s special in a warm friendly tone, slow push-in, soft window light, gentle cafe ambience.",
    "image_url": ["https://.../barista-headshot.jpg"],
    "duration": 8,
    "aspect_ratio": "9:16"
  }' \
  --output-dir <absolute/path>

Multi-modal (image + video + audio refs):

runcomfy run bytedance/seedance-v2/pro \
  --input '{
    "prompt": "Subject from image 1 walks through the café from video 1, voice tone matches audio 1.",
    "image_url": ["https://.../subject.jpg"],
    "video_url": ["https://.../cafe-locked-shot.mp4"],
    "audio_url": ["https://.../voice-ref.mp3"]
  }' \
  --output-dir <absolute/path>

The CLI submits, polls, fetches the result, downloads *.runcomfy.net/*.runcomfy.com URLs into --output-dir.

Prompting — what actually works

Image vs text division. This is the single most important rule. Stable identity (face, costume, brand mark, logo) → put in image_url. Evolving narrative (action, mood, lighting, camera) → put in prompt. Trying to verbally describe a face in detail wastes tokens and produces drift.

Camera + motion in plain language. "Medium close-up", "slow push-in", "handheld follow", "locked-off wide" all work as directives. Combine: "Medium close-up. Slow push-in over 3 seconds. Handheld, slight breathing motion."

Audio direction with generate_audio: true — say the tone: "warm friendly conversational", "calm instructional", "crisp newsroom delivery". For ambient: "gentle cafe chatter, distant traffic, no foreground music".

Reference media specs — videos must be 2–15s; audio must be ≤15MB and 2–15s. Out-of-range files reject. Match aspect ratio of refs to your output to avoid crops.

Anti-patterns:

  • Mixing radically different aesthetic refs (watercolor + photoreal) → confuses.
  • Conflicting style cues in prompt → simplify by removing contradictions.
  • Trying to describe stable identity verbally → use image_url instead.
  • Asking for >15s clips → 422; segment into multiple calls.

Where it shines

Use caseWhy Seedance 2.0 Pro
Spokesperson / dialogue adsNative in-pass lip-sync, no separate TTS step
Brand-consistent multi-language narrativesImage refs hold identity; text drives translation
Cinematic short-form film previsCamera-shot grammar + multi-modal refs
Ad creatives with reference music / VO toneAudio refs guide voice / mood without locking lip-sync
Reproducible variant testingSeed control + fixed schema

Sample prompts (verified to produce strong results)

Default playground example:

Golden hour on a quiet cafe terrace: a barista wipes the counter, then
looks up and explains today's special in a friendly tone, natural
lip-sync. Medium close-up, slow push-in; warm side light, soft bokeh
through glass, gentle cafe ambience and subtle film grain.

Multi-modal lip-sync (text + image):

Same person as image 1 in a softly-lit recording booth, leaning into
the mic, says: "We just shipped the biggest update of the year."
Calm conversational tone. Medium close-up, locked tripod, shallow DOF,
warm key light from camera-left.

Limitations

  • Duration 4–15s — no longer clips on this endpoint.
  • Resolution ceiling 720p on the playground variant.
  • Reference media specs — videos / audio must be 2–15s; audio < 15MB.
  • Lip-sync quality — depends on prompt clarity; not guaranteed perfect under all conditions.
  • No @-syntax for character binding — relies on image refs + prompt alignment.

Exit codes

codemeaning
0success
64bad CLI args
65bad input JSON / schema mismatch
69upstream 5xx
75retryable: timeout / 429
77not signed in or token rejected

Full reference: docs.runcomfy.com/cli/troubleshooting.

How it works

The skill invokes runcomfy run bytedance/seedance-v2/pro with a JSON body matching the schema. The CLI POSTs to https://model-api.runcomfy.net/v1/models/bytedance/seedance-v2/pro, polls the request, fetches the result, and downloads any .runcomfy.net/.runcomfy.com URL into --output-dir. Ctrl-C cancels the remote request before exit.

Security & Privacy

  • Token storage: runcomfy login writes the API token to ~/.config/runcomfy/token.json with mode 0600 (owner-only read/write). Set RUNCOMFY_TOKEN env var to bypass the file entirely in CI / containers.
  • Input boundary: the user prompt is passed as a JSON string to the CLI via --input. The CLI does NOT shell-expand the prompt; it transmits the JSON body directly to the Model API over HTTPS. No shell injection surface from prompt content.
  • Third-party content: image / mask / video URLs you pass are fetched by the RunComfy model server, not by the CLI on your machine. Treat external URLs as untrusted; image-based prompt injection is a known risk for any image-edit / video-edit model.
  • Outbound endpoints: only model-api.runcomfy.net (request submission) and *.runcomfy.net / *.runcomfy.com (download whitelist for generated outputs). No telemetry, no callbacks.
  • Generated-file size cap: the CLI aborts any single download > 2 GiB to prevent disk-fill from a malicious or runaway model output.

Lebih banyak skill dari runcomfy-com

video-inpainting
runcomfy-com
Region edits across video frames on RunComfy via the `runcomfy` CLI — remove an object that appears across many frames, clean up wires or watermarks, replace a region with matching motion. Routes across Wan 2-7 edit-video (default, prompt-driven region edits with spatial language), Lucy Edit Restyle (identity-stable region-aware restyle), and Seedream 4-0 edit-sequential (when treating the clip as a frame stack). Picks the right route based on whether the change is prose-driven,...
videocreativemedia
runcomfy-cli
runcomfy-com
We need to translate the given English text into Bahasa Indonesia. The text describes a CLI tool called runcomfy-cli (name preserved). The instruction says to translate only the text inside <text>, not include the name unless it appears in source. The name "runcomfy-cli" is not in the source text; the source mentions "runcomfy CLI" but not "runcomfy-cli". However, the directory item type is "agent skill" and name to preserve is "runcomfy-cli". But the instruction says "Do not include the name unless it appears in the source text." The source text has "runcomfy CLI" (with space) not "runcomfy-cli". So we should not add "runcomfy-cli" in the translation. We'll translate the text as is, preserving "runcomfy CLI" as a proper name. Also preserve URLs, numbers, technical terms. The text has "ControlNet", "LoRA", etc. Keep them as is. Translate: "Run any model on RunComfy
creativemediaapi
image-to-video
runcomfy-com
We need to translate the given text from English to Bahasa Indonesia. The text describes a skill called "image-to-video" but the instruction says not to include the name unless it appears in the source text. The name "image-to-video" does appear in the source? Actually the source text starts with "Animate any still image on RunComfy — this skill is a smart router..." The name is not explicitly written as "image-to-video" in the source text. The directory item type is "agent skill" and the name to preserve is "image-to-video". But the instruction says: "Do not include the name unless it appears in the source text." The name "image-to-video" does not appear in the source text. So we should not add it. We just translate the text inside <text>. We need to preserve product names, protocol names, URLs, numbers, technical terms. So "RunComfy", "HappyHorse 1.0 I2V", "Arena #1", "Wan 2.7",
creativevideomedia
flux-2-klein
runcomfy-com
Hasilkan gambar dengan Flux 2 Klein (varian cepat sulingan Flux 2 dari Black Forest Labs) di RunComfy — dibundel dengan pola prompt terdokumentasi model tersebut sehingga skill menghasilkan output yang lebih tajam dibandingkan prompt naif terhadap model yang sama. Mendokumentasikan keunggulan Flux 2 Klein (latensi sub-detik, gaya merek multi-referensi, prompt deklaratif subjek-pertama), strategi jumlah langkah (4–8 untuk iterasi cepat, ~25 untuk polesan), trade-off varian 9B vs 4B, dan kapan harus mengarahkan ke Flux 2 Pro /...
creativeimageresearch
ai-avatar-video
runcomfy-com
Create AI avatar, talking-head, and lip-sync videos on RunComfy via the `runcomfy` CLI. Routes across ByteDance OmniHuman (audio-driven full-body avatar), Wan-AI Wan 2-7 (audio-driven mouth sync via `audio_url` on a portrait), HappyHorse 1.0 (Arena #1 t2v / i2v with in-pass audio), and Seedance v2 Pro (multi-modal cinematic with reference audio + reference subject). Picks the right model for the user's actual intent — UGC voiceover, virtual presenter, dubbed product demo, lip-synced...
videocreativemedia
nano-banana-edit
runcomfy-com
Edit gambar dengan Google Nano Banana 2 (endpoint edit gambar-ke-gambar) di RunComfy. Mendokumentasikan kelebihan Nano Banana Edit (mempertahankan identitas subjek,
creativeimageapi
wan-2-7
runcomfy-com
We need to translate the given text from English to Bahasa Indonesia. The text describes an agent skill for generating text-to-video using Wan 2.7. We must preserve product names, protocol names, URLs, numbers, technical terms. The name "wan-2-7" is to be preserved if it appears in the source text. The instruction says: "Do not include the name unless it appears in the source text." The name appears multiple times: "Wan 2.7", "wan-2-7", "wan". So we keep those as is. Also "RunComfy", "HappyHorse 1.0", "Seedance 2.0", "Kling", "LTX 2", "CLI", "audio_url", "runcomfy run wan-ai/wan-2-7/text-to-video" are technical terms/names to preserve. We translate the rest naturally. Let's break down the text: "Generate text-to-video with Wan 2.7 (Wan-AI's flagship motion model
creativevideomedia
lipsync
runcomfy-com
We need to translate the given English text into Indonesian. The text describes a skill for lip-syncing. We must preserve product names, protocol names, URLs, numbers, technical terms. The name "lipsync" is not in the text? Actually it appears as "Lip-sync" at the beginning. The instruction says "Do not include the name unless it appears in the source text." The name "lipsync" appears as "Lip-sync" in the source text. So we should translate that as well? But the instruction says "Preserve product names" and "Name to preserve: lipsync". So we should keep "lipsync" as is? But the source has "Lip-sync" with hyphen and capital L. Probably we should keep it as "Lip-sync" or "lipsync"? The instruction says "Name to preserve: lipsync" (lowercase). But in the text it's "Lip-sync". I think we should preserve the exact form as in the source: "Lip-sync". However, the instruction says
creativevideomedia