image-to-video

bởi qu-skills

Hướng dẫn chuyển đổi ảnh tĩnh thành video: lựa chọn mô hình, tạo chuyển động và di chuyển camera. Bao gồm Wan 2.5 i2v, Seedance, Fabric, Grok Video với cách sử dụng từng loại. Dùng để: tạo hoạt ảnh cho ảnh, tạo video từ ảnh tĩnh, thêm chuyển động, tạo hoạt ảnh sản phẩm. Kích hoạt: ảnh thành video, i2v, tạo hoạt ảnh cho ảnh, ảnh tĩnh thành video, thêm chuyển động vào ảnh, hoạt ảnh ảnh, ảnh thành video, tạo hoạt ảnh cho ảnh tĩnh, wan i2v, image2video, làm ảnh sống động, tạo hoạt ảnh

npx skills add https://github.com/qu-skills/skills --skill image-to-video

Install the belt CLI skill: npx skills add belt-sh/cli

Image to Video

Convert still images to animated videos via inference.sh CLI.

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Generate a still image
belt app run falai/flux-dev-lora --input '{
  "prompt": "serene mountain lake at sunset, snow-capped peaks reflected in still water, golden hour light, landscape photography",
  "width": 1248,
  "height": 832
}'

# Animate it
belt app run falai/wan-2-5-i2v --input '{
  "prompt": "gentle ripples on the lake surface, clouds slowly drifting, warm light shifting, birds flying in the distance",
  "image": "path/to/lake-image.png"
}'

Model Selection

ModelApp IDBest ForMotion Style
Wan 2.5 i2vfalai/wan-2-5-i2vRealistic motion, natural movementPhotorealistic, subtle
WAN-I2V (Pruna)pruna/wan-i2vEconomical, fast, 480p/720pNatural, efficient
Seedance 2.0bytedance/seedance-2-0Up to 1080p, sync audio, all input typesVersatile, high quality
Seedance 2.0 Fastbytedance/seedance-2-0-fastFast variant, same capabilitiesVersatile, fast
Fabric 1.0falai/fabric-1-0Cloth, fabric, liquid, flowing materialsPhysics-based flow
Grok Imagine Videoxai/grok-imagine-videoGeneral animation, text-guidedVersatile

When to Use Each

ScenarioBest ModelWhy
Landscape with water/cloudsWan 2.5 i2vBest at natural, realistic motion
Portrait with subtle expressionWan 2.5 i2vMaintains face fidelity
Product with fabric/clothFabric 1.0Specialized in material physics
Flag waving, curtain flowingFabric 1.0Cloth simulation
Illustrated/artistic imageSeedance 2.0Matches stylized content
General "bring to life"Seedance 2.0Good all-rounder, up to 1080p
Quick test/iterationSeedance 2.0 FastFaster generation

Motion Types

Camera Movement

MovementPrompt KeywordEffect
Push in / Dolly forward"slow dolly forward", "camera pushes in"Increasing intimacy/focus
Pull out / Dolly back"camera pulls back", "slow zoom out"Reveal, context
Pan left/right"camera pans slowly to the right"Scanning, following
Tilt up/down"camera tilts upward"Revealing height
Orbit"camera orbits around the subject"3D exploration
Crane up"camera rises upward"Grand reveal
Static(no camera movement prompt)Subject motion only

Subject Motion

TypePrompt Examples
Natural elements"water rippling", "clouds drifting", "leaves rustling in breeze"
Hair/clothing"hair blowing gently in wind", "dress fabric flowing"
Atmospheric"fog slowly rolling", "dust particles floating in light beams"
Character"person slowly turns to camera", "subtle breathing motion"
Mechanical"gears turning", "clock hands moving"
Liquid"coffee steam rising", "paint dripping", "water pouring"

Prompting Best Practices

The Golden Rule: Subtle > Dramatic

AI video models produce better results with gentle, subtle motion than dramatic action. Requesting too much movement causes distortion and artifacts.

❌ "person running and jumping over obstacles while the camera spins"
✅ "person slowly walking forward, gentle breeze, camera follows alongside"

❌ "explosion with debris flying everywhere"
✅ "candle flame flickering gently, warm ambient light shifting"

❌ "fast zoom into the eyes with dramatic camera shake"
✅ "slow dolly forward toward the subject, subtle focus shift"

Prompt Structure

[Camera movement] + [Subject motion] + [Atmospheric effects] + [Mood/pace]

Examples by Scenario

# Landscape animation
belt app run falai/wan-2-5-i2v --input '{
  "prompt": "gentle camera pan right, water reflecting moving clouds, trees swaying slightly in breeze, warm golden light, peaceful and slow",
  "image": "landscape.png"
}'

# Portrait animation
belt app run falai/wan-2-5-i2v --input '{
  "prompt": "subtle breathing motion, slight head turn, natural eye blink, hair moving gently, soft ambient lighting shifts",
  "image": "portrait.png"
}'

# Product shot animation
belt app run bytedance/seedance-2-0 --input '{
  "prompt": "slow 360 degree orbit around the product, gentle spotlight movement, subtle reflections shifting, premium product showcase, smooth motion",
  "image": "product.png",
  "generate_audio": true
}'

# Fabric/cloth animation
belt app run falai/fabric-1-0 --input '{
  "prompt": "fabric flowing and rippling in gentle wind, natural cloth physics, soft movement",
  "image": "fabric-scene.png"
}'

# Architectural visualization
belt app run falai/wan-2-5-i2v --input '{
  "prompt": "slow dolly forward through the entrance, slight camera tilt upward, ambient light filtering through windows, dust particles in light beams",
  "image": "building-interior.png"
}'

Duration Guidelines

DurationQualityUse For
2-3 secondsHighest qualityGIFs, looping backgrounds, cinemagraphs
4-5 secondsHigh qualitySocial media posts, product reveals
6-8 secondsGood qualityShort clips, transitions
10+ secondsQuality degradesAvoid unless stitching shorter clips

Extending Duration

For longer videos, generate multiple short clips and stitch:

# Generate 3 clips from the same image with progressive motion
belt app run falai/wan-2-5-i2v --input '{
  "prompt": "slow pan left, gentle water motion",
  "image": "scene.png"
}' --no-wait

belt app run falai/wan-2-5-i2v --input '{
  "prompt": "continuing pan, clouds shifting, light changing",
  "image": "scene.png"
}' --no-wait

# Stitch together
belt app run infsh/media-merger --input '{
  "media": ["clip1.mp4", "clip2.mp4"]
}'

The Full Workflow

Still-to-Final-Video Pipeline

# 1. Generate source image (best quality)
belt app run bytedance/seedream-4-5 --input '{
  "prompt": "cinematic landscape, misty mountains at dawn, lake in foreground, dramatic clouds, golden hour, 4K quality, professional photography",
  "size": "2K"
}'

# 2. Animate the image
belt app run falai/wan-2-5-i2v --input '{
  "prompt": "gentle mist rolling through the valley, lake surface rippling, clouds slowly moving, birds in distance, warm light shifting",
  "image": "landscape.png"
}'

# 3. Upscale video if needed
belt app run falai/topaz-video-upscaler --input '{
  "video": "animated-landscape.mp4"
}'

# 4. Add ambient audio
belt app run infsh/hunyuanvideo-foley --input '{
  "video": "animated-landscape.mp4",
  "prompt": "gentle nature ambience, distant birds, soft wind, water lapping"
}'

# 5. Merge video with audio
belt app run infsh/video-audio-merger --input '{
  "video": "upscaled-landscape.mp4",
  "audio": "ambient-audio.mp3"
}'

Cinemagraph Effect

A cinemagraph is a still photo where only one element moves (e.g., waterfall moving in an otherwise frozen scene). To achieve this:

  1. Generate the still image with the motion element clearly defined
  2. Prompt for motion only in that specific element
  3. Keep to 2-4 seconds for seamless looping
belt app run falai/wan-2-5-i2v --input '{
  "prompt": "only the waterfall is moving, everything else remains perfectly still, water cascading smoothly, rest of scene frozen",
  "image": "waterfall-scene.png"
}'

Common Mistakes

MistakeProblemFix
Too much motion requestedDistortion, artifacts, warpingSubtle > dramatic, always
Wrong model for content typePoor resultsUse selection guide above
Clips too long (10s+)Quality degrades significantlyKeep to 3-5 seconds, stitch if needed
No camera movement specifiedRandom/unpredictable motionAlways specify camera behavior
Conflicting motion directionsChaotic, unnaturalOne primary motion direction
Low-res source imageLow-res video outputStart with highest quality source
Complex action scenesModels can't handleKeep motion simple and natural

Related Skills

npx skills add inference-sh/skills@ai-video-generation
npx skills add inference-sh/skills@ai-image-generation
npx skills add inference-sh/skills@p-video
npx skills add inference-sh/skills@video-prompting-guide
npx skills add inference-sh/skills@prompt-engineering

Browse all apps: belt app store

Thêm skills từ qu-skills

ai-video-generation
qu-skills
Tạo video AI với Google Veo, Seedance 2.0, HappyHorse, Wan, Grok và hơn 40 mô hình qua CLI inference.sh. Các mô hình: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Khả năng: văn bản thành video, hình ảnh thành video, tham chiếu thành video, chỉnh sửa video, đồng bộ môi, hoạt ảnh đại diện, nâng cấp video, âm thanh foley. Sử dụng cho: video mạng xã hội, nội dung tiếp thị, video giải thích, demo sản phẩm, đại diện AI. Kích hoạt: tạo video, video AI,...
videocreativemedia
remotion-render
qu-skills
Kết xuất video từ mã thành phần React/Remotion thông qua inference.sh. Nhập mã TSX, nhận MP4. Hỗ trợ tất cả API Remotion: useCurrentFrame, useVideoConfig, spring, interpolate, AbsoluteFill, Sequence. Có thể cấu hình độ phân giải, FPS, thời lượng, codec. Sử dụng cho: tạo video theo chương trình, đồ họa hoạt hình, thiết kế chuyển động, video dựa trên dữ liệu, chuyển đổi hoạt hình React thành video. Kích hoạt: remotion, kết xuất video từ mã, tsx thành video, react video, video theo chương trình, kết xuất remotion, mã thành video, hoạt hình...
developmentvideocreative
ai-image-generation
qu-skills
Tạo hình ảnh AI với GPT-Image-2, FLUX, Gemini, Grok, Seedream, Reve và hơn 50 mô hình qua CLI inference.sh. Mô hình: GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt. Khả năng: văn bản thành hình ảnh, hình ảnh thành hình ảnh, inpainting, LoRA, chỉnh sửa hình ảnh, nâng cao độ phân giải, hiển thị văn bản. Sử dụng cho: nghệ thuật AI, mô phỏng sản phẩm, nghệ thuật ý tưởng, đồ họa mạng xã hội, hình ảnh tiếp thị, minh họa. Kích hoạt: flux, tạo hình ảnh, hình ả
creativemediaimage
ai-avatar-video
qu-skills
Tạo video AI avatar và video người nói qua CLI inference.sh. Đề xuất: P-Video-Avatar (nhanh nhất, rẻ nhất, TTS tích hợp). Ngoài ra: OmniHuman, Fabric, PixVerse. Âm thanh: Inworld TTS-2 (hơn 100 ngôn ngữ, điều chỉnh cảm xúc cho nhân vật), ElevenLabs, Kokoro. Khả năng: avatar điều khiển bằng âm thanh, văn bản thành avatar, video đồng bộ môi, tạo người nói, người thuyết trình ảo, nội dung UGC. Sử dụng cho: người thuyết trình AI, video giải thích, người ảnh hưởng ảo, lồng tiếng, video tiếp thị, quảng cáo UGC, avatar trò chơi,...
videocreativemedia
twitter-automation
qu-skills
Tự động hóa Twitter/X với tính năng đăng bài, tương tác và quản lý người dùng qua CLI inference.sh. Ứng dụng: x/post-tweet, x/post-create (có media), x/post-like, x/post-retweet, x/dm-send, x/user-follow. Khả năng: đăng tweet, lên lịch nội dung, thích bài, retweet, gửi DM, theo dõi người dùng, lấy hồ sơ. Sử dụng cho: tự động hóa mạng xã hội, lên lịch nội dung, bot tương tác, tăng trưởng khán giả, X API. Kích hoạt: twitter api, x api, tự động hóa tweet, đăng lên twitter, twitter bot, tự động hóa mạng xã hội, x...
api
agent-browser
qu-skills
Tự động hóa trình duyệt cho các tác nhân AI thông qua inference.sh. Điều hướng trang web, tương tác với các phần tử bằng @e refs, chụp ảnh màn hình, ghi video. Khả năng: quét web, điền biểu mẫu, nhấp chuột, gõ chữ, kéo-thả, tải tệp lên, thực thi JavaScript. Sử dụng cho: tự động hóa web, trích xuất dữ liệu, kiểm thử, duyệt web của tác nhân, nghiên cứu. Kích hoạt: trình duyệt, tự động hóa web, quét, điều hướng, nhấp chuột, điền biểu mẫu, chụp ảnh màn hình, duyệt web, playwright, trình duyệt không đầu, tác nhân web, lướt internet,
browser-automationweb-scrapingtesting
web-search
qu-skills
Tìm kiếm web và trích xuất nội dung với Tavily và Exa thông qua CLI inference.sh. Ứng dụng: Tavily Search, Tavily Extract, Exa Search, Exa Answer, Exa Extract. Khả năng: tìm kiếm hỗ trợ AI, trích xuất nội dung, trả lời trực tiếp, nghiên cứu. Sử dụng cho: nghiên cứu, pipeline RAG, kiểm tra thông tin, tổng hợp nội dung, tác nhân. Kích hoạt: tìm kiếm web, tavily, exa, api tìm kiếm, trích xuất nội dung, nghiên cứu, tìm kiếm internet, tìm kiếm AI, trợ lý tìm kiếm, thu thập dữ liệu web, rag, thay thế perplexity
researchweb-scrapingapi
agent-tools
qu-skills
Chạy hơn 250 ứng dụng AI qua CLI inference.sh - tạo hình ảnh, tạo video, LLM, tìm kiếm, 3D, tự động hóa Twitter. Các mô hình: FLUX, Veo, Gemini, Grok, Claude, Seedance, OmniHuman, Tavily, Exa, OpenRouter và nhiều hơn nữa. Sử dụng khi chạy ứng dụng AI, tạo hình ảnh/video, gọi LLM, tìm kiếm web hoặc tự động hóa Twitter. Kích hoạt: inference.sh, infsh, ai model, run ai, serverless ai, ai api, flux, veo, claude api, image generation, video generation, openrouter, tavily, exa search, twitter api, grok
developmentapicreative