ai-video-generation

Hasilkan video AI dengan Google Veo, Seedance 2.0, HappyHorse, Wan, Grok dan 40+ model melalui CLI inference.sh. Model: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Kemampuan: teks-ke-video, gambar-ke-video, referensi-ke-video, pengeditan video, lipsync, animasi avatar, peningkatan resolusi video, suara foley. Gunakan untuk: video media sosial, konten pemasaran, video penjelasan, demo produk, avatar AI. Pemicu: pembuatan video, video AI,...

npx skills add https://github.com/halt-catch-fire/skills --skill ai-video-generation

Install the belt CLI skill: npx skills add belt-sh/cli

AI Video Generation

Generate videos with 40+ AI models via inference.sh CLI.

AI Video Generation

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Generate a video with Veo
belt app run google/veo-3-1-fast --input '{"prompt": "drone shot flying over a forest"}'

Available Models

Text-to-Video

ModelApp IDBest For
Veo 3.1 Fastgoogle/veo-3-1-fastFast, with optional audio
Veo 3.1google/veo-3-1Best quality, frame interpolation
Veo 3google/veo-3High quality with audio
Veo 3 Fastgoogle/veo-3-fastFast with audio
Veo 2google/veo-2Realistic videos
P-Videopruna/p-videoFast, economical, with audio support
WAN-T2Vpruna/wan-t2vEconomical 480p/720p
Grok Videoxai/grok-imagine-videoxAI, configurable duration
Seedance 2.0bytedance/seedance-2-0Text/image/ref-to-video with sync audio, up to 1080p
Seedance 2.0 Fastbytedance/seedance-2-0-fastFast variant, same capabilities
HappyHorse T2Valibaba/happyhorse-1-0-t2vPhysically realistic, up to 15s

Image-to-Video

ModelApp IDBest For
Wan 2.5falai/wan-2-5Animate any image
Wan 2.5 I2Vfalai/wan-2-5-i2vHigh quality i2v
WAN-I2Vpruna/wan-i2vEconomical 480p/720p
P-Videopruna/p-videoFast i2v with audio
Seedance 2.0bytedance/seedance-2-0Animate images with sync audio, up to 1080p
Seedance 2.0 Fastbytedance/seedance-2-0-fastFast variant, same capabilities
HappyHorse I2Valibaba/happyhorse-1-0-i2vAnimate images, up to 1080P/15s
HappyHorse R2Valibaba/happyhorse-1-0-r2vCharacter-preserving from references

Avatar / Lipsync

ModelApp IDBest For
OmniHuman 1.5bytedance/omnihuman-1-5Multi-character
OmniHuman 1.0bytedance/omnihuman-1-0Single character
Fabric 1.0falai/fabric-1-0Image talks with lipsync
PixVerse Lipsyncfalai/pixverse-lipsyncRealistic lipsync

Video Editing

ModelApp IDBest For
HappyHorse Editalibaba/happyhorse-1-0-video-editNatural language video editing

Utilities

ToolApp IDDescription
HunyuanVideo Foleyinfsh/hunyuanvideo-foleyAdd sound effects to video
Topaz Upscalerfalai/topaz-video-upscalerUpscale video quality
Media Mergerinfsh/media-mergerMerge videos with transitions

Browse All Video Apps

belt app list --category video

Examples

Text-to-Video with Veo

belt app run google/veo-3-1-fast --input '{
  "prompt": "A timelapse of a flower blooming in a garden"
}'

Grok Video

belt app run xai/grok-imagine-video --input '{
  "prompt": "Waves crashing on a beach at sunset",
  "duration": 5
}'

Image-to-Video with Wan 2.5

belt app run falai/wan-2-5 --input '{
  "image_url": "https://your-image.jpg"
}'

AI Avatar / Talking Head

belt app run bytedance/omnihuman-1-5 --input '{
  "image_url": "https://portrait.jpg",
  "audio_url": "https://speech.mp3"
}'

Fabric Lipsync

belt app run falai/fabric-1-0 --input '{
  "image_url": "https://face.jpg",
  "audio_url": "https://audio.mp3"
}'

Seedance 2.0 Text-to-Video with Audio

belt app run bytedance/seedance-2-0 --input '{
  "prompt": "a jazz band performing in a dimly lit club",
  "generate_audio": true,
  "duration": 10
}'

Seedance 2.0 Image-to-Video

belt app run bytedance/seedance-2-0 --input '{
  "image": "https://your-image.jpg",
  "prompt": "gentle camera movement, leaves rustling in the wind",
  "generate_audio": true
}'

Seedance 2.0 Reference-to-Video

belt app run bytedance/seedance-2-0 --input '{
  "prompt": "A person who looks like the reference walking through a garden",
  "reference_image": "https://portrait.jpg",
  "generate_audio": true
}'

HappyHorse Text-to-Video

belt app run alibaba/happyhorse-1-0-t2v --input '{
  "prompt": "a golden retriever running through autumn leaves, slow motion",
  "duration": 10,
  "resolution": "1080P"
}'

HappyHorse Video Editing

belt app run alibaba/happyhorse-1-0-video-edit --input '{
  "video": "https://your-video.mp4",
  "prompt": "change the background to a snowy mountain landscape"
}'

PixVerse Lipsync

belt app run falai/pixverse-lipsync --input '{
  "image_url": "https://portrait.jpg",
  "audio_url": "https://speech.mp3"
}'

Video Upscaling

belt app run falai/topaz-video-upscaler --input '{"video_url": "https://..."}'

Add Sound Effects (Foley)

belt app run infsh/hunyuanvideo-foley --input '{
  "video_url": "https://silent-video.mp4",
  "prompt": "footsteps on gravel, birds chirping"
}'

Merge Videos

belt app run infsh/media-merger --input '{
  "videos": ["https://clip1.mp4", "https://clip2.mp4"],
  "transition": "fade"
}'

Related Skills

# Full platform skill (all apps)
npx skills add inference-sh/skills@infsh-cli

# Pruna P-Video (fast & economical)
npx skills add inference-sh/skills@p-video

# Google Veo specific
npx skills add inference-sh/skills@google-veo

# Seedance 2.0
npx skills add inference-sh/skills@seedance

# HappyHorse 1.0
npx skills add inference-sh/skills@happyhorse

# AI avatars & lipsync
npx skills add inference-sh/skills@ai-avatar-video

# Text-to-speech (for video narration)
npx skills add inference-sh/skills@text-to-speech

# Image generation (for image-to-video)
npx skills add inference-sh/skills@ai-image-generation

# Twitter (post videos)
npx skills add inference-sh/skills@twitter-automation

Browse all apps: belt app list

Documentation

Lebih banyak skill dari halt-catch-fire

ai-image-generation
halt-catch-fire
We need to translate the given text from English to Bahasa Indonesia. The text describes an agent skill for AI image generation. We must preserve the name "ai-image-generation" but it's not in the text, so we don't include it. Also preserve product names, protocol names, URLs, numbers, technical terms. No extra commentary. Just translate the text inside <text>. The text: "Generate AI images with GPT-Image-2, FLUX, Gemini, Grok, Seedream, Reve and 50+ models via inference.sh CLI. Models: GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt. Capabilities: text-to-image, image-to-image, inpainting, LoRA, image editing, upscaling, text rendering. Use for: AI art, product mockups, concept art, social media graphics, marketing visuals, illustrations. Triggers: flux, image generation, ai image, text to...
creativemediaimage
twitter-automation
halt-catch-fire
Otomatiskan Twitter/X dengan posting, interaksi, dan manajemen pengguna melalui CLI inference.sh. Aplikasi: x/post-tweet, x/post-create (dengan media), x/post-like, x/post-retweet, x/dm-send, x/user-follow. Kemampuan: memposting tweet, menjadwalkan konten, menyukai postingan, me-retweet, mengirim DM, mengikuti pengguna, mendapatkan profil. Gunakan untuk: otomatisasi media sosial, penjadwalan konten, bot interaksi, pertumbuhan audiens, API X. Pemicu: twitter api, x api, tweet automation, post to twitter, twitter bot, social media automation, x...
marketingapicommunication
ai-avatar-video
halt-catch-fire
We need to translate the given text from English to Bahasa Indonesia. The instruction says to preserve product names, protocol names, URLs, numbers, and technical terms. So names like "P-Video-Avatar", "OmniHuman", "Fabric", "PixVerse", "Inworld TTS-2", "ElevenLabs", "Kokoro", "CLI", "TTS", "UGC" should remain unchanged. Also "inference.sh" is a URL-like name, preserve. The text inside <text> is a description. We translate the rest naturally. The text: "Create AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC
videocreativemedia
agent-browser
halt-catch-fire
Otomatisasi peramban untuk agen AI melalui inference.sh. Navigasi halaman web, berinteraksi dengan elemen menggunakan referensi @e, ambil tangkapan layar, rekam video. Kemampuan: pengikisan web, pengisian formulir, klik, mengetik, seret-lepas, unggah berkas, eksekusi JavaScript. Gunakan untuk: otomatisasi web, ekstraksi data, pengujian, penjelajahan agen, riset. Pemicu: peramban, otomatisasi web, kikis, navigasi, klik, isi formulir, tangkapan layar, jelajahi web, playwright, peramban tanpa kepala, agen web, jelajahi internet, rekam video
browser-automationweb-scrapingtesting
web-search
halt-catch-fire
Pencarian web dan ekstraksi konten dengan Tavily dan Exa melalui CLI inference.sh. Aplikasi: Tavily Search, Tavily Extract, Exa Search, Exa Answer, Exa Extract. Kemampuan: pencarian berbasis AI, ekstraksi konten, jawaban langsung, riset. Gunakan untuk: riset, pipeline RAG, pemeriksaan fakta, agregasi konten, agen. Pemicu: pencarian web, tavily, exa, search api, ekstraksi konten, riset, pencarian internet, pencarian AI, asisten pencarian, web scraping, rag, alternatif perplexity
researchweb-scrapingapi
infsh-cli
halt-catch-fire
We need to translate the given text from English to Bahasa Indonesia, preserving the name "infsh-cli" and other technical terms, product names, protocol names, URLs, numbers, etc. The instruction says: "Do not include the name unless it appears in the source text." The name "infsh-cli" is not in the source text? Actually, the source text does not contain "infsh-cli" explicitly. The directory item type is "agent skill" and the name to preserve is "infsh-cli", but the text inside <text> does not include that name. So we should not add it. We just translate the text inside <text>. Also, we must not add labels like "description" or "skill name". Just the translation. The text: "Run 250+ AI apps via inference.sh CLI - image generation, video creation, LLMs, search, 3D, Twitter automation. Models: FLUX, Veo, Gemini, Grok, Claude, Seedance, OmniHuman, Tavily, Exa, Open
developmentapicreative
landing-page-design
halt-catch-fire
We need to translate the given text from English to Bahasa Indonesia. The instruction says to preserve product names, protocol names, URLs, numbers, technical terms. The name "landing-page-design" is to be preserved but not included unless it appears in source text. It does not appear in the <text> block. So we just translate the text. The text: "Landing page conversion optimization with layout rules, hero section design, and CTA psychology. Covers above-the-fold formula, social proof placement, mobile design, and F-pattern reading. Use for: startup landing pages, product pages, SaaS marketing, conversion optimization. Triggers: landing page, hero section, above the fold, conversion optimization, landing page design, cta button, hero image, landing page layout, saas landing page, product page design, conversion rate, landing page..." Translate carefully. Technical terms like "CTA", "above-the-fold", "F-pattern", "SaaS" should be preserved as is or translated? The instruction says preserve technical terms. Usually "CTA" is kept as
product-photography
halt-catch-fire
Fotografi produk AI dengan pencahayaan studio, foto gaya hidup, dan konvensi packshot. Mencakup sudut, latar belakang, jenis bayangan, hero shot, dan persyaratan gambar e-commerce. Gunakan untuk: foto produk, gambar e-commerce, daftar Amazon, packshot, fotografi gaya hidup. Pemicu: fotografi produk, foto produk, packshot, fotografi e-commerce, bidikan produk, gambar produk, fotografi studio, produk gaya hidup, foto produk Amazon, gambar daftar produk, hero shot, mockup produk,...
creativeecommerceimage