ai-avatar-video

Tạo video AI avatar và video người nói qua CLI inference.sh. Đề xuất: P-Video-Avatar (nhanh nhất, rẻ nhất, TTS tích hợp). Ngoài ra: OmniHuman, Fabric, PixVerse. Âm thanh: Inworld TTS-2 (hơn 100 ngôn ngữ, điều chỉnh cảm xúc cho nhân vật), ElevenLabs, Kokoro. Khả năng: avatar điều khiển bằng âm thanh, văn bản thành avatar, video đồng bộ môi, tạo người nói, người thuyết trình ảo, nội dung UGC. Sử dụng cho: người thuyết trình AI, video giải thích, người có ảnh hưởng ảo, lồng tiếng, video tiếp thị, quảng cáo UGC, avatar trò chơi,...

npx skills add https://github.com/101-skills/skills --skill ai-avatar-video

Install the belt CLI skill: npx skills add belt-sh/cli

AI Avatar & Talking Head Videos

Create AI avatars and talking head videos via inference.sh CLI.

AI Avatar & Talking Head Videos

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS)
belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "voice_script": "Hello, welcome to our product demo!",
  "voice": "Zephyr (Female)"
}'

Available Models

Start with P-Video-Avatar — it's 18x faster and 6x cheaper than alternatives, with built-in TTS, dynamic backgrounds, and 1080p support.

ModelApp IDBest ForBuilt-in TTS
P-Video-Avatarpruna/p-video-avatarBest overall: speed, cost, quality, controlYes (30 voices, 10 languages)
OmniHuman 1.5bytedance/omnihuman-1-5Multi-character, audio-drivenNo
Fabric 1.0falai/fabric-1-0Image talks with lipsyncYes
PixVerse Lipsyncfalai/pixverse-lipsyncHighly realistic lipsyncNo

Cost & Speed Comparison

ModelSpeed (per sec of video)Cost per second
P-Video-Avatar~1.83s/s$0.025
OmniHuman 1.5~28s/s (15x slower)$0.16 (6.4x more)
Fabric 1.0~34s/s (18x slower)$0.14 (5.6x more)

Examples

P-Video-Avatar (Recommended)

Generate avatar from portrait + text script with built-in TTS:

belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "voice_script": "Welcome to our product walkthrough. Today I will show you three key features.",
  "voice": "Puck (Male)",
  "voice_language": "English (US)",
  "resolution": "720p"
}'

With custom style control:

belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "voice_script": "This is exciting news!",
  "voice": "Aoede (Female)",
  "voice_prompt": "Enthusiastic and energetic tone",
  "video_prompt": "The person is presenting on stage with dramatic lighting",
  "resolution": "1080p"
}'

With audio file instead of TTS:

belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "audio": "https://speech.mp3"
}'

Full Workflow: Generate Portrait + Avatar

Use Pruna P-Image to generate the portrait, then create the avatar:

# 1. Generate a portrait image
belt app run pruna/p-image --input '{
  "prompt": "professional headshot portrait of a young woman, neutral background, looking at camera, studio lighting, photorealistic",
  "aspect_ratio": "9:16"
}'

# 2. Create avatar video with built-in TTS
belt app run pruna/p-video-avatar --input '{
  "image": "<image-url-from-step-1>",
  "voice_script": "Hi there! Let me walk you through our latest features.",
  "voice": "Zephyr (Female)"
}'

OmniHuman 1.5 (Multi-Character)

belt app run bytedance/omnihuman-1-5 --input '{
  "image_url": "https://portrait.jpg",
  "audio_url": "https://speech.mp3"
}'

Supports specifying which character to drive in multi-person images.

Fabric 1.0 (Image Talks)

belt app run falai/fabric-1-0 --input '{
  "image_url": "https://face.jpg",
  "audio_url": "https://audio.mp3"
}'

PixVerse Lipsync

belt app run falai/pixverse-lipsync --input '{
  "image_url": "https://portrait.jpg",
  "audio_url": "https://speech.mp3"
}'

Full Workflow: TTS + Avatar (Non-TTS Models)

For models without built-in TTS (OmniHuman, PixVerse), generate speech first:

# 1. Generate speech — Inworld TTS-2 for expressive character voices
belt app run inworld/text-to-speech-2 --input '{
  "text": "[friendly] Welcome to our product demo! [excited] Let me show you three features that will change how you work.",
  "voice_id": "Sarah",
  "delivery_mode": "CREATIVE"
}' > speech.json

# 2. Create avatar video with the speech
belt app run bytedance/omnihuman-1-5 --input '{
  "image_url": "https://presenter-photo.jpg",
  "audio_url": "<audio-url-from-step-1>"
}'

Tip: For most use cases, P-Video-Avatar with built-in TTS is simpler — no separate audio step needed. Use this workflow only when you specifically need OmniHuman (multi-character) or PixVerse (realistic lipsync).

Full Workflow: Dub Video in Another Language

# 1. Transcribe original video
belt app run infsh/fast-whisper-large-v3 --input '{"audio_url": "https://video.mp4"}' > transcript.json

# 2. Translate text (manually or with an LLM)

# 3. Generate speech in new language
belt app run infsh/kokoro-tts --input '{"text": "<translated-text>"}' > new_speech.json

# 4. Lipsync the original video with new audio
belt app run infsh/latentsync-1-6 --input '{
  "video_url": "https://original-video.mp4",
  "audio_url": "<new-audio-url>"
}'

Avatar UGC Generation

Create UGC-style content with P-Video-Avatar — built-in TTS, no separate audio step needed:

# 1. Generate a relatable UGC-style portrait
belt app run pruna/p-image --input '{
  "prompt": "casual selfie-style photo of a young woman in a cozy room, natural lighting, looking at camera, warm smile, authentic feel",
  "aspect_ratio": "9:16"
}'

# 2. Create UGC avatar video with built-in TTS
belt app run pruna/p-video-avatar --input '{
  "image": "<image-url-from-step-1>",
  "voice_script": "Okay so I just tried this product and honestly? It is a game changer. I was not expecting to love it this much but here we are!",
  "voice": "Zephyr (Female)",
  "voice_prompt": "Excited, casual, authentic tone like talking to a friend",
  "video_prompt": "The person is talking casually to camera in their room, natural gestures",
  "resolution": "1080p"
}'

Why P-Video-Avatar for UGC

  • All-in-one — built-in TTS means no separate audio generation step
  • 30 voices, 10 languages — match your target audience
  • Voice + video prompts — control tone, emotion, body language, and background independently
  • 18x faster, 6x cheaper — produce UGC at scale vs. Fabric/OmniHuman/HeyGen
  • 1080p support — platform-ready vertical video from a single portrait image

Batch UGC: Same Product, Multiple Presenters

# Generate 3 different presenters
for voice in "Zephyr (Female)" "Puck (Male)" "Aoede (Female)"; do
  belt app run pruna/p-video-avatar --input "{
    \"image\": \"https://portrait.jpg\",
    \"voice_script\": \"This changed my morning routine completely. Five minutes and I am done.\",
    \"voice\": \"$voice\",
    \"voice_prompt\": \"Casual, authentic, like a real testimonial\",
    \"video_prompt\": \"Person talking to camera in a bright kitchen\",
    \"resolution\": \"1080p\"
  }"
done

Use Cases

  • UGC & Marketing: Product demos, UGC-style ads with AI presenters
  • Education: Course videos, explainers
  • Localization: Dub content across 10 languages from one image
  • Social Media: Consistent virtual influencer content
  • Corporate: Training videos, announcements
  • Gaming: Character avatars, NPC dialogue

Tips

  • Use high-quality portrait photos (front-facing, good lighting)
  • Audio should be clear with minimal background noise
  • P-Video-Avatar supports built-in TTS — no need for a separate speech generation step
  • P-Video-Avatar output aspect ratio matches the input image
  • Generate portraits with pruna/p-image using 9:16 aspect ratio for vertical videos
  • OmniHuman 1.5 supports multiple people in one image
  • LatentSync is best for syncing existing videos to new audio

Related Skills

# Dedicated P-Video-Avatar skill
npx skills add inference-sh/skills@p-video-avatar

# Full platform skill (all apps)
npx skills add inference-sh/skills@infsh-cli

# Text-to-speech (generate audio for non-TTS avatar models)
npx skills add inference-sh/skills@text-to-speech

# Speech-to-text (transcribe for dubbing)
npx skills add inference-sh/skills@speech-to-text

# Video generation
npx skills add inference-sh/skills@ai-video-generation

# Image generation (create avatar images)
npx skills add inference-sh/skills@ai-image-generation

Browse all video apps: belt app list --category video

Documentation

Thêm skills từ 101-skills

ai-video-generation
101-skills
Tạo video AI với Google Veo, Seedance 2.0, HappyHorse, Wan, Grok và hơn 40 mô hình qua CLI inference.sh. Các mô hình: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Khả năng: văn bản thành video, hình ảnh thành video, tham chiếu thành video, chỉnh sửa video, đồng bộ môi, hoạt ảnh đại diện, nâng cấp video, âm thanh foley. Sử dụng cho: video mạng xã hội, nội dung tiếp thị, video giải thích, demo sản phẩm, đại diện AI. Kích hoạt: tạo video, video AI,...
creativevideomedia
ai-image-generation
101-skills
Tạo hình ảnh AI với GPT-Image-2, FLUX, Gemini, Grok, Seedream, Reve và hơn 50 mô hình qua CLI inference.sh. Các mô hình: GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt. Khả năng: văn bản thành hình ảnh, hình ảnh thành hình ảnh, inpainting, LoRA, chỉnh sửa hình ảnh, nâng cấp độ phân giải, hiển thị văn bản. Sử dụng cho: nghệ thuật AI, mô phỏng sản phẩm, nghệ thuật ý tưởng, đồ họa mạng xã hội, hình ảnh tiếp thị, minh họa. Kích hoạt: flux, tạo hình ảnh, hình ảnh AI, văn bản thành...
creativemediaimage
remotion-render
101-skills
Kết xuất video từ mã React/Remotion component thông qua inference.sh. Nhập mã TSX, nhận MP4. Hỗ trợ tất cả API Remotion: useCurrentFrame, useVideoConfig, spring, interpolate, AbsoluteFill, Sequence. Có thể cấu hình độ phân giải, FPS, thời lượng, codec. Sử dụng cho: tạo video theo chương trình, đồ họa hoạt hình, thiết kế chuyển động, video dựa trên dữ liệu, chuyển đổi hoạt hình React thành video. Kích hoạt: remotion, kết xuất video từ mã, tsx thành video, react video, video theo chương trình, kết xuất remotion, mã thành video, hoạt hình...
videocreativedevelopment
web-search
101-skills
Tìm kiếm web và trích xuất nội dung với Tavily và Exa thông qua CLI inference.sh. Ứng dụng: Tavily Search, Tavily Extract, Exa Search, Exa Answer, Exa Extract. Khả năng: tìm kiếm hỗ trợ AI, trích xuất nội dung, trả lời trực tiếp, nghiên cứu. Sử dụng cho: nghiên cứu, pipeline RAG, kiểm tra thông tin, tổng hợp nội dung, tác nhân. Kích hoạt: tìm kiếm web, tavily, exa, api tìm kiếm, trích xuất nội dung, nghiên cứu, tìm kiếm internet, tìm kiếm AI, trợ lý tìm kiếm, thu thập dữ liệu web, rag, thay thế perplexity
researchweb-scrapingapi
agent-tools
101-skills
Chạy ứng dụng AI qua CLI inference.sh - tạo hình ảnh, tạo video, LLM, tìm kiếm, 3D, tự động hóa Twitter. Các mô hình: FLUX, Veo, Gemini, Grok, Claude, Seedance, OmniHuman, Tavily, Exa, OpenRouter và nhiều hơn nữa. Sử dụng khi chạy ứng dụng AI, tạo hình ảnh/video, gọi LLM, tìm kiếm web hoặc tự động hóa Twitter. Kích hoạt: inference.sh, infsh, ai model, run ai, serverless ai, ai api, flux, veo, claude api, image generation, video generation, openrouter, tavily, exa search, twitter api, grok
infsh-cli
101-skills
Chạy ứng dụng AI qua CLI inference.sh - tạo hình ảnh, tạo video, LLM, tìm kiếm, 3D, tự động hóa Twitter. Các mô hình: FLUX, Veo, Gemini, Grok, Claude, Seedance, OmniHuman, Tavily, Exa, OpenRouter và nhiều hơn nữa. Sử dụng khi chạy ứng dụng AI, tạo hình ảnh/video, gọi LLM, tìm kiếm web hoặc tự động hóa Twitter. Kích hoạt: inference.sh, infsh, mô hình AI, chạy AI, AI không máy chủ, API AI, flux, veo, claude api, tạo hình ảnh, tạo video, openrouter, tavily, tìm kiếm exa, twitter api, grok
landing-page-design
101-skills
Tối ưu hóa chuyển đổi trang đích với quy tắc bố cục, thiết kế phần hero và tâm lý học CTA. Bao gồm công thức above-the-fold, vị trí đặt bằng chứng xã hội, thiết kế di động và cách đọc theo mô hình F. Sử dụng cho: trang đích khởi nghiệp, trang sản phẩm, tiếp thị SaaS, tối ưu hóa chuyển đổi. Kích hoạt: trang đích, phần hero, above the fold, tối ưu hóa chuyển đổi, thiết kế trang đích, nút CTA, hình ảnh hero, bố cục trang đích, trang đích SaaS, thiết kế trang sản phẩm, tỷ lệ chuyển đổi, trang đích...
designmarketingcreative
product-photography
101-skills
Chụp ảnh sản phẩm bằng AI với ánh sáng studio, ảnh phong cách sống và quy ước ảnh packshot. Bao gồm góc chụp, phông nền, loại bóng, ảnh hero và yêu cầu hình ảnh thương mại điện tử. Sử dụng cho: ảnh sản phẩm, hình ảnh thương mại điện tử, danh sách Amazon, packshot, nhiếp ảnh phong cách sống. Kích hoạt: chụp ảnh sản phẩm, ảnh sản phẩm, packshot, nhiếp ảnh thương mại điện tử, chụp sản phẩm, hình ảnh sản phẩm, nhiếp ảnh studio, sản phẩm phong cách sống, ảnh sản phẩm Amazon, hình ảnh danh sách sản phẩm, ảnh hero, mockup sản phẩm,...
creativeecommerceimage