ai-avatar-video

โดย halt-catch-fire

We need to translate the given text from English to Thai. The instruction says to preserve product names, protocol names, URLs, numbers, and technical terms. So names like "P-Video-Avatar", "OmniHuman", "Fabric", "PixVerse", "Inworld TTS-2", "ElevenLabs", "Kokoro", "inference.sh CLI" should remain as is. Also "AI avatar", "talking head", "TTS", "lipsync", "UGC" etc. are technical terms that might be kept in English or translated? The instruction says "preserve product names, protocol names, URLs, numbers, and technical terms." Technical terms could be translated if common in Thai, but to be safe, we can keep them as is or use common Thai transliterations. However, the instruction says "preserve" meaning keep original. So we should keep all those terms in English. The rest of the text should be translated naturally into Thai. We need to output only the translated text, no extra labels.

npx skills add https://github.com/halt-catch-fire/skills --skill ai-avatar-video

Install the belt CLI skill: npx skills add belt-sh/cli

AI Avatar & Talking Head Videos

Create AI avatars and talking head videos via inference.sh CLI.

AI Avatar & Talking Head Videos

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS)
belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "voice_script": "Hello, welcome to our product demo!",
  "voice": "Zephyr (Female)"
}'

Available Models

Start with P-Video-Avatar — it's 18x faster and 6x cheaper than alternatives, with built-in TTS, dynamic backgrounds, and 1080p support.

ModelApp IDBest ForBuilt-in TTS
P-Video-Avatarpruna/p-video-avatarBest overall: speed, cost, quality, controlYes (30 voices, 10 languages)
OmniHuman 1.5bytedance/omnihuman-1-5Multi-character, audio-drivenNo
Fabric 1.0falai/fabric-1-0Image talks with lipsyncYes
PixVerse Lipsyncfalai/pixverse-lipsyncHighly realistic lipsyncNo

Cost & Speed Comparison

ModelSpeed (per sec of video)Cost per second
P-Video-Avatar~1.83s/s$0.025
OmniHuman 1.5~28s/s (15x slower)$0.16 (6.4x more)
Fabric 1.0~34s/s (18x slower)$0.14 (5.6x more)

Examples

P-Video-Avatar (Recommended)

Generate avatar from portrait + text script with built-in TTS:

belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "voice_script": "Welcome to our product walkthrough. Today I will show you three key features.",
  "voice": "Puck (Male)",
  "voice_language": "English (US)",
  "resolution": "720p"
}'

With custom style control:

belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "voice_script": "This is exciting news!",
  "voice": "Aoede (Female)",
  "voice_prompt": "Enthusiastic and energetic tone",
  "video_prompt": "The person is presenting on stage with dramatic lighting",
  "resolution": "1080p"
}'

With audio file instead of TTS:

belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "audio": "https://speech.mp3"
}'

Full Workflow: Generate Portrait + Avatar

Use Pruna P-Image to generate the portrait, then create the avatar:

# 1. Generate a portrait image
belt app run pruna/p-image --input '{
  "prompt": "professional headshot portrait of a young woman, neutral background, looking at camera, studio lighting, photorealistic",
  "aspect_ratio": "9:16"
}'

# 2. Create avatar video with built-in TTS
belt app run pruna/p-video-avatar --input '{
  "image": "<image-url-from-step-1>",
  "voice_script": "Hi there! Let me walk you through our latest features.",
  "voice": "Zephyr (Female)"
}'

OmniHuman 1.5 (Multi-Character)

belt app run bytedance/omnihuman-1-5 --input '{
  "image_url": "https://portrait.jpg",
  "audio_url": "https://speech.mp3"
}'

Supports specifying which character to drive in multi-person images.

Fabric 1.0 (Image Talks)

belt app run falai/fabric-1-0 --input '{
  "image_url": "https://face.jpg",
  "audio_url": "https://audio.mp3"
}'

PixVerse Lipsync

belt app run falai/pixverse-lipsync --input '{
  "image_url": "https://portrait.jpg",
  "audio_url": "https://speech.mp3"
}'

Full Workflow: TTS + Avatar (Non-TTS Models)

For models without built-in TTS (OmniHuman, PixVerse), generate speech first:

# 1. Generate speech — Inworld TTS-2 for expressive character voices
belt app run inworld/text-to-speech-2 --input '{
  "text": "[friendly] Welcome to our product demo! [excited] Let me show you three features that will change how you work.",
  "voice_id": "Sarah",
  "delivery_mode": "CREATIVE"
}' > speech.json

# 2. Create avatar video with the speech
belt app run bytedance/omnihuman-1-5 --input '{
  "image_url": "https://presenter-photo.jpg",
  "audio_url": "<audio-url-from-step-1>"
}'

Tip: For most use cases, P-Video-Avatar with built-in TTS is simpler — no separate audio step needed. Use this workflow only when you specifically need OmniHuman (multi-character) or PixVerse (realistic lipsync).

Full Workflow: Dub Video in Another Language

# 1. Transcribe original video
belt app run infsh/fast-whisper-large-v3 --input '{"audio_url": "https://video.mp4"}' > transcript.json

# 2. Translate text (manually or with an LLM)

# 3. Generate speech in new language
belt app run infsh/kokoro-tts --input '{"text": "<translated-text>"}' > new_speech.json

# 4. Lipsync the original video with new audio
belt app run infsh/latentsync-1-6 --input '{
  "video_url": "https://original-video.mp4",
  "audio_url": "<new-audio-url>"
}'

Avatar UGC Generation

Create UGC-style content with P-Video-Avatar — built-in TTS, no separate audio step needed:

# 1. Generate a relatable UGC-style portrait
belt app run pruna/p-image --input '{
  "prompt": "casual selfie-style photo of a young woman in a cozy room, natural lighting, looking at camera, warm smile, authentic feel",
  "aspect_ratio": "9:16"
}'

# 2. Create UGC avatar video with built-in TTS
belt app run pruna/p-video-avatar --input '{
  "image": "<image-url-from-step-1>",
  "voice_script": "Okay so I just tried this product and honestly? It is a game changer. I was not expecting to love it this much but here we are!",
  "voice": "Zephyr (Female)",
  "voice_prompt": "Excited, casual, authentic tone like talking to a friend",
  "video_prompt": "The person is talking casually to camera in their room, natural gestures",
  "resolution": "1080p"
}'

Why P-Video-Avatar for UGC

  • All-in-one — built-in TTS means no separate audio generation step
  • 30 voices, 10 languages — match your target audience
  • Voice + video prompts — control tone, emotion, body language, and background independently
  • 18x faster, 6x cheaper — produce UGC at scale vs. Fabric/OmniHuman/HeyGen
  • 1080p support — platform-ready vertical video from a single portrait image

Batch UGC: Same Product, Multiple Presenters

# Generate 3 different presenters
for voice in "Zephyr (Female)" "Puck (Male)" "Aoede (Female)"; do
  belt app run pruna/p-video-avatar --input "{
    \"image\": \"https://portrait.jpg\",
    \"voice_script\": \"This changed my morning routine completely. Five minutes and I am done.\",
    \"voice\": \"$voice\",
    \"voice_prompt\": \"Casual, authentic, like a real testimonial\",
    \"video_prompt\": \"Person talking to camera in a bright kitchen\",
    \"resolution\": \"1080p\"
  }"
done

Use Cases

  • UGC & Marketing: Product demos, UGC-style ads with AI presenters
  • Education: Course videos, explainers
  • Localization: Dub content across 10 languages from one image
  • Social Media: Consistent virtual influencer content
  • Corporate: Training videos, announcements
  • Gaming: Character avatars, NPC dialogue

Tips

  • Use high-quality portrait photos (front-facing, good lighting)
  • Audio should be clear with minimal background noise
  • P-Video-Avatar supports built-in TTS — no need for a separate speech generation step
  • P-Video-Avatar output aspect ratio matches the input image
  • Generate portraits with pruna/p-image using 9:16 aspect ratio for vertical videos
  • OmniHuman 1.5 supports multiple people in one image
  • LatentSync is best for syncing existing videos to new audio

Related Skills

# Dedicated P-Video-Avatar skill
npx skills add inference-sh/skills@p-video-avatar

# Full platform skill (all apps)
npx skills add inference-sh/skills@infsh-cli

# Text-to-speech (generate audio for non-TTS avatar models)
npx skills add inference-sh/skills@text-to-speech

# Speech-to-text (transcribe for dubbing)
npx skills add inference-sh/skills@speech-to-text

# Video generation
npx skills add inference-sh/skills@ai-video-generation

# Image generation (create avatar images)
npx skills add inference-sh/skills@ai-image-generation

Browse all video apps: belt app list --category video

Documentation

Skills เพิ่มเติมจาก halt-catch-fire

ai-image-generation
halt-catch-fire
สร้างภาพ AI ด้วย GPT-Image-2, FLUX, Gemini, Grok, Seedream, Reve และโมเดลอีก 50+ รุ่นผ่าน CLI inference.sh โมเดล: GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt ความสามารถ: ข้อความเป็นภาพ, ภาพเป็นภาพ, การเติมแต่งภาพ, LoRA, การแก้ไขภาพ, การเพิ่มความละเอียด, การเรนเดอร์ข้อความ ใช้สำหรับ: ศิลปะ AI, ตัวอย่างผลิตภัณฑ์, คอนเซปต์อาร์ต, กราฟิกโซเชียลมีเดีย, ภาพการตลาด, ภาพประกอบ ทริกเกอร์: flux, การสร้างภาพ, ภาพ ai, ข้อความเป็น...
creativemediaimage
ai-video-generation
halt-catch-fire
สร้างวิดีโอ AI ด้วย Google Veo, Seedance 2.0, HappyHorse, Wan, Grok และโมเดลอีกกว่า 40 รุ่นผ่าน CLI ของ inference.sh โมเดล: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo ความสามารถ: ข้อความเป็นวิดีโอ, รูปภาพเป็นวิดีโอ, อ้างอิงเป็นวิดีโอ, ตัดต่อวิดีโอ, ลิปซิงค์, แอนิเมชันอวาตาร์, ปรับขนาดวิดีโอ, เสียงประกอบฟอลีย์ ใช้สำหรับ: วิดีโอโซเชียลมีเดีย, เนื้อหาการตลาด, วิดีโออธิบาย, การสาธิตผลิตภัณฑ์, อวาตาร์ AI ทริกเกอร์: การสร้างวิดีโอ,
creativevideomedia
twitter-automation
halt-catch-fire
ทำให้ Twitter/X เป็นอัตโนมัติด้วยการโพสต์ การมีส่วนร่วม และการจัดการผู้ใช้ผ่าน CLI ของ inference.sh แอป: x/post-tweet, x/post-create (พร้อมสื่อ), x/post-like, x/post-retweet, x/dm-send, x/user-follow ความสามารถ: โพสต์ทวีต, จัดตารางเนื้อหา, กดไลก์โพสต์, รีทวีต, ส่ง DM, ติดตามผู้ใช้, ดูโปรไฟล์ ใช้สำหรับ: การทำให้โซเชียลมีเดียเป็นอัตโนมัติ, การจัดตารางเนื้อหา, บอทสร้างการมีส่วนร่วม, การเติบโตของผู้ชม, X API ทริกเกอร์: twitter api, x api, tweet automation, post to twitter, twitter bot, social media automation, x...
marketingapicommunication
agent-browser
halt-catch-fire
การทำงานอัตโนมัติของเบราว์เซอร์สำหรับเอเจนต์ AI ผ่าน inference.sh นำทางหน้าเว็บ โต้ตอบกับองค์ประกอบโดยใช้ @e refs จับภาพหน้าจอ บันทึกวิดีโอ ความสามารถ: การขูดข้อมูลเว็บ การกรอกฟอร์ม การคลิก การพิมพ์ การลากและวาง การอัปโหลดไฟล์ การเรียกใช้ JavaScript ใช้สำหรับ: ระบบอัตโนมัติทางเว็บ การดึงข้อมูล การทดสอบ การเรียกดูของเอเจนต์ การวิจัย ทริกเกอร์: เบราว์เซอร์, ระบบอัตโนมัติทางเว็บ, ขูดข้อมูล, นำทาง, คลิก, กรอกฟอร์ม, จับภาพหน้าจอ, เรียกดูเว็บ, playwright, เบราว์เซอร์ไร้หัว, เว็บเอเจนต์, ท่องอินเทอร์เน็ต, บันทึกวิดีโอ
browser-automationweb-scrapingtesting
web-search
halt-catch-fire
การค้นหาเว็บและดึงเนื้อหาด้วย Tavily และ Exa ผ่าน CLI ของ inference.sh แอปพลิเคชัน: Tavily Search, Tavily Extract, Exa Search, Exa Answer, Exa Extract ความสามารถ: การค้นหาด้วย AI, การดึงเนื้อหา, คำตอบโดยตรง, การวิจัย ใช้สำหรับ: การวิจัย, ไพพ์ไลน์ RAG, การตรวจสอบข้อเท็จจริง, การรวบรวมเนื้อหา, เอเจนต์ ทริกเกอร์: การค้นหาเว็บ, tavily, exa, search api, การดึงเนื้อหา, การวิจัย, การค้นหาทางอินเทอร์เน็ต, การค้นหาด้วย AI, ผู้ช่วยค้นหา, การขูดเว็บ, rag, ทางเลือกของ Perplexity
researchweb-scrapingapi
infsh-cli
halt-catch-fire
เรียกใช้แอป AI กว่า 250 รายการผ่าน CLI inference.sh - สร้างภาพ, สร้างวิดีโอ, LLM, ค้นหา, 3D, อัตโนมัติ Twitter โมเดล: FLUX, Veo, Gemini, Grok, Claude, Seedance, OmniHuman, Tavily, Exa, OpenRouter และอื่นๆ อีกมากมาย ใช้เมื่อเรียกใช้แอป AI, สร้างภาพ/วิดีโอ, เรียก LLM, ค้นหาเว็บ หรืออัตโนมัติ Twitter ทริกเกอร์: inference.sh, infsh, ai model, run ai, serverless ai, ai api, flux, veo, claude api, image generation, video generation, openrouter, tavily, exa search, twitter api, grok
developmentapicreative
landing-page-design
halt-catch-fire
We need to translate the given text from English to Thai. The text is about landing page conversion optimization. We must preserve the name "landing-page-design" but it's not in the text, so we ignore. We translate the entire text inside <text>. No extra commentary, no labels. Keep technical terms like "CTA", "SaaS", "F-pattern" as is or translate? The instruction says preserve technical terms, so we keep "CTA", "SaaS", "F-pattern" as English. Also "above-the-fold" is a term, we can translate or keep? Better to keep as is or translate? The instruction says "preserve product names, protocol names, URLs, numbers, and technical terms." So "above-the-fold" is a technical term, we can keep as "above-the-fold" or translate? I think it's common to keep in English in Thai context. But to be safe, we can translate the concept. However, the instruction says preserve technical terms, so I'll keep "above-the-fold" as is. Similarly
product-photography
halt-catch-fire
AI การถ่ายภาพผลิตภัณฑ์ด้วยแสงในสตูดิโอ ภาพไลฟ์สไตล์ และรูปแบบแพ็กช็อต ครอบคลุมมุมกล้อง พื้นหลัง ประเภทเงา ภาพฮีโร่ และข้อกำหนดภาพสำหรับอีคอมเมิร์ซ ใช้สำหรับ: ภาพผลิตภัณฑ์ ภาพอีคอมเมิร์ซ รายการสินค้าใน Amazon แพ็กช็อต ภาพถ่ายไลฟ์สไตล์ ทริกเกอร์: การถ่ายภาพผลิตภัณฑ์ ภาพผลิตภัณฑ์ แพ็กช็อต การถ่ายภาพอีคอมเมิร์ซ ภาพช็อตผลิตภัณฑ์ ภาพผลิตภัณฑ์ การถ่ายภาพในสตูดิโอ ผลิตภัณฑ์ไลฟ์สไตล์ ภาพผลิตภัณฑ์ Amazon ภาพรายการสินค้า ภาพฮีโร่ ภาพจำลองผลิตภัณฑ์...
creativeecommerceimage