ai-avatar-video

Crea videos de avatares IA y cabezas parlantes mediante la CLI de inference.sh. Recomendado: P-Video-Avatar (el más rápido, más económico, con TTS integrado). También: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (más de 100 idiomas, control de emociones para personajes), ElevenLabs, Kokoro. Capacidades: avatares controlados por audio, texto a avatar, videos con sincronización de labios, generación de cabezas parlantes, presentadores virtuales, contenido UGC. Uso para: presentadores IA, videos explicativos, influencers virtuales, doblaje, videos de marketing, anuncios UGC, avatares para juegos,...

npx skills add https://github.com/skills-101/superpowers --skill ai-avatar-video

Install the belt CLI skill: npx skills add belt-sh/cli

AI Avatar & Talking Head Videos

Create AI avatars and talking head videos via inference.sh CLI.

AI Avatar & Talking Head Videos

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS)
belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "voice_script": "Hello, welcome to our product demo!",
  "voice": "Zephyr (Female)"
}'

Available Models

Start with P-Video-Avatar — it's 18x faster and 6x cheaper than alternatives, with built-in TTS, dynamic backgrounds, and 1080p support.

ModelApp IDBest ForBuilt-in TTS
P-Video-Avatarpruna/p-video-avatarBest overall: speed, cost, quality, controlYes (30 voices, 10 languages)
OmniHuman 1.5bytedance/omnihuman-1-5Multi-character, audio-drivenNo
Fabric 1.0falai/fabric-1-0Image talks with lipsyncYes
PixVerse Lipsyncfalai/pixverse-lipsyncHighly realistic lipsyncNo

Cost & Speed Comparison

ModelSpeed (per sec of video)Cost per second
P-Video-Avatar~1.83s/s$0.025
OmniHuman 1.5~28s/s (15x slower)$0.16 (6.4x more)
Fabric 1.0~34s/s (18x slower)$0.14 (5.6x more)

Examples

P-Video-Avatar (Recommended)

Generate avatar from portrait + text script with built-in TTS:

belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "voice_script": "Welcome to our product walkthrough. Today I will show you three key features.",
  "voice": "Puck (Male)",
  "voice_language": "English (US)",
  "resolution": "720p"
}'

With custom style control:

belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "voice_script": "This is exciting news!",
  "voice": "Aoede (Female)",
  "voice_prompt": "Enthusiastic and energetic tone",
  "video_prompt": "The person is presenting on stage with dramatic lighting",
  "resolution": "1080p"
}'

With audio file instead of TTS:

belt app run pruna/p-video-avatar --input '{
  "image": "https://portrait.jpg",
  "audio": "https://speech.mp3"
}'

Full Workflow: Generate Portrait + Avatar

Use Pruna P-Image to generate the portrait, then create the avatar:

# 1. Generate a portrait image
belt app run pruna/p-image --input '{
  "prompt": "professional headshot portrait of a young woman, neutral background, looking at camera, studio lighting, photorealistic",
  "aspect_ratio": "9:16"
}'

# 2. Create avatar video with built-in TTS
belt app run pruna/p-video-avatar --input '{
  "image": "<image-url-from-step-1>",
  "voice_script": "Hi there! Let me walk you through our latest features.",
  "voice": "Zephyr (Female)"
}'

OmniHuman 1.5 (Multi-Character)

belt app run bytedance/omnihuman-1-5 --input '{
  "image_url": "https://portrait.jpg",
  "audio_url": "https://speech.mp3"
}'

Supports specifying which character to drive in multi-person images.

Fabric 1.0 (Image Talks)

belt app run falai/fabric-1-0 --input '{
  "image_url": "https://face.jpg",
  "audio_url": "https://audio.mp3"
}'

PixVerse Lipsync

belt app run falai/pixverse-lipsync --input '{
  "image_url": "https://portrait.jpg",
  "audio_url": "https://speech.mp3"
}'

Full Workflow: TTS + Avatar (Non-TTS Models)

For models without built-in TTS (OmniHuman, PixVerse), generate speech first:

# 1. Generate speech — Inworld TTS-2 for expressive character voices
belt app run inworld/text-to-speech-2 --input '{
  "text": "[friendly] Welcome to our product demo! [excited] Let me show you three features that will change how you work.",
  "voice_id": "Sarah",
  "delivery_mode": "CREATIVE"
}' > speech.json

# 2. Create avatar video with the speech
belt app run bytedance/omnihuman-1-5 --input '{
  "image_url": "https://presenter-photo.jpg",
  "audio_url": "<audio-url-from-step-1>"
}'

Tip: For most use cases, P-Video-Avatar with built-in TTS is simpler — no separate audio step needed. Use this workflow only when you specifically need OmniHuman (multi-character) or PixVerse (realistic lipsync).

Full Workflow: Dub Video in Another Language

# 1. Transcribe original video
belt app run infsh/fast-whisper-large-v3 --input '{"audio_url": "https://video.mp4"}' > transcript.json

# 2. Translate text (manually or with an LLM)

# 3. Generate speech in new language
belt app run infsh/kokoro-tts --input '{"text": "<translated-text>"}' > new_speech.json

# 4. Lipsync the original video with new audio
belt app run infsh/latentsync-1-6 --input '{
  "video_url": "https://original-video.mp4",
  "audio_url": "<new-audio-url>"
}'

Avatar UGC Generation

Create UGC-style content with P-Video-Avatar — built-in TTS, no separate audio step needed:

# 1. Generate a relatable UGC-style portrait
belt app run pruna/p-image --input '{
  "prompt": "casual selfie-style photo of a young woman in a cozy room, natural lighting, looking at camera, warm smile, authentic feel",
  "aspect_ratio": "9:16"
}'

# 2. Create UGC avatar video with built-in TTS
belt app run pruna/p-video-avatar --input '{
  "image": "<image-url-from-step-1>",
  "voice_script": "Okay so I just tried this product and honestly? It is a game changer. I was not expecting to love it this much but here we are!",
  "voice": "Zephyr (Female)",
  "voice_prompt": "Excited, casual, authentic tone like talking to a friend",
  "video_prompt": "The person is talking casually to camera in their room, natural gestures",
  "resolution": "1080p"
}'

Why P-Video-Avatar for UGC

  • All-in-one — built-in TTS means no separate audio generation step
  • 30 voices, 10 languages — match your target audience
  • Voice + video prompts — control tone, emotion, body language, and background independently
  • 18x faster, 6x cheaper — produce UGC at scale vs. Fabric/OmniHuman/HeyGen
  • 1080p support — platform-ready vertical video from a single portrait image

Batch UGC: Same Product, Multiple Presenters

# Generate 3 different presenters
for voice in "Zephyr (Female)" "Puck (Male)" "Aoede (Female)"; do
  belt app run pruna/p-video-avatar --input "{
    \"image\": \"https://portrait.jpg\",
    \"voice_script\": \"This changed my morning routine completely. Five minutes and I am done.\",
    \"voice\": \"$voice\",
    \"voice_prompt\": \"Casual, authentic, like a real testimonial\",
    \"video_prompt\": \"Person talking to camera in a bright kitchen\",
    \"resolution\": \"1080p\"
  }"
done

Use Cases

  • UGC & Marketing: Product demos, UGC-style ads with AI presenters
  • Education: Course videos, explainers
  • Localization: Dub content across 10 languages from one image
  • Social Media: Consistent virtual influencer content
  • Corporate: Training videos, announcements
  • Gaming: Character avatars, NPC dialogue

Tips

  • Use high-quality portrait photos (front-facing, good lighting)
  • Audio should be clear with minimal background noise
  • P-Video-Avatar supports built-in TTS — no need for a separate speech generation step
  • P-Video-Avatar output aspect ratio matches the input image
  • Generate portraits with pruna/p-image using 9:16 aspect ratio for vertical videos
  • OmniHuman 1.5 supports multiple people in one image
  • LatentSync is best for syncing existing videos to new audio

Related Skills

# Dedicated P-Video-Avatar skill
npx skills add inference-sh/skills@p-video-avatar

# Full platform skill (all apps)
npx skills add inference-sh/skills@infsh-cli

# Text-to-speech (generate audio for non-TTS avatar models)
npx skills add inference-sh/skills@text-to-speech

# Speech-to-text (transcribe for dubbing)
npx skills add inference-sh/skills@speech-to-text

# Video generation
npx skills add inference-sh/skills@ai-video-generation

# Image generation (create avatar images)
npx skills add inference-sh/skills@ai-image-generation

Browse all video apps: belt app list --category video

Documentation

Más skills de skills-101

ai-image-generation
skills-101
Genera imágenes de IA con GPT-Image-2, FLUX, Gemini, Grok, Seedream, Reve y más de 50 modelos mediante la CLI de inference.sh. Modelos: GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt. Capacidades: text-to-image, image-to-image, inpainting, LoRA, edición de imágenes, upscaling, renderizado de texto. Úsalo para: arte con IA, maquetas de productos, concept art, gráficos para redes sociales, visuales de marketing, ilustraciones. Desencadenantes: flux, image generation, ai image, text to...
agent-browser
skills-101
Automatización de navegador para agentes de IA mediante inference.sh. Navega páginas web, interactúa con elementos usando referencias @e, toma capturas de pantalla, graba video. Capacidades: extracción web, llenado de formularios, clics, escritura, arrastrar y soltar, carga de archivos, ejecución de JavaScript. Úsalo para: automatización web, extracción de datos, pruebas, navegación de agentes, investigación. Disparadores: navegador, automatización web, extraer, navegar, clic, llenar formulario, captura de pantalla, navegar web, playwright, navegador headless, agente web, navegar por internet, grabar video
agent-tools
skills-101
Ejecuta aplicaciones de IA mediante la CLI de inference.sh: generación de imágenes, creación de videos, LLMs, búsqueda, 3D, automatización de Twitter. Modelos: FLUX, Veo, Gemini, Grok, Claude, Seedance, OmniHuman, Tavily, Exa, OpenRouter y muchos más. Úsalo al ejecutar aplicaciones de IA, generar imágenes/videos, llamar a LLMs, búsqueda web o automatizar Twitter. Disparadores: inference.sh, infsh, ai model, run ai, serverless ai, ai api, flux, veo, claude api, image generation, video generation, openrouter, tavily, exa search, twitter api, grok
python-executor
skills-101
Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh). Pre-installed: NumPy, Pandas, Matplotlib, requests, BeautifulSoup, Selenium, Playwright, MoviePy, Pillow, OpenCV, trimesh, and 100+ more libraries. Use for: data processing, web scraping, image manipulation, video creation, 3D model processing, PDF generation, API calls, automation scripts. Triggers: python, execute code, run script, web scraping, data analysis, image processing, video editing, 3D...
remotion-render
skills-101
Renderiza videos a partir de código de componentes React/Remotion mediante inference.sh. Pasa código TSX, obtén MP4. Compatible con todas las APIs de Remotion: useCurrentFrame, useVideoConfig, spring, interpolate, AbsoluteFill, Sequence. Resolución, FPS, duración y códec configurables. Úsalo para: generación programática de videos, gráficos animados, motion design, videos basados en datos, animaciones de React a video. Disparadores: remotion, render video from code, tsx to video, react video, programmatic video, remotion render, code to video, animated...
infsh-cli
skills-101
Ejecuta aplicaciones de IA mediante la CLI de inference.sh: generación de imágenes, creación de videos, LLM, búsqueda, 3D, automatización de Twitter. Modelos: FLUX, Veo, Gemini, Grok, Claude, Seedance, OmniHuman, Tavily, Exa, OpenRouter y muchos más. Úsalo al ejecutar aplicaciones de IA, generar imágenes/videos, llamar a LLM, buscar en la web o automatizar Twitter. Desencadenantes: inference.sh, infsh, modelo de IA, ejecutar IA, IA sin servidor, API de IA, flux, veo, API de claude, generación de imágenes, generación de videos, openrouter, tavily, búsqueda de exa, API de twitter, grok
landing-page-design
skills-101
Optimización de conversión de landing pages con reglas de diseño, diseño de sección hero y psicología del CTA. Cubre la fórmula del above-the-fold, la colocación de pruebas sociales, el diseño móvil y la lectura en patrón F. Útil para: landing pages de startups, páginas de producto, marketing SaaS, optimización de conversión. Disparadores: landing page, sección hero, above the fold, optimización de conversión, diseño de landing page, botón CTA, imagen hero, layout de landing page, landing page SaaS, diseño de página de producto, tasa de conversión, landing page...
designmarketing
product-photography
skills-101
Fotografía de productos con IA que incluye iluminación de estudio, tomas de estilo de vida y convenciones de packshot. Cubre ángulos, fondos, tipos de sombra, tomas principales y requisitos de imágenes para comercio electrónico. Útil para: fotos de productos, imágenes de comercio electrónico, listados de Amazon, packshots, fotografía de estilo de vida. Activadores: fotografía de productos, foto de producto, packshot, fotografía de comercio electrónico, toma de producto, imagen de producto, fotografía de estudio, producto de estilo de vida, foto de producto de Amazon, imagen de listado de producto, toma principal, mockup de producto,...