video-ad-specs

Pembuatan iklan video dengan spesifikasi tepat sesuai platform untuk TikTok, Instagram, YouTube, Facebook, LinkedIn. Mencakup dimensi, batas durasi, kerangka AIDA, dan persyaratan teks. Gunakan untuk: iklan video, iklan media sosial, kreatif media berbayar, pemasaran video, produksi iklan. Pemicu: iklan video, iklan media sosial, iklan tiktok, iklan instagram, iklan youtube, iklan facebook, iklan linkedin, kreatif video, spesifikasi iklan, media berbayar, pemasaran video, produksi iklan, iklan reels, iklan stories, pre roll, bumper ad

npx skills add https://github.com/halt-catch-fire/skills --skill video-ad-specs

Install the belt CLI skill: npx skills add belt-sh/cli

Video Ad Specs

Create platform-specific video ads via inference.sh CLI.

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Generate a vertical video ad scene
belt app run bytedance/seedance-2-0 --input '{
  "prompt": "vertical video, person excitedly unboxing a product, clean modern room, bright natural lighting, social media ad style, authentic feeling",
  "ratio": "9:16",
  "generate_audio": true
}'

Platform Specifications

TikTok

SpecValue
Aspect ratio9:16 (vertical)
Resolution1080 x 1920 px
Duration5-60 seconds (15-30s recommended)
File sizeMax 500 MB
FormatMP4, MOV
SoundOn by default (design with sound)
Text safe zone150px from all edges
Hook window1 second — first frame must grab attention

Instagram Reels

SpecValue
Aspect ratio9:16 (vertical)
Resolution1080 x 1920 px
DurationUp to 90 seconds (15-30s for ads)
Cover imageSeparate upload, shows in grid
SoundOn by default
Caption areaBottom 20% reserved for text overlay

Instagram Stories

SpecValue
Aspect ratio9:16
Resolution1080 x 1920 px
DurationUp to 15 seconds per segment
Swipe-up/LinkAvailable for ads
Top/bottom14% top and 20% bottom = unsafe for key content

YouTube

FormatAspectDurationSkip
Bumper16:96 seconds exactlyNon-skippable
Non-skippable16:915 secondsNon-skippable
Skippable (TrueView)16:9Any lengthSkip after 5 seconds
Shorts9:16Up to 60 secondsN/A

Resolution: 1920 x 1080 (16:9) or 1080 x 1920 (Shorts)

Facebook Feed

SpecValue
Aspect ratio1:1 (square) or 4:5 (recommended for mobile)
Resolution1080 x 1080 or 1080 x 1350
DurationUp to 240 min (15-30s recommended)
AutoplaySilent — captions are essential
Sound85% of Facebook video is watched without sound

LinkedIn

SpecValue
Aspect ratio1:1 or 16:9
Resolution1080 x 1080 or 1920 x 1080
Duration3 seconds to 10 minutes (15-30s for ads)
ToneProfessional
AutoplaySilent in feed

AIDA Framework for Video Ads

PhaseTimeGoalTechnique
Attention0-3sStop the scrollPattern interrupt, bold visual, question
Interest3-10sKeep watchingState the problem, show relevance
Desire10-20sWant the solutionShow the product/outcome, social proof
ActionFinal 3-5sClick/buy/sign upClear CTA, urgency, offer

Hook Techniques (First 3 Seconds)

TechniqueExample
Bold statement"This tool replaced my entire marketing team"
Question"Why are you still doing this manually?"
Surprising visualUnexpected transformation, before/after reveal
Pattern interruptStart mid-action, unusual angle, bright color
Social proof"2 million people switched to this"
Pain point"If you hate [common frustration], watch this"

Creating Video Ads

Vertical (TikTok, Reels, Stories, Shorts)

# Hook scene (0-3s)
belt app run google/veo-3-1-fast --input '{
  "prompt": "vertical 9:16 video, close-up of hands struggling with tangled cables and messy desk, frustrated energy, shaky handheld camera, authentic social media style, bright lighting"
}'

# Solution reveal (3-15s)
belt app run bytedance/seedance-2-0 --input '{
  "prompt": "vertical video, smooth product reveal, clean wireless charging station on minimalist desk, satisfying organization transformation, bright modern room, social media ad aesthetic",
  "ratio": "9:16",
  "generate_audio": true
}'

# Add voiceover
belt app run falai/dia-tts --input '{
  "prompt": "[S1] Stop wasting time with this mess. This one product changed my entire setup. Everything charges. Everything is organized. Link in bio."
}'

# Merge video + audio
belt app run infsh/video-audio-merger --input '{
  "video": "solution-reveal.mp4",
  "audio": "voiceover.mp3"
}'

# Add captions (critical for silent autoplay)
belt app run infsh/caption-videos --input '{
  "video": "ad-with-audio.mp4",
  "caption_file": "captions.srt"
}'

Square (Facebook, LinkedIn Feed)

belt app run google/veo-3-1-fast --input '{
  "prompt": "square 1:1 video, professional person at desk discovering a new software tool, laptop screen showing clean dashboard, natural office lighting, corporate commercial style, satisfied expression"
}'

YouTube Bumper (6 Seconds)

# 6-second bumper: one message, one visual, one CTA
belt app run google/veo-3-1-fast --input '{
  "prompt": "6 second product ad, quick montage of a sleek app being used on phone, fast cuts, modern, energetic, brand logo reveal at end, punchy and dynamic, wide 16:9"
}'

# Keep it tight
belt app run falai/dia-tts --input '{
  "prompt": "[S1] Your reports. Automated. Try DataFlow free."
}'

Captions Are Mandatory

85% of Facebook and 40%+ of Instagram video is watched on mute.

Caption Best Practices

RuleReason
Always add captionsSilent viewing is the default on most platforms
Large, readable fontSmall text is invisible on mobile
High contrastWhite text with dark outline/background
Centered or bottom-thirdStandard viewing position
Max 2 lines at a timeMore text = can't be read fast enough
Key words in bold/colorDraws eye to important words
# Generate captions from audio
# (create SRT file from your script, then burn in)
belt app run infsh/caption-videos --input '{
  "video": "ad-video.mp4",
  "caption_file": "ad-captions.srt"
}'

Ad Structure Templates

Testimonial Ad (15-30s)

TimeContent
0-3sCustomer states the problem they had
3-15sHow they discovered and tried the product
15-25sThe specific result they achieved
25-30sProduct name + CTA

Demo Ad (15-30s)

TimeContent
0-3sThe problem (text or visual)
3-20sProduct demo showing the solution
20-25sKey result/benefit
25-30sCTA + offer

Before/After Ad (15s)

TimeContent
0-3s"Before" state (messy, slow, frustrating)
3-5sTransition / product introduction
5-12s"After" state (clean, fast, satisfying)
12-15sCTA

Common Mistakes

MistakeProblemFix
No hook in first 1-3sViewer scrolls pastOpen with pattern interrupt
Landscape video on TikTok/ReelsLetterboxed, looks amateurUse 9:16 for vertical platforms
No captionsMost viewers watch silentAlways add captions
CTA too lateViewers already leftClear CTA within last 5 seconds
Too long for platformForced skip or dropoutMatch platform duration norms
Same ad for all platformsWrong specs, wrong toneCreate platform-specific versions
Logo in first 3sFeels like a commercial, gets skippedSave branding for the end
Text in unsafe zonesCut off by platform UICheck safe zone per platform

Checklist

  • Correct aspect ratio for target platform
  • Hook in first 1-3 seconds
  • Captions added (readable, high contrast)
  • CTA clear and within final 5 seconds
  • Duration matches platform norms
  • Text outside platform unsafe zones
  • Audio designed for both sound-on and sound-off
  • Platform-specific version (not one-size-fits-all)

Related Skills

npx skills add inference-sh/skills@ai-video-generation
npx skills add inference-sh/skills@video-prompting-guide
npx skills add inference-sh/skills@text-to-speech
npx skills add inference-sh/skills@prompt-engineering

Browse all apps: belt app list

Lebih banyak skill dari halt-catch-fire

ai-image-generation
halt-catch-fire
We need to translate the given text from English to Bahasa Indonesia. The text describes an agent skill for AI image generation. We must preserve the name "ai-image-generation" but it's not in the text, so we don't include it. Also preserve product names, protocol names, URLs, numbers, technical terms. No extra commentary. Just translate the text inside <text>. The text: "Generate AI images with GPT-Image-2, FLUX, Gemini, Grok, Seedream, Reve and 50+ models via inference.sh CLI. Models: GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt. Capabilities: text-to-image, image-to-image, inpainting, LoRA, image editing, upscaling, text rendering. Use for: AI art, product mockups, concept art, social media graphics, marketing visuals, illustrations. Triggers: flux, image generation, ai image, text to...
creativemediaimage
ai-video-generation
halt-catch-fire
Hasilkan video AI dengan Google Veo, Seedance 2.0, HappyHorse, Wan, Grok dan 40+ model melalui CLI inference.sh. Model: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Kemampuan: teks-ke-video, gambar-ke-video, referensi-ke-video, pengeditan video, lipsync, animasi avatar, peningkatan resolusi video, suara foley. Gunakan untuk: video media sosial, konten pemasaran, video penjelasan, demo produk, avatar AI. Pemicu: pembuatan video, video AI,...
creativevideomedia
twitter-automation
halt-catch-fire
Otomatiskan Twitter/X dengan posting, interaksi, dan manajemen pengguna melalui CLI inference.sh. Aplikasi: x/post-tweet, x/post-create (dengan media), x/post-like, x/post-retweet, x/dm-send, x/user-follow. Kemampuan: memposting tweet, menjadwalkan konten, menyukai postingan, me-retweet, mengirim DM, mengikuti pengguna, mendapatkan profil. Gunakan untuk: otomatisasi media sosial, penjadwalan konten, bot interaksi, pertumbuhan audiens, API X. Pemicu: twitter api, x api, tweet automation, post to twitter, twitter bot, social media automation, x...
marketingapicommunication
ai-avatar-video
halt-catch-fire
We need to translate the given text from English to Bahasa Indonesia. The instruction says to preserve product names, protocol names, URLs, numbers, and technical terms. So names like "P-Video-Avatar", "OmniHuman", "Fabric", "PixVerse", "Inworld TTS-2", "ElevenLabs", "Kokoro", "CLI", "TTS", "UGC" should remain unchanged. Also "inference.sh" is a URL-like name, preserve. The text inside <text> is a description. We translate the rest naturally. The text: "Create AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC
videocreativemedia
agent-browser
halt-catch-fire
Otomatisasi peramban untuk agen AI melalui inference.sh. Navigasi halaman web, berinteraksi dengan elemen menggunakan referensi @e, ambil tangkapan layar, rekam video. Kemampuan: pengikisan web, pengisian formulir, klik, mengetik, seret-lepas, unggah berkas, eksekusi JavaScript. Gunakan untuk: otomatisasi web, ekstraksi data, pengujian, penjelajahan agen, riset. Pemicu: peramban, otomatisasi web, kikis, navigasi, klik, isi formulir, tangkapan layar, jelajahi web, playwright, peramban tanpa kepala, agen web, jelajahi internet, rekam video
browser-automationweb-scrapingtesting
web-search
halt-catch-fire
Pencarian web dan ekstraksi konten dengan Tavily dan Exa melalui CLI inference.sh. Aplikasi: Tavily Search, Tavily Extract, Exa Search, Exa Answer, Exa Extract. Kemampuan: pencarian berbasis AI, ekstraksi konten, jawaban langsung, riset. Gunakan untuk: riset, pipeline RAG, pemeriksaan fakta, agregasi konten, agen. Pemicu: pencarian web, tavily, exa, search api, ekstraksi konten, riset, pencarian internet, pencarian AI, asisten pencarian, web scraping, rag, alternatif perplexity
researchweb-scrapingapi
infsh-cli
halt-catch-fire
We need to translate the given text from English to Bahasa Indonesia, preserving the name "infsh-cli" and other technical terms, product names, protocol names, URLs, numbers, etc. The instruction says: "Do not include the name unless it appears in the source text." The name "infsh-cli" is not in the source text? Actually, the source text does not contain "infsh-cli" explicitly. The directory item type is "agent skill" and the name to preserve is "infsh-cli", but the text inside <text> does not include that name. So we should not add it. We just translate the text inside <text>. Also, we must not add labels like "description" or "skill name". Just the translation. The text: "Run 250+ AI apps via inference.sh CLI - image generation, video creation, LLMs, search, 3D, Twitter automation. Models: FLUX, Veo, Gemini, Grok, Claude, Seedance, OmniHuman, Tavily, Exa, Open
developmentapicreative
landing-page-design
halt-catch-fire
We need to translate the given text from English to Bahasa Indonesia. The instruction says to preserve product names, protocol names, URLs, numbers, technical terms. The name "landing-page-design" is to be preserved but not included unless it appears in source text. It does not appear in the <text> block. So we just translate the text. The text: "Landing page conversion optimization with layout rules, hero section design, and CTA psychology. Covers above-the-fold formula, social proof placement, mobile design, and F-pattern reading. Use for: startup landing pages, product pages, SaaS marketing, conversion optimization. Triggers: landing page, hero section, above the fold, conversion optimization, landing page design, cta button, hero image, landing page layout, saas landing page, product page design, conversion rate, landing page..." Translate carefully. Technical terms like "CTA", "above-the-fold", "F-pattern", "SaaS" should be preserved as is or translated? The instruction says preserve technical terms. Usually "CTA" is kept as