ai-image-generation

作者: halt-catch-fire

透過 inference.sh CLI 使用 GPT-Image-2、FLUX、Gemini、Grok、Seedream、Reve 及 50 多種模型生成 AI 圖像。模型包括:GPT-Image-2、FLUX Dev LoRA、FLUX.2 Klein LoRA、Gemini 3 Pro Image、Grok Imagine、Seedream 4.5、Reve、ImagineArt。功能:文字轉圖像、圖像轉圖像、修補、LoRA、圖像編輯、放大、文字渲染。適用於:AI 藝術、產品模型、概念藝術、社交媒體圖形、行銷視覺、插圖。觸發詞:flux、圖像生成、AI 圖像、文字轉...

npx skills add https://github.com/halt-catch-fire/skills --skill ai-image-generation

Install the belt CLI skill: npx skills add belt-sh/cli

AI Image Generation

Generate images with 50+ AI models via inference.sh CLI.

AI Image Generation

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Generate an image with FLUX
belt app run falai/flux-dev-lora --input '{"prompt": "a cat astronaut in space"}'

Available Models

ModelApp IDBest For
GPT-Image-2openai/gpt-image-2Text-to-image, editing, inpainting
FLUX Dev LoRAfalai/flux-dev-loraHigh quality with custom styles
FLUX.2 Klein LoRAfalai/flux-2-klein-loraFast with LoRA support (4B/9B)
P-Imagepruna/p-imageFast, economical, multiple aspects
P-Image-LoRApruna/p-image-loraFast with preset LoRA styles
P-Image-Editpruna/p-image-editFast image editing
Gemini 3 Progoogle/gemini-3-pro-image-previewGoogle's latest
Gemini 2.5 Flashgoogle/gemini-2-5-flash-imageFast Google model
Grok Imaginexai/grok-imagine-imagexAI's model, multiple aspects
Seedream 4.5bytedance/seedream-4-52K-4K cinematic quality
Seedream 4.0bytedance/seedream-4-0High quality 2K-4K
Seedream 3.0bytedance/seedream-3-0-t2iAccurate text rendering
Revefalai/reveNatural language editing, text rendering
ImagineArt 1.5 Profalai/imagine-art-1-5-pro-previewUltra-high-fidelity 4K
FLUX Klein 4Bpruna/flux-klein-4bUltra-cheap ($0.0001/image)
Topaz Upscalerfalai/topaz-image-upscalerProfessional upscaling

Browse All Image Apps

belt app list --category image

Examples

GPT-Image-2

belt app run openai/gpt-image-2 --input '{
  "prompt": "professional product photo of sneakers, studio lighting",
  "quality": "high"
}'

GPT-Image-2 Editing

belt app run openai/gpt-image-2 --input '{
  "prompt": "change the background to a beach at sunset",
  "images": ["https://your-image.jpg"]
}'

Text-to-Image with FLUX

belt app run falai/flux-dev-lora --input '{
  "prompt": "professional product photo of a coffee mug, studio lighting"
}'

Fast Generation with FLUX Klein

belt app run falai/flux-2-klein-lora --input '{"prompt": "sunset over mountains"}'

Google Gemini 3 Pro

belt app run google/gemini-3-pro-image-preview --input '{
  "prompt": "photorealistic landscape with mountains and lake"
}'

Grok Imagine

belt app run xai/grok-imagine-image --input '{
  "prompt": "cyberpunk city at night",
  "aspect_ratio": "16:9"
}'

Reve (with Text Rendering)

belt app run falai/reve --input '{
  "prompt": "A poster that says HELLO WORLD in bold letters"
}'

Seedream 4.5 (4K Quality)

belt app run bytedance/seedream-4-5 --input '{
  "prompt": "cinematic portrait of a woman, golden hour lighting"
}'

Image Upscaling

belt app run falai/topaz-image-upscaler --input '{"image_url": "https://..."}'

Stitch Multiple Images

belt app run infsh/stitch-images --input '{
  "images": ["https://img1.jpg", "https://img2.jpg"],
  "direction": "horizontal"
}'

Related Skills

# Full platform skill (all apps)
npx skills add inference-sh/skills@infsh-cli

# Pruna P-Image (fast & economical)
npx skills add inference-sh/skills@p-image

# GPT-Image-2 (OpenAI)
npx skills add inference-sh/skills@gpt-image

# FLUX-specific skill
npx skills add inference-sh/skills@flux-image

# Upscaling & enhancement
npx skills add inference-sh/skills@image-upscaling

# Background removal
npx skills add inference-sh/skills@background-removal

# Video generation
npx skills add inference-sh/skills@ai-video-generation

# AI avatars from images
npx skills add inference-sh/skills@ai-avatar-video

Browse all apps: belt app list

Documentation

來自 halt-catch-fire 的更多技能

ai-video-generation
halt-catch-fire
We need to translate the given text from English to Traditional Chinese. The instruction says to preserve product names, protocol names, URLs, numbers, and technical terms. So names like Google Veo, Seedance 2.0, HappyHorse, Wan, Grok, inference.sh CLI, Veo 3.1, etc. should remain as is. Also preserve terms like text-to-video, image-to-video, etc. but translate the surrounding Chinese text. The text inside <text> is a description of an agent skill. We need to output only the translated text, no extra labels. Let's translate paragraph by paragraph: "Generate AI videos with Google Veo, Seedance 2.0, HappyHorse, Wan, Grok and 40+ models via inference.sh CLI." -> "通過 inference.sh CLI 使用 Google Veo、Seedance 2.0、HappyHorse、Wan、Grok 及 40 多個模型生成 AI 影片。" "Models: Veo 3.1, Veo 3,
creativevideomedia
twitter-automation
halt-catch-fire
We need to translate the given text from English to Traditional Chinese. The instruction says to preserve product names, protocol names, URLs, numbers, and technical terms. The name "twitter-automation" is to be preserved but not included unless it appears in the source text. The source text does not include that name, so we don't add it. We translate only the text inside <text>. No extra labels or commentary. The text: "Automate Twitter/X with posting, engagement, and user management via inference.sh CLI. Apps: x/post-tweet, x/post-create (with media), x/post-like, x/post-retweet, x/dm-send, x/user-follow. Capabilities: post tweets, schedule content, like posts, retweet, send DMs, follow users, get profiles. Use for: social media automation, content scheduling, engagement bots, audience growth, X API. Triggers: twitter api, x api, tweet automation, post to twitter, twitter bot, social media automation, x..." We need to translate terms like "Automate",
marketingapicommunication
ai-avatar-video
halt-catch-fire
透過 inference.sh CLI 建立 AI 虛擬角色與說話頭像影片。推薦:P-Video-Avatar(最快、最便宜、內建 TTS)。另可選:OmniHuman、Fabric、PixVerse。音訊:Inworld TTS-2(支援 100 多種語言、角色情感引導)、ElevenLabs、Kokoro。功能:音訊驅動虛擬角色、文字轉虛擬角色、唇形同步影片、說話頭像生成、虛擬主持人、UGC 內容。適用於:AI 主持人、解說影片、虛擬網紅、配音、行銷影片、UGC 廣告、遊戲虛擬角色等。
videocreativemedia
agent-browser
halt-catch-fire
透過 inference.sh 為 AI 代理提供瀏覽器自動化功能。可導覽網頁、使用 @e 參考與元素互動、擷取螢幕截圖、錄製影片。功能包括:網頁抓取、表單填寫、點擊、打字、拖放、檔案上傳、執行 JavaScript。適用於:網頁自動化、資料擷取、測試、代理瀏覽、研究。觸發詞:瀏覽器、網頁自動化、抓取、導覽、點擊、填寫表單、螢幕截圖、瀏覽網頁、playwright、無頭瀏覽器、網頁代理、上網、錄製影片
browser-automationweb-scrapingtesting
web-search
halt-catch-fire
透過 inference.sh CLI 使用 Tavily 和 Exa 進行網路搜尋與內容擷取。應用:Tavily 搜尋、Tavily 擷取、Exa 搜尋、Exa 回答、Exa 擷取。功能:AI 驅動搜尋、內容擷取、直接回答、研究。用途:研究、RAG 管線、事實查核、內容彙整、代理程式。觸發詞:網路搜尋、tavily、exa、搜尋 API、內容擷取、研究、網際網路搜尋、AI 搜尋、搜尋助理、網頁抓取、rag、perplexity 替代方案
researchweb-scrapingapi
infsh-cli
halt-catch-fire
透過 inference.sh CLI 執行 250 多個 AI 應用程式,包括圖片生成、影片創作、大型語言模型、搜尋、3D、Twitter 自動化。模型包含:FLUX、Veo、Gemini、Grok、Claude、Seedance、OmniHuman、Tavily、Exa、OpenRouter 等眾多選擇。適用於執行 AI 應用程式、生成圖片/影片、呼叫大型語言模型、網路搜尋或自動化 Twitter 時使用。觸發詞:inference.sh、infsh、ai model、run ai、serverless ai、ai api、flux、veo、claude api、image generation、video generation、openrouter、tavily、exa search、twitter api、grok
developmentapicreative
landing-page-design
halt-catch-fire
著陸頁轉換優化,包含佈局規則、英雄區塊設計與CTA心理學。涵蓋首屏公式、社會證明放置、行動設計與F型閱讀模式。適用於:新創著陸頁、產品頁面、SaaS行銷、轉換優化。觸發詞:著陸頁、英雄區塊、首屏、轉換優化、著陸頁設計、CTA按鈕、英雄圖片、著陸頁佈局、SaaS著陸頁、產品頁面設計、轉換率、著陸頁...
product-photography
halt-catch-fire
AI產品攝影,包含棚拍燈光、生活風格照及商品照慣例。涵蓋角度、背景、陰影類型、主視覺照及電商圖片需求。適用於:產品照片、電商圖片、Amazon商品頁、商品照、生活風格攝影。觸發詞:產品攝影、產品照片、商品照、電商攝影、產品拍攝、產品圖片、棚拍攝影、生活風格產品、亞馬遜產品照片、商品列表圖片、主視覺照、產品示意圖...
creativeecommerceimage