ai-image-generation

作者: halt-catch-fire

通过 inference.sh CLI 使用 GPT-Image-2、FLUX、Gemini、Grok、Seedream、Reve 及 50 多个模型生成 AI 图像。模型包括:GPT-Image-2、FLUX Dev LoRA、FLUX.2 Klein LoRA、Gemini 3 Pro Image、Grok Imagine、Seedream 4.5、Reve、ImagineArt。功能:文本转图像、图像转图像、图像修复、LoRA、图像编辑、图像放大、文本渲染。用途:AI 艺术、产品模型、概念艺术、社交媒体图形、营销视觉、插图。触发词:flux、图像生成、AI 图像、文本转...

npx skills add https://github.com/halt-catch-fire/skills --skill ai-image-generation

Install the belt CLI skill: npx skills add belt-sh/cli

AI Image Generation

Generate images with 50+ AI models via inference.sh CLI.

AI Image Generation

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Generate an image with FLUX
belt app run falai/flux-dev-lora --input '{"prompt": "a cat astronaut in space"}'

Available Models

ModelApp IDBest For
GPT-Image-2openai/gpt-image-2Text-to-image, editing, inpainting
FLUX Dev LoRAfalai/flux-dev-loraHigh quality with custom styles
FLUX.2 Klein LoRAfalai/flux-2-klein-loraFast with LoRA support (4B/9B)
P-Imagepruna/p-imageFast, economical, multiple aspects
P-Image-LoRApruna/p-image-loraFast with preset LoRA styles
P-Image-Editpruna/p-image-editFast image editing
Gemini 3 Progoogle/gemini-3-pro-image-previewGoogle's latest
Gemini 2.5 Flashgoogle/gemini-2-5-flash-imageFast Google model
Grok Imaginexai/grok-imagine-imagexAI's model, multiple aspects
Seedream 4.5bytedance/seedream-4-52K-4K cinematic quality
Seedream 4.0bytedance/seedream-4-0High quality 2K-4K
Seedream 3.0bytedance/seedream-3-0-t2iAccurate text rendering
Revefalai/reveNatural language editing, text rendering
ImagineArt 1.5 Profalai/imagine-art-1-5-pro-previewUltra-high-fidelity 4K
FLUX Klein 4Bpruna/flux-klein-4bUltra-cheap ($0.0001/image)
Topaz Upscalerfalai/topaz-image-upscalerProfessional upscaling

Browse All Image Apps

belt app list --category image

Examples

GPT-Image-2

belt app run openai/gpt-image-2 --input '{
  "prompt": "professional product photo of sneakers, studio lighting",
  "quality": "high"
}'

GPT-Image-2 Editing

belt app run openai/gpt-image-2 --input '{
  "prompt": "change the background to a beach at sunset",
  "images": ["https://your-image.jpg"]
}'

Text-to-Image with FLUX

belt app run falai/flux-dev-lora --input '{
  "prompt": "professional product photo of a coffee mug, studio lighting"
}'

Fast Generation with FLUX Klein

belt app run falai/flux-2-klein-lora --input '{"prompt": "sunset over mountains"}'

Google Gemini 3 Pro

belt app run google/gemini-3-pro-image-preview --input '{
  "prompt": "photorealistic landscape with mountains and lake"
}'

Grok Imagine

belt app run xai/grok-imagine-image --input '{
  "prompt": "cyberpunk city at night",
  "aspect_ratio": "16:9"
}'

Reve (with Text Rendering)

belt app run falai/reve --input '{
  "prompt": "A poster that says HELLO WORLD in bold letters"
}'

Seedream 4.5 (4K Quality)

belt app run bytedance/seedream-4-5 --input '{
  "prompt": "cinematic portrait of a woman, golden hour lighting"
}'

Image Upscaling

belt app run falai/topaz-image-upscaler --input '{"image_url": "https://..."}'

Stitch Multiple Images

belt app run infsh/stitch-images --input '{
  "images": ["https://img1.jpg", "https://img2.jpg"],
  "direction": "horizontal"
}'

Related Skills

# Full platform skill (all apps)
npx skills add inference-sh/skills@infsh-cli

# Pruna P-Image (fast & economical)
npx skills add inference-sh/skills@p-image

# GPT-Image-2 (OpenAI)
npx skills add inference-sh/skills@gpt-image

# FLUX-specific skill
npx skills add inference-sh/skills@flux-image

# Upscaling & enhancement
npx skills add inference-sh/skills@image-upscaling

# Background removal
npx skills add inference-sh/skills@background-removal

# Video generation
npx skills add inference-sh/skills@ai-video-generation

# AI avatars from images
npx skills add inference-sh/skills@ai-avatar-video

Browse all apps: belt app list

Documentation

来自 halt-catch-fire 的更多技能

ai-video-generation
halt-catch-fire
通过 inference.sh CLI,使用 Google Veo、Seedance 2.0、HappyHorse、Wan、Grok 及 40 多个模型生成 AI 视频。模型包括:Veo 3.1、Veo 3、Seedance 2.0、HappyHorse 1.0、Wan 2.5、Grok Imagine Video、OmniHuman、Fabric、HunyuanVideo。功能涵盖:文生视频、图生视频、参考视频生成、视频编辑、唇形同步、虚拟人动画、视频增强、拟音音效。适用于:社交媒体视频、营销内容、解说视频、产品演示、AI 虚拟人。触发词:视频生成、AI 视频……
creativevideomedia
twitter-automation
halt-catch-fire
通过inference.sh CLI实现Twitter/X的自动化发帖、互动和用户管理。应用:x/post-tweet、x/post-create(支持媒体)、x/post-like、x/post-retweet、x/dm-send、x/user-follow。功能:发布推文、安排内容、点赞推文、转发、发送私信、关注用户、获取个人资料。用途:社交媒体自动化、内容排期、互动机器人、受众增长、X API。触发词:twitter api、x api、推文自动化、发布到twitter、twitter机器人、社交媒体自动化、x...
marketingapicommunication
ai-avatar-video
halt-catch-fire
通过inference.sh CLI创建AI虚拟形象和说话头像视频。推荐:P-Video-Avatar(最快、最便宜、内置TTS)。其他选项:OmniHuman、Fabric、PixVerse。音频:Inworld TTS-2(支持100+语言、角色情感控制)、ElevenLabs、Kokoro。功能:音频驱动虚拟形象、文本转虚拟形象、唇形同步视频、说话头像生成、虚拟主持人、UGC内容。用途:AI主持人、解说视频、虚拟网红、配音、营销视频、UGC广告、游戏虚拟形象……
videocreativemedia
agent-browser
halt-catch-fire
通过inference.sh为AI代理提供浏览器自动化功能。可导航网页、使用@e引用与元素交互、截图、录制视频。能力包括:网页抓取、表单填写、点击、输入、拖放、文件上传、JavaScript执行。适用于:网页自动化、数据提取、测试、代理浏览、研究。触发词:浏览器、网页自动化、抓取、导航、点击、填写表单、截图、浏览网页、playwright、无头浏览器、网页代理、上网、录制视频
browser-automationweb-scrapingtesting
web-search
halt-catch-fire
通过inference.sh CLI使用Tavily和Exa进行网络搜索与内容提取。应用:Tavily搜索、Tavily提取、Exa搜索、Exa回答、Exa提取。功能:AI驱动搜索、内容提取、直接回答、研究。用途:研究、RAG管道、事实核查、内容聚合、智能体。触发词:网络搜索、tavily、exa、搜索API、内容提取、研究、互联网搜索、AI搜索、搜索助手、网页抓取、rag、perplexity替代方案
researchweb-scrapingapi
infsh-cli
halt-catch-fire
通过inference.sh CLI运行250+款AI应用——图像生成、视频创作、大语言模型、搜索、3D、Twitter自动化。支持模型:FLUX、Veo、Gemini、Grok、Claude、Seedance、OmniHuman、Tavily、Exa、OpenRouter等。适用于运行AI应用、生成图像/视频、调用大语言模型、网络搜索或自动化Twitter操作。触发词:inference.sh、infsh、ai model、run ai、serverless ai、ai api、flux、veo、claude api、image generation、video generation、openrouter、tavily、exa search、twitter api、grok
developmentapicreative
landing-page-design
halt-catch-fire
着陆页转化优化,涵盖布局规则、首屏设计及CTA心理学。包含首屏公式、社交证明放置、移动端设计及F型阅读模式。适用场景:创业公司着陆页、产品页面、SaaS营销、转化优化。触发词:着陆页、首屏、首屏以上、转化优化、着陆页设计、CTA按钮、首屏图片、着陆页布局、SaaS着陆页、产品页面设计、转化率、着陆页……
product-photography
halt-catch-fire
We need to translate the given English text into Simplified Chinese. The instruction says to preserve product names, protocol names, URLs, numbers, and technical terms. The name "product-photography" is to be preserved only if it appears in the source text. It does not appear in the provided text, so we don't include it. We must not add any labels like "description" or "skill name". Just translate the text inside <text> directly. The text: "AI product photography with studio lighting, lifestyle shots, and packshot conventions. Covers angles, backgrounds, shadow types, hero shots, and e-commerce image requirements. Use for: product photos, e-commerce images, Amazon listings, packshots, lifestyle photography. Triggers: product photography, product photo, packshot, e-commerce photography, product shot, product image, studio photography, lifestyle product, amazon product photo, product listing image, hero shot, product mockup,..." We need to translate this into natural Chinese. Keep technical terms like "studio lighting", "lifestyle shots", "packshot",
creativeecommerceimage