agent-browser

tarafından halt-catch-fire

AI ajanları için inference.sh üzerinden tarayıcı otomasyonu. Web sayfalarında gezinme, @e referanslarıyla öğelerle etkileşim, ekran görüntüsü alma, video kaydetme. Yetenekler: web kazıma, form doldurma, tıklama, yazma, sürükle-bırak, dosya yükleme, JavaScript çalıştırma. Kullanım alanları: web otomasyonu, veri çıkarma, test, ajan taraması, araştırma. Tetikleyiciler: tarayıcı, web otomasyonu, kazıma, gezinme, tıklama, form doldurma, ekran görüntüsü, web'de gezinme, playwright, başsız tarayıcı, web ajanı, internette gezinme, video kaydetme

npx skills add https://github.com/halt-catch-fire/skills --skill agent-browser

Install the belt CLI skill: npx skills add belt-sh/cli

Agentic Browser

Browser automation for AI agents via inference.sh. Uses Playwright under the hood with a simple @e ref system for element interaction.

Agentic Browser

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Open a page and get interactive elements
belt app run agent-browser --function open --input '{"url": "https://example.com"}' --session new

Core Workflow

Every browser automation follows this pattern:

  1. Open - Navigate to URL, get @e refs for elements
  2. Interact - Use refs to click, fill, drag, etc.
  3. Re-snapshot - After navigation/changes, get fresh refs
  4. Close - End session (returns video if recording)
# 1. Start session
RESULT=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://example.com/login"
}')
SESSION_ID=$(echo $RESULT | jq -r '.session_id')
# Elements: @e1 [input] "Email", @e2 [input] "Password", @e3 [button] "Sign In"

# 2. Fill and submit
belt app run agent-browser --function interact --session $SESSION_ID --input '{
  "action": "fill", "ref": "@e1", "text": "user@example.com"
}'
belt app run agent-browser --function interact --session $SESSION_ID --input '{
  "action": "fill", "ref": "@e2", "text": "password123"
}'
belt app run agent-browser --function interact --session $SESSION_ID --input '{
  "action": "click", "ref": "@e3"
}'

# 3. Re-snapshot after navigation
belt app run agent-browser --function snapshot --session $SESSION_ID --input '{}'

# 4. Close when done
belt app run agent-browser --function close --session $SESSION_ID --input '{}'

Functions

FunctionDescription
openNavigate to URL, configure browser (viewport, proxy, video recording)
snapshotRe-fetch page state with @e refs after DOM changes
interactPerform actions using @e refs (click, fill, drag, upload, etc.)
screenshotTake page screenshot (viewport or full page)
executeRun JavaScript code on the page
closeClose session, returns video if recording was enabled

Interact Actions

ActionDescriptionRequired Fields
clickClick elementref
dblclickDouble-click elementref
fillClear and type textref, text
typeType text (no clear)text
pressPress key (Enter, Tab, etc.)text
selectSelect dropdown optionref, text
hoverHover over elementref
checkCheck checkboxref
uncheckUncheck checkboxref
dragDrag and dropref, target_ref
uploadUpload file(s)ref, file_paths
scrollScroll pagedirection (up/down/left/right), scroll_amount
backGo back in history-
waitWait millisecondswait_ms
gotoNavigate to URLurl

Element Refs

Elements are returned with @e refs:

@e1 [a] "Home" href="/"
@e2 [input type="text"] placeholder="Search"
@e3 [button] "Submit"
@e4 [select] "Choose option"
@e5 [input type="checkbox"] name="agree"

Important: Refs are invalidated after navigation. Always re-snapshot after:

  • Clicking links/buttons that navigate
  • Form submissions
  • Dynamic content loading

Features

Video Recording

Record browser sessions for debugging or documentation:

# Start with recording enabled (optionally show cursor indicator)
SESSION=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://example.com",
  "record_video": true,
  "show_cursor": true
}' | jq -r '.session_id')

# ... perform actions ...

# Close to get the video file
belt app run agent-browser --function close --session $SESSION --input '{}'
# Returns: {"success": true, "video": <File>}

Cursor Indicator

Show a visible cursor in screenshots and video (useful for demos):

belt app run agent-browser --function open --session new --input '{
  "url": "https://example.com",
  "show_cursor": true,
  "record_video": true
}'

The cursor appears as a red dot that follows mouse movements and shows click feedback.

Proxy Support

Route traffic through a proxy server:

belt app run agent-browser --function open --session new --input '{
  "url": "https://example.com",
  "proxy_url": "http://proxy.example.com:8080",
  "proxy_username": "user",
  "proxy_password": "pass"
}'

File Upload

Upload files to file inputs:

belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "upload",
  "ref": "@e5",
  "file_paths": ["/path/to/file.pdf"]
}'

Drag and Drop

Drag elements to targets:

belt app run agent-browser --function interact --session $SESSION --input '{
  "action": "drag",
  "ref": "@e1",
  "target_ref": "@e2"
}'

JavaScript Execution

Run custom JavaScript:

belt app run agent-browser --function execute --session $SESSION --input '{
  "code": "document.querySelectorAll(\"h2\").length"
}'
# Returns: {"result": "5", "screenshot": <File>}

Deep-Dive Documentation

ReferenceDescription
references/commands.mdFull function reference with all options
references/snapshot-refs.mdRef lifecycle, invalidation rules, troubleshooting
references/session-management.mdSession persistence, parallel sessions
references/authentication.mdLogin flows, OAuth, 2FA handling
references/video-recording.mdRecording workflows for debugging
references/proxy-support.mdProxy configuration, geo-testing

Ready-to-Use Templates

TemplateDescription
templates/form-automation.shForm filling with validation
templates/authenticated-session.shLogin once, reuse session
templates/capture-workflow.shContent extraction with screenshots

Examples

Form Submission

SESSION=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://example.com/contact"
}' | jq -r '.session_id')

# Get elements: @e1 [input] "Name", @e2 [input] "Email", @e3 [textarea], @e4 [button] "Send"

belt app run agent-browser --function interact --session $SESSION --input '{"action": "fill", "ref": "@e1", "text": "John Doe"}'
belt app run agent-browser --function interact --session $SESSION --input '{"action": "fill", "ref": "@e2", "text": "john@example.com"}'
belt app run agent-browser --function interact --session $SESSION --input '{"action": "fill", "ref": "@e3", "text": "Hello!"}'
belt app run agent-browser --function interact --session $SESSION --input '{"action": "click", "ref": "@e4"}'

belt app run agent-browser --function snapshot --session $SESSION --input '{}'
belt app run agent-browser --function close --session $SESSION --input '{}'

Search and Extract

SESSION=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://google.com"
}' | jq -r '.session_id')

belt app run agent-browser --function interact --session $SESSION --input '{"action": "fill", "ref": "@e1", "text": "weather today"}'
belt app run agent-browser --function interact --session $SESSION --input '{"action": "press", "text": "Enter"}'
belt app run agent-browser --function interact --session $SESSION --input '{"action": "wait", "wait_ms": 2000}'

belt app run agent-browser --function snapshot --session $SESSION --input '{}'
belt app run agent-browser --function close --session $SESSION --input '{}'

Screenshot with Video

SESSION=$(belt app run agent-browser --function open --session new --input '{
  "url": "https://example.com",
  "record_video": true
}' | jq -r '.session_id')

# Take full page screenshot
belt app run agent-browser --function screenshot --session $SESSION --input '{
  "full_page": true
}'

# Close and get video
RESULT=$(belt app run agent-browser --function close --session $SESSION --input '{}')
echo $RESULT | jq '.video'

Sessions

Browser state persists within a session. Always:

  1. Start with --session new on first call
  2. Use returned session_id for subsequent calls
  3. Close session when done

Related Skills

# Web search (for research + browse)
npx skills add inference-sh/skills@web-search

# LLM models (analyze extracted content)
npx skills add inference-sh/skills@llm-models

Documentation

halt-catch-fire tarafından daha fazla skill

ai-image-generation
halt-catch-fire
GPT-Image-2, FLUX, Gemini, Grok, Seedream, Reve ve inference.sh CLI üzerinden 50'den fazla model ile AI görselleri oluşturun. Modeller: GPT-Image-2, FLUX Dev LoRA, FLUX.2 Klein LoRA, Gemini 3 Pro Image, Grok Imagine, Seedream 4.5, Reve, ImagineArt. Yetenekler: metinden görsele, görselden görsele, iç boyama, LoRA, görsel düzenleme, yükseltme, metin oluşturma. Kullanım alanları: AI sanatı, ürün maketleri, konsept sanatı, sosyal medya grafikleri, pazarlama görselleri, illüstrasyonlar. Tetikleyiciler: flux, görsel oluşturma, ai görsel, metinden...
creativemediaimage
ai-video-generation
halt-catch-fire
Google Veo, Seedance 2.0, HappyHorse, Wan, Grok ve 40'tan fazla model ile inference.sh CLI üzerinden yapay zeka videoları oluşturun. Modeller: Veo 3.1, Veo 3, Seedance 2.0, HappyHorse 1.0, Wan 2.5, Grok Imagine Video, OmniHuman, Fabric, HunyuanVideo. Yetenekler: metinden videoya, görüntüden videoya, referanstan videoya, video düzenleme, dudak senkronizasyonu, avatar animasyonu, video yükseltme, foley ses. Kullanım alanları: sosyal medya videoları, pazarlama içerikleri, açıklayıcı videolar, ürün tanıtımları, yapay zeka avatarları. Tetikleyiciler: video oluşturma
creativevideomedia
twitter-automation
halt-catch-fire
Twitter/X otomasyonu: inference.sh CLI ile gönderi, etkileşim ve kullanıcı yönetimi. Uygulamalar: x/post-tweet, x/post-create (medya ile), x/post-like, x/post-retweet, x/dm-send, x/user-follow. Yetenekler: tweet gönderme, içerik planlama, gönderi beğenme, retweet yapma, DM gönderme, kullanıcı takip etme, profil alma. Kullanım alanları: sosyal medya otomasyonu, içerik planlama, etkileşim botları, kitle büyütme, X API. Tetikleyiciler: twitter api, x api, tweet otomasyonu, twitter'a gönderi, twitter botu, sosyal medya otomasyonu, x...
marketingapicommunication
ai-avatar-video
halt-catch-fire
We need to translate the given text from English to Turkish. The target language is Türkçe. The directory item type is agent skill, and the name to preserve is "ai-avatar-video". The instruction says: "Translate only the text inside <text>. Do not include the name unless it appears in the source text." The name "ai-avatar-video" does not appear in the source text, so we should not include it. Also, do not include labels like "description", "server name", or "skill name". Just translate the content. The text: "Create AI avatar and talking head videos via inference.sh CLI. Recommended: P-Video-Avatar (fastest, cheapest, built-in TTS). Also: OmniHuman, Fabric, PixVerse. Audio: Inworld TTS-2 (100+ languages, emotion steering for characters), ElevenLabs, Kokoro. Capabilities: audio-driven avatars, text-to-avatar, lipsync videos, talking head generation, virtual presenters, UGC content. Use for: AI
videocreativemedia
web-search
halt-catch-fire
We need to translate the given text from English to Turkish, preserving specific terms like "web-search", "Tavily", "Exa", "inference.sh CLI", "RAG", etc. The instruction says to preserve product names, protocol names, URLs, numbers, and technical terms. So "Tavily", "Exa", "inference.sh CLI", "RAG", "AI", "API" should remain as is. Also "web-search" is the name to preserve but it's not in the text? Actually the name is "web-search" but the text doesn't contain that exact string. The text has "web search" (two words). The instruction says "Do not include the name unless it appears in the source text." So we translate "web search" as "web araması" or similar? But careful: "web search" appears multiple times. We should translate it as "web araması" but keep technical terms like "Tavily Search" as is? The instruction says preserve product names. "Tavily Search
researchweb-scrapingapi
infsh-cli
halt-catch-fire
inference.sh CLI üzerinden 250'den fazla AI uygulamasını çalıştırın - görüntü oluşturma, video oluşturma, LLM'ler, arama, 3D, Twitter otomasyonu. Modeller: FLUX, Veo, Gemini, Grok, Claude, Seedance, OmniHuman, Tavily, Exa, OpenRouter ve daha fazlası. AI uygulamalarını çalıştırırken, görüntü/video oluştururken, LLM'leri çağırırken, web araması yaparken veya Twitter'ı otomatikleştirirken kullanın. Tetikleyiciler: inference.sh, infsh, ai model, run ai, serverless ai, ai api, flux, veo, claude api, image generation, video generation, openrouter, tavily, exa search, twitter api, grok
developmentapicreative
landing-page-design
halt-catch-fire
Açılış sayfası dönüşüm optimizasyonu; düzen kuralları, kahraman bölümü tasarımı ve CTA psikolojisi ile. Ekran üstü formülü, sosyal kanıt yerleşimi, mobil tasarım ve F-deseni okumayı kapsar. Kullanım alanları: startup açılış sayfaları, ürün sayfaları, SaaS pazarlama, dönüşüm optimizasyonu. Tetikleyiciler: açılış sayfası, kahraman bölümü, ekran üstü, dönüşüm optimizasyonu, açılış sayfası tasarımı, cta butonu, kahraman görseli, açılış sayfası düzeni, saas açılış sayfası, ürün sayfası tasarımı, dönüşüm oranı
product-photography
halt-catch-fire
Yapay zeka ile stüdyo aydınlatmalı ürün fotoğrafçılığı, yaşam tarzı çekimleri ve paket görüntüsü kuralları. Açılar, arka planlar, gölge türleri, kahraman çekimleri ve e-ticaret görsel gereksinimlerini kapsar. Kullanım alanları: ürün fotoğrafları, e-ticaret görselleri, Amazon listeleme görselleri, paket görüntüleri, yaşam tarzı fotoğrafçılığı. Tetikleyiciler: ürün fotoğrafçılığı, ürün fotoğrafı, paket görüntüsü, e-ticaret fotoğrafçılığı, ürün çekimi, ürün görseli, stüdyo fotoğrafçılığı, yaşam tarzı
creativeecommerceimage