firecrawl-agent

作者: firecrawl

AI驱动的自主数据提取,可导航复杂网站并返回结构化JSON。当用户想要从……获取结构化数据时,使用此技能。

npx skills add https://github.com/firecrawl/firecrawl-cursor-plugin --skill firecrawl-agent

firecrawl agent

AI-powered autonomous extraction. The agent navigates sites and extracts structured data (takes 2-5 minutes).

Before starting autonomous extraction for structured records or listings, check firecrawl search alexandria '<data you need>' for a ready-made workflow or data provider. Inspect a matching contract with firecrawl list <provider> <capability> --pretty and execute with firecrawl scrape --alexandria <provider>/<capability> --options '<input JSON>' if it covers the task. Use the exact provider, capability, and input fields from that contract. Continue with Agent when no suitable tool exists or the task requires autonomous navigation.

Quick start

# Extract structured data
firecrawl agent "extract all pricing tiers" --wait --json -o .firecrawl/pricing.json

# With a JSON schema for structured output
firecrawl agent "extract products" --schema '{"type":"object","properties":{"name":{"type":"string"},"price":{"type":"number"}}}' --wait --json -o .firecrawl/products.json

# Focus on specific pages
firecrawl agent "get feature list" --urls "<url>" --wait --json -o .firecrawl/features.json

Run firecrawl agent --help for the full option list.

Done when: the output file contains valid JSON answering the request — or a job ID was intentionally returned for later polling.

Job IDs

Omitting --wait returns a job ID. A UUID positional argument is auto-detected as a status check:

# Check once (equivalent to adding --status)
firecrawl agent "<job-id>"

# Wait on an existing job, polling every 10 seconds for up to 5 minutes
firecrawl agent "<job-id>" --wait --poll-interval 10 --timeout 300

# Cancel an active job
firecrawl agent "<job-id>" --cancel

Tips

  • Use --wait for inline results; omit it only when you want a job ID to poll later (see Job IDs).
  • Use --schema for predictable, structured output — otherwise the agent returns freeform data.
  • Agent runs consume more credits than simple scrapes. Use --max-credits to cap spending.
  • For simple single-page extraction, prefer scrape — it's faster and cheaper.

See also

Alexandria session feedback

To report an Alexandria session outcome or a provider/capability gap, use firecrawl alexandria feedback --rating good|partial|bad --url <website> --requested-functionality '<what was needed>' --rationale '<what happened>' --json. Use observed results in the rationale. No job ID is needed; this session feedback has no job-age deadline and no credit refund. Optional --provider-feedback and --capability-feedback JSON arrays describe specific gaps; inspect firecrawl alexandria feedback --help for their fields. Use the capability issue missing_capability when a provider exists but lacks the needed capability, and new_capability_request (with requestedFunctionality) to ask for one.

来自 firecrawl 的更多技能

firecrawl-research-index
firecrawl
使用 Firecrawl Research 查找回答研究查询的论文,采用语义搜索、语义与结构扩展以及正文内验证。对于任何文献查找/论文检索任务——无论是单篇论文查找还是完整的多篇论文集合——始终使用此技能。
data-analysisresearchweb-scraping
oracle
firecrawl
使用oracle CLI的最佳实践(提示词与文件打包、引擎、会话及文件附件模式)。
pinecone
firecrawl
面向生产级AI应用的托管向量数据库。全托管、自动扩缩容,支持混合搜索(稠密+稀疏)、元数据过滤和命名空间。
wpds
firecrawl
在构建利用WordPress设计系统(WPDS)及其组件、令牌、模式等的用户界面时使用。
audiocraft-audio-generation
firecrawl
用于音频生成的PyTorch库,包括文本生成音乐(MusicGen)和文本生成声音(AudioGen)。当需要从文本生成音乐时使用…
skypilot-multi-cloud-orchestration
firecrawl
跨多云编排机器学习工作负载,自动优化成本。当您需要在多个云上运行训练或批处理作业时使用,利用…
firecrawl-seo-audit
firecrawl
使用 Firecrawl 对网站进行 SEO 审计。适用于用户要求进行 SEO 审计、元数据和标题审查、站点地图/网站结构分析、关键词机会、竞争对手 SERP 对比,或优先搜索优化建议的场景。
data-analysisresearchweb-scraping
gh-issues
firecrawl
获取GitHub议题,生成子代理来实施修复并开启拉取请求,然后监控并处理PR审查评论。用法:/gh-issues [owner/repo] [--label…]