firecrawl-agent

bởi firecrawl

Trích xuất tự động được hỗ trợ bởi AI. Agent điều hướng các trang web và trích xuất dữ liệu có cấu trúc (mất 2-5 phút).

npx skills add https://github.com/firecrawl/firecrawl-claude-plugin --skill firecrawl-agent

firecrawl agent

AI-powered autonomous extraction. The agent navigates sites and extracts structured data (takes 2-5 minutes).

Before starting autonomous extraction for structured records or listings, check firecrawl search alexandria '<data you need>' for a ready-made workflow or data provider. Inspect a matching contract with firecrawl list <provider> <capability> --pretty and execute with firecrawl scrape --alexandria <provider>/<capability> --options '<input JSON>' if it covers the task. Use the exact provider, capability, and input fields from that contract. Continue with Agent when no suitable tool exists or the task requires autonomous navigation.

Quick start

# Extract structured data
firecrawl agent "extract all pricing tiers" --wait --json -o .firecrawl/pricing.json

# With a JSON schema for structured output
firecrawl agent "extract products" --schema '{"type":"object","properties":{"name":{"type":"string"},"price":{"type":"number"}}}' --wait --json -o .firecrawl/products.json

# Focus on specific pages
firecrawl agent "get feature list" --urls "<url>" --wait --json -o .firecrawl/features.json

Run firecrawl agent --help for the full option list.

Done when: the output file contains valid JSON answering the request — or a job ID was intentionally returned for later polling.

Job IDs

Omitting --wait returns a job ID. A UUID positional argument is auto-detected as a status check:

# Check once (equivalent to adding --status)
firecrawl agent "<job-id>"

# Wait on an existing job, polling every 10 seconds for up to 5 minutes
firecrawl agent "<job-id>" --wait --poll-interval 10 --timeout 300

# Cancel an active job
firecrawl agent "<job-id>" --cancel

Tips

  • Use --wait for inline results; omit it only when you want a job ID to poll later (see Job IDs).
  • Use --schema for predictable, structured output — otherwise the agent returns freeform data.
  • Agent runs consume more credits than simple scrapes. Use --max-credits to cap spending.
  • For simple single-page extraction, prefer scrape — it's faster and cheaper.

See also

Alexandria session feedback

To report an Alexandria session outcome or a provider/capability gap, use firecrawl alexandria feedback --rating good|partial|bad --url <website> --requested-functionality '<what was needed>' --rationale '<what happened>' --json. Use observed results in the rationale. No job ID is needed; this session feedback has no job-age deadline and no credit refund. Optional --provider-feedback and --capability-feedback JSON arrays describe specific gaps; inspect firecrawl alexandria feedback --help for their fields. Use the capability issue missing_capability when a provider exists but lacks the needed capability, and new_capability_request (with requestedFunctionality) to ask for one.

Thêm skills từ firecrawl

firecrawl-research-index
firecrawl
Tìm các bài báo trả lời truy vấn nghiên cứu với Firecrawl Research, sử dụng tìm kiếm ngữ nghĩa, mở rộng ngữ nghĩa và cấu trúc, cùng xác minh trong nội dung. Luôn sử dụng kỹ năng này cho bất kỳ nhiệm vụ tìm kiếm tài liệu/truy xuất bài báo nào — tra cứu một bài báo đơn lẻ hoặc toàn bộ bộ nhiều bài báo.
data-analysisresearchweb-scraping
oracle
firecrawl
Các phương pháp hay nhất khi sử dụng CLI oracle (gộp lời nhắc + tệp, engine, phiên và các mẫu đính kèm tệp).
pinecone
firecrawl
Cơ sở dữ liệu vector được quản lý cho các ứng dụng AI sản xuất. Được quản lý hoàn toàn, tự động mở rộng, với tìm kiếm kết hợp (dense + sparse), lọc metadata và không gian tên.…
wpds
firecrawl
Sử dụng khi xây dựng giao diện người dùng dựa trên Hệ thống Thiết kế WordPress (WPDS) và các thành phần, token, mẫu thiết kế, v.v. của nó.
audiocraft-audio-generation
firecrawl
Thư viện PyTorch để tạo âm thanh bao gồm chuyển văn bản thành nhạc (MusicGen) và chuyển văn bản thành âm thanh (AudioGen). Sử dụng khi bạn cần tạo nhạc từ văn bản…
skypilot-multi-cloud-orchestration
firecrawl
Điều phối đa đám mây cho khối lượng công việc ML với tối ưu hóa chi phí tự động. Sử dụng khi bạn cần chạy các công việc đào tạo hoặc xử lý hàng loạt trên nhiều đám mây, tận dụng…
firecrawl-seo-audit
firecrawl
Kiểm tra SEO của một trang web với Firecrawl. Sử dụng khi người dùng yêu cầu kiểm tra SEO, đánh giá metadata và tiêu đề, phân tích sitemap/cấu trúc trang web, cơ hội từ khóa, so sánh SERP đối thủ, hoặc các đề xuất tối ưu hóa tìm kiếm được ưu tiên.
data-analysisresearchweb-scraping
gh-issues
firecrawl
Lấy các issue GitHub, tạo sub-agent để triển khai sửa lỗi và mở PR, sau đó theo dõi và xử lý các nhận xét đánh giá PR. Cách dùng: /gh-issues [owner/repo] [--label…