firecrawl-crawl
Trích xuất nội dung hàng loạt từ toàn bộ trang web hoặc các phần của trang web với bộ lọc độ sâu và đường dẫn. Thu thập các trang theo liên kết đến giới hạn độ sâu và số trang có thể cấu hình, với bộ lọc bao gồm/loại trừ đường dẫn để phạm vi trích xuất. Hỗ trợ thăm dò công việc không đồng bộ hoặc chờ đồng bộ với hiển thị tiến trình qua cờ --wait và --progress. Cung cấp kiểm soát đồng thời, độ trễ yêu cầu và định dạng đầu ra JSON để tích hợp vào quy trình làm việc của tác nhân. Là một phần của mô hình leo thang bốn bước: tìm kiếm → cạo →...
npx skills add https://github.com/firecrawl/cli --skill firecrawl-crawlfirecrawl crawl
Bulk extract content from a website. Crawls pages following links up to a depth/limit.
Prerequisite: crawl requires authentication (no keyless free tier); without credentials the CLI prompts an interactive login.
Quick start
# Crawl a docs section
firecrawl crawl "<url>" --include-paths /docs --limit 50 --wait -o .firecrawl/crawl.json
# Full crawl with depth limit
firecrawl crawl "<url>" --max-depth 3 --wait --progress -o .firecrawl/crawl.json
# Check status of a running crawl
firecrawl crawl <job-id>
Run firecrawl crawl --help for the full option list.
Done when: the crawl reaches a terminal status and the saved output under .firecrawl/ contains the expected pages.
Tips
- Use
--waitwhen you need the results immediately. It has no default timeout; use--timeout <seconds>to bound polling. Without--wait, crawl returns a job ID for async polling. - Scope crawls with
--include-pathswhenever the request names a section — crawl only the pages you need. - Crawl consumes credits per page. Check
firecrawl credit-usagebefore large crawls (credit-usagerequires authentication).
See also
- firecrawl-scrape — scrape individual pages
- firecrawl-map — discover URLs before deciding to crawl
- firecrawl-download — download site to local files (uses map + scrape)
- firecrawl-build-scrape — building bulk extraction into an app instead of running it here