firecrawl-crawl

bởi firecrawl

Trích xuất nội dung hàng loạt từ toàn bộ trang web hoặc các phần của trang web với bộ lọc độ sâu và đường dẫn. Thu thập các trang theo liên kết đến giới hạn độ sâu và số trang có thể cấu hình, với bộ lọc bao gồm/loại trừ đường dẫn để phạm vi trích xuất. Hỗ trợ thăm dò công việc không đồng bộ hoặc chờ đồng bộ với hiển thị tiến trình qua cờ --wait và --progress. Cung cấp kiểm soát đồng thời, độ trễ yêu cầu và định dạng đầu ra JSON để tích hợp vào quy trình làm việc của tác nhân. Là một phần của mô hình leo thang bốn bước: tìm kiếm → cạo →...

npx skills add https://github.com/firecrawl/cli --skill firecrawl-crawl

firecrawl crawl

Bulk extract content from a website. Crawls pages following links up to a depth/limit.

When to use

  • You need content from many pages on a site (e.g., all /docs/)
  • You want to extract an entire site section
  • Step 4 in the workflow escalation pattern: search → scrape → map → crawl → interact

Quick start

# Crawl a docs section
firecrawl crawl "<url>" --include-paths /docs --limit 50 --wait -o .firecrawl/crawl.json

# Full crawl with depth limit
firecrawl crawl "<url>" --max-depth 3 --wait --progress -o .firecrawl/crawl.json

# Check status of a running crawl
firecrawl crawl <job-id>

Options

OptionDescription
--waitWait for crawl to complete before returning
--progressShow progress while waiting
--limit <n>Max pages to crawl
--max-depth <n>Max link depth to follow
--include-paths <paths>Only crawl URLs matching these paths
--exclude-paths <paths>Skip URLs matching these paths
--delay <ms>Delay between requests
--max-concurrency <n>Max parallel crawl workers
--prettyPretty print JSON output
-o, --output <path>Output file path

Tips

  • Always use --wait when you need the results immediately. Without it, crawl returns a job ID for async polling.
  • Use --include-paths to scope the crawl — don't crawl an entire site when you only need one section.
  • Crawl consumes credits per page. Check firecrawl credit-usage before large crawls.

See also

Thêm skills từ firecrawl

oracle
firecrawl
Các phương pháp hay nhất khi sử dụng CLI oracle (gộp lời nhắc + tệp, engine, phiên và các mẫu đính kèm tệp).
official
pinecone
firecrawl
Cơ sở dữ liệu vector được quản lý cho các ứng dụng AI sản xuất. Được quản lý hoàn toàn, tự động mở rộng, với tìm kiếm kết hợp (dense + sparse), lọc metadata và không gian tên.…
official
sentence-transformers
firecrawl
Khung cho các embedding câu, văn bản và hình ảnh tiên tiến nhất. Cung cấp hơn 5000 mô hình được huấn luyện sẵn cho độ tương đồng ngữ nghĩa, phân cụm và truy xuất.
official
wp-playground
firecrawl
Sử dụng cho quy trình làm việc WordPress Playground: các phiên bản WP dùng một lần nhanh trong trình duyệt hoặc cục bộ qua @wp-playground/cli (server, run-blueprint, build-snapshot),…
official
wp-plugin-development
firecrawl
Sử dụng khi phát triển plugin WordPress: kiến trúc và hooks, kích hoạt/hủy kích hoạt/gỡ cài đặt, giao diện quản trị và Settings API, lưu trữ dữ liệu, cron/tác vụ, bảo mật…
official
wp-project-triage
firecrawl
Sử dụng khi bạn cần kiểm tra xác định một kho lưu trữ WordPress (plugin/theme/block theme/WP core/Gutenberg/full site) bao gồm công cụ/kiểm tra/phiên bản…
official
wp-rest-api
firecrawl
Sử dụng khi xây dựng, mở rộng hoặc gỡ lỗi các endpoint/route của WordPress REST API: register_rest_route, các lớp WP_REST_Controller/controller, schema/argument…
official
wp-wpcli-and-ops
firecrawl
Sử dụng khi làm việc với WP-CLI (wp) cho các thao tác WordPress: tìm kiếm-thay thế an toàn, xuất/nhập db, quản lý plugin/chủ đề/người dùng/nội dung, cron, xóa bộ nhớ đệm,…
official