parallel-web-extract

Trích xuất nội dung từ nhiều URL song song, tiết kiệm token. Xử lý trang web, bài báo, PDF và các trang nặng JavaScript chỉ với một lệnh. Chạy trong ngữ cảnh fork để giảm thiểu chi phí token so với WebFetch tích hợp sẵn. Hỗ trợ trích xuất hàng loạt nhiều URL với mục tiêu tập trung tùy chọn. Yêu cầu cài đặt và xác thực parallel-cli; xuất nội dung đã trích xuất dưới dạng markdown vào tệp cục bộ cho các truy vấn tiếp theo.

npx skills add https://github.com/parallel-web/parallel-agent-skills --skill parallel-web-extract

URL Extraction

Extract content from: $ARGUMENTS

Command

Choose a short, descriptive filename based on the URL or content (e.g., vespa-docs, react-hooks-api). Use lowercase with hyphens, no spaces. Substitute it into the command inline$FILENAME is a placeholder, not a shell variable.

parallel-cli extract "$ARGUMENTS" --json -o "/tmp/$FILENAME.json"

Concrete example:

parallel-cli extract "https://docs.parallel.ai" --json -o "/tmp/parallel-docs.json"

Note: -o always saves JSON. The extension must be .json.

Options if needed:

  • --objective "focus area" to focus extraction on a specific goal (also silences the "neither objective nor search_queries" warning that V1 emits when neither is set)
  • -q "keyword" (repeatable) to prioritize keywords in excerpts
  • --full-content to include the complete page body (for long articles, PDFs, or when excerpts may not capture what you need)
  • --full-content-max-chars N to cap full-content size per result
  • --no-excerpts to strip excerpts when you only want full content

Handling failed extractions

If the response has an errors field, an empty results array, or a 404/timeout for the URL, do NOT fabricate content. Tell the user the extraction failed, surface the upstream status, and suggest:

  • Verifying the URL (the page may have moved)
  • Retrying with --full-content if excerpts came back empty but the page exists
  • Using parallel-cli search to locate the current URL if the page was renamed

Response format

Return content as:

Page Title

Then the extracted content verbatim, with these rules:

  • Keep content verbatim - do not paraphrase or summarize
  • Parse lists exhaustively - extract EVERY numbered/bulleted item
  • Strip only obvious noise: nav menus, footers, ads
  • Preserve all facts, names, numbers, dates, quotes

After the response, mention the output file path (/tmp/$FILENAME.json) so the user knows it's available for follow-up questions.

Setup

If parallel-cli is not found, install and authenticate:

/parallel:parallel-cli-setup

If parallel-cli extract returns 403, tell the user balance is likely required. Offer to run parallel-cli balance get, and if needed ask for explicit confirmation before running parallel-cli balance add <amount_cents>. Then retry the original extract command.

Thêm skills từ parallel-web

parallel-monitor
parallel-web
Liên tục theo dõi web để phát hiện thay đổi theo chu kỳ lặp lại. Sử dụng khi người dùng yêu cầu 'giám sát', 'theo dõi thay đổi của', 'xem xét', hoặc 'cảnh báo tôi khi' một điều gì đó…
migrate-to-parallel
parallel-web
Di chuyển hoàn toàn các tích hợp dữ liệu web từ Exa, Tavily, Perplexity hoặc Firecrawl sang các sản phẩm Parallel phù hợp trong khi vẫn bảo toàn hành vi của ứng dụng. Sử dụng…
parallel-data-enrichment
parallel-web
Làm giàu hàng loạt dữ liệu công ty, con người hoặc sản phẩm với các trường lấy từ web như tên CEO, thông tin tài trợ và liên hệ. Chấp nhận dữ liệu JSON nội tuyến hoặc tệp CSV; xuất kết quả đã làm giàu ra CSV. Chạy không đồng bộ với theo dõi tiến độ qua URL giám sát và lệnh thăm dò. Yêu cầu công cụ parallel-cli và kết nối internet; xử lý tập dữ liệu lớn với thời gian chờ có thể cấu hình. Hỗ trợ yêu cầu trường linh hoạt thông qua mô tả ý định ngôn ngữ tự nhiên (ví dụ: "tên CEO và năm thành lập").
parallel-deep-research
parallel-web
Nghiên cứu toàn diện với các tùy chọn điều chỉnh độ sâu, độ trễ và chi phí cho các chủ đề phức tạp. Ba cấp xử lý (pro-fast, ultra-fast, ultra) từ 30 giây đến 25 phút, với chi phí từ 1x đến 3x so với mức cơ bản. Thực thi bất đồng bộ với cơ chế polling: khởi tạo nghiên cứu ngay lập tức, theo dõi tiến độ qua URL, lấy kết quả khi sẵn sàng mà không bị chặn. Đầu ra bao gồm báo cáo định dạng markdown và siêu dữ liệu JSON; tóm tắt điều hành được in ra stdout để xem nhanh. Được thiết kế cho các...
parallel-findall
parallel-web
Khám phá các thực thể (công ty, con người, sản phẩm, v.v.) khớp với mô tả ngôn ngữ tự nhiên. Sử dụng khi người dùng yêu cầu 'tìm tất cả X' hoặc 'liệt kê mọi Y mà…' —…
parallel-memory
parallel-web
Nhớ lại các lần chạy Parallel Task, Monitor và FindAll trước đây khi chúng có thể hữu ích; loại bỏ các lần chạy hoặc xóa bộ nhớ khi được yêu cầu.
parallel-web-search
parallel-web
Tìm kiếm web nhanh để lấy thông tin hiện tại, nghiên cứu và tra cứu thực tế trên internet. Thực hiện các truy vấn dựa trên một mục tiêu duy nhất hoặc nhiều tìm kiếm từ khóa song song, trả về tối đa 10 kết quả kèm trích đoạn và siêu dữ liệu. Hỗ trợ lọc nhạy cảm với thời gian qua --after-date và tìm kiếm theo tên miền cụ thể với --include-domains. Xuất ra JSON có cấu trúc với tiêu đề, URL, ngày xuất bản và trích đoạn để dễ dàng phân tích và truy vấn tiếp theo. Yêu cầu trích dẫn nội tuyến cho mọi tuyên bố bằng markdown...
setup
parallel-web
Thiết lập plugin Parallel (cài đặt CLI)