parallel-data-enrichment

Làm giàu hàng loạt dữ liệu công ty, con người hoặc sản phẩm với các trường lấy từ web như tên CEO, thông tin tài trợ và liên hệ. Chấp nhận dữ liệu JSON nội tuyến hoặc tệp CSV; xuất kết quả đã làm giàu ra CSV. Chạy không đồng bộ với theo dõi tiến độ qua URL giám sát và lệnh thăm dò. Yêu cầu công cụ parallel-cli và kết nối internet; xử lý tập dữ liệu lớn với thời gian chờ có thể cấu hình. Hỗ trợ yêu cầu trường linh hoạt thông qua mô tả ý định ngôn ngữ tự nhiên (ví dụ: "tên CEO và năm thành lập").

npx skills add https://github.com/parallel-web/parallel-agent-skills --skill parallel-data-enrichment

Data Enrichment

Enrich: $ARGUMENTS

Before starting

Inform the user that enrichment may take several minutes depending on the number of rows and fields requested.

Optional: Suggest output columns

If the user gave a vague intent ("enrich these companies with useful info") and you're not sure what columns to add, ask the API for a suggestion before kicking off the run:

parallel-cli enrich suggest "Find CEO and recent funding info" --json

The response is an envelope: {title, processor, enriched_columns, warnings}. Extract just the enriched_columns array (not the whole envelope) and pass it as the value of --enriched-columns on enrich run, in place of --intent — the two flags are alternative ways to specify what to enrich, not combined. If suggest returned a processor, pass it through explicitly via --processor on the run call (it's a tuned recommendation for the schema). Skip this whole section if the user already specified the fields they want.

enrich suggest requires parallel-cli ≥ 0.3.0. If it errors with anything resembling no such command / No such command / unknown command, do not bail — skip the suggestion step, fall through to step 1 with --intent, complete the run, and mention parallel-cli update (or pipx upgrade parallel-web-tools) in the final response so the user picks up the feature next time.

Step 1: Start the enrichment

Use ONE of these command patterns (substitute user's actual data):

For inline data:

parallel-cli enrich run --data '[{"company": "Google"}, {"company": "Microsoft"}]' --intent "CEO name and founding year" --target "output.csv" --no-wait --json

For CSV file:

parallel-cli enrich run --source-type csv --source "input.csv" --target "output.csv" --source-columns '[{"name": "company", "description": "Company name"}]' --intent "CEO name and founding year" --no-wait --json

If this is a follow-up to a previous research task and you have its interaction_id, add context chaining:

parallel-cli enrich run --data '...' --intent "..." --target "output.csv" --no-wait --json --previous-interaction-id "$INTERACTION_ID"

The enrichment will run with the full context of that prior research — so you can enrich entities discovered earlier without restating what was already found. Note: enrichment does not itself produce a new interaction_id, so you cannot chain a further follow-up off of an enrichment.

IMPORTANT: Always include --no-wait so the command returns immediately instead of blocking.

Parse the --json output to extract taskgroup_id and url. The output is {taskgroup_id, url, num_runs} — there is no interaction_id field, do not look for one. Immediately tell the user:

  • Enrichment has been kicked off
  • The monitoring URL where they can track progress

Tell them they can background the polling step to continue working while it runs.

Step 2: Poll for results

Pick a concrete output path (e.g., /tmp/enrichment-acme.json). Note: the file is JSON regardless of the extension you choose — it's an array of {input, output} objects, not a CSV. Name it .json to avoid confusing yourself or the user.

parallel-cli enrich poll "$TASKGROUP_ID" --timeout 540 --output "/tmp/enrichment-<descriptive-name>.json"

Important:

  • Use --timeout 540 (9 minutes) to stay within tool execution limits
  • The --target from step 1 is unused in --no-wait mode — only --output here determines where results are saved, and the file is always JSON

If the poll times out

Enrichment of large datasets can take longer than 9 minutes. If the poll exits without completing:

  1. Tell the user the enrichment is still running server-side
  2. Re-run the same parallel-cli enrich poll command to continue waiting

Response format

After step 1: Share the monitoring URL (for tracking progress).

After step 2:

  1. Report number of rows enriched
  2. Preview first few rows from the output file (it's a JSON array of {input, output} objects)
  3. Tell the user the full path to the output file

Do NOT re-share the monitoring URL after completion — the results are in the output file.

Setup

If parallel-cli is not found, install and authenticate:

/parallel:parallel-cli-setup

If any parallel-cli enrich command returns 403, tell the user balance is likely required. Offer to run parallel-cli balance get, and if needed ask for explicit confirmation before running parallel-cli balance add <amount_cents>. Then retry the original enrichment command.

Thêm skills từ parallel-web

parallel-monitor
parallel-web
Liên tục theo dõi web để phát hiện thay đổi theo chu kỳ lặp lại. Sử dụng khi người dùng yêu cầu 'giám sát', 'theo dõi thay đổi của', 'xem xét', hoặc 'cảnh báo tôi khi' một điều gì đó…
migrate-to-parallel
parallel-web
Di chuyển hoàn toàn các tích hợp dữ liệu web từ Exa, Tavily, Perplexity hoặc Firecrawl sang các sản phẩm Parallel phù hợp trong khi vẫn bảo toàn hành vi của ứng dụng. Sử dụng…
parallel-deep-research
parallel-web
Nghiên cứu toàn diện với các tùy chọn điều chỉnh độ sâu, độ trễ và chi phí cho các chủ đề phức tạp. Ba cấp xử lý (pro-fast, ultra-fast, ultra) từ 30 giây đến 25 phút, với chi phí từ 1x đến 3x so với mức cơ bản. Thực thi bất đồng bộ với cơ chế polling: khởi tạo nghiên cứu ngay lập tức, theo dõi tiến độ qua URL, lấy kết quả khi sẵn sàng mà không bị chặn. Đầu ra bao gồm báo cáo định dạng markdown và siêu dữ liệu JSON; tóm tắt điều hành được in ra stdout để xem nhanh. Được thiết kế cho các...
parallel-findall
parallel-web
Khám phá các thực thể (công ty, con người, sản phẩm, v.v.) khớp với mô tả ngôn ngữ tự nhiên. Sử dụng khi người dùng yêu cầu 'tìm tất cả X' hoặc 'liệt kê mọi Y mà…' —…
parallel-memory
parallel-web
Nhớ lại các lần chạy Parallel Task, Monitor và FindAll trước đây khi chúng có thể hữu ích; loại bỏ các lần chạy hoặc xóa bộ nhớ khi được yêu cầu.
parallel-web-extract
parallel-web
Trích xuất nội dung từ nhiều URL song song, tiết kiệm token. Xử lý trang web, bài báo, PDF và các trang nặng JavaScript chỉ với một lệnh. Chạy trong ngữ cảnh fork để giảm thiểu chi phí token so với WebFetch tích hợp sẵn. Hỗ trợ trích xuất hàng loạt nhiều URL với mục tiêu tập trung tùy chọn. Yêu cầu cài đặt và xác thực parallel-cli; xuất nội dung đã trích xuất dưới dạng markdown vào tệp cục bộ cho các truy vấn tiếp theo.
parallel-web-search
parallel-web
Tìm kiếm web nhanh để lấy thông tin hiện tại, nghiên cứu và tra cứu thực tế trên internet. Thực hiện các truy vấn dựa trên một mục tiêu duy nhất hoặc nhiều tìm kiếm từ khóa song song, trả về tối đa 10 kết quả kèm trích đoạn và siêu dữ liệu. Hỗ trợ lọc nhạy cảm với thời gian qua --after-date và tìm kiếm theo tên miền cụ thể với --include-domains. Xuất ra JSON có cấu trúc với tiêu đề, URL, ngày xuất bản và trích đoạn để dễ dàng phân tích và truy vấn tiếp theo. Yêu cầu trích dẫn nội tuyến cho mọi tuyên bố bằng markdown...
setup
parallel-web
Thiết lập plugin Parallel (cài đặt CLI)