parallel-data-enrichment

작성자: parallel-web

회사, 인물 또는 제품 데이터를 CEO 이름, 자금 정보, 연락처 정보 등 웹에서 수집한 필드로 대량 보강합니다. 인라인 JSON 데이터 또는 CSV 파일을 입력받아 보강된 결과를 CSV로 출력합니다. 모니터링 URL 및 폴링 명령어를 통한 진행 상황 추적과 함께 비동기적으로 실행됩니다. parallel-cli 도구와 인터넷 접속이 필요하며, 구성 가능한 타임아웃으로 대규모 데이터셋을 처리합니다. 자연어 의도 설명(예: "CEO 이름 및 설립 연도")을 통해 유연한 필드 요청을 지원합니다.

npx skills add https://github.com/parallel-web/parallel-agent-skills --skill parallel-data-enrichment

Data Enrichment

Enrich: $ARGUMENTS

Before starting

Inform the user that enrichment may take several minutes depending on the number of rows and fields requested.

Optional: Suggest output columns

If the user gave a vague intent ("enrich these companies with useful info") and you're not sure what columns to add, ask the API for a suggestion before kicking off the run:

parallel-cli enrich suggest "Find CEO and recent funding info" --json

The response is an envelope: {title, processor, enriched_columns, warnings}. Extract just the enriched_columns array (not the whole envelope) and pass it as the value of --enriched-columns on enrich run, in place of --intent — the two flags are alternative ways to specify what to enrich, not combined. If suggest returned a processor, pass it through explicitly via --processor on the run call (it's a tuned recommendation for the schema). Skip this whole section if the user already specified the fields they want.

enrich suggest requires parallel-cli ≥ 0.3.0. If it errors with anything resembling no such command / No such command / unknown command, do not bail — skip the suggestion step, fall through to step 1 with --intent, complete the run, and mention parallel-cli update (or pipx upgrade parallel-web-tools) in the final response so the user picks up the feature next time.

Step 1: Start the enrichment

Use ONE of these command patterns (substitute user's actual data):

For inline data:

parallel-cli enrich run --data '[{"company": "Google"}, {"company": "Microsoft"}]' --intent "CEO name and founding year" --target "output.csv" --no-wait --json

For CSV file:

parallel-cli enrich run --source-type csv --source "input.csv" --target "output.csv" --source-columns '[{"name": "company", "description": "Company name"}]' --intent "CEO name and founding year" --no-wait --json

If this is a follow-up to a previous research task and you have its interaction_id, add context chaining:

parallel-cli enrich run --data '...' --intent "..." --target "output.csv" --no-wait --json --previous-interaction-id "$INTERACTION_ID"

The enrichment will run with the full context of that prior research — so you can enrich entities discovered earlier without restating what was already found. Note: enrichment does not itself produce a new interaction_id, so you cannot chain a further follow-up off of an enrichment.

IMPORTANT: Always include --no-wait so the command returns immediately instead of blocking.

Parse the --json output to extract taskgroup_id and url. The output is {taskgroup_id, url, num_runs} — there is no interaction_id field, do not look for one. Immediately tell the user:

  • Enrichment has been kicked off
  • The monitoring URL where they can track progress

Tell them they can background the polling step to continue working while it runs.

Step 2: Poll for results

Pick a concrete output path (e.g., /tmp/enrichment-acme.json). Note: the file is JSON regardless of the extension you choose — it's an array of {input, output} objects, not a CSV. Name it .json to avoid confusing yourself or the user.

parallel-cli enrich poll "$TASKGROUP_ID" --timeout 540 --output "/tmp/enrichment-<descriptive-name>.json"

Important:

  • Use --timeout 540 (9 minutes) to stay within tool execution limits
  • The --target from step 1 is unused in --no-wait mode — only --output here determines where results are saved, and the file is always JSON

If the poll times out

Enrichment of large datasets can take longer than 9 minutes. If the poll exits without completing:

  1. Tell the user the enrichment is still running server-side
  2. Re-run the same parallel-cli enrich poll command to continue waiting

Response format

After step 1: Share the monitoring URL (for tracking progress).

After step 2:

  1. Report number of rows enriched
  2. Preview first few rows from the output file (it's a JSON array of {input, output} objects)
  3. Tell the user the full path to the output file

Do NOT re-share the monitoring URL after completion — the results are in the output file.

Setup

If parallel-cli is not found, install and authenticate:

/parallel:parallel-cli-setup

If any parallel-cli enrich command returns 403, tell the user balance is likely required. Offer to run parallel-cli balance get, and if needed ask for explicit confirmation before running parallel-cli balance add <amount_cents>. Then retry the original enrichment command.

parallel-web의 다른 스킬

parallel-monitor
parallel-web
지속적으로 웹을 추적하여 정해진 주기로 변경 사항을 감지합니다. 사용자가 '모니터링', '변경 추적', '감시', '알림' 등을 요청할 때 사용하세요.
migrate-to-parallel
parallel-web
Exa, Tavily, Perplexity 또는 Firecrawl 웹데이터 통합을 애플리케이션 동작을 유지하면서 적절한 Parallel 제품으로 완전히 마이그레이션합니다. 사용…
parallel-deep-research
parallel-web
복잡한 주제에 대해 구성 가능한 깊이, 지연 시간, 비용 트레이드오프를 제공하는 철저한 연구. 30초에서 25분까지의 세 가지 프로세서 계층(pro-fast, ultra-fast, ultra)과 1배에서 3배까지의 기본 비용 스케일링. 폴링을 통한 비동기 실행: 연구를 즉시 시작하고, URL을 통해 진행 상황을 모니터링하며, 준비 완료 시 차단 없이 결과를 검색. 출력은 포맷된 마크다운 보고서와 JSON 메타데이터로 제공되며, 빠른 개요를 위해 실행 요약이 stdout에 출력됨. 명시적...
parallel-findall
parallel-web
자연어 설명과 일치하는 엔터티(회사, 인물, 제품 등)를 발견합니다. 사용자가 '모든 X를 찾아줘' 또는 '…하는 모든 Y를 나열해줘'라고 요청할 때 사용하세요.
parallel-memory
parallel-web
과거 Parallel Task, Monitor, FindAll 실행이 도움이 될 때 이를 회상하고, 요청 시 실행을 제거하거나 메모리를 비웁니다.
parallel-web-extract
parallel-web
여러 URL에서 병렬로 콘텐츠를 추출하며, 토큰 효율적으로 처리합니다. 단일 명령어로 웹페이지, 기사, PDF, JavaScript 중심 사이트를 처리합니다. 포크된 컨텍스트에서 실행되어 내장 WebFetch보다 토큰 오버헤드를 최소화합니다. 선택적 초점 목표와 함께 여러 URL의 배치 추출을 지원합니다. parallel-cli 설치 및 인증이 필요하며, 추출된 콘텐츠를 마크다운 형식으로 로컬 파일에 출력하여 후속 질의에 활용할 수 있습니다.
parallel-web-search
parallel-web
인터넷 전반에 걸쳐 최신 정보, 연구, 사실 확인을 위한 빠른 웹 검색. 단일 목표 기반 쿼리 또는 여러 키워드 검색을 병렬로 실행하여 최대 10개의 결과를 발췌문 및 메타데이터와 함께 반환합니다. --after-date를 통한 시간 기반 필터링과 --include-domains를 통한 도메인별 검색을 지원합니다. 제목, URL, 게시 날짜, 발췌문이 포함된 구조화된 JSON을 출력하여 쉽게 구문 분석하고 후속 쿼리를 수행할 수 있습니다. 모든 주장에 대해 마크다운을 사용한 인라인 인용이 필요합니다...
setup
parallel-web
Parallel 플러그인 설정 (CLI 설치)