parallel-data-enrichment

Enriquecimento em lote de dados de empresas, pessoas ou produtos com campos obtidos da web, como nomes de CEOs, financiamento e informações de contato. Aceita dados JSON inline ou arquivos CSV; gera resultados enriquecidos em CSV. Executa de forma assíncrona com rastreamento de progresso via URL de monitoramento e comandos de polling. Requer a ferramenta parallel-cli e acesso à internet; lida com grandes conjuntos de dados com timeouts configuráveis. Suporta solicitações flexíveis de campos por meio de descrições de intenção em linguagem natural (ex.: "nome do CEO e ano de fundação").

npx skills add https://github.com/parallel-web/parallel-agent-skills --skill parallel-data-enrichment

Data Enrichment

Enrich: $ARGUMENTS

Before starting

Inform the user that enrichment may take several minutes depending on the number of rows and fields requested.

Optional: Suggest output columns

If the user gave a vague intent ("enrich these companies with useful info") and you're not sure what columns to add, ask the API for a suggestion before kicking off the run:

parallel-cli enrich suggest "Find CEO and recent funding info" --json

The response is an envelope: {title, processor, enriched_columns, warnings}. Extract just the enriched_columns array (not the whole envelope) and pass it as the value of --enriched-columns on enrich run, in place of --intent — the two flags are alternative ways to specify what to enrich, not combined. If suggest returned a processor, pass it through explicitly via --processor on the run call (it's a tuned recommendation for the schema). Skip this whole section if the user already specified the fields they want.

enrich suggest requires parallel-cli ≥ 0.3.0. If it errors with anything resembling no such command / No such command / unknown command, do not bail — skip the suggestion step, fall through to step 1 with --intent, complete the run, and mention parallel-cli update (or pipx upgrade parallel-web-tools) in the final response so the user picks up the feature next time.

Step 1: Start the enrichment

Use ONE of these command patterns (substitute user's actual data):

For inline data:

parallel-cli enrich run --data '[{"company": "Google"}, {"company": "Microsoft"}]' --intent "CEO name and founding year" --target "output.csv" --no-wait --json

For CSV file:

parallel-cli enrich run --source-type csv --source "input.csv" --target "output.csv" --source-columns '[{"name": "company", "description": "Company name"}]' --intent "CEO name and founding year" --no-wait --json

If this is a follow-up to a previous research task and you have its interaction_id, add context chaining:

parallel-cli enrich run --data '...' --intent "..." --target "output.csv" --no-wait --json --previous-interaction-id "$INTERACTION_ID"

The enrichment will run with the full context of that prior research — so you can enrich entities discovered earlier without restating what was already found. Note: enrichment does not itself produce a new interaction_id, so you cannot chain a further follow-up off of an enrichment.

IMPORTANT: Always include --no-wait so the command returns immediately instead of blocking.

Parse the --json output to extract taskgroup_id and url. The output is {taskgroup_id, url, num_runs} — there is no interaction_id field, do not look for one. Immediately tell the user:

  • Enrichment has been kicked off
  • The monitoring URL where they can track progress

Tell them they can background the polling step to continue working while it runs.

Step 2: Poll for results

Pick a concrete output path (e.g., /tmp/enrichment-acme.json). Note: the file is JSON regardless of the extension you choose — it's an array of {input, output} objects, not a CSV. Name it .json to avoid confusing yourself or the user.

parallel-cli enrich poll "$TASKGROUP_ID" --timeout 540 --output "/tmp/enrichment-<descriptive-name>.json"

Important:

  • Use --timeout 540 (9 minutes) to stay within tool execution limits
  • The --target from step 1 is unused in --no-wait mode — only --output here determines where results are saved, and the file is always JSON

If the poll times out

Enrichment of large datasets can take longer than 9 minutes. If the poll exits without completing:

  1. Tell the user the enrichment is still running server-side
  2. Re-run the same parallel-cli enrich poll command to continue waiting

Response format

After step 1: Share the monitoring URL (for tracking progress).

After step 2:

  1. Report number of rows enriched
  2. Preview first few rows from the output file (it's a JSON array of {input, output} objects)
  3. Tell the user the full path to the output file

Do NOT re-share the monitoring URL after completion — the results are in the output file.

Setup

If parallel-cli is not found, install and authenticate:

/parallel:parallel-cli-setup

If any parallel-cli enrich command returns 403, tell the user balance is likely required. Offer to run parallel-cli balance get, and if needed ask for explicit confirmation before running parallel-cli balance add <amount_cents>. Then retry the original enrichment command.

Mais skills de parallel-web

parallel-monitor
parallel-web
Acompanhe continuamente a web em busca de mudanças em uma cadência recorrente. Use quando o usuário pedir para 'monitorar', 'rastrear mudanças em', 'observar' ou 'me alertar quando' algo…
migrate-to-parallel
parallel-web
Migre integrações de dados web da Exa, Tavily, Perplexity ou Firecrawl completamente para os produtos Parallel apropriados, preservando o comportamento do aplicativo. Use…
parallel-deep-research
parallel-web
Pesquisa exaustiva com profundidade, latência e compensações de custo configuráveis para tópicos complexos. Três níveis de processamento (pro-fast, ultra-fast, ultra) variando de 30 segundos a 25 minutos, com escalonamento de custo de 1x a 3x da linha de base. Execução assíncrona com polling: inicie a pesquisa instantaneamente, monitore o progresso via URL, recupere os resultados quando estiver pronto sem bloqueio. Saídas formatadas em relatório markdown e metadados JSON; resumo executivo impresso no stdout para visão geral rápida. Projetado para explícito...
parallel-findall
parallel-web
Descubra entidades (empresas, pessoas, produtos, etc.) que correspondam a uma descrição em linguagem natural. Use quando o usuário pedir para 'encontrar todos os X' ou 'listar todos os Y que…' —…
parallel-memory
parallel-web
Lembre-se de execuções passadas de Parallel Task, Monitor e FindAll quando elas puderem ajudar; remova execuções ou limpe a memória quando solicitado.
parallel-web-extract
parallel-web
Extrai conteúdo de múltiplas URLs em paralelo, de forma eficiente em tokens. Lida com páginas web, artigos, PDFs e sites com muito JavaScript em um único comando. Executa em um contexto bifurcado para minimizar a sobrecarga de tokens em comparação com o WebFetch integrado. Suporta extração em lote de várias URLs com objetivos de foco opcionais. Requer instalação e autenticação do parallel-cli; gera o conteúdo extraído como markdown em um arquivo local para consultas de acompanhamento.
parallel-web-search
parallel-web
Pesquisa web rápida para informações atuais, pesquisas e verificação de fatos na internet. Executa consultas baseadas em um único objetivo ou múltiplas pesquisas por palavras-chave em paralelo, retornando até 10 resultados com trechos e metadados. Suporta filtragem sensível ao tempo via --after-date e pesquisas específicas de domínio com --include-domains. Gera JSON estruturado com títulos, URLs, datas de publicação e trechos para fácil análise e consultas de acompanhamento. Requer citações inline para cada afirmação usando markdown...
setup
parallel-web
Configurar o plugin Parallel (instalar CLI)