parallel-data-enrichment

Массовое обогащение данных. Добавляет поля из веб-источников (имена CEO, финансирование, контактная информация) в списки компаний, людей или продуктов. Используется для обогащения CSV-файлов или…

npx skills add https://github.com/parallel-web/parallel-cursor-plugin --skill parallel-data-enrichment

Data Enrichment

Enrich: $ARGUMENTS

Before starting

Inform the user that enrichment may take several minutes depending on the number of rows and fields requested.

Optional: Suggest output columns

If the user gave a vague intent ("enrich these companies with useful info") and you're not sure what columns to add, ask the API for a suggestion before kicking off the run:

parallel-cli enrich suggest "Find CEO and recent funding info" --json

The response is an envelope: {title, processor, enriched_columns, warnings}. Extract just the enriched_columns array (not the whole envelope) and pass it as the value of --enriched-columns on enrich run, in place of --intent — the two flags are alternative ways to specify what to enrich, not combined. If suggest returned a processor, pass it through explicitly via --processor on the run call (it's a tuned recommendation for the schema). Skip this whole section if the user already specified the fields they want.

enrich suggest requires parallel-cli ≥ 0.3.0. If it errors with anything resembling no such command / No such command / unknown command, do not bail — skip the suggestion step, fall through to step 1 with --intent, complete the run, and mention parallel-cli update (or pipx upgrade parallel-web-tools) in the final response so the user picks up the feature next time.

Step 1: Start the enrichment

Use ONE of these command patterns (substitute user's actual data):

For inline data:

parallel-cli enrich run --data '[{"company": "Google"}, {"company": "Microsoft"}]' --intent "CEO name and founding year" --target "output.csv" --no-wait --json

For CSV file:

parallel-cli enrich run --source-type csv --source "input.csv" --target "output.csv" --source-columns '[{"name": "company", "description": "Company name"}]' --intent "CEO name and founding year" --no-wait --json

If this is a follow-up to a previous research task and you have its interaction_id, add context chaining:

parallel-cli enrich run --data '...' --intent "..." --target "output.csv" --no-wait --json --previous-interaction-id "$INTERACTION_ID"

The enrichment will run with the full context of that prior research — so you can enrich entities discovered earlier without restating what was already found. Note: enrichment does not itself produce a new interaction_id, so you cannot chain a further follow-up off of an enrichment.

IMPORTANT: Always include --no-wait so the command returns immediately instead of blocking.

Parse the --json output to extract taskgroup_id and url. The output is {taskgroup_id, url, num_runs} — there is no interaction_id field, do not look for one. Immediately tell the user:

  • Enrichment has been kicked off
  • The monitoring URL where they can track progress

Tell them they can background the polling step to continue working while it runs.

Step 2: Poll for results

Pick a concrete output path (e.g., /tmp/enrichment-acme.json). Note: the file is JSON regardless of the extension you choose — it's an array of {input, output} objects, not a CSV. Name it .json to avoid confusing yourself or the user.

parallel-cli enrich poll "$TASKGROUP_ID" --timeout 540 --output "/tmp/enrichment-<descriptive-name>.json"

Important:

  • Use --timeout 540 (9 minutes) to stay within tool execution limits
  • The --target from step 1 is unused in --no-wait mode — only --output here determines where results are saved, and the file is always JSON

If the poll times out

Enrichment of large datasets can take longer than 9 minutes. If the poll exits without completing:

  1. Tell the user the enrichment is still running server-side
  2. Re-run the same parallel-cli enrich poll command to continue waiting

Response format

After step 1: Share the monitoring URL (for tracking progress).

After step 2:

  1. Report number of rows enriched
  2. Preview first few rows from the output file (it's a JSON array of {input, output} objects)
  3. Tell the user the full path to the output file

Do NOT re-share the monitoring URL after completion — the results are in the output file.

If the parallel-cli binary is not installed

If the shell reports command not found: parallel-cli (i.e. the binary itself is missing — distinct from a No such command error from a stale CLI, which the in-body guidance above covers), stop immediately. Do NOT search the web yourself, do NOT use any built-in search tools, and do NOT try to answer the query from your own knowledge. Instead, tell the user:

  1. parallel-cli is not installed
  2. Run /parallel-setup to install it
  3. Then retry their request

Больше skills от parallel-web

parallel-monitor
parallel-web
Непрерывно отслеживайте изменения в интернете с заданной периодичностью. Используйте, когда пользователь просит «отслеживать», «следить за изменениями», «наблюдать» или «уведомить меня, когда» что-то…
migrate-to-parallel
parallel-web
Полностью перенесите интеграции веб-данных Exa, Tavily, Perplexity или Firecrawl на соответствующие продукты Parallel, сохраняя поведение приложения. Используйте…
parallel-data-enrichment
parallel-web
Массовое обогащение данных о компаниях, людях или продуктах веб-полями, такими как имена CEO, финансирование и контактная информация. Принимает встроенные JSON-данные или CSV-файлы; выводит обогащённые результаты в CSV. Работает асинхронно с отслеживанием прогресса через URL мониторинга и команды опроса. Требует инструмент parallel-cli и доступ в интернет; обрабатывает большие наборы данных с настраиваемыми тайм-аутами. Поддерживает гибкие запросы полей через описания на естественном языке (например, «имя CEO и год основания»).
parallel-deep-research
parallel-web
Исчерпывающее исследование с настраиваемыми компромиссами по глубине, задержке и стоимости для сложных тем. Три уровня процессоров (pro-fast, ultra-fast, ultra) от 30 секунд до 25 минут, с масштабированием стоимости от 1x до 3x от базовой. Асинхронное выполнение с опросом: запуск исследования мгновенно, отслеживание прогресса через URL, получение результатов по готовности без блокировки. Вывод отформатированного отчёта в Markdown и метаданных в JSON; исполнительное резюме выводится в stdout для быстрого обзора. Разработано для явного...
parallel-findall
parallel-web
Обнаруживает сущности (компании, людей, продукты и т.д.), соответствующие описанию на естественном языке. Используйте, когда пользователь просит «найти все X» или «перечислить все Y, которые…» —…
parallel-memory
parallel-web
Вспоминать прошлые запуски Parallel Task, Monitor и FindAll, когда они могут помочь; удалять запуски или очищать память по запросу.
parallel-web-extract
parallel-web
Извлекает содержимое из нескольких URL-адресов параллельно, эффективно расходуя токены. Обрабатывает веб-страницы, статьи, PDF-файлы и сайты с интенсивным использованием JavaScript одной командой. Работает в форкнутом контексте, чтобы минимизировать затраты токенов по сравнению со встроенным WebFetch. Поддерживает пакетное извлечение нескольких URL-адресов с опциональными целевыми задачами. Требует установки parallel-cli и аутентификации; выводит извлечённое содержимое в формате markdown в локальный файл для последующих запросов.
parallel-web-search
parallel-web
Быстрый веб-поиск для получения актуальной информации, исследований и проверки фактов в интернете. Выполняет одиночные целевые запросы или множественные параллельные поиски по ключевым словам, возвращая до 10 результатов с выдержками и метаданными. Поддерживает фильтрацию по времени через --after-date и поиск по конкретным доменам с помощью --include-domains. Выводит структурированный JSON с заголовками, URL, датами публикации и выдержками для удобного анализа и последующих запросов. Требует встроенных цитат для каждого утверждения с использованием Markdown...