parallel-findall

작성자: parallel-web

자연어 설명과 일치하는 엔터티(회사, 인물, 제품 등)를 발견합니다. 사용자가 '모든 X를 찾아줘' 또는 '…하는 모든 Y를 나열해줘'라고 요청할 때 사용하세요.

npx skills add https://github.com/parallel-web/parallel-agent-skills --skill parallel-findall

FindAll: Entity Discovery

Find: $ARGUMENTS

Requires parallel-cli ≥ 0.6.0 (the findall entity-search command was added in 0.6.0; the broader findall command was added in 0.3.0). If either errors with no such command or similar, tell the user to run parallel-cli update (or pipx upgrade parallel-web-tools if installed via pipx), then retry.

When to use this skill

Use FindAll when the user wants a structured list of entities matching a description, not webpages or a narrative answer.

User asks for…Use
"Find all X that…" / "List every Y…"parallel-findall (this skill)
Webpage results / quick answers / current infoparallel-web-search
Narrative report / analysis / "research X"parallel-deep-research
Add fields to a list you already haveparallel-data-enrichment

If the user already has a list and just wants to add fields, this is the wrong skill — use parallel-data-enrichment.

FindAll has two paths: the comprehensive, asynchronous findall run (Steps 1–2) and the fast, synchronous entity-search (final section).

  • entity-search — very fast (few seconds), only supports people or company search. Supports a more limited set of query arguments. Optimized for recall over precision; results are not individually verified.
  • findall run — Provides comprehensive coverage, complex, match conditions, exclusions, enrichment, citations, or a type other than people/companies.

If it's ambiguous, ask the user which they'd prefer and offer a default. Remember entity search limits: companies/people only, no exclusions/generator/enrichment, and entity_set_id can't be used with enrich/extend (re-run via findall run if needed).

Switch to entity-search only when the user explicitly signals they want a fast, throwaway list. entity-search is also strictly more limited: it only supports companies or people entity types, no exclusions, no generator choice, no enrichment, and the returned entity_set_id is not usable with findall enrich/extend. If you start there and the user later asks to enrich or extend, you'll have to re-run via findall run.

Step 1: Start the run

parallel-cli findall run "$ARGUMENTS" --no-wait --json

Defaults: generator core, match limit 10. Stick with core unless the user has a reason to escalate:

  • -g pro — most thorough generator (slower, costlier). Use when the user asks for "comprehensive" coverage or matches are sparse on core
  • -g base — fastest, but markedly lower quality. Often returns query-echo entities (e.g., directory pages, the literal query string), entries with no URL, or category placeholders. Only use if the user explicitly asks for a quick scan and accepts noise; otherwise prefer core
  • -n 50 — return up to 50 matched entities (5–1000 allowed)

If the user wants to exclude known entities (e.g., "find competitors but not Google or OpenAI"):

parallel-cli findall run "$ARGUMENTS" --no-wait --json \
    --exclude '[{"name":"Google","url":"google.com"},{"name":"OpenAI","url":"openai.com"}]'

Tip — preview the schema first if the objective is ambiguous: parallel-cli findall ingest "$ARGUMENTS" --json shows the entity type and match conditions the API inferred, so you can refine wording before paying for a run.

Parse the JSON output to extract the findall_id and any monitoring URL. Tell the user:

  • A FindAll run has been started
  • Approximate cadence (minutes for core, longer for pro)
  • They can keep working while it runs

Step 2: Poll for results

Choose a descriptive filename (e.g., series-a-ai-2026, charlotte-roofers). Use lowercase with hyphens, no spaces.

parallel-cli findall poll "$FINDALL_ID" -o "/tmp/$FILENAME.json" --timeout 540

Important:

  • Use --timeout 540 (9 minutes) to stay within tool execution limits
  • Do NOT pass --json for large result sets — it will flood context. -o saves the full results to disk

If the poll times out

Re-run the same parallel-cli findall poll command to continue waiting. Server-side the run continues regardless.

Response format

Before presenting matches, filter the results for obvious noise:

  • Drop entries with empty/missing url
  • Drop entries whose name echoes the user's query (e.g., literal "YC W25 batch companies in developer tools") — those are search-result placeholders, not real entities
  • Drop entries whose url is a third-party directory or profile page rather than the entity's own domain. The URL should be something the entity itself owns (its product site, docs, or marketing site)

If filtering removes a meaningful share of matches, mention this to the user and suggest re-running with -g pro or a higher -n.

Sanity-check -g base results. The base generator can hallucinate categorical attributes (e.g., return a YC S22 company as a YC W25 match). The filter rules above only catch URL/name shape, not factual correctness. If the user's query has a falsifiable attribute (a specific batch, year, geography, etc.), spot-check the kept entries against the source URL and flag any that don't fit. Recommend re-running with -g core (or higher) if either multiple kept entries fail the spot-check or noise filtering dropped a meaningful share of the matched set (say, ≥40%) — both indicate base isn't producing reliable results for this query.

Present the remaining (real) entities as a markdown table or list. Lead with the count, then list each entity with its name, URL, and a one-line description if available. Cite each entity with its source URL.

Tell the user:

  • How many entities were matched (and how many were filtered as noise, if any)
  • The full results path (/tmp/$FILENAME.json)
  • That they can:
    • Add fields to these results, e.g.:

      parallel-cli findall enrich $FINDALL_ID '{"properties":{"ceo":{"type":"string"},"employee_count":{"type":"number"}}}'
      

      The schema is a JSON Schema-style object with properties mapping field names → {type, description?}.

    • Get more matches: parallel-cli findall extend $FINDALL_ID 50

Fast entity search

Use this path only when the user explicitly signals they want a quick/rough/preview list — do not pick it just because the entity type happens to be companies or people.

Synchronous call. No polling, no findall_id. Pick a descriptive $FILENAME (lowercase, hyphens, no spaces), as in Step 2.

parallel-cli findall entity-search "$ARGUMENTS" -t companies -n 100 -o "/tmp/$FILENAME.json"

Flags:

  • -t companies|people — entity type (required). The endpoint only supports these two; for anything else, use findall run
  • -n 5..1000 — match limit (default 10). When possible, request more than the user needs (e.g. -n 100) and select after filtering — results are ranked but not individually verified, and a low limit can omit relevant entities
  • Do NOT pass --json for large result sets — it will flood context. -o saves the full results to disk

Avoid highly restrictive objectives on this path: the API fills toward the limit, so relevance declines toward the tail. Keep the core criterion in the objective and filter the rest downstream, or use findall run.

Response shape:

{ "entity_set_id": "entity_set_…", "entities": [ {"name": "...", "url": "...", "description": "..."},
… ] }

Unlike the full path, the url returned by entity-search is usually a directory/profile link — expected, not noise. Don't drop them; only filter out entries with an empty url or a name that echoes the query.

Present the kept entities as a markdown table or list, lead with the count, and cite each with its source URL. Tell the user:

  • How many entities came back (and how many were filtered as noise)
  • The full results path (/tmp/$FILENAME.json) if -o was used

Setup

Requires parallel-cli (installed and authenticated). If parallel-cli --version fails, or if a later command fails with an authentication error, tell the user to see https://docs.parallel.ai/integrations/cli and stop.

parallel-web의 다른 스킬

parallel-monitor
parallel-web
지속적으로 웹을 추적하여 정해진 주기로 변경 사항을 감지합니다. 사용자가 '모니터링', '변경 추적', '감시', '알림' 등을 요청할 때 사용하세요.
migrate-to-parallel
parallel-web
Exa, Tavily, Perplexity 또는 Firecrawl 웹데이터 통합을 애플리케이션 동작을 유지하면서 적절한 Parallel 제품으로 완전히 마이그레이션합니다. 사용…
parallel-data-enrichment
parallel-web
회사, 인물 또는 제품 데이터를 CEO 이름, 자금 정보, 연락처 정보 등 웹에서 수집한 필드로 대량 보강합니다. 인라인 JSON 데이터 또는 CSV 파일을 입력받아 보강된 결과를 CSV로 출력합니다. 모니터링 URL 및 폴링 명령어를 통한 진행 상황 추적과 함께 비동기적으로 실행됩니다. parallel-cli 도구와 인터넷 접속이 필요하며, 구성 가능한 타임아웃으로 대규모 데이터셋을 처리합니다. 자연어 의도 설명(예: "CEO 이름 및 설립 연도")을 통해 유연한 필드 요청을 지원합니다.
parallel-deep-research
parallel-web
복잡한 주제에 대해 구성 가능한 깊이, 지연 시간, 비용 트레이드오프를 제공하는 철저한 연구. 30초에서 25분까지의 세 가지 프로세서 계층(pro-fast, ultra-fast, ultra)과 1배에서 3배까지의 기본 비용 스케일링. 폴링을 통한 비동기 실행: 연구를 즉시 시작하고, URL을 통해 진행 상황을 모니터링하며, 준비 완료 시 차단 없이 결과를 검색. 출력은 포맷된 마크다운 보고서와 JSON 메타데이터로 제공되며, 빠른 개요를 위해 실행 요약이 stdout에 출력됨. 명시적...
parallel-memory
parallel-web
과거 Parallel Task, Monitor, FindAll 실행이 도움이 될 때 이를 회상하고, 요청 시 실행을 제거하거나 메모리를 비웁니다.
parallel-web-extract
parallel-web
여러 URL에서 병렬로 콘텐츠를 추출하며, 토큰 효율적으로 처리합니다. 단일 명령어로 웹페이지, 기사, PDF, JavaScript 중심 사이트를 처리합니다. 포크된 컨텍스트에서 실행되어 내장 WebFetch보다 토큰 오버헤드를 최소화합니다. 선택적 초점 목표와 함께 여러 URL의 배치 추출을 지원합니다. parallel-cli 설치 및 인증이 필요하며, 추출된 콘텐츠를 마크다운 형식으로 로컬 파일에 출력하여 후속 질의에 활용할 수 있습니다.
parallel-web-search
parallel-web
인터넷 전반에 걸쳐 최신 정보, 연구, 사실 확인을 위한 빠른 웹 검색. 단일 목표 기반 쿼리 또는 여러 키워드 검색을 병렬로 실행하여 최대 10개의 결과를 발췌문 및 메타데이터와 함께 반환합니다. --after-date를 통한 시간 기반 필터링과 --include-domains를 통한 도메인별 검색을 지원합니다. 제목, URL, 게시 날짜, 발췌문이 포함된 구조화된 JSON을 출력하여 쉽게 구문 분석하고 후속 쿼리를 수행할 수 있습니다. 모든 주장에 대해 마크다운을 사용한 인라인 인용이 필요합니다...
setup
parallel-web
Parallel 플러그인 설정 (CLI 설치)