apify-verified-email-finder

작성자: apify

Google Maps, Google SERPs 또는 사용자가 제공한 URL 목록에서 검증된 비즈니스 이메일 목록을 구축합니다. 검증은 동일한 Apify 실행 내에서 이루어집니다 — 추가...

npx skills add https://github.com/apify/awesome-skills --skill apify-verified-email-finder

Verified Email Finder

Return a list of verified business emails by routing the user's input to the right Apify Actor and turning on the leads enrichment + email verification add-ons in a single run. No third-party verifier (Hunter, NeverBounce, Apollo) needed — verification happens inside the same Actor run.

Prerequisites

(No need to check it upfront)

The skill supports two execution paths. Pick the one that matches your environment — Steps 4 and 5 show commands for both.

MCP path (default in Claude sessions, recommended). If the Apify MCP server is connected, no setup is needed — auth runs through the user's Apify account. Use the call-actor and get-dataset-items MCP tools.

Script path (CLI / scheduled / non-Claude execution). Requires:

  • .env file with APIFY_TOKEN
  • Node.js 20.6+ (for native --env-file support)

Workflow

Copy this checklist and track progress:

Task Progress:
- [ ] Step 1: Collect the six required anchor inputs
- [ ] Step 2: Route to the correct Actor (confirm if ambiguous)
- [ ] Step 3: Build the Actor input (verification always ON)
- [ ] Step 4: Run the Actor and wait
- [ ] Step 5: Apply the result-scope filter, deduplicate, and render

Step 1: Collect the Six Required Anchor Inputs

Ask all six as one block before any Actor call. Don't bundle Actor-specific optional fields (country code, language, max pages) into this round — surface those as follow-ups.

  1. What do you have to start with? — location query / SERP keyword / URL list. This drives the routing decision.
  2. The actual input — the location string, the keyword(s), or the URLs themselves.
  3. Department filter — one or more of: c_suite, product, engineering_technical, design, education, finance, human_resources, information_technology, legal, marketing, medical_health, operations, sales, consulting. Default is any (leave the array empty), but ask every time.
  4. Max contacts per domain / business — passed as maximumLeadsEnrichmentRecords. Default 3, but ask every time.
  5. Output format — CSV or JSON. Ask every time.
  6. Result scope — which leads to keep in the deliverable. The Actor always runs the same way (verification always on); this only controls post-run filtering. Pick one:
    • verified-only (default) — only leads with emailVerification.result == "ok". Safest for cold email.
    • verified-plus-catchall — ok plus catch_all. Catch-all is often deliverable but unprovable.
    • all-emails — any lead with a non-empty email, regardless of verification.
    • with-phone — any lead with a non-empty phone number, regardless of email status. Use for call campaigns.
    • everything — every lead the Actor returned, even incomplete ones.

Step 2: Route to the Correct Actor

Inspect anchor #1 and pick the Actor.

User has to start withActor IDUse when
Location + business type ("dentists in Berlin")compass/crawler-google-placesLocal leads list from Maps listings; best when user wants address / phone / hours too
Keyword / search query ("best CRM software")apify/google-search-scraperContacts from whichever sites Google ranks for a topic
Pre-existing URL list (pasted, file path)vdrmota/contact-info-scraperUser already has domains; cheapest route since no discovery step

All three Actors share the same three add-on fields, so verification behavior is identical across routes.

Decision examples

User saysRoute
"Dentists in Munich" / "Lawyers in Prague"Maps
"Marketing contacts at the top results for 'AI agent builder'"Search
"Find emails for these 5 URLs: acme-co.example, demo-co.example..."URL list
"Find HR contacts at Fortune 500 companies"Ask: SERP for "Fortune 500 HR" or a URL list?
"Find contacts at SaaS companies in Berlin"Ask: Maps for "SaaS companies in Berlin" or SERP for "SaaS companies Berlin"? Maps works best when businesses are Google-Maps-listed.
(User pastes both a SERP keyword AND a URL list)Ask: run one route, the other, or both as separate deliverables?

Ambiguity rule: if anchor #1 is unclear, ask one follow-up before running. Never burn Actor compute on a guessed route.

Mixed deliverables: if the user explicitly asks for two routes in one deliverable, run both Actors and concatenate. The Source column makes the mix clear; dedupe by email across the combined output.

Step 3: Build the Actor Input

Always set these three fields, regardless of which Actor is selected.

FieldValue
maximumLeadsEnrichmentRecordsanchor #4 (default 3, min 1)
leadsEnrichmentDepartmentsanchor #3 as array, or [] if "any"
verifyLeadsEnrichmentEmailstrue (always) — guard rail, never set to false

Full per-Actor input parameters and example payloads are in reference/apify-actor-usage.md.

URL-list pre-validation: before submitting URLs to vdrmota/contact-info-scraper, parse each one and check it is http/https and parseable. Skipped entries must appear in the output as skipped — invalid URL, never silently dropped.

Step 4: Run the Actor

Maps and SERP runs with leads enrichment can take several minutes per query. Raise the timeout for large jobs.

MCP path (default in Claude sessions):

Call the call-actor tool:

  • actor: one of the three Actor IDs (compass/crawler-google-places, apify/google-search-scraper, vdrmota/contact-info-scraper)
  • input: the JSON payload from Step 3
  • callOptions: {"timeout": 1800, "memory": 4096} for a generous budget

The tool returns runId and datasetId. If status is still RUNNING, poll with get-actor-run (waitSecs up to 45) until SUCCEEDED. Capture both IDs for the run_metadata.json sidecar.

Script path (CLI / scheduled use):

node --env-file=.env ${CLAUDE_PLUGIN_ROOT}/reference/scripts/run_actor.js \
  --actor "ACTOR_ID" \
  --input 'JSON_INPUT' \
  --output YYYY-MM-DD_verified-emails.csv \
  --format csv \
  --timeout 900

Use --format json for JSON. The script writes the raw dataset to disk; Step 5 still applies the spurious-match + scope filters on top.

Step 5: Filter, Deduplicate, and Render

Pull the dataset:

  • MCP path: call get-dataset-items with the datasetId from Step 4. Use the fields parameter (e.g., title,searchString,countryCode,city,address,phone,website,leadsEnrichment) and clean: true to keep the response small. For datasets that still exceed the response cap, fetch directly via curl https://api.apify.com/v2/datasets/<id>/items?fields=...&clean=true and pipe through jq.
  • Script path: the raw dataset is already on disk in the file from Step 4.

Each record contains business fields plus a leadsEnrichment array (Maps, SERP) or top-level lead fields (URL list). Each lead has a departments array, a companyWebsite, and an emailVerification object with result (ok / invalid / disposable / catch_all / unknown / error) and quality (good / risky / bad).

  • Spurious-match filter (mandatory, always on). Apply this first, before any other filter. The lead-enrichment service can return global-fallback leads when no local match exists (real case observed: a single US-zoo CFO whose companyWebsite=zoo.org was attributed to 8 unrelated Polish zoos because the matcher latched onto the zoo substring). Drop any lead whose companyWebsite hostname doesn't equal the source URL's hostname (strip https?://, leading www., anything after /; lowercase). Count drops in run_metadata.json and call them out in the deliverable header if non-zero.

  • Filter by result scope (anchor #6). Applied second.

    ScopeRow-keep logic
    verified-onlyemailVerification.result == "ok"
    verified-plus-catchallemailVerification.result in {"ok", "catch_all"}
    all-emailsemail is non-empty (any result, including missing verification)
    with-phonephone (or company phone) is non-empty (regardless of email)
    everythingkeep every lead, no filter
  • Dedupe: group by lowercased email; keep the first occurrence and merge Source Query or URL if the same email appears from multiple sources. For with-phone rows that have no email, dedupe by lowercased phone instead.

  • Empty-result surfacing: if the department filter (anchor #3) produces zero leads for a given domain, include a row for that domain with Email = "" and Email Verification Status = "no leads matched filter". Do not silently drop it. This is separate from the result-scope filter above — empty-domain rows are inserted before scope-filtering and always shown.

Output row schema (16 columns, including Departments) and per-format rendering details are in reference/output-formats.md.

Worked Examples

Quality Rules (always enforce)

  • Guard rail: never submit a run with verifyLeadsEnrichmentEmails: false.
  • Provenance & traceability: populate the Source column on every row; carry Apify runId + datasetId in run_metadata.json.
  • No fabrication: missing dataset fields stay blank.
  • Deliverable header transparency: state the active result scope and the spurious-match drop count; offer to re-render under a different scope.
  • Ambiguity confirm: if anchor #1 is unclear, ask before running.

Cost & Pricing

Email verification is charged only for decisive results (ok / invalid / disposable); catch_all / unknown / error are free. Leads enrichment is charged per successfully extracted lead. Check the Apify console for live rates (they vary by subscription tier and change over time).

Error Handling

See reference/troubleshooting.md.

apify의 다른 스킬

apify-influencer-brand-collabs
apify
인스타그램 브랜드-크리에이터 파트너십을 Apify 액터를 연결하여 발견하세요. 사용자가 브랜드와 협업하는 사람, 크리에이터가 유료로 진행한 브랜드 등을 물을 때 사용하세요.
apify-actor-development
apify
서버리스 클라우드 프로그램을 생성, 디버깅 및 배포하여 웹 스크래핑, 자동화 및 데이터 처리를 수행합니다. JavaScript, TypeScript 및 Python 템플릿을 지원하며, HTTP 및 브라우저 기반 크롤링을 위한 통합 Crawlee, Playwright 및 Cheerio 라이브러리를 포함합니다. 격리된 스토리지와 함께 apify run을 통한 로컬 테스트, 입력/출력에 대한 스키마 검증, apify push를 통한 Apify 플랫폼 배포를 포함합니다. Apify CLI 인증 및 AI를 위한 .actor/actor.json의 필수 generatedBy 메타데이터가 필요합니다...
apify-actorization
apify
기존 프로젝트를 언어별 SDK 통합을 통해 서버리스 Apify Actor로 변환합니다. JavaScript/TypeScript(Actor.init() / Actor.exit() 사용), Python(비동기 컨텍스트 매니저), CLI 래퍼를 통한 모든 언어를 지원합니다. 구조화된 워크플로우를 제공합니다: apify init으로 스캐폴딩, SDK 래핑 적용, 입출력 스키마 구성, apify run으로 로컬 테스트, apify push로 배포. 입출력 스키마 검증, Docker 컨테이너화, 선택적 이벤트당 과금을 포함합니다.
apify-content-analytics
apify
Apify Actors를 통한 Instagram, Facebook, YouTube, TikTok의 멀티 플랫폼 콘텐츠 분석. 네 플랫폼의 게시물, 릴스, 스토리, 댓글, 해시태그, 팔로워, 광고를 포함한 17개 이상의 특화 Actors를 지원합니다. mcpc CLI를 사용하여 Actor 스키마를 동적으로 가져와 필요한 입력과 사용 가능한 출력 필드를 결정합니다. 빠른 채팅 표시, CSV 내보내기, JSON 내보내기(결과 수 사용자 지정 가능)의 세 가지 형식으로 결과를 출력합니다. .env 파일에 Apify 토큰이 필요하며 Node.js 20.6+가 필요합니다...
apify-ecommerce
apify
50개 이상의 전자상거래 마켓플레이스에서 제품 데이터, 가격, 리뷰, 판매자 정보를 추출합니다. 세 가지 워크플로우 모드: 제품 및 가격(가격 추적, 경쟁사 분석), 고객 리뷰(감정 분석, 품질 문제), 판매자 인텔리전스(Google Shopping을 통한 공급업체 발견). Amazon(20개 이상 지역), Walmart, eBay, IKEA, Costco, 유럽 소매업체 지원; 제품 URL, 카테고리 URL 또는 키워드 검색을 통해 입력. 선택적 AI 기반 분석으로 가격에 대한 인사이트를 생성합니다...
apify-generate-output-schema
apify
Apify Actor의 소스 코드를 분석하여 출력 스키마(dataset_schema.json, output_schema.json, key_value_store_schema.json)를 생성합니다. 다음과 같은 경우에 사용하세요…
apify-influencer-discovery
apify
Instagram, Facebook, YouTube, TikTok에서 Apify Actors를 사용하여 인플루언서를 발견하고 평가합니다. 발견 요청을 15개 이상의 전문 Actors로 라우팅하여 프로필 스크래핑, 해시태그 검색, 참여도 분석, 모든 주요 플랫폼의 틈새 발견을 다룹니다. 실행 전에 mcpc를 통해 Actor 스키마를 동적으로 가져와 필요한 입력과 사용 가능한 출력 필드를 결정합니다. 인라인 채팅 표시, CSV 또는 JSON 파일 출력의 세 가지 내보내기 모드를 지원하며 결과 수를 사용자 지정할 수 있습니다...
apify-ultimate-scraper
apify
Instagram, TikTok, YouTube, Facebook, Google Maps 등 55개 이상의 플랫폼에 최적의 Actor를 선택하는 자동화된 웹 스크래퍼. 8개 주요 플랫폼에 걸쳐 55개 이상의 사전 구성된 Actor를 포함하며, 사용 사례별 선택 가이드(리드 생성, 인플루언서 발굴, 브랜드 모니터링, 경쟁사 분석, 트렌드 조사)를 제공합니다. 빠른 채팅 표시, CSV 내보내기, 또는 사용자 정의 가능한 결과 제한이 있는 JSON 내보내기의 세 가지 출력 형식을 지원합니다. 복잡한 작업을 위한 다중 Actor 워크플로 패턴을 포함합니다...