apify-ultimate-scraper

โดย apify

เครื่องมือขูดเว็บที่ขับเคลื่อนด้วย AI แบบสากลสำหรับทุกแพลตฟอร์ม ขูดข้อมูลจาก Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Google Search,…

npx skills add https://github.com/apify/apify-claude-code-plugin --skill apify-ultimate-scraper

Universal web scraper

AI-driven data extraction from ~100 Actors across 15+ platforms via the Apify CLI.

Rule: Pass --json and redirect stderr with 2>/dev/null on data-returning commands (actors call, actors start, actors info, actors search, datasets get-items, runs info). JSON output is stable across CLI versions. stderr contains progress messages and version warnings that break JSON parsers if not redirected.

This rule does not apply to status/auth commands (apify info, apify --version, apify login). For those, use 2>&1 so authentication and version errors are visible.

Exception: if --input returns no data, re-run with 2>&1 to confirm whether the cause is a missing schema vs. a network/auth error.

Rule: pass --user-agent apify-claude-code-plugin/apify-ultimate-scraper only on actor start and actor run commands (apify actors start, apify actors call). Do not add it to login, info, search, or other CLI commands.

Prerequisites

  • Apify CLI v1.5.0+ (npm install -g apify-cli) — older versions reject the user-agent flag
  • Authenticated session (see below)

Authentication

If a CLI command fails with an auth error, authenticate using one of these methods:

  1. OAuth (interactive): apify login (opens browser)
  2. Environment variable: export APIFY_TOKEN=your_token_here
  3. From .env file: source .env (if the file contains APIFY_TOKEN=...)

Generate token: https://console.apify.com/settings/integrations

Workflow

Step 0: Verify CLI readiness before doing anything else

Before using the Apify CLI, always verify the local environment:

  1. Check that the CLI is installed and is v1.5.0 or newer:
    apify --help
    apify --version

If this fails, install the CLI first:

       npm install -g apify-cli
  1. Check that the CLI is authenticated:
    # Auth check — do NOT pipe to /dev/null, you need to see errors
    apify info 2>&1

If this shows the user is not logged in, instruct them to authenticate with a token:

    apify login --token TOKEN
  1. Authenticated Apify CLI commands need file access to ~/.apify/, where the CLI keeps its credentials. A host that sandboxes file access can deny this even when the login is valid; that is a sandbox problem, not a login problem.

  2. Assume many Apify commands block with zero output until completion, so allow at least 60 seconds before treating one as stuck. If your shell tool takes a timeout, raise it accordingly.

  3. For long or unknown-duration runs, prefer the async pattern:

    apify actors start "ACTOR_ID" -i 'JSON_INPUT' --user-agent apify-claude-code-plugin/apify-ultimate-scraper --json 2>/dev/null

Then poll the run status:

    apify runs info RUN_ID --json

Check .status for SUCCEEDED or FAILED.

Step 1: Understand goal and select Actor

Identify the target platform and use case. Read references/actor-index.md to find the right Actor. Prefer apify-tier actors; use community-tier only when no apify actor covers the task. For input schemas, fetch dynamically: apify actors info "ACTOR_ID" --input --json 2>/dev/null If the output is empty, re-run without the redirect (2>&1) to surface auth or network errors before proceeding.

If the task involves a multi-step pipeline, also read the matching workflow guide:

Task involves...Read
leads, contacts, emails, B2Breferences/workflows/lead-generation.md
competitor, ads, pricingreferences/workflows/competitive-intel.md
influencer, creatorreferences/workflows/influencer-vetting.md
brand, mentions, sentimentreferences/workflows/brand-monitoring.md
reviews, ratings, reputationreferences/workflows/review-analysis.md
SEO, SERP, crawl, content, RAGreferences/workflows/content-and-seo.md
analytics, engagement, performancereferences/workflows/social-media-analytics.md
trends, keywords, hashtagsreferences/workflows/trend-research.md
jobs, recruiting, candidatesreferences/workflows/job-market-and-recruitment.md
real estate, listings, hotelsreferences/workflows/real-estate-and-hospitality.md
price monitoring, e-commerce, productsreferences/workflows/ecommerce-price-monitoring.md
contact enrichment, email extractionreferences/workflows/contact-enrichment.md
knowledge base, RAG, LLM data feedreferences/workflows/knowledge-base-and-rag.md
company research, due diligencereferences/workflows/company-research.md

If no Actor matches in the index, search dynamically:

apify actors search "KEYWORDS" --json --limit 10 2>/dev/null

From results: items[].username/items[].name (Actor ID), items[].title, items[].stats.totalUsers30Days, items[].currentPricingInfo.pricingModel.

Step 2: Fetch Actor schema and check gotchas

Some Actors don't register an input schema with the platform (their schema lives in code). Try schema sources in this order — fall through on empty/error:

  1. Input schema (human-readable):
    apify actors info "ACTOR_ID" --input 2>/dev/null

If output is Error: No input schema found for this Actor, skip to source 2.

  1. Input schema (JSON keys only):
    apify actors info "ACTOR_ID" --input --json 2>/dev/null | jq '.input.schema.properties // empty | keys'

Empty result means no registered schema — fall through to source 3. To drill into a specific field:

    apify actors info "ACTOR_ID" --input --json 2>/dev/null | jq '.input.schema.properties.FIELD_NAME'
  1. README fallback (always works, contains usage examples):
    apify actors info "ACTOR_ID" --readme 2>/dev/null

Grep the README for an "Input" / "Example input" section to copy the JSON shape.

  1. Last resort — call with minimal known input (e.g. {"startUrls":[{"url":"..."}]} for crawlers) and let the Actor surface validation errors that reveal required fields. See references/gotchas.md for known-good minimal inputs for common Actors.

Also read references/gotchas.md to check for common pitfalls and cost guardrails for the selected Actor.

Step 3: Configure and run

Skip user preferences for simple lookups (e.g., "Nike's follower count"). Go straight to running with quick answer mode.

For larger tasks, confirm output format (quick answer / CSV / JSON) and result count.

Before starting the run, double-check whether the task is short enough for a blocking call or should use the async pattern from Step 0.

Standard run (blocking):

    apify actors call "ACTOR_ID" -i 'JSON_INPUT' --user-agent apify-claude-code-plugin/apify-ultimate-scraper --json 2>/dev/null

From output: .id (run ID), .status, .defaultDatasetId, .stats.durationMillis

Fetch results:

    apify datasets get-items DATASET_ID --format json

For CSV: apify datasets get-items DATASET_ID --format csv

Quick answer mode: Fetch results as JSON, pick top 5, present formatted in chat.

Save to file: Fetch results, use Write tool to save as YYYY-MM-DD_descriptive-name.csv or .json.

Large/long-running scrapes:

    apify actors start "ACTOR_ID" -i 'JSON_INPUT' --user-agent apify-claude-code-plugin/apify-ultimate-scraper --json 2>/dev/null

Poll: apify runs info RUN_ID --json (check .status for SUCCEEDED or FAILED).

Step 4: Deliver results

Report: result count, file location (if saved), key data fields, and links:

  • Dataset: https://console.apify.com/storage/datasets/DATASET_ID
  • Run: https://console.apify.com/actors/runs/RUN_ID

For multi-step workflows: suggest the next pipeline step from the workflow guide.

Troubleshooting

Common errors and pitfalls are documented in references/gotchas.md. Read it before running PPE (pay-per-event) Actors.

Skills เพิ่มเติมจาก apify

apify-influencer-brand-collabs
apify
ค้นหาความร่วมมือระหว่างแบรนด์และครีเอเตอร์บน Instagram โดยการเชื่อมต่อ Apify Actors ใช้เมื่อผู้ใช้ถามว่าใครร่วมงานกับแบรนด์ แบรนด์ใดที่ครีเอเตอร์ได้รับค่าตอบแทน...
apify-actor-development
apify
สร้าง, ดีบัก, และปรับใช้โปรแกรมคลาวด์แบบไร้เซิร์ฟเวอร์สำหรับการขูดเว็บ, ระบบอัตโนมัติ, และการประมวลผลข้อมูล รองรับเทมเพลต JavaScript, TypeScript, และ Python พร้อมไลบรารี Crawlee, Playwright, และ Cheerio ในตัวสำหรับการรวบรวมข้อมูลผ่าน HTTP และเบราว์เซอร์ รวมถึงการทดสอบในเครื่องผ่าน apify run พร้อมพื้นที่จัดเก็บแบบแยกส่วน, การตรวจสอบความถูกต้องของสคีมาสำหรับอินพุต/เอาต์พุต, และการปรับใช้ไปยังแพลตฟอร์ม Apify ผ่าน apify push ต้องมีการรับรองความถูกต้องของ Apify CLI และข้อมูลเมตา generatedBy ที่จำเป็นใน .actor/actor.json สำหรับ AI...
apify-actorization
apify
แปลงโปรเจกต์ที่มีอยู่ให้เป็น Apify Actors แบบไร้เซิร์ฟเวอร์ พร้อมการผสานรวม SDK เฉพาะภาษา รองรับ JavaScript/TypeScript (ด้วย Actor.init() / Actor.exit()), Python (ตัวจัดการบริบทแบบอะซิงก์) และภาษาอื่นๆ ผ่าน CLI wrapper มีเวิร์กโฟลว์ที่มีโครงสร้าง: apify init เพื่อสร้างโครงร่าง, ใช้ SDK wrapping, กำหนดค่า schemas อินพุต/เอาต์พุต, ทดสอบในเครื่องด้วย apify run, จากนั้นปรับใช้ด้วย apify push รวมถึงการตรวจสอบความถูกต้องของ schema อินพุตและเอาต์พุต, การทำ Docker containerization, และตัวเลือกการจ่ายต่อเหตุการณ์...
apify-content-analytics
apify
การวิเคราะห์เนื้อหาหลายแพลตฟอร์มผ่าน Apify Actors สำหรับ Instagram, Facebook, YouTube และ TikTok รองรับ Actors เฉพาะทางมากกว่า 17 รายการครอบคลุมโพสต์ รีล สตอรี่ คอมเมนต์ แฮชแท็ก ผู้ติดตาม และโฆษณาทั่วทั้งสี่แพลตฟอร์ม ดึงข้อมูลสคีมาของ Actor แบบไดนามิกโดยใช้ mcpc CLI เพื่อกำหนดอินพุตที่จำเป็นและฟิลด์เอาต์พุตที่มีอยู่ แสดงผลลัพธ์ในสามรูปแบบ: การแสดงผลแชทด่วน การส่งออก CSV หรือการส่งออก JSON พร้อมจำนวนผลลัพธ์ที่ปรับแต่งได้ ต้องใช้โทเค็น Apoken ในไฟล์ .env และ Node.js 20.6+...
apify-ecommerce
apify
ดึงข้อมูลสินค้า ราคา รีวิว และข้อมูลผู้ขายจากตลาดอีคอมเมิร์ซกว่า 50 แห่ง มีโหมดการทำงานสามแบบ: สินค้าและราคา (ติดตามราคา วิเคราะห์คู่แข่ง), รีวิวลูกค้า (วิเคราะห์ความรู้สึก ปัญหาคุณภาพ), และข้อมูลผู้ขาย (ค้นหาผู้ขายผ่าน Google Shopping) รองรับ Amazon (กว่า 20 ภูมิภาค), Walmart, eBay, IKEA, Costco และร้านค้าปลีกในยุโรป; ป้อนข้อมูลผ่าน URL สินค้า, URL หมวดหมู่ หรือค้นหาด้วยคำสำคัญ มีการวิเคราะห์ด้วย AI แบบเสริมเพื่อสร้างข้อมูลเชิงลึกเกี่ยวกับราคา...
apify-generate-output-schema
apify
สร้างสคีมาเอาต์พุต (dataset_schema.json, output_schema.json, key_value_store_schema.json) สำหรับ Apify Actor โดยการวิเคราะห์ซอร์สโค้ดของมัน ใช้เมื่อ...
apify-influencer-discovery
apify
ค้นหาและประเมินอินฟลูเอนเซอร์บน Instagram, Facebook, YouTube และ TikTok โดยใช้ Apify Actors เส้นทางคำขอค้นหาไปยัง Actors เฉพาะทางมากกว่า 15 รายการที่ครอบคลุมการขูดข้อมูลโปรไฟล์ การค้นหาแฮชแท็ก การวิเคราะห์การมีส่วนร่วม และการค้นหากลุ่มเฉพาะบนแพลตฟอร์มหลักทั้งหมด ดึงข้อมูลสคีมาของ Actor แบบไดนามิกผ่าน mcpc เพื่อกำหนดอินพุตที่จำเป็นและฟิลด์เอาต์พุตที่มีก่อนดำเนินการ รองรับโหมดการส่งออกสามแบบ: แสดงผลในแชทแบบอินไลน์, ไฟล์ CSV หรือ JSON พร้อมกำหนดจำนวนผลลัพธ์ที่ปรับแต่งได้...
apify-ultimate-scraper
apify
เว็บสแครปเปอร์อัตโนมัติที่เลือก Actor ที่เหมาะสมที่สุดสำหรับ 55+ แพลตฟอร์ม รวมถึง Instagram, TikTok, YouTube, Facebook, Google Maps และอื่นๆ ครอบคลุม Actor ที่กำหนดค่าไว้ล่วงหน้ากว่า 55 ตัวใน 8 แพลตฟอร์มหลัก พร้อมคำแนะนำการเลือกตามกรณีการใช้งานเฉพาะ (การสร้างลีด, การค้นหาอินฟลูเอนเซอร์, การตรวจสอบแบรนด์, การวิเคราะห์คู่แข่ง, การวิจัยเทรนด์) รองรับรูปแบบเอาต์พุตสามแบบ: การแสดงผลแชทด่วน, การส่งออก CSV หรือการส่งออก JSON พร้อมขีดจำกัดผลลัพธ์ที่ปรับแต่งได้ รวมถึงรูปแบบเวิร์กโฟลว์แบบหลาย Actor สำหรับการทำงานที่ซับซ้อน...