apify-ultimate-scraper

von apify

Universeller KI-gestützter Web-Scraper für jede Plattform. Scrapen Sie Daten von Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Google Search,…

npx skills add https://github.com/apify/apify-claude-code-plugin --skill apify-ultimate-scraper

Universal web scraper

AI-driven data extraction from ~100 Actors across 15+ platforms via the Apify CLI.

Rule: Pass --json and redirect stderr with 2>/dev/null on data-returning commands (actors call, actors start, actors info, actors search, datasets get-items, runs info). JSON output is stable across CLI versions. stderr contains progress messages and version warnings that break JSON parsers if not redirected.

This rule does not apply to status/auth commands (apify info, apify --version, apify login). For those, use 2>&1 so authentication and version errors are visible.

Exception: if --input returns no data, re-run with 2>&1 to confirm whether the cause is a missing schema vs. a network/auth error.

Rule: pass --user-agent apify-claude-code-plugin/apify-ultimate-scraper only on actor start and actor run commands (apify actors start, apify actors call). Do not add it to login, info, search, or other CLI commands.

Prerequisites

  • Apify CLI v1.5.0+ (npm install -g apify-cli) — older versions reject the user-agent flag
  • Authenticated session (see below)

Authentication

If a CLI command fails with an auth error, authenticate using one of these methods:

  1. OAuth (interactive): apify login (opens browser)
  2. Environment variable: export APIFY_TOKEN=your_token_here
  3. From .env file: source .env (if the file contains APIFY_TOKEN=...)

Generate token: https://console.apify.com/settings/integrations

Workflow

Step 0: Verify CLI readiness before doing anything else

Before using the Apify CLI, always verify the local environment:

  1. Check that the CLI is installed and is v1.5.0 or newer:
    apify --help
    apify --version

If this fails, install the CLI first:

       npm install -g apify-cli
  1. Check that the CLI is authenticated:
    # Auth check — do NOT pipe to /dev/null, you need to see errors
    apify info 2>&1

If this shows the user is not logged in, instruct them to authenticate with a token:

    apify login --token TOKEN
  1. Authenticated Apify CLI commands need file access to ~/.apify/, where the CLI keeps its credentials. A host that sandboxes file access can deny this even when the login is valid; that is a sandbox problem, not a login problem.

  2. Assume many Apify commands block with zero output until completion, so allow at least 60 seconds before treating one as stuck. If your shell tool takes a timeout, raise it accordingly.

  3. For long or unknown-duration runs, prefer the async pattern:

    apify actors start "ACTOR_ID" -i 'JSON_INPUT' --user-agent apify-claude-code-plugin/apify-ultimate-scraper --json 2>/dev/null

Then poll the run status:

    apify runs info RUN_ID --json

Check .status for SUCCEEDED or FAILED.

Step 1: Understand goal and select Actor

Identify the target platform and use case. Read references/actor-index.md to find the right Actor. Prefer apify-tier actors; use community-tier only when no apify actor covers the task. For input schemas, fetch dynamically: apify actors info "ACTOR_ID" --input --json 2>/dev/null If the output is empty, re-run without the redirect (2>&1) to surface auth or network errors before proceeding.

If the task involves a multi-step pipeline, also read the matching workflow guide:

Task involves...Read
leads, contacts, emails, B2Breferences/workflows/lead-generation.md
competitor, ads, pricingreferences/workflows/competitive-intel.md
influencer, creatorreferences/workflows/influencer-vetting.md
brand, mentions, sentimentreferences/workflows/brand-monitoring.md
reviews, ratings, reputationreferences/workflows/review-analysis.md
SEO, SERP, crawl, content, RAGreferences/workflows/content-and-seo.md
analytics, engagement, performancereferences/workflows/social-media-analytics.md
trends, keywords, hashtagsreferences/workflows/trend-research.md
jobs, recruiting, candidatesreferences/workflows/job-market-and-recruitment.md
real estate, listings, hotelsreferences/workflows/real-estate-and-hospitality.md
price monitoring, e-commerce, productsreferences/workflows/ecommerce-price-monitoring.md
contact enrichment, email extractionreferences/workflows/contact-enrichment.md
knowledge base, RAG, LLM data feedreferences/workflows/knowledge-base-and-rag.md
company research, due diligencereferences/workflows/company-research.md

If no Actor matches in the index, search dynamically:

apify actors search "KEYWORDS" --json --limit 10 2>/dev/null

From results: items[].username/items[].name (Actor ID), items[].title, items[].stats.totalUsers30Days, items[].currentPricingInfo.pricingModel.

Step 2: Fetch Actor schema and check gotchas

Some Actors don't register an input schema with the platform (their schema lives in code). Try schema sources in this order — fall through on empty/error:

  1. Input schema (human-readable):
    apify actors info "ACTOR_ID" --input 2>/dev/null

If output is Error: No input schema found for this Actor, skip to source 2.

  1. Input schema (JSON keys only):
    apify actors info "ACTOR_ID" --input --json 2>/dev/null | jq '.input.schema.properties // empty | keys'

Empty result means no registered schema — fall through to source 3. To drill into a specific field:

    apify actors info "ACTOR_ID" --input --json 2>/dev/null | jq '.input.schema.properties.FIELD_NAME'
  1. README fallback (always works, contains usage examples):
    apify actors info "ACTOR_ID" --readme 2>/dev/null

Grep the README for an "Input" / "Example input" section to copy the JSON shape.

  1. Last resort — call with minimal known input (e.g. {"startUrls":[{"url":"..."}]} for crawlers) and let the Actor surface validation errors that reveal required fields. See references/gotchas.md for known-good minimal inputs for common Actors.

Also read references/gotchas.md to check for common pitfalls and cost guardrails for the selected Actor.

Step 3: Configure and run

Skip user preferences for simple lookups (e.g., "Nike's follower count"). Go straight to running with quick answer mode.

For larger tasks, confirm output format (quick answer / CSV / JSON) and result count.

Before starting the run, double-check whether the task is short enough for a blocking call or should use the async pattern from Step 0.

Standard run (blocking):

    apify actors call "ACTOR_ID" -i 'JSON_INPUT' --user-agent apify-claude-code-plugin/apify-ultimate-scraper --json 2>/dev/null

From output: .id (run ID), .status, .defaultDatasetId, .stats.durationMillis

Fetch results:

    apify datasets get-items DATASET_ID --format json

For CSV: apify datasets get-items DATASET_ID --format csv

Quick answer mode: Fetch results as JSON, pick top 5, present formatted in chat.

Save to file: Fetch results, use Write tool to save as YYYY-MM-DD_descriptive-name.csv or .json.

Large/long-running scrapes:

    apify actors start "ACTOR_ID" -i 'JSON_INPUT' --user-agent apify-claude-code-plugin/apify-ultimate-scraper --json 2>/dev/null

Poll: apify runs info RUN_ID --json (check .status for SUCCEEDED or FAILED).

Step 4: Deliver results

Report: result count, file location (if saved), key data fields, and links:

  • Dataset: https://console.apify.com/storage/datasets/DATASET_ID
  • Run: https://console.apify.com/actors/runs/RUN_ID

For multi-step workflows: suggest the next pipeline step from the workflow guide.

Troubleshooting

Common errors and pitfalls are documented in references/gotchas.md. Read it before running PPE (pay-per-event) Actors.

Mehr Skills von apify

apify-influencer-brand-collabs
apify
Entdecken Sie Instagram-Marken-Creator-Partnerschaften durch Verkettung von Apify-Actors. Verwenden Sie dies, wenn der Benutzer fragt, wer mit einer Marke zusammenarbeitet, mit welchen Marken ein Creator bezahlte Kooperationen hatte…
apify-actor-development
apify
Erstellen, debuggen und bereitstellen von serverlosen Cloud-Programmen für Web Scraping, Automatisierung und Datenverarbeitung. Unterstützt JavaScript-, TypeScript- und Python-Vorlagen mit integrierten Crawlee-, Playwright- und Cheerio-Bibliotheken für HTTP- und browserbasiertes Crawling. Beinhaltet lokale Tests über apify run mit isoliertem Speicher, Schema-Validierung für Ein-/Ausgaben und Bereitstellung auf der Apify-Plattform über apify push. Erfordert Apify CLI-Authentifizierung und zwingend erforderliche generatedBy-Metadaten in .actor/actor.json für KI...
apify-actorization
apify
Konvertieren Sie bestehende Projekte in serverlose Apify Actors mit sprachspezifischer SDK-Integration. Unterstützt JavaScript/TypeScript (mit Actor.init() / Actor.exit()), Python (asynchroner Kontextmanager) und jede Sprache über CLI-Wrapper. Bietet strukturierten Workflow: apify init zum Erstellen des Grundgerüsts, Anwenden von SDK-Wrapping, Konfigurieren von Eingabe-/Ausgabeschemata, lokales Testen mit apify run, dann Bereitstellung mit apify push. Enthält Eingabe- und Ausgabeschemavalidierung, Docker-Containerisierung und optionales Pay-per-Event...
apify-content-analytics
apify
Multiplattform-Inhaltsanalysen über Apify Actors für Instagram, Facebook, YouTube und TikTok. Unterstützt 17+ spezialisierte Actors für Beiträge, Reels, Stories, Kommentare, Hashtags, Follower und Werbung auf allen vier Plattformen. Ruft dynamisch Actor-Schemas über die mcpc CLI ab, um erforderliche Eingaben und verfügbare Ausgabefelder zu ermitteln. Gibt Ergebnisse in drei Formaten aus: Kurzanzeige im Chat, CSV-Export oder JSON-Export mit anpassbaren Ergebnisanzahlen. Erfordert Apify-Token in der .env-Datei und Node.js 20.6+...
apify-ecommerce
apify
Extrahiere Produktdaten, Preise, Bewertungen und Verkäuferinformationen von über 50 E-Commerce-Marktplätzen. Drei Workflow-Modi: Produkte & Preise (Preisverfolgung, Wettbewerbsanalyse), Kundenbewertungen (Sentimentanalyse, Qualitätsprobleme) und Verkäufer-Intelligenz (Anbietererkennung über Google Shopping). Unterstützt Amazon (über 20 Regionen), Walmart, eBay, IKEA, Costco und europäische Einzelhändler; Eingabe über Produkt-URLs, Kategorie-URLs oder Stichwortsuche. Optionale KI-gestützte Analyse generiert Erkenntnisse zu Preisen...
apify-generate-output-schema
apify
Generieren Sie Ausgabeschemata (dataset_schema.json, output_schema.json, key_value_store_schema.json) für einen Apify Actor durch Analyse seines Quellcodes. Verwenden Sie, wenn…
apify-influencer-discovery
apify
Entdecken und bewerten Sie Influencer auf Instagram, Facebook, YouTube und TikTok mit Apify Actors. Leitet Discovery-Anfragen an über 15 spezialisierte Actors weiter, die Profil-Scraping, Hashtag-Suche, Engagement-Analyse und Nischen-Discovery auf allen großen Plattformen abdecken. Ruft dynamisch Actor-Schemas über mcpc ab, um erforderliche Eingaben und verfügbare Ausgabefelder vor der Ausführung zu ermitteln. Unterstützt drei Exportmodi: Inline-Chat-Anzeige, CSV- oder JSON-Dateiausgabe mit anpassbaren Ergebnisanzahlen...
apify-ultimate-scraper
apify
Automatisierter Web-Scraper, der optimale Actors für über 55 Plattformen auswählt, darunter Instagram, TikTok, YouTube, Facebook, Google Maps und mehr. Umfasst über 55 vorkonfigurierte Actors auf 8 großen Plattformen mit anwendungsfallspezifischer Auswahlhilfe (Lead-Generierung, Influencer-Entdeckung, Markenüberwachung, Wettbewerbsanalyse, Trendforschung). Unterstützt drei Ausgabeformate: schnelle Chat-Anzeige, CSV-Export oder JSON-Export mit anpassbaren Ergebnislimits. Enthält Multi-Actor-Workflow-Muster für komplexe...