apify-ultimate-scraper

от apify

Универсальный веб-скрапер на базе ИИ для любой платформы. Извлекайте данные из Instagram, Facebook, TikTok, YouTube, LinkedIn, X/Twitter, Google Maps, Google Search,…

npx skills add https://github.com/apify/apify-claude-code-plugin --skill apify-ultimate-scraper

Universal web scraper

AI-driven data extraction from ~100 Actors across 15+ platforms via the Apify CLI.

Rule: Pass --json and redirect stderr with 2>/dev/null on data-returning commands (actors call, actors start, actors info, actors search, datasets get-items, runs info). JSON output is stable across CLI versions. stderr contains progress messages and version warnings that break JSON parsers if not redirected.

This rule does not apply to status/auth commands (apify info, apify --version, apify login). For those, use 2>&1 so authentication and version errors are visible.

Exception: if --input returns no data, re-run with 2>&1 to confirm whether the cause is a missing schema vs. a network/auth error.

Rule: pass --user-agent apify-claude-code-plugin/apify-ultimate-scraper only on actor start and actor run commands (apify actors start, apify actors call). Do not add it to login, info, search, or other CLI commands.

Prerequisites

  • Apify CLI v1.5.0+ (npm install -g apify-cli) — older versions reject the user-agent flag
  • Authenticated session (see below)

Authentication

If a CLI command fails with an auth error, authenticate using one of these methods:

  1. OAuth (interactive): apify login (opens browser)
  2. Environment variable: export APIFY_TOKEN=your_token_here
  3. From .env file: source .env (if the file contains APIFY_TOKEN=...)

Generate token: https://console.apify.com/settings/integrations

Workflow

Step 0: Verify CLI readiness before doing anything else

Before using the Apify CLI, always verify the local environment:

  1. Check that the CLI is installed and is v1.5.0 or newer:
    apify --help
    apify --version

If this fails, install the CLI first:

       npm install -g apify-cli
  1. Check that the CLI is authenticated:
    # Auth check — do NOT pipe to /dev/null, you need to see errors
    apify info 2>&1

If this shows the user is not logged in, instruct them to authenticate with a token:

    apify login --token TOKEN
  1. Authenticated Apify CLI commands need file access to ~/.apify/, where the CLI keeps its credentials. A host that sandboxes file access can deny this even when the login is valid; that is a sandbox problem, not a login problem.

  2. Assume many Apify commands block with zero output until completion, so allow at least 60 seconds before treating one as stuck. If your shell tool takes a timeout, raise it accordingly.

  3. For long or unknown-duration runs, prefer the async pattern:

    apify actors start "ACTOR_ID" -i 'JSON_INPUT' --user-agent apify-claude-code-plugin/apify-ultimate-scraper --json 2>/dev/null

Then poll the run status:

    apify runs info RUN_ID --json

Check .status for SUCCEEDED or FAILED.

Step 1: Understand goal and select Actor

Identify the target platform and use case. Read references/actor-index.md to find the right Actor. Prefer apify-tier actors; use community-tier only when no apify actor covers the task. For input schemas, fetch dynamically: apify actors info "ACTOR_ID" --input --json 2>/dev/null If the output is empty, re-run without the redirect (2>&1) to surface auth or network errors before proceeding.

If the task involves a multi-step pipeline, also read the matching workflow guide:

Task involves...Read
leads, contacts, emails, B2Breferences/workflows/lead-generation.md
competitor, ads, pricingreferences/workflows/competitive-intel.md
influencer, creatorreferences/workflows/influencer-vetting.md
brand, mentions, sentimentreferences/workflows/brand-monitoring.md
reviews, ratings, reputationreferences/workflows/review-analysis.md
SEO, SERP, crawl, content, RAGreferences/workflows/content-and-seo.md
analytics, engagement, performancereferences/workflows/social-media-analytics.md
trends, keywords, hashtagsreferences/workflows/trend-research.md
jobs, recruiting, candidatesreferences/workflows/job-market-and-recruitment.md
real estate, listings, hotelsreferences/workflows/real-estate-and-hospitality.md
price monitoring, e-commerce, productsreferences/workflows/ecommerce-price-monitoring.md
contact enrichment, email extractionreferences/workflows/contact-enrichment.md
knowledge base, RAG, LLM data feedreferences/workflows/knowledge-base-and-rag.md
company research, due diligencereferences/workflows/company-research.md

If no Actor matches in the index, search dynamically:

apify actors search "KEYWORDS" --json --limit 10 2>/dev/null

From results: items[].username/items[].name (Actor ID), items[].title, items[].stats.totalUsers30Days, items[].currentPricingInfo.pricingModel.

Step 2: Fetch Actor schema and check gotchas

Some Actors don't register an input schema with the platform (their schema lives in code). Try schema sources in this order — fall through on empty/error:

  1. Input schema (human-readable):
    apify actors info "ACTOR_ID" --input 2>/dev/null

If output is Error: No input schema found for this Actor, skip to source 2.

  1. Input schema (JSON keys only):
    apify actors info "ACTOR_ID" --input --json 2>/dev/null | jq '.input.schema.properties // empty | keys'

Empty result means no registered schema — fall through to source 3. To drill into a specific field:

    apify actors info "ACTOR_ID" --input --json 2>/dev/null | jq '.input.schema.properties.FIELD_NAME'
  1. README fallback (always works, contains usage examples):
    apify actors info "ACTOR_ID" --readme 2>/dev/null

Grep the README for an "Input" / "Example input" section to copy the JSON shape.

  1. Last resort — call with minimal known input (e.g. {"startUrls":[{"url":"..."}]} for crawlers) and let the Actor surface validation errors that reveal required fields. See references/gotchas.md for known-good minimal inputs for common Actors.

Also read references/gotchas.md to check for common pitfalls and cost guardrails for the selected Actor.

Step 3: Configure and run

Skip user preferences for simple lookups (e.g., "Nike's follower count"). Go straight to running with quick answer mode.

For larger tasks, confirm output format (quick answer / CSV / JSON) and result count.

Before starting the run, double-check whether the task is short enough for a blocking call or should use the async pattern from Step 0.

Standard run (blocking):

    apify actors call "ACTOR_ID" -i 'JSON_INPUT' --user-agent apify-claude-code-plugin/apify-ultimate-scraper --json 2>/dev/null

From output: .id (run ID), .status, .defaultDatasetId, .stats.durationMillis

Fetch results:

    apify datasets get-items DATASET_ID --format json

For CSV: apify datasets get-items DATASET_ID --format csv

Quick answer mode: Fetch results as JSON, pick top 5, present formatted in chat.

Save to file: Fetch results, use Write tool to save as YYYY-MM-DD_descriptive-name.csv or .json.

Large/long-running scrapes:

    apify actors start "ACTOR_ID" -i 'JSON_INPUT' --user-agent apify-claude-code-plugin/apify-ultimate-scraper --json 2>/dev/null

Poll: apify runs info RUN_ID --json (check .status for SUCCEEDED or FAILED).

Step 4: Deliver results

Report: result count, file location (if saved), key data fields, and links:

  • Dataset: https://console.apify.com/storage/datasets/DATASET_ID
  • Run: https://console.apify.com/actors/runs/RUN_ID

For multi-step workflows: suggest the next pipeline step from the workflow guide.

Troubleshooting

Common errors and pitfalls are documented in references/gotchas.md. Read it before running PPE (pay-per-event) Actors.

Больше skills от apify

apify-influencer-brand-collabs
apify
Обнаруживайте партнёрства брендов и создателей в Instagram, объединяя Apify Actors. Используйте, когда пользователь спрашивает, кто сотрудничает с брендом, с какими брендами работал создатель на платной основе…
apify-actor-development
apify
Создавайте, отлаживайте и развертывайте серверные облачные программы для веб-скрапинга, автоматизации и обработки данных. Поддерживает шаблоны JavaScript, TypeScript и Python с интегрированными библиотеками Crawlee, Playwright и Cheerio для HTTP- и браузерного краулинга. Включает локальное тестирование через apify run с изолированным хранилищем, проверку схемы для входных/выходных данных и развертывание на платформе Apify через apify push. Требуется аутентификация Apify CLI и обязательные метаданные generatedBy в .actor/actor.json для AI...
apify-actorization
apify
Преобразуйте существующие проекты в бессерверные Apify Actors с интеграцией SDK для конкретного языка. Поддерживает JavaScript/TypeScript (с Actor.init() / Actor.exit()), Python (асинхронный контекстный менеджер) и любой язык через CLI-обёртку. Предоставляет структурированный рабочий процесс: apify init для создания каркаса, применение SDK-обёртки, настройка схем ввода/вывода, локальное тестирование с apify run, затем развёртывание с apify push. Включает валидацию схем ввода и вывода, контейнеризацию Docker и опциональную оплату за событие...
apify-content-analytics
apify
Мультиплатформенный анализ контента через Apify Actors для Instagram, Facebook, YouTube и TikTok. Поддерживает 17+ специализированных Actors, охватывающих посты, рилсы, истории, комментарии, хештеги, подписчиков и рекламу на всех четырех платформах. Динамически получает схемы Actors с помощью mcpc CLI для определения необходимых входных данных и доступных полей вывода. Выводит результаты в трех форматах: быстрый чат-дисплей, экспорт в CSV или экспорт в JSON с настраиваемым количеством результатов. Требует токен Apify в файле .env и Node.js 20.6+...
apify-ecommerce
apify
Извлекайте данные о товарах, ценах, отзывах и продавцах с более чем 50 торговых площадок. Три режима работы: «Товары и цены» (отслеживание цен, анализ конкурентов), «Отзывы покупателей» (анализ тональности, проблемы с качеством) и «Информация о продавцах» (поиск поставщиков через Google Shopping). Поддерживает Amazon (более 20 регионов), Walmart, eBay, IKEA, Costco и европейских ритейлеров; ввод через URL товаров, URL категорий или поиск по ключевым словам. Опциональный AI-анализ формирует выводы о ценах...
apify-generate-output-schema
apify
Генерирует схемы вывода (dataset_schema.json, output_schema.json, key_value_store_schema.json) для Apify Actor путем анализа его исходного кода. Используйте, когда…
apify-influencer-discovery
apify
Обнаружение и оценка инфлюенсеров в Instagram, Facebook, YouTube и TikTok с помощью Apify Actors. Маршрутизирует запросы на поиск к 15+ специализированным Actors, охватывающим сбор профилей, поиск по хештегам, анализ вовлеченности и поиск по нишам на всех основных платформах. Динамически получает схемы Actors через mcpc для определения необходимых входных данных и доступных полей вывода перед выполнением. Поддерживает три режима экспорта: встроенное отображение в чате, вывод в CSV или JSON-файл с настраиваемым количеством результатов...
apify-ultimate-scraper
apify
Автоматизированный веб-скрапер, выбирающий оптимальные Акторы для 55+ платформ, включая Instagram, TikTok, YouTube, Facebook, Google Maps и другие. Охватывает 55+ предварительно настроенных Акторов на 8 основных платформах с рекомендациями по выбору в зависимости от сценария использования (генерация лидов, поиск инфлюенсеров, мониторинг бренда, анализ конкурентов, исследование трендов). Поддерживает три формата вывода: быстрый чат-дисплей, экспорт в CSV или JSON с настраиваемыми лимитами результатов. Включает многоАкторные шаблоны рабочих процессов для сложных...