apify-ecommerce

par apify

Extraire des données e-commerce pour les prix, les avis, les meilleures ventes et la découverte de vendeurs sur plus de 30 plateformes, dont Amazon, Walmart, eBay, Shopify, WooCommerce, et…

npx skills add https://github.com/apify/awesome-skills --skill apify-ecommerce

E-Commerce Cluster

Answer natural language e-commerce questions by routing to the right Apify Actor and delivering a synthesized answer via the apify CLI.

CLI rules: Always pass --user-agent apify-awesome-skills/apify-ecommerce, --json (or the relevant --format flag on datasets get-items), and 2>/dev/null. The --user-agent flag is critical for telemetry — never omit it.

Prerequisites

(No need to check it upfront)

  • Apify CLI v1.5.0+ (npm install -g apify-cli)
  • jq (recommended for quick extraction and filtering; brew install jq on macOS, apt install jq on Linux)
  • Authentication via one of:

Verify auth: apify info --user-agent apify-awesome-skills/apify-ecommerce — should show username and userId.

Workflow

Copy this checklist and track progress:

Task Progress:
- [ ] Step 1: Detect intent and select Actor
- [ ] Step 2: Fetch Actor schema
- [ ] Step 3: Ask user preferences (format, result count)
- [ ] Step 4: Run the Actor and fetch results
- [ ] Step 5: Analyze results and deliver synthesized answer

Step 1: Detect Intent and Select Actor

Classify the user's message into an intent, then pick the right Actor.

Intent signals:

Signals in user messageIntent
price, cost, cheapest, compare prices, pricingpricing
review, rating, sentiment, stars, feedbackreviews
bestseller, top selling, most popular, trendingbestsellers
seller, vendor, reseller, who sellssellers
all products from, scrape store, full catalogstore-scrape
what platform, built on, tech stack, Shopify or WooCommercetech-stack
SEO, listing quality, product page auditseo-audit
competitor funnel, competitor pricing, conversion elementscompetitor
search intent, keyword intent, SERP intentsearch-intent
match products, same product on different platformsproduct-matching
restaurant, food delivery, DoorDash, UberEats, TheForkfood-delivery
enrich store, store metadata, store liststore-enrichment
event, concert, ticket, Eventbriteevents
property, real estate, house listing, Realtorreal-estate
Facebook ads, Meta ads, ad library, competitor adsads-intelligence
classified, Craigslist, used item for saleclassifieds
car, used car, vehicle, automotive, Webmotorsautomotive
pins, inspiration, Pinterest boards, visual search, Pinterest trendscontent-discovery
TikTok Shop, TikTok store, TikTok creatortiktok-shop
website for sale, domain for sale, Flippawebsite-marketplace

If multiple intents are detected, ask: "Do you want [intent A] or [intent B]?"

Actor routing table — always try Primary first, switch to Fallback only if it fails or returns 0 results. The Primary actor (apify/e-commerce-scraping-tool) handles most intents once you feed the right input mode:

  • Have target URLs (a listing, profile, or category page) → detailsUrls / listingUrls.
  • Have a keyword/marketplace → keyword + marketplaces (product-details mode).
  • Broad discovery (competitor, search-intent, classifieds, automotive, real-estate, website-marketplace, events) → use searchEngineKeyword (search-engine mode) or keyword/detailsUrls depending on whether you have a query or URLs.

Exception — skip the Primary and go straight to the Fallback for intents the Primary genuinely can't do (different data source or specialized analysis): tech-stack, seo-audit, store-enrichment, product-matching, ads-intelligence, content-discovery (Pinterest), and tiktok-shop. Routing these to the Primary wastes a run and credits.

IntentPlatformPrimary ActorFallback Actor
pricingAmazon / Walmart / genericapify/e-commerce-scraping-tool—
pricingeBayapify/e-commerce-scraping-toolivanvs/ebay-scraper-pay-per-result
pricingEtsyapify/e-commerce-scraping-toolepctex/etsy-scraper
pricingGoogle Shoppingapify/e-commerce-scraping-toolepctex/google-shopping-scraper
pricingFacebook Marketplaceapify/e-commerce-scraping-toolapify/facebook-marketplace-scraper
pricingSHEINapify/e-commerce-scraping-toolseamless_coffer/shein-product-scraper
pricingLazadaapify/e-commerce-scraping-toolfatihtahta/lazada-scraper
pricingCanadian Tireapify/e-commerce-scraping-toolazzouzana/canadiantire-ca-scraper
pricingTescoapify/e-commerce-scraping-toolradeance/tesco-scraper
pricingShopifyapify/e-commerce-scraping-tooltrovevault/shopify-products-scraper
pricingWooCommerceapify/e-commerce-scraping-tooltrovevault/woocommerce-products-scraper
reviewsAmazon / Walmart / genericapify/e-commerce-scraping-tooljunglee/amazon-reviews-scraper
reviewsTrustpilotapify/e-commerce-scraping-toolcasper11515/trustpilot-reviews-scraper
reviewsTheForkapify/e-commerce-scraping-tooljdtpnjtp/thefork-restaurant-scraper-advanced
bestsellersAmazonapify/e-commerce-scraping-tooljunglee/amazon-bestsellers
sellersAmazonapify/e-commerce-scraping-tooljunglee/amazon-seller-scraper
sellerseBayapify/e-commerce-scraping-toolivanvs/ebay-scraper-pay-per-result
store-scrapeShopifyapify/e-commerce-scraping-tooltrovevault/shopify-products-scraper
store-scrapeWooCommerceapify/e-commerce-scraping-tooltrovevault/woocommerce-products-scraper
store-scrapeAmazonapify/e-commerce-scraping-tooljunglee/Amazon-crawler
store-scrapeFlippaapify/e-commerce-scraping-toolscraped/flippa-scraper
tech-stackanyapify/e-commerce-scraping-tooltrovevault/e-commerce-tech-stack-detector
seo-auditanyapify/e-commerce-scraping-tooltrovevault/product-listing-seo-auditor
competitoranyapify/e-commerce-scraping-tooltrovevault/competitor-intelligence-scraper---funnel-pricing-conversion
search-intentanyapify/e-commerce-scraping-tooltrovevault/ai-serp-intent-extractor---search-intent-classifier
product-matchinganyapify/e-commerce-scraping-tool—
store-enrichmentanyapify/e-commerce-scraping-tooltrovevault/e-commerce-store-data-enricher
food-deliveryDoorDashapify/e-commerce-scraping-tooltri_angle/doordash-store-details-scraper
food-deliveryUberEatsapify/e-commerce-scraping-toole-commerce/ubereats-reviews-scraper
food-deliveryTheForkapify/e-commerce-scraping-tooljdtpnjtp/thefork-restaurant-scraper-advanced
ads-intelligenceFacebook / Metaapify/e-commerce-scraping-toolapify/facebook-ads-scraper
classifiedsCraigslistapify/e-commerce-scraping-toolivanvs/craigslist-scraper-pay-per-result
automotiveWebmotorsapify/e-commerce-scraping-toolstealth_mode/webmotors-auto-search-scraper
eventsEventbriteapify/e-commerce-scraping-toolaitorsm/eventbrite
real-estateRealtor.comapify/e-commerce-scraping-toolpowerai/realtor-properties-search-scraper
content-discoveryPinterestapify/e-commerce-scraping-toolfatihtahta/pinterest-scraper-search
tiktok-shopTikTok Shopapify/e-commerce-scraping-toollemur/tiktok-shop-creators
website-marketplaceFlippaapify/e-commerce-scraping-toolscraped/flippa-scraper

Escalation — if both Primary and Fallback fail or return 0 results, discover a current alternative live instead of guessing an ID:

# Find relevant, well-rated, pay-per-event Actors for the platform/intent.
# Keep the default relevance sort — `--sort-by popularity` surfaces generic
# big-name scrapers over the platform you actually asked for.
apify actors search "PLATFORM or INTENT keywords" \
  --pricing-model PAY_PER_EVENT --limit 10 --json \
  --user-agent apify-awesome-skills/apify-ecommerce 2>/dev/null \
  | jq '[.items[]
      | select(.stats.totalUsers > 100 and .actorReviewRating > 4.5)
      | {id: (.username + "/" + .name), users: .stats.totalUsers,
         rating: (.actorReviewRating | (. * 100 | round / 100)),
         pricing: .currentPricingInfo.pricingModel}]'

Pick the top match. Before running it, confirm it requests only limited permissions (check the Actor's Store page / README — prefer Actors that don't require full account access). If the PAY_PER_EVENT filter returns nothing, drop the --pricing-model flag and re-run, keeping the ≥100-users and ≥4.5-rating bar.

Step 2: Fetch Actor Schema

Fetch the Actor summary, input schema, and README:

# Summary (title, description, pricing, stats)
apify actors info "ACTOR_ID" --user-agent apify-awesome-skills/apify-ecommerce --json 2>/dev/null

# Input schema — use --input WITHOUT --json to get the clean schema directly.
# (Adding --json returns the full ~250 KB actor object instead, with the schema
#  buried as an escaped string under .taggedBuilds.latest.build.inputSchema.)
apify actors info "ACTOR_ID" --user-agent apify-awesome-skills/apify-ecommerce --input 2>/dev/null

# README (capabilities, examples, gotchas)
apify actors info "ACTOR_ID" --user-agent apify-awesome-skills/apify-ecommerce --readme 2>/dev/null

Replace ACTOR_ID with the selected Actor (e.g., apify/e-commerce-scraping-tool).

Primary actor input cheat-sheet. apify/e-commerce-scraping-tool is mode-driven — pick fields by intent (always set the matching max…Results cap):

IntentMinimal input
pricing (keyword){"keyword": "wireless earbuds", "marketplaces": ["www.amazon.com"], "maxProductResults": 50}
pricing (specific URLs){"detailsUrls": [{"url": "https://…"}], "maxProductResults": 50}
store-scrape (category){"listingUrls": [{"url": "https://…/category"}], "maxProductResults": 500}
reviews{"keywordReviews": "echo dot", "marketplacesReviews": ["www.amazon.com"], "sortReview": "Most recent", "maxReviewResults": 200}
sellers{"sellerUrls": [{"url": "https://…"}], "maxSellerResults": 50}
pricing (Google Shopping){"searchEngineKeyword": "ps5", "countryCode": "us", "maxSearchEngineResults": 50}
food-delivery{"keywordDelivery": "pizza", "marketplacesDelivery": ["www.doordash.com"], "addressDelivery": "New York, NY", "maxDeliveryResults": 50}

For any other actor (or fields not listed), fetch the schema with the --input command above.

Step 3: Ask User Preferences

Before running, ask:

  1. Output format:
    • Quick answer (default) — synthesized answer in chat, no file saved
    • CSV — full export saved to disk
    • JSON — full export saved to disk
  2. Result count — suggest defaults by intent:
IntentDefault
pricing50 products
reviews200 reviews
bestsellers100 items
sellers50 sellers
store-scrapeall (unlimited)
food-delivery50 restaurants
all others20–50

Cost safety: Always set a sensible result limit in the Actor input. For the Primary actor the cap field is mode-specific — maxProductResults, maxReviewResults, maxSellerResults, maxSearchEngineResults, or maxDeliveryResults (there is no single maxResults). For Fallback actors, use whatever the schema exposes (maxResults, resultsLimit, maxItems, maxCrawledPages, etc.). Default to the per-intent values above unless the user explicitly asks for more. Warn the user before running large scrapes (1000+ results) as they consume more Apify credits.

Step 4: Run the Actor and Fetch Results

Two steps: run the Actor (blocks until done), then fetch dataset items in the requested format.

Run the Actor — returns run metadata as JSON; extract defaultDatasetId for the next step:

apify actors call "ACTOR_ID" -i 'JSON_INPUT' \
  --user-agent apify-awesome-skills/apify-ecommerce --json 2>/dev/null

From the output use .id (run ID), .status (should be SUCCEEDED), and .defaultDatasetId.

Fetch results — pick the variant based on the user's preference:

# Quick answer: total count + fields + top 5 in chat (no file)
apify datasets info DATASET_ID --json \
  --user-agent apify-awesome-skills/apify-ecommerce 2>/dev/null \
  | jq '{itemCount, fields, consoleUrl}'
apify datasets get-items DATASET_ID --limit 5 \
  --user-agent apify-awesome-skills/apify-ecommerce --format json 2>/dev/null

# CSV file
apify datasets get-items DATASET_ID \
  --user-agent apify-awesome-skills/apify-ecommerce --format csv 2>/dev/null > YYYY-MM-DD_OUTPUT_FILE.csv

# JSON file
apify datasets get-items DATASET_ID \
  --user-agent apify-awesome-skills/apify-ecommerce --format json 2>/dev/null > YYYY-MM-DD_OUTPUT_FILE.json

Other --format options: jsonl, xlsx, xml, rss, html. Use --offset N to paginate large datasets.

Tip: for anything more than a quick peek, save the dataset to a local file first (with > file.json / > file.csv) and run further analysis from disk. apify datasets get-items always streams over the network, so piping it straight into jq re-downloads the whole thing every iteration.

Combining with jq for quick extraction:

Treat jq as a complement to apify datasets get-items, not a replacement: server-side --limit / --offset / --format keeps cost and bandwidth down. Use jq on a sample item or on a file you already saved.

# Discover real field names from one sample item (Actor outputs vary —
# use this before composing further jq queries)
apify datasets get-items DATASET_ID --limit 1 --format json \
  --user-agent apify-awesome-skills/apify-ecommerce 2>/dev/null \
  | jq '.[0]'

# Quick aggregation from a JSON file you already saved with the commands above
jq '[.[] | select(.rating != null and .rating >= 4.5)] | length' YYYY-MM-DD_OUTPUT_FILE.json

Step 5: Analyze Results and Deliver Answer

After the run completes, deliver a direct synthesized answer — not a data dump:

  • Pricing: price range, average, top 5 cheapest with URLs
  • Reviews: average rating, top 3 positive and negative themes, recent snippets
  • Bestsellers: top 10 by rank with name, price, rating, URL
  • Sellers: total sellers, price range per seller, unauthorized seller flags
  • Store-scrape: total products, category breakdown, price range, stock summary
  • Tech-stack: platform detected, confidence level, notable plugins
  • Food delivery: restaurant count, average rating, price tier breakdown
  • Ads intelligence: total ads, active/inactive split, top creative formats

Error Handling

  • Auth error → run apify login, or set APIFY_TOKEN env var
  • Actor not found → check Actor ID spelling in the routing table
  • Run status FAILED → open the console URL (.consoleUrl from run metadata) for logs
  • Timeout / very long run → pass --timeout <seconds> to apify actors call
  • No results → broaden the keyword, switch to the Fallback Actor, then use the Escalation discovery command (under Step 1) if both fail
  • proxy is required → add "proxy": {"useApifyProxy": true} to the Actor input
  • Platform not detected → default to apify/e-commerce-scraping-tool with generic intent

Gotchas

  • --input --json is a trap. It returns the full ~250 KB actor object, not the schema. Use apify actors info ID --input --user-agent apify-awesome-skills/apify-ecommerce 2>/dev/null (no --json) for the clean schema; only dig into .taggedBuilds.latest.build.inputSchema if you specifically need it as JSON.
  • The Primary actor has no maxResults field. Its caps are mode-specific (maxProductResults, maxReviewResults, maxSellerResults, maxSearchEngineResults, maxDeliveryResults). Setting maxResults does nothing and the run scrapes unbounded.
  • The Primary handles most intents via the right input mode (URLs → detailsUrls/listingUrls; query → keyword or searchEngineKeyword), including competitor, search-intent, classifieds, automotive, real-estate, website-marketplace, and events. It genuinely can't do tech-stack, seo-audit, store-enrichment, product-matching, ads-intelligence, content-discovery (Pinterest), or tiktok-shop — route those straight to the Fallback.
  • apify actors call -i expects valid JSON on one line. For inputs with URL arrays or quotes, write a file and pass -i @input.json instead of inlining — shell quoting silently corrupts complex inputs.
  • datasets get-items always streams over the network. Save to a file once (> file.json), then run jq against the file — don't re-pipe the command into jq repeatedly or you re-download every time.
  • apify actors search --sort-by popularity ignores relevance. It returns the biggest-name scrapers regardless of your query (an "etsy" search surfaces Instagram/Google Maps Actors). For escalation discovery keep the default relevance sort and filter on stats.totalUsers/actorReviewRating instead.
  • marketplaces values are full domain slugs, e.g. ["www.amazon.com", "www.ebay.com"] — not "amazon" or display names. Delivery mode is even narrower: marketplacesDelivery only accepts ["www.doordash.com", "www.instacart.com"] (no UberEats — use the e-commerce/ubereats-reviews-scraper fallback for that). Always confirm accepted values from the --input schema's enum before guessing.

Plus de skills de apify

apify-influencer-brand-collabs
apify
Découvrez les partenariats entre marques et créateurs Instagram en enchaînant les Apify Actors. Utilisez lorsque l'utilisateur demande qui collabore avec une marque, quelles marques un créateur a faites payantes…
apify-actor-development
apify
Créez, déboguez et déployez des programmes cloud serverless pour le scraping web, l'automatisation et le traitement de données. Prend en charge les modèles JavaScript, TypeScript et Python avec les bibliothèques intégrées Crawlee, Playwright et Cheerio pour le crawling HTTP et basé sur navigateur. Inclut des tests locaux via apify run avec stockage isolé, validation de schéma pour les entrées/sorties, et déploiement sur la plateforme Apify via apify push. Nécessite l'authentification Apify CLI et les métadonnées obligatoires generatedBy dans .actor/actor.json pour l'IA...
apify-actorization
apify
Convertissez des projets existants en Apify Actors serverless avec intégration SDK spécifique au langage. Prend en charge JavaScript/TypeScript (avec Actor.init() / Actor.exit()), Python (gestionnaire de contexte asynchrone) et tout langage via un wrapper CLI. Fournit un flux de travail structuré : apify init pour générer la structure, appliquer le wrapping SDK, configurer les schémas d'entrée/sortie, tester localement avec apify run, puis déployer avec apify push. Inclut la validation des schémas d'entrée et de sortie, la conteneurisation Docker et une option de paiement par événement...
apify-content-analytics
apify
Analytique de contenu multiplateforme via les Acteurs Apify pour Instagram, Facebook, YouTube et TikTok. Prend en charge plus de 17 Acteurs spécialisés couvrant les publications, reels, stories, commentaires, hashtags, abonnés et publicités sur les quatre plateformes. Récupère dynamiquement les schémas des Acteurs à l'aide de l'interface CLI mcpc pour déterminer les entrées requises et les champs de sortie disponibles. Produit les résultats en trois formats : affichage rapide dans le chat, export CSV ou export JSON avec des nombres de résultats personnalisables. Nécessite un jeton Apify dans le fichier .env et Node.js 20.6+...
apify-ecommerce
apify
Extrayez les données produits, les prix, les avis et les informations vendeurs de plus de 50 places de marché e-commerce. Trois modes de workflow : Produits & Tarification (suivi des prix, analyse concurrentielle), Avis Clients (analyse des sentiments, problèmes de qualité) et Renseignement Vendeurs (découverte de fournisseurs via Google Shopping). Prend en charge Amazon (plus de 20 régions), Walmart, eBay, IKEA, Costco et les détaillants européens ; saisie via URL de produits, URL de catégories ou recherche par mot-clé. Analyse optionnelle basée sur l'IA génère des insights sur les prix...
apify-generate-output-schema
apify
Générer des schémas de sortie (dataset_schema.json, output_schema.json, key_value_store_schema.json) pour un acteur Apify en analysant son code source. Utiliser lorsque…
apify-influencer-discovery
apify
Découvrez et évaluez des influenceurs sur Instagram, Facebook, YouTube et TikTok à l'aide des Apify Actors. Achemine les demandes de découverte vers plus de 15 Actors spécialisés couvrant le scraping de profils, la recherche par hashtag, l'analyse d'engagement et la découverte de niches sur toutes les grandes plateformes. Récupère dynamiquement les schémas des Actors via mcpc pour déterminer les entrées requises et les champs de sortie disponibles avant l'exécution. Prend en charge trois modes d'exportation : affichage en ligne dans le chat, sortie en fichier CSV ou JSON, avec des nombres de résultats personnalisables...
apify-ultimate-scraper
apify
Grattoir web automatisé sélectionnant les meilleurs Actors pour plus de 55 plateformes, dont Instagram, TikTok, YouTube, Facebook, Google Maps et autres. Couvre plus de 55 Actors préconfigurés sur 8 plateformes majeures avec des conseils de sélection spécifiques aux cas d'usage (génération de leads, découverte d'influenceurs, surveillance de marque, analyse concurrentielle, recherche de tendances). Prend en charge trois formats de sortie : affichage rapide dans le chat, export CSV ou export JSON avec limites de résultats personnalisables. Inclut des schémas de workflow multi-Actors pour des tâches complexes...