apify-ecommerce

作者: apify

抓取超過30個平台(包括Amazon、Walmart、eBay、Shopify、WooCommerce等)的電子商務資料,涵蓋價格、評論、暢銷商品及賣家探索。

npx skills add https://github.com/apify/awesome-skills --skill apify-ecommerce

E-Commerce Cluster

Answer natural language e-commerce questions by routing to the right Apify Actor and delivering a synthesized answer via the apify CLI.

CLI rules: Always pass --user-agent apify-awesome-skills/apify-ecommerce, --json (or the relevant --format flag on datasets get-items), and 2>/dev/null. The --user-agent flag is critical for telemetry — never omit it.

Prerequisites

(No need to check it upfront)

  • Apify CLI v1.5.0+ (npm install -g apify-cli)
  • jq (recommended for quick extraction and filtering; brew install jq on macOS, apt install jq on Linux)
  • Authentication via one of:

Verify auth: apify info --user-agent apify-awesome-skills/apify-ecommerce — should show username and userId.

Workflow

Copy this checklist and track progress:

Task Progress:
- [ ] Step 1: Detect intent and select Actor
- [ ] Step 2: Fetch Actor schema
- [ ] Step 3: Ask user preferences (format, result count)
- [ ] Step 4: Run the Actor and fetch results
- [ ] Step 5: Analyze results and deliver synthesized answer

Step 1: Detect Intent and Select Actor

Classify the user's message into an intent, then pick the right Actor.

Intent signals:

Signals in user messageIntent
price, cost, cheapest, compare prices, pricingpricing
review, rating, sentiment, stars, feedbackreviews
bestseller, top selling, most popular, trendingbestsellers
seller, vendor, reseller, who sellssellers
all products from, scrape store, full catalogstore-scrape
what platform, built on, tech stack, Shopify or WooCommercetech-stack
SEO, listing quality, product page auditseo-audit
competitor funnel, competitor pricing, conversion elementscompetitor
search intent, keyword intent, SERP intentsearch-intent
match products, same product on different platformsproduct-matching
restaurant, food delivery, DoorDash, UberEats, TheForkfood-delivery
enrich store, store metadata, store liststore-enrichment
event, concert, ticket, Eventbriteevents
property, real estate, house listing, Realtorreal-estate
Facebook ads, Meta ads, ad library, competitor adsads-intelligence
classified, Craigslist, used item for saleclassifieds
car, used car, vehicle, automotive, Webmotorsautomotive
pins, inspiration, Pinterest boards, visual search, Pinterest trendscontent-discovery
TikTok Shop, TikTok store, TikTok creatortiktok-shop
website for sale, domain for sale, Flippawebsite-marketplace

If multiple intents are detected, ask: "Do you want [intent A] or [intent B]?"

Actor routing table — always try Primary first, switch to Fallback only if it fails or returns 0 results. The Primary actor (apify/e-commerce-scraping-tool) handles most intents once you feed the right input mode:

  • Have target URLs (a listing, profile, or category page) → detailsUrls / listingUrls.
  • Have a keyword/marketplace → keyword + marketplaces (product-details mode).
  • Broad discovery (competitor, search-intent, classifieds, automotive, real-estate, website-marketplace, events) → use searchEngineKeyword (search-engine mode) or keyword/detailsUrls depending on whether you have a query or URLs.

Exception — skip the Primary and go straight to the Fallback for intents the Primary genuinely can't do (different data source or specialized analysis): tech-stack, seo-audit, store-enrichment, product-matching, ads-intelligence, content-discovery (Pinterest), and tiktok-shop. Routing these to the Primary wastes a run and credits.

IntentPlatformPrimary ActorFallback Actor
pricingAmazon / Walmart / genericapify/e-commerce-scraping-tool—
pricingeBayapify/e-commerce-scraping-toolivanvs/ebay-scraper-pay-per-result
pricingEtsyapify/e-commerce-scraping-toolepctex/etsy-scraper
pricingGoogle Shoppingapify/e-commerce-scraping-toolepctex/google-shopping-scraper
pricingFacebook Marketplaceapify/e-commerce-scraping-toolapify/facebook-marketplace-scraper
pricingSHEINapify/e-commerce-scraping-toolseamless_coffer/shein-product-scraper
pricingLazadaapify/e-commerce-scraping-toolfatihtahta/lazada-scraper
pricingCanadian Tireapify/e-commerce-scraping-toolazzouzana/canadiantire-ca-scraper
pricingTescoapify/e-commerce-scraping-toolradeance/tesco-scraper
pricingShopifyapify/e-commerce-scraping-tooltrovevault/shopify-products-scraper
pricingWooCommerceapify/e-commerce-scraping-tooltrovevault/woocommerce-products-scraper
reviewsAmazon / Walmart / genericapify/e-commerce-scraping-tooljunglee/amazon-reviews-scraper
reviewsTrustpilotapify/e-commerce-scraping-toolcasper11515/trustpilot-reviews-scraper
reviewsTheForkapify/e-commerce-scraping-tooljdtpnjtp/thefork-restaurant-scraper-advanced
bestsellersAmazonapify/e-commerce-scraping-tooljunglee/amazon-bestsellers
sellersAmazonapify/e-commerce-scraping-tooljunglee/amazon-seller-scraper
sellerseBayapify/e-commerce-scraping-toolivanvs/ebay-scraper-pay-per-result
store-scrapeShopifyapify/e-commerce-scraping-tooltrovevault/shopify-products-scraper
store-scrapeWooCommerceapify/e-commerce-scraping-tooltrovevault/woocommerce-products-scraper
store-scrapeAmazonapify/e-commerce-scraping-tooljunglee/Amazon-crawler
store-scrapeFlippaapify/e-commerce-scraping-toolscraped/flippa-scraper
tech-stackanyapify/e-commerce-scraping-tooltrovevault/e-commerce-tech-stack-detector
seo-auditanyapify/e-commerce-scraping-tooltrovevault/product-listing-seo-auditor
competitoranyapify/e-commerce-scraping-tooltrovevault/competitor-intelligence-scraper---funnel-pricing-conversion
search-intentanyapify/e-commerce-scraping-tooltrovevault/ai-serp-intent-extractor---search-intent-classifier
product-matchinganyapify/e-commerce-scraping-tool—
store-enrichmentanyapify/e-commerce-scraping-tooltrovevault/e-commerce-store-data-enricher
food-deliveryDoorDashapify/e-commerce-scraping-tooltri_angle/doordash-store-details-scraper
food-deliveryUberEatsapify/e-commerce-scraping-toole-commerce/ubereats-reviews-scraper
food-deliveryTheForkapify/e-commerce-scraping-tooljdtpnjtp/thefork-restaurant-scraper-advanced
ads-intelligenceFacebook / Metaapify/e-commerce-scraping-toolapify/facebook-ads-scraper
classifiedsCraigslistapify/e-commerce-scraping-toolivanvs/craigslist-scraper-pay-per-result
automotiveWebmotorsapify/e-commerce-scraping-toolstealth_mode/webmotors-auto-search-scraper
eventsEventbriteapify/e-commerce-scraping-toolaitorsm/eventbrite
real-estateRealtor.comapify/e-commerce-scraping-toolpowerai/realtor-properties-search-scraper
content-discoveryPinterestapify/e-commerce-scraping-toolfatihtahta/pinterest-scraper-search
tiktok-shopTikTok Shopapify/e-commerce-scraping-toollemur/tiktok-shop-creators
website-marketplaceFlippaapify/e-commerce-scraping-toolscraped/flippa-scraper

Escalation — if both Primary and Fallback fail or return 0 results, discover a current alternative live instead of guessing an ID:

# Find relevant, well-rated, pay-per-event Actors for the platform/intent.
# Keep the default relevance sort — `--sort-by popularity` surfaces generic
# big-name scrapers over the platform you actually asked for.
apify actors search "PLATFORM or INTENT keywords" \
  --pricing-model PAY_PER_EVENT --limit 10 --json \
  --user-agent apify-awesome-skills/apify-ecommerce 2>/dev/null \
  | jq '[.items[]
      | select(.stats.totalUsers > 100 and .actorReviewRating > 4.5)
      | {id: (.username + "/" + .name), users: .stats.totalUsers,
         rating: (.actorReviewRating | (. * 100 | round / 100)),
         pricing: .currentPricingInfo.pricingModel}]'

Pick the top match. Before running it, confirm it requests only limited permissions (check the Actor's Store page / README — prefer Actors that don't require full account access). If the PAY_PER_EVENT filter returns nothing, drop the --pricing-model flag and re-run, keeping the ≥100-users and ≥4.5-rating bar.

Step 2: Fetch Actor Schema

Fetch the Actor summary, input schema, and README:

# Summary (title, description, pricing, stats)
apify actors info "ACTOR_ID" --user-agent apify-awesome-skills/apify-ecommerce --json 2>/dev/null

# Input schema — use --input WITHOUT --json to get the clean schema directly.
# (Adding --json returns the full ~250 KB actor object instead, with the schema
#  buried as an escaped string under .taggedBuilds.latest.build.inputSchema.)
apify actors info "ACTOR_ID" --user-agent apify-awesome-skills/apify-ecommerce --input 2>/dev/null

# README (capabilities, examples, gotchas)
apify actors info "ACTOR_ID" --user-agent apify-awesome-skills/apify-ecommerce --readme 2>/dev/null

Replace ACTOR_ID with the selected Actor (e.g., apify/e-commerce-scraping-tool).

Primary actor input cheat-sheet. apify/e-commerce-scraping-tool is mode-driven — pick fields by intent (always set the matching max…Results cap):

IntentMinimal input
pricing (keyword){"keyword": "wireless earbuds", "marketplaces": ["www.amazon.com"], "maxProductResults": 50}
pricing (specific URLs){"detailsUrls": [{"url": "https://…"}], "maxProductResults": 50}
store-scrape (category){"listingUrls": [{"url": "https://…/category"}], "maxProductResults": 500}
reviews{"keywordReviews": "echo dot", "marketplacesReviews": ["www.amazon.com"], "sortReview": "Most recent", "maxReviewResults": 200}
sellers{"sellerUrls": [{"url": "https://…"}], "maxSellerResults": 50}
pricing (Google Shopping){"searchEngineKeyword": "ps5", "countryCode": "us", "maxSearchEngineResults": 50}
food-delivery{"keywordDelivery": "pizza", "marketplacesDelivery": ["www.doordash.com"], "addressDelivery": "New York, NY", "maxDeliveryResults": 50}

For any other actor (or fields not listed), fetch the schema with the --input command above.

Step 3: Ask User Preferences

Before running, ask:

  1. Output format:
    • Quick answer (default) — synthesized answer in chat, no file saved
    • CSV — full export saved to disk
    • JSON — full export saved to disk
  2. Result count — suggest defaults by intent:
IntentDefault
pricing50 products
reviews200 reviews
bestsellers100 items
sellers50 sellers
store-scrapeall (unlimited)
food-delivery50 restaurants
all others20–50

Cost safety: Always set a sensible result limit in the Actor input. For the Primary actor the cap field is mode-specific — maxProductResults, maxReviewResults, maxSellerResults, maxSearchEngineResults, or maxDeliveryResults (there is no single maxResults). For Fallback actors, use whatever the schema exposes (maxResults, resultsLimit, maxItems, maxCrawledPages, etc.). Default to the per-intent values above unless the user explicitly asks for more. Warn the user before running large scrapes (1000+ results) as they consume more Apify credits.

Step 4: Run the Actor and Fetch Results

Two steps: run the Actor (blocks until done), then fetch dataset items in the requested format.

Run the Actor — returns run metadata as JSON; extract defaultDatasetId for the next step:

apify actors call "ACTOR_ID" -i 'JSON_INPUT' \
  --user-agent apify-awesome-skills/apify-ecommerce --json 2>/dev/null

From the output use .id (run ID), .status (should be SUCCEEDED), and .defaultDatasetId.

Fetch results — pick the variant based on the user's preference:

# Quick answer: total count + fields + top 5 in chat (no file)
apify datasets info DATASET_ID --json \
  --user-agent apify-awesome-skills/apify-ecommerce 2>/dev/null \
  | jq '{itemCount, fields, consoleUrl}'
apify datasets get-items DATASET_ID --limit 5 \
  --user-agent apify-awesome-skills/apify-ecommerce --format json 2>/dev/null

# CSV file
apify datasets get-items DATASET_ID \
  --user-agent apify-awesome-skills/apify-ecommerce --format csv 2>/dev/null > YYYY-MM-DD_OUTPUT_FILE.csv

# JSON file
apify datasets get-items DATASET_ID \
  --user-agent apify-awesome-skills/apify-ecommerce --format json 2>/dev/null > YYYY-MM-DD_OUTPUT_FILE.json

Other --format options: jsonl, xlsx, xml, rss, html. Use --offset N to paginate large datasets.

Tip: for anything more than a quick peek, save the dataset to a local file first (with > file.json / > file.csv) and run further analysis from disk. apify datasets get-items always streams over the network, so piping it straight into jq re-downloads the whole thing every iteration.

Combining with jq for quick extraction:

Treat jq as a complement to apify datasets get-items, not a replacement: server-side --limit / --offset / --format keeps cost and bandwidth down. Use jq on a sample item or on a file you already saved.

# Discover real field names from one sample item (Actor outputs vary —
# use this before composing further jq queries)
apify datasets get-items DATASET_ID --limit 1 --format json \
  --user-agent apify-awesome-skills/apify-ecommerce 2>/dev/null \
  | jq '.[0]'

# Quick aggregation from a JSON file you already saved with the commands above
jq '[.[] | select(.rating != null and .rating >= 4.5)] | length' YYYY-MM-DD_OUTPUT_FILE.json

Step 5: Analyze Results and Deliver Answer

After the run completes, deliver a direct synthesized answer — not a data dump:

  • Pricing: price range, average, top 5 cheapest with URLs
  • Reviews: average rating, top 3 positive and negative themes, recent snippets
  • Bestsellers: top 10 by rank with name, price, rating, URL
  • Sellers: total sellers, price range per seller, unauthorized seller flags
  • Store-scrape: total products, category breakdown, price range, stock summary
  • Tech-stack: platform detected, confidence level, notable plugins
  • Food delivery: restaurant count, average rating, price tier breakdown
  • Ads intelligence: total ads, active/inactive split, top creative formats

Error Handling

  • Auth error → run apify login, or set APIFY_TOKEN env var
  • Actor not found → check Actor ID spelling in the routing table
  • Run status FAILED → open the console URL (.consoleUrl from run metadata) for logs
  • Timeout / very long run → pass --timeout <seconds> to apify actors call
  • No results → broaden the keyword, switch to the Fallback Actor, then use the Escalation discovery command (under Step 1) if both fail
  • proxy is required → add "proxy": {"useApifyProxy": true} to the Actor input
  • Platform not detected → default to apify/e-commerce-scraping-tool with generic intent

Gotchas

  • --input --json is a trap. It returns the full ~250 KB actor object, not the schema. Use apify actors info ID --input --user-agent apify-awesome-skills/apify-ecommerce 2>/dev/null (no --json) for the clean schema; only dig into .taggedBuilds.latest.build.inputSchema if you specifically need it as JSON.
  • The Primary actor has no maxResults field. Its caps are mode-specific (maxProductResults, maxReviewResults, maxSellerResults, maxSearchEngineResults, maxDeliveryResults). Setting maxResults does nothing and the run scrapes unbounded.
  • The Primary handles most intents via the right input mode (URLs → detailsUrls/listingUrls; query → keyword or searchEngineKeyword), including competitor, search-intent, classifieds, automotive, real-estate, website-marketplace, and events. It genuinely can't do tech-stack, seo-audit, store-enrichment, product-matching, ads-intelligence, content-discovery (Pinterest), or tiktok-shop — route those straight to the Fallback.
  • apify actors call -i expects valid JSON on one line. For inputs with URL arrays or quotes, write a file and pass -i @input.json instead of inlining — shell quoting silently corrupts complex inputs.
  • datasets get-items always streams over the network. Save to a file once (> file.json), then run jq against the file — don't re-pipe the command into jq repeatedly or you re-download every time.
  • apify actors search --sort-by popularity ignores relevance. It returns the biggest-name scrapers regardless of your query (an "etsy" search surfaces Instagram/Google Maps Actors). For escalation discovery keep the default relevance sort and filter on stats.totalUsers/actorReviewRating instead.
  • marketplaces values are full domain slugs, e.g. ["www.amazon.com", "www.ebay.com"] — not "amazon" or display names. Delivery mode is even narrower: marketplacesDelivery only accepts ["www.doordash.com", "www.instacart.com"] (no UberEats — use the e-commerce/ubereats-reviews-scraper fallback for that). Always confirm accepted values from the --input schema's enum before guessing.

來自 apify 的更多技能

apify-influencer-brand-collabs
apify
探索Instagram品牌與創作者的合作關係,透過串聯Apify Actors。當使用者詢問某品牌與誰合作、某創作者曾與哪些品牌進行付費合作時使用…
apify-actor-development
apify
建立、除錯及部署無伺服器雲端程式,用於網頁爬取、自動化及資料處理。支援 JavaScript、TypeScript 及 Python 範本,內建 Crawlee、Playwright 與 Cheerio 函式庫,適用於 HTTP 及瀏覽器爬取。包含透過 apify run 進行本地測試(具備隔離儲存)、輸入/輸出結構驗證,以及透過 apify push 部署至 Apify 平台。需進行 Apify CLI 驗證,並在 .actor/actor.json 中強制加入 generatedBy 元資料以供 AI 使用...
apify-actorization
apify
將現有專案轉換為無伺服器 Apify Actors,並整合語言專屬 SDK。支援 JavaScript/TypeScript(使用 Actor.init() / Actor.exit())、Python(非同步上下文管理器),以及透過 CLI 包裝器的任何語言。提供結構化工作流程:使用 apify init 建立專案骨架、套用 SDK 包裝、設定輸入/輸出架構、以 apify run 進行本地測試,再透過 apify push 部署。包含輸入與輸出架構驗證、Docker 容器化,以及可選的按事件付費...
apify-content-analytics
apify
透過 Apify Actors 進行多平台內容分析,支援 Instagram、Facebook、YouTube 及 TikTok。涵蓋 17 種以上專用 Actors,可處理貼文、Reels、限時動態、留言、Hashtag、粉絲及廣告等內容,並動態使用 mcpc CLI 擷取 Actor 架構,以判斷所需輸入與可用輸出欄位。結果提供三種格式:快速聊天顯示、CSV 匯出或 JSON 匯出,並可自訂結果數量。需在 .env 檔案中設定 Apify Token,並使用 Node.js 20.6+...
apify-ecommerce
apify
從50多個電子商務平台提取產品數據、價格、評論及賣家資訊。三種工作流程模式:產品與定價(價格追蹤、競爭對手分析)、客戶評論(情感分析、品質問題)及賣家情報(透過Google Shopping發現供應商)。支援Amazon(20多個地區)、Walmart、eBay、IKEA、Costco及歐洲零售商;可透過產品網址、分類網址或關鍵字搜尋輸入。可選AI驅動分析,生成價格洞察...
apify-generate-output-schema
apify
為 Apify Actor 分析其原始碼,生成輸出結構(dataset_schema.json、output_schema.json、key_value_store_schema.json)。用於…
apify-influencer-discovery
apify
使用Apify Actors在Instagram、Facebook、YouTube和TikTok上發現並評估網紅。將發現請求路由至15個以上專門的Actors,涵蓋所有主要平台的個人資料抓取、標籤搜尋、互動分析及利基發現。透過mcpc動態獲取Actor架構,以在執行前確定所需輸入與可用輸出欄位。支援三種匯出模式:內嵌聊天顯示、CSV或JSON檔案輸出,並可自訂結果數量...
apify-ultimate-scraper
apify
自動化網頁爬蟲,為55多個平台選擇最佳Actor,包括Instagram、TikTok、YouTube、Facebook、Google地圖等。涵蓋8大主要平台的55多個預配置Actor,並提供針對特定使用案例的選擇指引(潛在客戶開發、網紅發現、品牌監控、競爭對手分析、趨勢研究)。支援三種輸出格式:快速聊天顯示、CSV匯出或JSON匯出,並可自訂結果數量限制。包含多Actor工作流程模式,適用於複雜...