firecrawl-scrape

Ekstrak markdown bersih dari URL mana pun, termasuk SPA yang dirender JavaScript. Gunakan skill ini setiap kali pengguna memberikan URL dan menginginkan kontennya, mengatakan "scrape",…

npx skills add https://github.com/firecrawl/firecrawl-codex-plugin --skill firecrawl-scrape

firecrawl scrape

Scrape one or more URLs. Returns clean, LLM-optimized markdown. Multiple URLs are scraped concurrently.

When to use

  • You have a specific URL and want its content
  • The page is static or JS-rendered (SPA)
  • Step 2 in the workflow escalation pattern: search → scrape → map → crawl → interact

Quick start

# Basic markdown extraction
firecrawl scrape "<url>" -o .firecrawl/page.md

# Main content only, no nav/footer
firecrawl scrape "<url>" --only-main-content -o .firecrawl/page.md

# Wait for JS to render, then scrape
firecrawl scrape "<url>" --wait-for 3000 -o .firecrawl/page.md

# Multiple URLs (each saved to .firecrawl/)
firecrawl scrape https://example.com https://example.com/blog https://example.com/docs

# Get markdown and links together
firecrawl scrape "<url>" --format markdown,links -o .firecrawl/page.json

# Ask a question about the page
firecrawl scrape "https://example.com/pricing" --query "What is the enterprise plan price?"

Options

OptionDescription
-f, --format <formats>Output formats: markdown, html, rawHtml, links, screenshot, json
-Q, --query <prompt>Ask a question about the page content (5 credits)
-HInclude HTTP headers in output
--only-main-contentStrip nav, footer, sidebar — main content only
--wait-for <ms>Wait for JS rendering before scraping
--include-tags <tags>Only include these HTML tags
--exclude-tags <tags>Exclude these HTML tags
-o, --output <path>Output file path

Tips

  • Prefer plain scrape over --query. Scrape to a file, then use grep, head, or read the markdown directly — you can search and reason over the full content yourself. Use --query only when you want a single targeted answer without saving the page (costs 5 extra credits).
  • Try scrape before interact. Scrape handles static pages and JS-rendered SPAs. Only escalate to interact when you need interaction (clicks, form fills, pagination).
  • Multiple URLs are scraped concurrently — check firecrawl --status for your concurrency limit.
  • Single format outputs raw content. Multiple formats (e.g., --format markdown,links) output JSON.
  • Always quote URLs — shell interprets ? and & as special characters.
  • Naming convention: .firecrawl/{site}-{path}.md

See also

Lebih banyak skill dari firecrawl

firecrawl-research-index
firecrawl
Temukan makalah yang menjawab pertanyaan riset dengan Firecrawl Research, menggunakan pencarian semantik, ekspansi semantik dan struktural, serta verifikasi dalam tubuh. Selalu gunakan keterampilan ini untuk tugas pencarian literatur atau pengambilan makalah — baik pencarian satu makalah maupun kumpulan multi-makalah.
data-analysisresearchweb-scraping
oracle
firecrawl
Praktik terbaik dalam menggunakan CLI oracle (penggabungan prompt dan file, mesin, sesi, dan pola lampiran file).
pinecone
firecrawl
Basis data vektor terkelola untuk aplikasi AI produksi. Sepenuhnya terkelola, penskalaan otomatis, dengan pencarian hibrida (padat + jarang), pemfilteran metadata, dan ruang nama.…
wpds
firecrawl
Gunakan saat membangun antarmuka pengguna yang memanfaatkan WordPress Design System (WPDS) dan komponen, token, pola, dll.
audiocraft-audio-generation
firecrawl
Perpustakaan PyTorch untuk pembuatan audio termasuk teks-ke-musik (MusicGen) dan teks-ke-suara (AudioGen). Gunakan saat Anda perlu menghasilkan musik dari teks…
skypilot-multi-cloud-orchestration
firecrawl
Orkestrasi multi-cloud untuk beban kerja ML dengan optimasi biaya otomatis. Gunakan saat Anda perlu menjalankan pelatihan atau pekerjaan batch di berbagai cloud, memanfaatkan…
firecrawl-seo-audit
firecrawl
Audit SEO situs web dengan Firecrawl. Gunakan saat pengguna meminta audit SEO, tinjauan metadata dan heading, analisis peta situs/struktur situs, peluang kata kunci, perbandingan SERP pesaing, atau rekomendasi optimasi pencarian yang diprioritaskan.
data-analysisresearchweb-scraping
gh-issues
firecrawl
Ambil isu GitHub, hasilkan sub-agen untuk menerapkan perbaikan dan membuka PR, lalu pantau serta tangani komentar tinjauan PR. Penggunaan: /gh-issues [owner/repo] [--label…