firecrawl-parse

โดย firecrawl

สกัดและแปลงเนื้อหาของไฟล์ในเครื่องใดๆ เช่น PDF, DOCX, DOC, ODT, RTF, XLSX, XLS หรือ HTML อย่างมีประสิทธิภาพ ให้เป็นมาร์กดาวน์ที่สะอาดและจัดรูปแบบอย่างดี บันทึกไว้…

npx skills add https://github.com/firecrawl/firecrawl-codex-plugin --skill firecrawl-parse

firecrawl parse

Turn a local document into clean markdown on disk. Supports PDF, DOCX, DOC, ODT, RTF, XLSX, XLS, HTML/HTM/XHTML.

When to use

  • You have a file on disk (not a URL) and want its text as markdown
  • User drops a PDF/DOCX and asks what it says, or to summarize it
  • Use scrape instead when the source is a URL

Quick start

Always save to .firecrawl/ with -o — parsed docs can be hundreds of KB and blow up context if streamed to stdout. Add .firecrawl/ to .gitignore.

mkdir -p .firecrawl

# File → markdown
firecrawl parse ./paper.pdf -o .firecrawl/paper.md

# AI summary
firecrawl parse ./paper.pdf -S -o .firecrawl/paper-summary.md

# Ask a question about the doc
firecrawl parse ./paper.pdf -Q "What are the main conclusions?" \
  -o .firecrawl/paper-qa.md

Then head, grep, rg etc., or incrementally read the file - don't load the whole thing at once.

Options

OptionDescription
-S, --summaryAI-generated summary
-Q, --query <prompt>Ask a question about the parsed content
-o, --output <path>Output file path — always use this
-f, --format <fmt>markdown (default), html, summary
--timeout <ms>Timeout for the parse job
--timingShow request duration

Tips

  • Quote paths with spaces: firecrawl parse "./My Doc.pdf" -o .firecrawl/mydoc.md.
  • Max upload size: 50 MB per file.
  • Credits: ~1 per PDF page; HTML is 1 flat.
  • Check .firecrawl/ before re-parsing the same file.
  • To check your credit balance (recommended for batch processing and similar workflows), use the firecrawl credit-usage command.

See also

Skills เพิ่มเติมจาก firecrawl

firecrawl-research-index
firecrawl
ค้นหาเอกสารที่ตอบคำถามการวิจัยด้วย Firecrawl Research โดยใช้การค้นหาเชิงความหมาย การขยายผลเชิงความหมายและโครงสร้าง และการตรวจสอบภายในเนื้อหา ใช้ทักษะนี้เสมอสำหรับงานค้นหาวรรณกรรมหรือดึงเอกสาร ไม่ว่าจะเป็นการค้นหาเอกสารเดี่ยวหรือชุดเอกสารหลายชิ้น
data-analysisresearchweb-scraping
oracle
firecrawl
แนวทางปฏิบัติที่ดีที่สุดสำหรับการใช้ oracle CLI (การรวม prompt และไฟล์, เอ็นจิน, เซสชัน, และรูปแบบการแนบไฟล์)
pinecone
firecrawl
ฐานข้อมูลเวกเตอร์ที่จัดการแล้วสำหรับแอปพลิเคชัน AI ในระบบผลิต จัดการเต็มรูปแบบ ปรับขนาดอัตโนมัติ พร้อมการค้นหาแบบไฮบริด (dense + sparse) การกรองเมตาดาต้า และเนมสเปซ…
wpds
firecrawl
ใช้เมื่อสร้าง UI ที่ใช้ประโยชน์จาก WordPress Design System (WPDS) และส่วนประกอบ โทเค็น รูปแบบ ฯลฯ
audiocraft-audio-generation
firecrawl
ไลบรารี PyTorch สำหรับการสร้างเสียง รวมถึงการแปลงข้อความเป็นเพลง (MusicGen) และการแปลงข้อความเป็นเสียง (AudioGen) ใช้เมื่อคุณต้องการสร้างเพลงจากข้อความ…
skypilot-multi-cloud-orchestration
firecrawl
การจัดระเบียบการทำงานข้ามคลาวด์สำหรับภาระงาน ML พร้อมการปรับต้นทุนอัตโนมัติ ใช้เมื่อคุณต้องการรันงานฝึกอบรมหรืองานแบตช์ข้ามคลาวด์หลายแห่ง ใช้ประโยชน์จาก…
firecrawl-seo-audit
firecrawl
ตรวจสอบ SEO ของเว็บไซต์ด้วย Firecrawl ใช้เมื่อผู้ใช้ขอการตรวจสอบ SEO การตรวจสอบข้อมูลเมตาและหัวข้อ การวิเคราะห์แผนผังเว็บไซต์/โครงสร้างเว็บไซต์ โอกาสของคำสำคัญ การเปรียบเทียบ SERP ของคู่แข่ง หรือคำแนะนำการปรับแต่งการค้นหาที่จัดลำดับความสำคัญ
data-analysisresearchweb-scraping
gh-issues
firecrawl
ดึงข้อมูล Issue จาก GitHub สร้างเอเยนต์ย่อยเพื่อดำเนินการแก้ไขและเปิด Pull Request จากนั้นติดตามและจัดการกับความคิดเห็นในการตรวจสอบ PR การใช้งาน: /gh-issues [owner/repo] [--label…