firecrawl-parse
bởi firecrawl
Trích xuất và chuyển đổi hiệu quả nội dung của bất kỳ tệp cục bộ nào—như PDF, DOCX, DOC, ODT, RTF, XLSX, XLS hoặc HTML—thành markdown sạch, được định dạng tốt và được lưu…
npx skills add https://github.com/firecrawl/firecrawl-cursor-plugin --skill firecrawl-parsefirecrawl parse
Turn a local document into clean markdown on disk. Supports PDF, DOCX, DOC, ODT, RTF, XLSX, XLS, HTML/HTM.
Quick start
Always save to .firecrawl/ with -o — parsed docs can be hundreds of KB and blow up context if streamed to stdout. Add .firecrawl/ to .gitignore.
mkdir -p .firecrawl
# File → markdown
firecrawl parse ./paper.pdf -o .firecrawl/paper.md
# AI summary
firecrawl parse ./paper.pdf -S -o .firecrawl/paper-summary.md
# Ask a question about the doc
firecrawl parse ./paper.pdf -Q "What are the main conclusions?" \
-o .firecrawl/paper-qa.md
Then read the output incrementally with head, grep, or rg.
Run firecrawl parse --help for the full option list.
Done when: the markdown, summary, or answer is written under .firecrawl/ and you have inspected it with bounded reads.
Tips
- Quote paths with spaces:
firecrawl parse "./My Doc.pdf" -o .firecrawl/mydoc.md. - Max upload size: 50 MB per file.
- Credits: ~1 per PDF page; HTML is 1 flat.
- Check
.firecrawl/before re-parsing the same file. - To check your credit balance (recommended for batch processing and similar workflows), use
firecrawl credit-usage(requires authentication).
See also
- firecrawl-scrape — same idea for URLs
- firecrawl-build-scrape — building document extraction into an app instead of running it here