firecrawl-knowledge-ingest

bởi firecrawl

Thu thập các cơ sở kiến thức và cổng tài liệu công khai hoặc có xác thực bằng trình duyệt Firecrawl. Sử dụng cho tài liệu nặng về JS, cổng yêu cầu đăng nhập, trung tâm trợ giúp phân trang, cơ sở kiến thức hỗ trợ hoặc trích xuất JSON/markdown có cấu trúc từ các trang tài liệu.

npx skills add https://github.com/firecrawl/firecrawl-workflows --skill firecrawl-knowledge-ingest

Tải ZIP GitHub

107

Firecrawl Knowledge Ingest

Use this when a docs portal needs browser navigation, auth, pagination, or JS rendering.

Onboarding Interview

Infer the portal URL, output format, auth needs, and page limit from context. If the portal is clear, proceed immediately.

Ask at most 1-3 concise questions only if blocked, such as the portal URL, whether authentication is required, or the desired output format.

Firecrawl Collection Plan

Use Firecrawl browser to:

open the portal and inspect navigation
identify sections, categories, sidebar links, and article URLs
follow sidebar navigation, next links, pagination, load-more controls, or search
scrape article content as markdown
extract metadata such as title, section, last updated date, author, and tags

Try Firecrawl map as a supplement for public URLs, but use browser navigation for auth-gated or JS-heavy content.

Final Deliverable

# Knowledge Ingest: [Portal]

## Summary
[Pages extracted, sections covered, limitations]

## Output
[JSON/markdown/merged file path or content]

## Sections
[Section names and article counts]

## Failed Or Restricted Pages
[Any access/loading issues]

## Sources
[URLs extracted]

## Rerun Inputs
workflow: firecrawl-knowledge-ingest
url: [portal url]
format: [json/markdown/merged]
max_pages: [number]