firecrawl-knowledge-ingest

โดย firecrawl

ใช้ Firecrawl browser เพื่อดึงข้อมูลจากฐานความรู้และพอร์ทัลเอกสารที่เปิดเผยต่อสาธารณะหรือต้องยืนยันตัวตน ใช้สำหรับเอกสารที่ใช้ JavaScript หนัก พอร์ทัลที่ต้องล็อกอิน ศูนย์ช่วยเหลือแบบแบ่งหน้า ฐานความรู้สนับสนุน หรือการดึงข้อมูล JSON/markdown ที่มีโครงสร้างจากเว็บไซต์เอกสาร

npx skills add https://github.com/firecrawl/firecrawl-workflows --skill firecrawl-knowledge-ingest

Firecrawl Knowledge Ingest

Use this when a docs portal needs browser navigation, auth, pagination, or JS rendering.

Onboarding Interview

Infer the portal URL, output format, auth needs, and page limit from context. If the portal is clear, proceed immediately.

Ask at most 1-3 concise questions only if blocked, such as the portal URL, whether authentication is required, or the desired output format.

Firecrawl Collection Plan

Use Firecrawl browser to:

  • open the portal and inspect navigation
  • identify sections, categories, sidebar links, and article URLs
  • follow sidebar navigation, next links, pagination, load-more controls, or search
  • scrape article content as markdown
  • extract metadata such as title, section, last updated date, author, and tags

Try Firecrawl map as a supplement for public URLs, but use browser navigation for auth-gated or JS-heavy content.

Final Deliverable

# Knowledge Ingest: [Portal]

## Summary
[Pages extracted, sections covered, limitations]

## Output
[JSON/markdown/merged file path or content]

## Sections
[Section names and article counts]

## Failed Or Restricted Pages
[Any access/loading issues]

## Sources
[URLs extracted]

## Rerun Inputs
workflow: firecrawl-knowledge-ingest
url: [portal url]
format: [json/markdown/merged]
max_pages: [number]

JSON Shape

Use source, url, extractedAt, totalArticles, and sections[] with article title, url, section, content, and metadata.

Quality Bar

  • Preserve code examples, tables, and formatting.
  • Strip nav chrome, headers, and footers.
  • Track extraction progress and page failures.
  • Respect authentication boundaries.

Skills เพิ่มเติมจาก firecrawl

oracle
firecrawl
แนวทางปฏิบัติที่ดีที่สุดสำหรับการใช้ oracle CLI (การรวม prompt และไฟล์, เอ็นจิน, เซสชัน, และรูปแบบการแนบไฟล์)
official
firecrawl-demo-walkthrough
firecrawl
ดำเนินการตามขั้นตอนหลักของผลิตภัณฑ์ด้วยเบราว์เซอร์ Firecrawl และสร้างคู่มือการใช้งาน UX/ผลิตภัณฑ์ที่มีโครงสร้าง ใช้สำหรับการสมัครสมาชิก การเริ่มต้นใช้งาน ราคา เอกสาร แดชบอร์ด การเตรียมสาธิตผลิตภัณฑ์ การวิเคราะห์ UX และการวิเคราะห์ประสบการณ์การใช้งานครั้งแรก
browser-automationofficialresearch
sag
firecrawl
ElevenLabs แปลงข้อความเป็นเสียง พร้อมประสบการณ์การใช้งานแบบ say สไตล์ Mac
official
pinecone
firecrawl
ฐานข้อมูลเวกเตอร์ที่จัดการแล้วสำหรับแอปพลิเคชัน AI ในระบบผลิต จัดการเต็มรูปแบบ ปรับขนาดอัตโนมัติ พร้อมการค้นหาแบบไฮบริด (dense + sparse) การกรองเมตาดาต้า และเนมสเปซ…
official
sentence-transformers
firecrawl
เฟรมเวิร์กสำหรับการฝังประโยค ข้อความ และรูปภาพที่ทันสมัยที่สุด มีโมเดลที่ผ่านการฝึกอบรมล่วงหน้ามากกว่า 5000 โมเดลสำหรับความคล้ายคลึงทางความหมาย การจัดกลุ่ม และการดึงข้อมูล
official
model-merging
firecrawl
รวมโมเดลที่ปรับแต่งหลายตัวโดยใช้ mergekit เพื่อรวมความสามารถโดยไม่ต้องฝึกซ้ำ ใช้เมื่อสร้างโมเดลเฉพาะทางโดยการผสมผสานโมเดลที่เชี่ยวชาญเฉพาะด้าน...
official
deepspeed
firecrawl
คำแนะนำจากผู้เชี่ยวชาญสำหรับการฝึกอบรมแบบกระจายด้วย DeepSpeed - ขั้นตอนการปรับแต่ง ZeRO, การทำงานแบบขนานของไปป์ไลน์, FP16/BF16/FP8, 1-bit Adam, การสนใจแบบกระจัดกระจาย
official
discord
firecrawl
การดำเนินการ Discord ผ่านเครื่องมือข้อความ (channel=discord)
official