structured-extraction

작성자: firecrawl

웹사이트에서 JSON 스키마와 일치하는 구조화된 데이터를 추출합니다. 복잡한 중첩 스키마, 배열, 페이지 매김 및 검증을 처리합니다. 항상 다음을 통해 출력합니다…

npx skills add https://github.com/firecrawl/web-agent --skill structured-extraction

Structured Extraction

Use this skill when extracting data that must match a specific JSON schema.

Strategy by task type

Simple query (single fact or small object)

  1. Search for relevant results.
  2. Scrape promising results with a targeted query.
  3. Build the result object and call formatOutput immediately.

Single target research (one entity, multiple fields)

  1. Search for relevant URLs.
  2. Scrape to extract data — stay in the orchestrator unless you have many independent sources (roughly 5+) where parallel workers clearly help.
  3. Compile findings and call formatOutput.

List of items (array in schema)

  1. Search/scrape to get the list of items.
  2. Are all requested details included in the list?
    • Yes: Build the result and call formatOutput.
    • No: If there are many items (roughly 5+), use spawnAgents so each worker gets the item and fields; otherwise fetch details sequentially in the orchestrator.
  3. Aggregate all results and call formatOutput.

All items from a website

  1. Check sitemaps (sitemap.xml, robots.txt) for an easy route to all pages.
  2. Scrape the entry page. Determine: pagination? Categories? Subcategories?
  3. For pagination, use interact to click through every page.
  4. For categories, scrape each category — use spawnAgents only when many independent categories warrant parallel fan-out.
  5. Aggregate and call formatOutput.

Scraping for structured data

  • PREFER scrape with a targeted query over raw page dumps. It keeps context lean.
  • When scraping lists, ALWAYS ask about pagination in your query: "How many total results? Is there a next page?"
  • For many independent URLs (roughly 5+), spawnAgents can help — each worker gets specific URLs and fields. Fewer URLs: handle in the orchestrator.
  • If a scrape returns a 404 or bot-check, do NOT retry. Move on to alternative sources.

Building the output

  • Match the schema EXACTLY. Every required field must be present.
  • Use null for missing fields — never omit keys.
  • Arrays must be arrays even for single items.
  • Numbers must be actual numbers, not strings (10.99 not "$10.99").
  • Use bashExec with jq to merge data from multiple sources:
    jq -s '.[0] * .[1]' /data/part1.json /data/part2.json > /data/merged.json
    

Validation before output

Before calling formatOutput, verify:

  1. All required fields from the schema are present.
  2. Types match (numbers are numbers, arrays are arrays).
  3. No duplicate entries in arrays.
  4. Source URLs are included where the schema has citation fields.

CRITICAL: Always call formatOutput

When you have gathered ALL data, call formatOutput with format "json" and the structured data. Do NOT stream data inline as markdown tables or JSON code blocks. Do NOT skip formatOutput — downstream systems depend on the structured output.

firecrawl의 다른 스킬

firecrawl-research-index
firecrawl
Firecrawl Research를 사용하여 연구 질문에 답하는 논문을 찾습니다. 의미론적 검색, 의미론적 및 구조적 확장, 본문 내 검증을 활용합니다. 단일 논문 조회나 전체 다중 논문 세트 등 논문 검색/문
data-analysisresearchweb-scraping
oracle
firecrawl
oracle CLI 사용 모범 사례 (프롬프트 + 파일 번들링, 엔진, 세션 및 파일 첨부 패턴)
pinecone
firecrawl
프로덕션 AI 애플리케이션을 위한 관리형 벡터 데이터베이스입니다. 완전 관리형, 자동 확장, 하이브리드 검색(밀집 + 희소), 메타데이터 필터링, 네임스페이스를 지원합니다.
wpds
firecrawl
WordPress 디자인 시스템(WPDS)과 그 컴포넌트, 토큰, 패턴 등을 활용하여 UI를 구축할 때 사용합니다.
audiocraft-audio-generation
firecrawl
PyTorch 라이브러리로, 텍스트-음악(MusicGen) 및 텍스트-사운드(AudioGen)를 포함한 오디오 생성을 지원합니다. 텍스트로부터 음악을 생성해야 할 때 사용합니다…
skypilot-multi-cloud-orchestration
firecrawl
다중 클라우드에서 ML 워크로드를 오케스트레이션하며 자동 비용 최적화를 제공합니다. 여러 클라우드에 걸쳐 학습 또는 배치 작업을 실행해야 하거나, 활용해야 할 때 사용하세요.
firecrawl-seo-audit
firecrawl
Firecrawl을 사용하여 웹사이트의 SEO를 감사합니다. 사용자가 SEO 감사, 메타데이터 및 헤딩 검토, 사이트맵/사이트 구조 분석, 키워드 기회, 경쟁사 SERP 비교, 또는 우선순위가 지정된 검색 최적화 추천을 요청할 때 사용하세요.
data-analysisresearchweb-scraping
gh-issues
firecrawl
GitHub 이슈를 가져오고, 수정을 구현할 하위 에이전트를 생성한 후 PR을 열고, PR 리뷰 코멘트를 모니터링하고 대응합니다. 사용법: /gh-issues [소유자/저장소] [--레이블…]