doc-extract

โดย anthropic

แยกข้อความธรรมดาจากไฟล์เอกสาร - PDF, DOCX, XLSX, PPTX, RTF หรือข้อความธรรมดา/markdown/HTML ใช้เมื่อต้องการแปลงเอกสารไบนารีให้เป็นข้อความ สำหรับ…

npx skills add https://github.com/anthropics/healthcare --skill doc-extract

doc-extract

Shared document-to-text extraction. One script, no state: reads an input file, prints JSON to stdout, writes nothing to disk (PHI-safe — no caches, no temp files; callers own any caching).

Setup (once)

cd <this skill dir> && bun install

This pulls liteparse (the lit bin, used for PDF/DOCX/XLSX/PPTX, OCR included) and rtf-to-text (RTF). Without it, PDFs still work via a pdftotext -layout fallback if poppler is installed; other binary formats require liteparse.

Use

bun <this skill dir>/scripts/extract.ts <input-file> [--content-type <mime>]

Output on stdout:

{ "text": "...", "method": "liteparse | pdftotext | rtf-to-text | passthrough", "pages": 12 }
  • text is page-anchored for paged formats: === [page N] === markers between pages.
  • pages is present when page markers exist.
  • method is the extractor that actually produced the text.
  • Format is taken from the file extension; pass --content-type (e.g. application/pdf) when the file has no useful extension, as with downloaded EHR attachments. Note liteparse refuses extension-less files, so those PDFs go through the pdftotext fallback.
  • Errors print {"error": "..."} to stderr and exit 1.

Table caveat

Tables with multiple value columns (option A vs option B, in-tier vs out-of-tier) can interleave columns line-by-line in the extracted text: fragments of adjacent cells alternate, and a cell's text can even land mid-sentence inside a neighboring column. Values usually survive, but which column a value belongs to can become ambiguous. When an answer comes from one column of a multi-column table and the document has no redundant restatement of the value elsewhere, verify it by reading the original page directly before treating it as ground truth. The extracted text's === [page N] === anchor tells you which page: pass it to the Read tool's pages parameter (e.g. pages: "37") to render just that page to vision instead of the whole document.

For other skills

Import the functions instead of shelling out when you're already in bun TS:

import { extract, resolveLit } from "../doc-extract/scripts/extract";
const lit = resolveLit([myRoot]); // also checks myRoot/node_modules/.bin/lit
const text = extract(lit, "/path/to/file.pdf"); // string | null

The contracts skill consumes it this way (its ingest caching stays on the contracts side).

Skills เพิ่มเติมจาก anthropic

analyzing-financial-statements
anthropic
ทักษะนี้คำนวณอัตราส่วนทางการเงินและตัวชี้วัดสำคัญจากข้อมูลงบการเงินเพื่อการวิเคราะห์การลงทุน
applying-brand-guidelines
anthropic
ทักษะนี้ใช้การสร้างแบรนด์และสไตล์องค์กรที่สอดคล้องกันกับเอกสารที่สร้างขึ้นทั้งหมด รวมถึงสี แบบอักษร เค้าโครง และข้อความ
creating-financial-models
anthropic
ทักษะนี้มีชุดเครื่องมือสร้างแบบจำลองทางการเงินขั้นสูง พร้อมการวิเคราะห์ DCF การทดสอบความไว การจำลองแบบมอนติคาร์โล และการวางแผนสถานการณ์สำหรับการลงทุน…
board-minutes
anthropic
ร่างรายงานการประชุมคณะกรรมการหรือคณะอนุกรรมการในรูปแบบขององค์กรของคุณ ตรวจจับการประชุมคณะกรรมการและคณะอนุกรรมการที่กำลังจะมาถึงจากปฏิทินของคุณโดยอัตโนมัติ สอบถามวาระการประชุมและ…
crm-cleanup
anthropic
สแกน HubSpot เพื่อหาดีลที่ค้างอยู่ คอนแทคที่ซ้ำกัน และฟิลด์ที่ขาดหาย จากนั้นแก้ไขตามที่เจ้าของอนุมัติ รองรับอาร์กิวเมนต์ขอบเขตแบบเลือกได้สำหรับดีล คอนแทค…
redshift-api
anthropic
รัน SQL กับ Amazon Redshift — ส่งคำสั่ง ตรวจสอบสถานะ เรียกดูผลลัพธ์แบบแบ่งหน้า และเรียกดูฐานข้อมูล/สคีมา/ตาราง ใช้สิ่งนี้เมื่อผู้ใช้ต้องการ…
ticket-deflector
anthropic
อ่านอีเมลหรือตั๋วของลูกค้าที่ถูกส่งต่อ ดึงสถานะคำสั่งซื้อ/การคืนเงินจาก PayPal และประวัติบัญชีจาก HubSpot ร่างคำตอบที่ปรับโทนเสียงให้ตรงกับเจ้าของ...
reg-feed-watcher
anthropic
ตรวจสอบฟีดข้อบังคับตอนนี้และรายงานสิ่งใหม่ตั้งแต่การตรวจสอบครั้งล่าสุด โดยกรองตามเกณฑ์ความสำคัญที่คุณกำหนด ใช้เมื่อผู้ใช้พูดว่า "ตรวจสอบฟีด"…