pdf-parsing

โดย google-gemini

เปิดเผยยูทิลิตี้การแยก PDF แบบแบตช์ที่ทำงานในเครื่อง 100% แบบออฟไลน์ (extract_to_markdown.py) ซึ่งแยกใบแจ้งหนี้ภายใต้ invoices/ และแปลง PDF เป็นข้อความที่สะอาด…

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill pdf-parsing

PDF Invoice Local Extraction

This skill exposes a 100% offline, local PDF batch extraction utility (extract_to_markdown.py) that isolates all invoice documents inside the invoices/ directory and exports them into clean, readable Markdown files (.md).

This runs entirely offline without any external network, server, or API key dependencies. This decouples raw text extraction from structuring, allowing you (the LLM agent) to perform robust, flexible data extraction natively without relying on fragile regular expressions.

Batch Extraction Tool: extract_to_markdown.py

Processes a directory of PDF invoices sequentially, isolating all PDFs into a dedicated invoices/ subdirectory and converting each to a matching .md file.

python3 skills/pdf-parsing/scripts/extract_to_markdown.py --workspace .agents/workspace
  • Standard PDF Files: Uses pypdf locally to instantly extract text and save it as <filename>.md under .agents/workspace/invoices/.

Agent Orchestration Guidelines

As the LLM agent, you should coordinate the invoice structuring workflow as follows:

  1. Move and Batch Extract: Execute the batch extraction tool:

    python3 /.agents/skills/pdf-parsing/scripts/extract_to_markdown.py --workspace .agents/workspace
    

    This isolates all invoices inside .agents/workspace/invoices/ and populates the folder with matching <invoice_name>.md files.

  2. LLM-Native Structuring: Read the generated .md files under .agents/workspace/invoices/. Use your own agent reasoning (LLM context) to natively extract the structured fields from the markdown:

    • merchant_name
    • date (format: YYYY-MM-DD)
    • amount (float)
    • invoice_number
    • source_file
  3. Compile Structured Database: Combine all structured invoice objects into a single JSON list and save it directly as .agents/workspace/parsed_invoices.json using your file creation tools, matching this schema:

    [
      {
        "date": "2026-05-15",
        "merchant_name": "Google",
        "amount": 150.00,
        "invoice_number": "INV-GCP-1029",
        "source_file": "google_invoice.png"
      }
    ]
    

Dependencies

  • pypdf (>= 4.0.0)

Skills เพิ่มเติมจาก google-gemini

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
คู่มือการใช้เครื่องมือ CLI ของ Gemini API ใช้เมื่อคุณต้องการโต้ตอบกับ Gemini API ผ่านทางบรรทัดคำสั่ง จัดการเอเจนต์ หรือสร้างสื่อ (รูปภาพ, …)
behavioral-evals
google-gemini
คำแนะนำสำหรับการสร้าง การรัน การแก้ไข และการส่งเสริมการประเมินพฤติกรรม ใช้เมื่อตรวจสอบตรรกะการตัดสินใจของตัวแทน การแก้ไขข้อบกพร่อง การดีบักพรอมต์…
gemini-live-api-dev
google-gemini
การสตรีมแบบสองทิศทางแบบเรียลไทม์กับ Gemini ผ่าน WebSockets สำหรับการสนทนาด้วยเสียง วีดีโอ และข้อความ รองรับการป้อน/ส่งออกเสียง (16 kHz PCM), เฟรมวีดีโอ, ข้อความ และการถอดความอัตโนมัติพร้อมการตรวจจับกิจกรรมเสียงเพื่อจัดการการขัดจังหวะ รวมถึงคุณสมบัติเสียงแบบเนทีฟ: การสนทนาที่มีอารมณ์, เสียงเชิงรุก และโหมดการคิด; การเรียกใช้ฟังก์ชันสำหรับการใช้เครื่องมือแบบซิงโครนัสและอะซิงโครนัส; และการอ้างอิง Google Search มีการจัดการเซสชันด้วยการบีบอัดบริบท, การกลับมาดำเนินการต่อ และ...
gemini-omni-flash-api
google-gemini
ใช้ทักษะนี้สำหรับการตัดต่อวิดีโอเชิงสร้างสรรค์ การสร้างวิดีโอจากข้อความ การสร้างวิดีโอโดยอ้างอิงจากภาพ และแอนิเมชันเปลี่ยนผ่านจากเฟรมแรกสู่วิดีโอ โดยใช้…
gemini-api-dev
google-gemini
สร้างแอปพลิเคชันด้วยโมเดล Gemini ของ Google รองรับเนื้อหาหลายรูปแบบ การเรียกใช้ฟังก์ชัน และผลลัพธ์ที่มีโครงสร้างใน Python, JavaScript, Go และ Java เข้าถึงโมเดล Gemini 3 ปัจจุบัน (Pro, Flash, Pro Image) พร้อมบริบท 1M โทเค็น; โมเดล Gemini 2.x และ 1.5 รุ่นเก่าถูกเลิกใช้งานแล้ว รองรับการสร้างข้อความ การทำความเข้าใจรูปภาพ/เสียง/วิดีโอ การเรียกใช้ฟังก์ชัน ผลลัพธ์ JSON ที่มีโครงสร้าง การเรียกใช้โค้ด การแคชบริบท และการฝังเวกเตอร์ SDK อย่างเป็นทางการ: google-genai (Python),...
gemini-interactions-api
google-gemini
อินเทอร์เฟซแบบรวมสำหรับโมเดล Gemini และเอเจนต์ พร้อมสถานะฝั่งเซิร์ฟเวอร์ การสตรีม และการจัดระเบียบเครื่องมือ รองรับโมเดลปัจจุบันหลายรุ่น (gemini-3-flash-preview, gemini-3-pro-preview, gemini-2.5-flash/pro) และเอเจนต์ Deep Research; แทนที่รหัสโมเดลที่เลิกใช้งานโดยอัตโนมัติด้วยทางเลือกปัจจุบัน ถ่ายโอนประวัติการสนทนาไปยังเซิร์ฟเวอร์ผ่าน previous_interaction_id สำหรับการโต้ตอบแบบหลายเทิร์นที่มีสถานะโดยไม่ต้องจัดการประวัติด้วยตนเอง มีการจัดระเบียบเครื่องมือในตัวรวมถึง...
deliver
google-gemini
โพสต์เวอร์ชันย่อของสรุปข้อมูลไปยัง Google Chat หรือ Slack incoming webhook เพื่อให้การส่งรายงานประจำวันเกิดขึ้นได้เอง — จะข้ามอย่างเงียบ ๆ เมื่อไม่มี webhook…