pdf-parsing

Cung cấp tiện ích trích xuất PDF hàng loạt ngoại tuyến, hoàn toàn cục bộ (extract_to_markdown.py), giúp tách biệt hóa đơn trong thư mục invoices/ và chuyển đổi PDF thành văn bản sạch…

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill pdf-parsing

PDF Invoice Local Extraction

This skill exposes a 100% offline, local PDF batch extraction utility (extract_to_markdown.py) that isolates all invoice documents inside the invoices/ directory and exports them into clean, readable Markdown files (.md).

This runs entirely offline without any external network, server, or API key dependencies. This decouples raw text extraction from structuring, allowing you (the LLM agent) to perform robust, flexible data extraction natively without relying on fragile regular expressions.

Batch Extraction Tool: extract_to_markdown.py

Processes a directory of PDF invoices sequentially, isolating all PDFs into a dedicated invoices/ subdirectory and converting each to a matching .md file.

python3 skills/pdf-parsing/scripts/extract_to_markdown.py --workspace .agents/workspace
  • Standard PDF Files: Uses pypdf locally to instantly extract text and save it as <filename>.md under .agents/workspace/invoices/.

Agent Orchestration Guidelines

As the LLM agent, you should coordinate the invoice structuring workflow as follows:

  1. Move and Batch Extract: Execute the batch extraction tool:

    python3 /.agents/skills/pdf-parsing/scripts/extract_to_markdown.py --workspace .agents/workspace
    

    This isolates all invoices inside .agents/workspace/invoices/ and populates the folder with matching <invoice_name>.md files.

  2. LLM-Native Structuring: Read the generated .md files under .agents/workspace/invoices/. Use your own agent reasoning (LLM context) to natively extract the structured fields from the markdown:

    • merchant_name
    • date (format: YYYY-MM-DD)
    • amount (float)
    • invoice_number
    • source_file
  3. Compile Structured Database: Combine all structured invoice objects into a single JSON list and save it directly as .agents/workspace/parsed_invoices.json using your file creation tools, matching this schema:

    [
      {
        "date": "2026-05-15",
        "merchant_name": "Google",
        "amount": 150.00,
        "invoice_number": "INV-GCP-1029",
        "source_file": "google_invoice.png"
      }
    ]
    

Dependencies

  • pypdf (>= 4.0.0)

Thêm skills từ google-gemini

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
Hướng dẫn sử dụng công cụ CLI Gemini API. Sử dụng khi bạn cần tương tác với Gemini API qua dòng lệnh, quản lý tác nhân hoặc tạo phương tiện (hình ảnh,…)
behavioral-evals
google-gemini
Hướng dẫn tạo, chạy, sửa lỗi và thúc đẩy các đánh giá hành vi. Sử dụng khi xác minh logic quyết định của tác nhân, gỡ lỗi lỗi, gỡ lỗi prompt…
gemini-live-api-dev
google-gemini
Truyền phát hai chiều thời gian thực với Gemini qua WebSockets cho các cuộc hội thoại âm thanh, video và văn bản. Hỗ trợ đầu vào/đầu ra âm thanh (PCM 16 kHz), khung hình video, văn bản và phiên âm tự động với phát hiện hoạt động giọng nói để xử lý ngắt quãng. Bao gồm các tính năng âm thanh gốc: hội thoại cảm xúc, âm thanh chủ động và chế độ suy nghĩ; gọi hàm để sử dụng công cụ đồng bộ và không đồng bộ; và truy vấn Google Search. Cung cấp quản lý phiên với nén ngữ cảnh, tiếp tục và...
gemini-omni-flash-api
google-gemini
Sử dụng kỹ năng này để chỉnh sửa video tạo sinh, chuyển văn bản thành video, tạo video tham chiếu hình ảnh và hoạt ảnh chuyển tiếp từ khung hình đầu tiên sang video bằng…
gemini-api-dev
google-gemini
Xây dựng ứng dụng với các mô hình Gemini của Google, hỗ trợ nội dung đa phương thức, gọi hàm và đầu ra có cấu trúc trên Python, JavaScript, Go và Java. Truy cập các mô hình Gemini 3 hiện tại (Pro, Flash, Pro Image) với ngữ cảnh 1 triệu token; các mô hình Gemini 2.x và 1.5 cũ đã ngừng hỗ trợ. Hỗ trợ tạo văn bản, hiểu hình ảnh/âm thanh/video, gọi hàm, đầu ra JSON có cấu trúc, thực thi mã, lưu trữ ngữ cảnh và nhúng. Các SDK chính thức có sẵn: google-genai (Python),...
gemini-interactions-api
google-gemini
Giao diện thống nhất cho các mô hình và tác nhân Gemini với trạng thái phía máy chủ, phát trực tuyến và điều phối công cụ. Hỗ trợ nhiều mô hình hiện tại (gemini-3-flash-preview, gemini-3-pro-preview, gemini-2.5-flash/pro) và tác nhân Deep Research; tự động thay thế các ID mô hình đã lỗi thời bằng các lựa chọn thay thế hiện tại. Giảm tải lịch sử hội thoại lên máy chủ qua previous_interaction_id cho các tương tác đa lượt có trạng thái mà không cần quản lý lịch sử thủ công. Điều phối công cụ tích hợp bao gồm...
deliver
google-gemini
Đăng phiên bản tóm tắt của bản briefing tới một webhook đến của Google Chat hoặc Slack, để bản chạy hàng ngày tự giao — bỏ qua một cách im lặng khi không có webhook nào…