pdf-parsing

作成者: google-gemini

100%ローカルでオフライン動作するPDFバッチ抽出ユーティリティ(extract_to_markdown.py)を公開し、invoices/配下の請求書を分離し、PDFをクリーンな形式に変換します…

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill pdf-parsing

PDF Invoice Local Extraction

This skill exposes a 100% offline, local PDF batch extraction utility (extract_to_markdown.py) that isolates all invoice documents inside the invoices/ directory and exports them into clean, readable Markdown files (.md).

This runs entirely offline without any external network, server, or API key dependencies. This decouples raw text extraction from structuring, allowing you (the LLM agent) to perform robust, flexible data extraction natively without relying on fragile regular expressions.

Batch Extraction Tool: extract_to_markdown.py

Processes a directory of PDF invoices sequentially, isolating all PDFs into a dedicated invoices/ subdirectory and converting each to a matching .md file.

python3 skills/pdf-parsing/scripts/extract_to_markdown.py --workspace .agents/workspace
  • Standard PDF Files: Uses pypdf locally to instantly extract text and save it as <filename>.md under .agents/workspace/invoices/.

Agent Orchestration Guidelines

As the LLM agent, you should coordinate the invoice structuring workflow as follows:

  1. Move and Batch Extract: Execute the batch extraction tool:

    python3 /.agents/skills/pdf-parsing/scripts/extract_to_markdown.py --workspace .agents/workspace
    

    This isolates all invoices inside .agents/workspace/invoices/ and populates the folder with matching <invoice_name>.md files.

  2. LLM-Native Structuring: Read the generated .md files under .agents/workspace/invoices/. Use your own agent reasoning (LLM context) to natively extract the structured fields from the markdown:

    • merchant_name
    • date (format: YYYY-MM-DD)
    • amount (float)
    • invoice_number
    • source_file
  3. Compile Structured Database: Combine all structured invoice objects into a single JSON list and save it directly as .agents/workspace/parsed_invoices.json using your file creation tools, matching this schema:

    [
      {
        "date": "2026-05-15",
        "merchant_name": "Google",
        "amount": 150.00,
        "invoice_number": "INV-GCP-1029",
        "source_file": "google_invoice.png"
      }
    ]
    

Dependencies

  • pypdf (>= 4.0.0)

google-geminiのその他のスキル

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
Gemini API CLIツールの使用ガイド。コマンドラインからGemini APIとやり取りする必要がある場合、エージェントを管理する場合、またはメディア(画像など)を生成する場合に使用します。
behavioral-evals
google-gemini
行動評価の作成、実行、修正、促進に関するガイダンス。エージェントの意思決定ロジックの検証、障害のデバッグ、プロンプトのデバッグなどに使用します。
gemini-live-api-dev
google-gemini
Geminiを介したWebSocket上のリアルタイム双方向ストリーミングにより、音声、動画、テキストの会話を実現。音声入出力(16kHz PCM)、動画フレーム、テキスト、および割り込み処理のための音声アクティビティ検出による自動文字起こしをサポート。ネイティブ音声機能(感情対話、プロアクティブ音声、思考モード)、同期・非同期ツール使用のための関数呼び出し、Google Searchグラウンディングを搭載。コンテキスト圧縮、再開などを備えたセッション管理を提供。
gemini-omni-flash-api
google-gemini
このスキルは、生成型ビデオ編集、テキストからビデオ、画像参照によるビデオ生成、および最初のフレームからビデオへのトランジションアニメーションに使用します…
gemini-api-dev
google-gemini
GoogleのGeminiモデルを使用してアプリケーションを構築します。マルチモーダルコンテンツ、関数呼び出し、構造化出力をサポートし、Python、JavaScript、Go、Javaに対応しています。現在のGemini 3モデル(Pro、Flash、Pro Image)にアクセス可能で、100万トークンのコンテキストを備えています。レガシーのGemini 2.xおよび1.5モデルは非推奨です。テキスト生成、画像/音声/動画の理解、関数呼び出し、構造化JSON出力、コード実行、コンテキストキャッシング、埋め込みをサポートしています。公式SDKとしてgoogle-genai(Python)などが利用可能です。
gemini-interactions-api
google-gemini
Geminiモデルとエージェントのための統合インターフェース。サーバーサイドの状態、ストリーミング、ツールオーケストレーションを備えています。複数の現行モデル(gemini-3-flash-preview、gemini-3-pro-preview、gemini-2.5-flash/pro)およびDeep Researchエージェントをサポート。非推奨のモデルIDを現行の代替モデルに自動的に置き換えます。previous_interaction_idを介して会話履歴をサーバーにオフロードし、手動での履歴管理なしでステートフルなマルチターン対話を実現。組み込みのツールオーケストレーションを含む...
deliver
google-gemini
ブリーフィングの要約版をGoogle ChatまたはSlackの受信ウェブフックに投稿するので、毎日の実行が自動的に配信される — ウェブフックが設定されていない場合は静かにスキップする…