pdf-parsing

tarafından google-gemini

%100 yerel, çevrimdışı bir PDF toplu çıkarma aracını (extract_to_markdown.py) kullanıma sunar; bu araç, invoices/ altındaki faturaları ayırır ve PDF'leri temiz bir şekilde çevirir…

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill pdf-parsing

PDF Invoice Local Extraction

This skill exposes a 100% offline, local PDF batch extraction utility (extract_to_markdown.py) that isolates all invoice documents inside the invoices/ directory and exports them into clean, readable Markdown files (.md).

This runs entirely offline without any external network, server, or API key dependencies. This decouples raw text extraction from structuring, allowing you (the LLM agent) to perform robust, flexible data extraction natively without relying on fragile regular expressions.

Batch Extraction Tool: extract_to_markdown.py

Processes a directory of PDF invoices sequentially, isolating all PDFs into a dedicated invoices/ subdirectory and converting each to a matching .md file.

python3 skills/pdf-parsing/scripts/extract_to_markdown.py --workspace .agents/workspace
  • Standard PDF Files: Uses pypdf locally to instantly extract text and save it as <filename>.md under .agents/workspace/invoices/.

Agent Orchestration Guidelines

As the LLM agent, you should coordinate the invoice structuring workflow as follows:

  1. Move and Batch Extract: Execute the batch extraction tool:

    python3 /.agents/skills/pdf-parsing/scripts/extract_to_markdown.py --workspace .agents/workspace
    

    This isolates all invoices inside .agents/workspace/invoices/ and populates the folder with matching <invoice_name>.md files.

  2. LLM-Native Structuring: Read the generated .md files under .agents/workspace/invoices/. Use your own agent reasoning (LLM context) to natively extract the structured fields from the markdown:

    • merchant_name
    • date (format: YYYY-MM-DD)
    • amount (float)
    • invoice_number
    • source_file
  3. Compile Structured Database: Combine all structured invoice objects into a single JSON list and save it directly as .agents/workspace/parsed_invoices.json using your file creation tools, matching this schema:

    [
      {
        "date": "2026-05-15",
        "merchant_name": "Google",
        "amount": 150.00,
        "invoice_number": "INV-GCP-1029",
        "source_file": "google_invoice.png"
      }
    ]
    

Dependencies

  • pypdf (>= 4.0.0)

google-gemini tarafından daha fazla skill

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
Gemini API CLI aracını kullanma kılavuzu. Gemini API ile komut satırı üzerinden etkileşim kurmanız, aracıları yönetmeniz veya medya (görseller, …) oluşturmanız gerektiğinde kullanın.
behavioral-evals
google-gemini
Davranışsal değerlendirmeler oluşturma, çalıştırma, düzeltme ve teşvik etme rehberi. Ajan karar mantığını doğrularken, hataları ayıklarken, prompt hatalarını ayıklarken kullanın…
gemini-live-api-dev
google-gemini
Gerçek zamanlı çift yönlü akış, WebSockets üzerinden Gemini ile ses, video ve metin konuşmaları için. Ses giriş/çıkışı (16 kHz PCM), video kareleri, metin ve kesinti yönetimi için ses etkinliği algılamalı otomatik transkripsiyonları destekler. Yerel ses özelliklerini içerir: duygusal diyalog, proaktif ses ve düşünme modu; senkron ve asenkron araç kullanımı için fonksiyon çağrısı; ve Google Search grounding. Oturum yönetimi, bağlam sıkıştırma, devam ettirme ve...
gemini-omni-flash-api
google-gemini
Üretken video düzenleme, metinden videoya, görsel referanslı video oluşturma ve ilk kareden videoya geçiş animasyonları için bu beceriyi kullanın…
gemini-api-dev
google-gemini
Google'ın Gemini modelleriyle uygulamalar geliştirin; çok modlu içerik, fonksiyon çağırma ve yapılandırılmış çıktıları Python, JavaScript, Go ve Java'da destekler. 1M token bağlamıyla mevcut Gemini 3 modellerine (Pro, Flash, Pro Image) erişin; eski Gemini 2.x ve 1.5 modelleri kullanımdan kaldırılmıştır. Metin oluşturma, görüntü/ses/video anlama, fonksiyon çağırma, yapılandırılmış JSON çıktısı, kod yürütme, bağlam önbelleğe alma ve gömme işlemlerini destekler. Resmi SDK'lar mevcuttur: google-genai (Python),...
gemini-interactions-api
google-gemini
Gemini modelleri ve ajanları için sunucu tarafı durumu, akış ve araç orkestrasyonu ile birleşik arayüz. Birden çok güncel modeli (gemini-3-flash-preview, gemini-3-pro-preview, gemini-2.5-flash/pro) ve Deep Research ajanını destekler; kullanımdan kaldırılan model kimliklerini otomatik olarak güncel alternatiflerle değiştirir. previous_interaction_id aracılığıyla konuşma geçmişini sunucuya aktararak manuel geçmiş yönetimi olmadan durum bilgisi olan çok turlu etkileşimler sağlar. Dahili araç orkestrasyonu dahil...
deliver
google-gemini
Özet brifingin kısaltılmış bir sürümünü bir Google Chat veya Slack gelen webhook'una gönderir, böylece günlük çalıştırma kendini teslim eder — webhook yoksa sessizce atlar…