gpu-document-processing

作者: langchain-ai

在處理大型PDF、文件集合或大量文字擷取任務時使用,這些任務可受益於GPU加速處理。當使用者…時觸發。

npx skills add https://github.com/langchain-ai/deepagents --skill gpu-document-processing

GPU Document Processing Skill

Process large documents and document collections using GPU-accelerated tools. This skill uses the sandbox-as-tool pattern: the agent runs on CPU for reasoning, and sends document processing work to a GPU-equipped environment.

When to Use This Skill

Use this skill when:

  • Processing large PDF files (50+ pages)
  • Analyzing collections of documents (10+ files)
  • Extracting structured data from unstructured documents
  • Performing bulk text extraction and chunking
  • Generating embeddings for large document sets
  • The user uploads or references large documents for analysis

Architecture: Sandbox as Tool

This skill follows the sandbox-as-tool pattern for GPU execution:

  1. Agent reasons on CPU - planning, synthesis, report writing
  2. Processing sent to GPU sandbox - document parsing, embedding, extraction
  3. Results returned to agent - structured output for further analysis

This separation ensures:

  • API keys stay outside the sandbox (security)
  • Agent state persists independently of processing jobs
  • Processing can be parallelized across documents
  • Cost-efficient: GPU used only during processing, not during reasoning

Capabilities

PDF Text Extraction

Extract text content from PDF documents with layout preservation:

  • Headers, paragraphs, lists, and tables detected separately
  • Page numbers and section boundaries preserved
  • Multi-column layout handling

Tabular Data Extraction

Extract tables from documents into structured formats:

  • PDF tables to CSV/DataFrames using GPU-accelerated parsing
  • Automatic column type detection
  • Handles merged cells and multi-row headers

Document Chunking

Split large documents into meaningful chunks for analysis:

  • Semantic chunking (by topic/section boundaries)
  • Fixed-size chunking with overlap for embedding
  • Configurable chunk sizes (default: 512 tokens)

Embedding Generation

Generate vector embeddings for document chunks:

  • Uses NVIDIA NeMo Retriever NIM for GPU-accelerated embedding
  • Supports batch processing for large document sets
  • Compatible with standard vector stores (Milvus, ChromaDB)

Workflow

  1. Receive document reference from the orchestrator
  2. Determine processing type (extraction, analysis, embedding)
  3. Send to GPU sandbox for processing
  4. Collect structured results (text, tables, embeddings)
  5. Write findings to /shared/ for the orchestrator to synthesize

Processing Large Document Collections

For multiple documents:

  1. Process documents in parallel batches (3-5 concurrent)
  2. Extract key metadata first (title, date, author, page count)
  3. Generate per-document summaries
  4. Cross-reference findings across documents
  5. Write consolidated findings with per-document citations

Output Format

When reporting document processing results:

  • Include document metadata (filename, pages, size)
  • Structure extracted content by section/chapter
  • Format tables as markdown tables
  • Include page references for all extracted content
  • Note any extraction quality issues (scanned images, corrupted pages)

Integration with NVIDIA NIM

For production deployments, GPU document processing can leverage:

  • NVIDIA NeMo Retriever: GPU-accelerated embedding and retrieval
  • NVIDIA RAPIDS cuDF: Tabular data processing from extracted tables
  • NVIDIA Triton: Scalable inference for document classification models

See NVIDIA's NIM documentation for self-hosted deployment options.

來自 langchain-ai 的更多技能

deepagents-thread-inspector
langchain-ai
檢查並解釋本地 Deep Agents Code SQLite 工作階段儲存庫中的對話。當 LangSmith 追蹤工具不可用時作為備用方案,用於…
deepagents-python-quickstart
langchain-ai
按照官方快速入門,在 Python 中搭建一個最小的本地 Deep Agent,使用提供者原生的網路搜尋而非 Tavily。當使用者想要……時使用。
deepagents-typescript-quickstart
langchain-ai
按照官方快速入門指南,以 TypeScript 搭建一個最小的本地 Deep Agent,使用供應商原生的網路搜尋而非 Tavily。當使用者……時使用
eval-engineering
langchain-ai
反覆檢查 agent 儲存庫與使用者提供的可選追蹤資料,訪談使用者,並逐一建立、執行及稽核 Harbor evals。用於……
LangChain RAG Pipeline
langchain-ai
在構建任何檢索增強生成(RAG)系統時,請調用此技能。涵蓋文檔加載器、遞迴字符文本分割器、嵌入(OpenAI)等。
LangChain Structured Output & HITL
langchain-ai
langchain-structured-output-&-hitl — 一個可安裝的 AI 代理技能,由 langchain-ai/langchain-skills 發布。
LangSmith Datasets
langchain-ai
當從追蹤建立評估資料集,或將資料集上傳至 LangSmith,或查詢資料集時,請調用此技能。涵蓋資料集類型(final_response、…)
langsmith-evaluator
langchain-ai
在為 LangSmith 建立評估管道時,請調用此技能。涵蓋三個核心組件:(1) 建立評估器 - LLM 作為評審、自訂程式碼;(2)…