gpu-document-processing

作者: langchain-ai

在处理大型PDF、文档集合或批量文本提取任务时使用,这些任务受益于GPU加速处理。当用户…时触发。

npx skills add https://github.com/langchain-ai/deepagents --skill gpu-document-processing

GPU Document Processing Skill

Process large documents and document collections using GPU-accelerated tools. This skill uses the sandbox-as-tool pattern: the agent runs on CPU for reasoning, and sends document processing work to a GPU-equipped environment.

When to Use This Skill

Use this skill when:

  • Processing large PDF files (50+ pages)
  • Analyzing collections of documents (10+ files)
  • Extracting structured data from unstructured documents
  • Performing bulk text extraction and chunking
  • Generating embeddings for large document sets
  • The user uploads or references large documents for analysis

Architecture: Sandbox as Tool

This skill follows the sandbox-as-tool pattern for GPU execution:

  1. Agent reasons on CPU - planning, synthesis, report writing
  2. Processing sent to GPU sandbox - document parsing, embedding, extraction
  3. Results returned to agent - structured output for further analysis

This separation ensures:

  • API keys stay outside the sandbox (security)
  • Agent state persists independently of processing jobs
  • Processing can be parallelized across documents
  • Cost-efficient: GPU used only during processing, not during reasoning

Capabilities

PDF Text Extraction

Extract text content from PDF documents with layout preservation:

  • Headers, paragraphs, lists, and tables detected separately
  • Page numbers and section boundaries preserved
  • Multi-column layout handling

Tabular Data Extraction

Extract tables from documents into structured formats:

  • PDF tables to CSV/DataFrames using GPU-accelerated parsing
  • Automatic column type detection
  • Handles merged cells and multi-row headers

Document Chunking

Split large documents into meaningful chunks for analysis:

  • Semantic chunking (by topic/section boundaries)
  • Fixed-size chunking with overlap for embedding
  • Configurable chunk sizes (default: 512 tokens)

Embedding Generation

Generate vector embeddings for document chunks:

  • Uses NVIDIA NeMo Retriever NIM for GPU-accelerated embedding
  • Supports batch processing for large document sets
  • Compatible with standard vector stores (Milvus, ChromaDB)

Workflow

  1. Receive document reference from the orchestrator
  2. Determine processing type (extraction, analysis, embedding)
  3. Send to GPU sandbox for processing
  4. Collect structured results (text, tables, embeddings)
  5. Write findings to /shared/ for the orchestrator to synthesize

Processing Large Document Collections

For multiple documents:

  1. Process documents in parallel batches (3-5 concurrent)
  2. Extract key metadata first (title, date, author, page count)
  3. Generate per-document summaries
  4. Cross-reference findings across documents
  5. Write consolidated findings with per-document citations

Output Format

When reporting document processing results:

  • Include document metadata (filename, pages, size)
  • Structure extracted content by section/chapter
  • Format tables as markdown tables
  • Include page references for all extracted content
  • Note any extraction quality issues (scanned images, corrupted pages)

Integration with NVIDIA NIM

For production deployments, GPU document processing can leverage:

  • NVIDIA NeMo Retriever: GPU-accelerated embedding and retrieval
  • NVIDIA RAPIDS cuDF: Tabular data processing from extracted tables
  • NVIDIA Triton: Scalable inference for document classification models

See NVIDIA's NIM documentation for self-hosted deployment options.

来自 langchain-ai 的更多技能

langgraph-docs
langchain-ai
访问LangGraph文档,构建有状态代理和多代理工作流。获取官方LangGraph Python文档,涵盖状态机、基于图的代理设计以及人机协同模式。根据查询类型优先提供相关文档:实现指南解答操作问题,概念页面讲解理论,教程提供端到端示例,API参考提供技术细节。自动选择2–4个最相关的文档URL并检索其内容以回答...
official
langgraph-human-in-the-loop
langchain-ai
暂停图执行以进行人工审查、批准或验证,随后根据其输入恢复执行。需要三个组件:检查点存储器(InMemorySaver 或 PostgresSaver)、配置中的线程 ID 以及 JSON 可序列化的中断负载。interrupt(value) 暂停执行并展示数据;Command(resume=value) 恢复执行并将该值返回给暂停的节点。恢复时,interrupt() 之前的所有代码会重新执行,因此副作用必须具有幂等性(使用 upsert 而非 insert)。支持审批工作流,...
official
web-research
langchain-ai
用于处理与网络研究相关的请求;它提供了一种结构化的方法来进行全面的网络研究
official
langchain-oss-primer
langchain-ai
任何LangChain、Deep Agents或LangGraph代理构建项目都请始终从这里开始。在选择其他技能或编写任何内容之前,这是必需的起点。
official
skill-creator
langchain-ai
创建有效技能的指南,通过专业知识、工作流程或工具集成来扩展代理能力。当用户……时使用此技能。
official
social-media
langchain-ai
根据研究内容起草特定平台的社交媒体帖子,并生成配套图片。支持领英帖子(1300字符,专业语气)和推特/X话题(每条推文280字符,采用1/🧵格式)。需在撰写前将研究任务委托给子代理,随后阅读研究结果以确保准确性和相关性。使用generate_social_image工具自动生成引人注目的社交图片,采用粗体高对比度构图,针对小屏幕进行优化...
official
deep-agents-memory
langchain-ai
为Deep Agents提供可插拔的内存与文件后端,支持临时、持久化和混合路由选项。四种后端类型:StateBackend(线程作用域,临时)、StoreBackend(跨会话持久化)、FilesystemBackend(本地开发时真实磁盘访问)和CompositeBackend(将不同路径路由到不同后端)。FilesystemMiddleware提供六种文件操作工具:ls、read_file、write_file、edit_file、glob、grep。CompositeBackend使用最长前缀匹配进行路由...
official
deep-agents-orchestration
langchain-ai
编排子代理,规划多步骤任务,并对敏感操作要求人工审批。通过任务工具将工作委派给专业子代理;自定义子代理支持独立的工具集和系统提示,而默认的“通用”子代理继承主代理配置。使用write_todos规划并跟踪复杂工作流,将任务组织为待处理、进行中和已完成状态;需要thread_id以实现跨调用的持久化。实现...
official