gpu-document-processing

โดย langchain-ai

Use when processing large PDFs, document collections, or bulk text extraction tasks that benefit from GPU-accelerated processing. Triggers when the user…

npx skills add https://github.com/langchain-ai/deepagents --skill gpu-document-processing

GPU Document Processing Skill

Process large documents and document collections using GPU-accelerated tools. This skill uses the sandbox-as-tool pattern: the agent runs on CPU for reasoning, and sends document processing work to a GPU-equipped environment.

When to Use This Skill

Use this skill when:

  • Processing large PDF files (50+ pages)
  • Analyzing collections of documents (10+ files)
  • Extracting structured data from unstructured documents
  • Performing bulk text extraction and chunking
  • Generating embeddings for large document sets
  • The user uploads or references large documents for analysis

Architecture: Sandbox as Tool

This skill follows the sandbox-as-tool pattern for GPU execution:

  1. Agent reasons on CPU - planning, synthesis, report writing
  2. Processing sent to GPU sandbox - document parsing, embedding, extraction
  3. Results returned to agent - structured output for further analysis

This separation ensures:

  • API keys stay outside the sandbox (security)
  • Agent state persists independently of processing jobs
  • Processing can be parallelized across documents
  • Cost-efficient: GPU used only during processing, not during reasoning

Capabilities

PDF Text Extraction

Extract text content from PDF documents with layout preservation:

  • Headers, paragraphs, lists, and tables detected separately
  • Page numbers and section boundaries preserved
  • Multi-column layout handling

Tabular Data Extraction

Extract tables from documents into structured formats:

  • PDF tables to CSV/DataFrames using GPU-accelerated parsing
  • Automatic column type detection
  • Handles merged cells and multi-row headers

Document Chunking

Split large documents into meaningful chunks for analysis:

  • Semantic chunking (by topic/section boundaries)
  • Fixed-size chunking with overlap for embedding
  • Configurable chunk sizes (default: 512 tokens)

Embedding Generation

Generate vector embeddings for document chunks:

  • Uses NVIDIA NeMo Retriever NIM for GPU-accelerated embedding
  • Supports batch processing for large document sets
  • Compatible with standard vector stores (Milvus, ChromaDB)

Workflow

  1. Receive document reference from the orchestrator
  2. Determine processing type (extraction, analysis, embedding)
  3. Send to GPU sandbox for processing
  4. Collect structured results (text, tables, embeddings)
  5. Write findings to /shared/ for the orchestrator to synthesize

Processing Large Document Collections

For multiple documents:

  1. Process documents in parallel batches (3-5 concurrent)
  2. Extract key metadata first (title, date, author, page count)
  3. Generate per-document summaries
  4. Cross-reference findings across documents
  5. Write consolidated findings with per-document citations

Output Format

When reporting document processing results:

  • Include document metadata (filename, pages, size)
  • Structure extracted content by section/chapter
  • Format tables as markdown tables
  • Include page references for all extracted content
  • Note any extraction quality issues (scanned images, corrupted pages)

Integration with NVIDIA NIM

For production deployments, GPU document processing can leverage:

  • NVIDIA NeMo Retriever: GPU-accelerated embedding and retrieval
  • NVIDIA RAPIDS cuDF: Tabular data processing from extracted tables
  • NVIDIA Triton: Scalable inference for document classification models

See NVIDIA's NIM documentation for self-hosted deployment options.

Skills เพิ่มเติมจาก langchain-ai

langgraph-docs
langchain-ai
เข้าถึงเอกสาร LangGraph เพื่อสร้างเอเจนต์ที่มีสถานะและเวิร์กโฟลว์แบบหลายเอเจนต์ ดึงข้อมูลเอกสาร Python อย่างเป็นทางการของ LangGraph ซึ่งครอบคลุมเครื่องจักรสถานะ การออกแบบเอเจนต์แบบกราฟ และรูปแบบมนุษย์ในวงจร จัดลำดับความสำคัญของเอกสารที่เกี่ยวข้องตามประเภทคำถาม: คู่มือการใช้งานสำหรับคำถามวิธีทำ หน้าคอนเซปต์สำหรับทฤษฎี บทช่วยสอนสำหรับตัวอย่างแบบครบวงจร และเอกสารอ้างอิง API สำหรับรายละเอียดทางเทคนิค เลือก URL เอกสารที่เกี่ยวข้องมากที่สุด 2–4 รายการโดยอัตโนมัติและดึงเนื้อหามาเพื่อตอบ...
official
langgraph-human-in-the-loop
langchain-ai
หยุดการทำงานของกราฟเพื่อให้มนุษย์ตรวจสอบ อนุมัติ หรือตรวจสอบความถูกต้อง จากนั้นดำเนินการต่อด้วยข้อมูลที่มนุษย์ป้อนเข้าไป ต้องมีสามองค์ประกอบ: ตัวตรวจสอบจุด (InMemorySaver หรือ PostgresSaver), ID เธรดใน config, และเพย์โหลดขัดจังหวะที่แปลงเป็น JSON ได้ interrupt(value) จะหยุดและแสดงข้อมูล; Command(resume=value) จะดำเนินการต่อและส่งคืนค่านั้นไปยังโหนดที่ถูกหยุด โค้ดทั้งหมดก่อน interrupt() จะถูกเรียกใช้ใหม่เมื่อดำเนินการต่อ ดังนั้นผลข้างเคียงต้องเป็น idempotent (ใช้ upsert ไม่ใช่ insert) รองรับเวิร์กโฟลว์การอนุมัติ...
official
web-research
langchain-ai
ใช้ทักษะนี้สำหรับคำขอที่เกี่ยวข้องกับการค้นคว้าทางเว็บ โดยมีแนวทางที่มีโครงสร้างเพื่อดำเนินการค้นคว้าทางเว็บอย่างครอบคลุม
official
langchain-oss-primer
langchain-ai
เริ่มต้นที่นี่เสมอสำหรับโปรเจกต์สร้างเอเจนต์ LangChain, Deep Agents หรือ LangGraph ใดๆ จุดเริ่มต้นที่จำเป็นก่อนเลือกสกิลอื่นหรือเขียนอะไรก็ตาม…
official
skill-creator
langchain-ai
คู่มือสำหรับสร้างสกิลที่มีประสิทธิภาพเพื่อขยายความสามารถของเอเจนต์ด้วยความรู้เฉพาะทาง เวิร์กโฟลว์ หรือการรวมเครื่องมือ ใช้สกิลนี้เมื่อผู้ใช้…
official
social-media
langchain-ai
ร่างโพสต์โซเชียลมีเดียเฉพาะแพลตฟอร์มพร้อมเนื้อหาที่มีงานวิจัยรองรับและภาพประกอบที่สร้างขึ้น รองรับโพสต์ LinkedIn (1,300 ตัวอักษรด้วยน้ำเสียงมืออาชีพ) และเธรด Twitter/X (280 ตัวอักษรต่อทวีตในรูปแบบ 1/🧵) ต้องมอบหมายการวิจัยให้กับซับเอเจนต์ก่อนเขียน จากนั้นอ่านผลลัพธ์เพื่อให้แน่ใจว่าถูกต้องและเกี่ยวข้อง สร้างภาพโซเชียลที่สะดุดตาโดยอัตโนมัติโดยใช้เครื่องมือ generate_social_image ด้วยองค์ประกอบที่โดดเด่นและคอนทราสต์สูงซึ่งปรับให้เหมาะสมสำหรับขนาดเล็ก...
official
deep-agents-memory
langchain-ai
ปลั๊กอินหน่วยความจำและแบ็กเอนด์ไฟล์สำหรับ Deep Agents พร้อมตัวเลือกการกำหนดเส้นทางแบบชั่วคราว ถาวร และแบบผสม แบ็กเอนด์สี่ประเภท: StateBackend (ขอบเขตเธรด, ชั่วคราว), StoreBackend (คงอยู่ข้ามเซสชัน), FilesystemBackend (เข้าถึงดิสก์จริงสำหรับการพัฒนาในเครื่อง) และ CompositeBackend (กำหนดเส้นทางพาธที่แตกต่างไปยังแบ็กเอนด์ที่แตกต่างกัน) FilesystemMiddleware มีเครื่องมือปฏิบัติการไฟล์หกอย่าง: ls, read_file, write_file, edit_file, glob, grep CompositeBackend ใช้การจับคู่คำนำหน้าที่ยาวที่สุดเพื่อกำหนดเส้นทาง...
official
deep-agents-orchestration
langchain-ai
จัดระเบียบเอเจนต์ย่อย วางแผนงานหลายขั้นตอน และต้องได้รับการอนุมัติจากมนุษย์สำหรับการดำเนินการที่ละเอียดอ่อน มอบหมายงานให้กับเอเจนต์ย่อยเฉพาะทางผ่านเครื่องมืองาน เอเจนต์ย่อยแบบกำหนดเองรองรับชุดเครื่องมือและพรอมต์ระบบที่แยกออกจากกัน ในขณะที่เอเจนต์ย่อย "วัตถุประสงค์ทั่วไป" เริ่มต้นจะสืบทอดการกำหนดค่าเอเจนต์หลัก วางแผนและติดตามเวิร์กโฟลว์ที่ซับซ้อนด้วย write_todos โดยจัดระเบียบงานในสถานะรอดำเนินการ กำลังดำเนินการ และเสร็จสมบูรณ์ ต้องใช้ thread_id เพื่อความต่อเนื่องในการเรียกใช้หลายครั้ง ดำเนินการ...
official