convert-pdf-to-md

bởi github

Chuyển đổi tài liệu PDF (.pdf) thành Markdown để nội dung của chúng có thể được phân tích, tóm tắt, tìm kiếm hoặc trích xuất một cách chính xác. Sử dụng kỹ năng này bất cứ khi nào…

npx skills add https://github.com/github/awesome-copilot --skill convert-pdf-to-md

Convert PDF to Markdown

When to use this skill

Trigger this skill any time there is a .pdf file that needs to be understood or processed — for example, a user attaches a PDF and asks questions about it, wants a summary, wants specific data or tables pulled out, or wants multiple PDFs in a folder processed together. PDF is a layout/print format, not reliably readable as plain text, so always convert it to Markdown first using the script in this skill rather than trying to open or parse the file directly.

This skill only supports .pdf — that's MarkItDown's only PDF-family format, so there's no legacy format to worry about here (unlike Word's .doc or Excel's .xls).

Mixed file types: When the user references a folder or set of documents containing multiple supported file types (.pdf, .docx, .xlsx), this skill handles only .pdf files. The agent MUST also invoke the sibling skills in parallel:

  • convert-word-to-md for any .docx files
  • convert-excel-to-md for any .xlsx files

Never process a folder and silently skip a supported file type. All three skills must be invoked together when mixed types are present.

Setup (once per environment)

Before the first conversion in a given environment, follow references/setup.md step by step to ensure Python, pip, markitdown, and pymupdf (for image extraction) are installed. Do this proactively rather than guessing whether the environment is ready — the script itself will also fail with a clear pointer back to that file if a dependency turns out to be missing, so it's safe to just try the conversion first if you're reasonably confident setup was already done.

Usage

The conversion script lives at scripts/convert_pdf_to_md.py.

Output structure: MarkItDown's PDF converter extracts text and tables only — it has no concept of embedded images at all. This script separately extracts real embedded images via PyMuPDF and writes a self-contained folder per document:

<name>/
    img/
        page001_img001.<ext>
        page002_img001.<ext>
        ...
    <name>.md

Because MarkItDown's PDF text does not preserve reliable per-page markers, there's no safe way to know exactly where inline an image belongs. Rather than risk misplacing images next to the wrong paragraph, the script appends a ## Extracted Images section at the end of the Markdown, with a ### Page N subheading per page that has images — read this section separately from the main body text. If the document has no embedded images, no img/ folder or Extracted Images section is created.

Single file:

python scripts\convert_pdf_to_md.py "C:\path\to\document.pdf"

This creates a document\ folder next to the source file (containing document.md and, if present, document\img\). To control the destination folder explicitly:

python scripts\convert_pdf_to_md.py "C:\path\to\document.pdf" -o "C:\path\to\output_folder"

A folder of PDFs (batch mode):

python scripts\convert_pdf_to_md.py "C:\path\to\folder"

Add --recursive to also include subfolders:

python scripts\convert_pdf_to_md.py "C:\path\to\folder" --recursive

Each .pdf found gets its own <name>\ output folder next to it by default. Pass -o "C:\path\to\output_parent" to collect all the generated <name>\ folders under a separate parent directory instead (subfolder structure is preserved when combined with --recursive).

After conversion, read the resulting .md file(s) to perform the actual analysis the user asked for — the script's job is only to produce accurate Markdown (and images), not to interpret the content.

Deciding where output goes

Default — always output next to the source file. The <name>/ folder is created in the same directory as the source .pdf. This is the required default for every case. Do NOT override it unless the user explicitly asks for a different location.

Only use -o when the user explicitly provides an output path (e.g., "save the output to C:\output", "put the results in D:\work"). Do NOT pass -o based on the agent's current working directory, the session state folder, or any implied location.

If the source file path cannot be fully resolved — for example, the user provides only a filename with no directory, or the path is ambiguous — use ask_user to confirm the full absolute path before running the conversion. Never guess or assume the directory.

Troubleshooting

SymptomLikely causeFix
ModuleNotFoundError: No module named 'markitdown' or 'fitz' / exit code 2MarkItDown or PyMuPDF not installedFollow references/setup.md
ERROR: Unsupported file type '...' / exit code 3Not a .pdf fileAsk the user for the correct file, or if it's .doc/.docx/.xlsx, use the matching sibling skill instead
ERROR: Input path not found / exit code 3Wrong path, or file movedConfirm the correct path with the user
FAILED <file> -> ... in batch outputThat specific file is corrupt, password-protected, or otherwise unreadableReport which file(s) failed; other files in the batch still succeed
NOTE: skipped N non-.pdf file(s)Folder contains non-PDF filesExpected — those files are intentionally ignored
Markdown body is empty or near-empty despite images being extractedThe PDF is scanned/image-only with no embedded text layer; MarkItDown does not perform OCRTell the user OCR isn't supported — the extracted page images are still available for them to view
Images appear in an appendix instead of inline with the textDeliberate limitation — MarkItDown's PDF text has no reliable per-page markers to place images inlineExpected behavior; cross-reference the ### Page N heading with the surrounding text context if needed

Thêm skills từ github

debugging-workflows
github
Hướng dẫn gỡ lỗi các quy trình tác nhân GitHub - phân tích nhật ký, kiểm tra lần chạy và khắc phục sự cố
go-codemod
github
Triển khai và kiểm thử các codemod Go cho lệnh gh aw fix.
acreadiness-policy
github
Giúp người dùng chọn, viết hoặc áp dụng chính sách AgentRC. Chính sách tùy chỉnh điểm sẵn sàng bằng cách tắt các kiểm tra không liên quan, ghi đè mức độ tác động/cấp độ, thiết lập…
ai-ready
github
Biến bất kỳ kho lưu trữ nào thành sẵn sàng cho AI — phân tích mã nguồn của bạn và tạo ra AGENTS.md, copilot-instructions.md, quy trình CI, mẫu issue, và nhiều hơn nữa. Khai thác đánh giá PR của bạn…
create-oo-component-documentation
github
Tạo tài liệu toàn diện, chuẩn hóa cho các thành phần hướng đối tượng, tuân theo các phương pháp thực hành tốt nhất trong ngành và tiêu chuẩn tài liệu kiến trúc.
dependabot
github
Dependabot là công cụ quản lý phụ thuộc tích hợp sẵn của GitHub với ba khả năng cốt lõi:
doublecheck
github
Quy trình xác minh ba lớp cho đầu ra AI. Trích xuất các tuyên bố có thể kiểm chứng, tìm nguồn hỗ trợ hoặc mâu thuẫn qua tìm kiếm web, thực hiện đánh giá đối kháng…
foundry-agent-sync
github
Tạo và đồng bộ hóa các tác nhân AI dựa trên prompt trực tiếp trong Azure AI Foundry thông qua REST API, từ một tệp kê khai JSON cục bộ. Không giống như các kỹ năng scaffolding chỉ…