lintlab PDF to Markdown

lintlabによるPDF to Markdown:公開PDFのURLを、リンク、メタデータ、簡易テーブル付きのページ対応Markdownまたはプレーンテキストに変換し、RAGや検索インデックスに利用可能。画像のみのページをフラグ付け(OCRなし)。従量課金制のApify Actorで、Apifyのホスト型MCPサーバーを介してMCPツールとして使用。Apifyトークンが必要。

ホスト型 MCP サーバー

npx add-mcp 'https://mcp.apify.com?tools=lintlab/pdf-to-markdown'

Claude Code、Codex、Cursor などにインストールできます

ドキュメント

PDF to Markdown & Text Extractor (lintlab/pdf-to-markdown) Actor

PDF to Markdown converter and PDF text extractor for RAG and AI agents, including scanned PDFs via OCR (7 Latin-script languages). Convert a public PDF to text or page-aware Markdown with metadata, links and simple tables; OCR reads image-only pages instead of only flagging them.

Pricing

from $2.40 / 1,000 pdf processeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events. Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes. In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours, and optionally produces a well-defined JSON output, datasets with results, or files in key-value store. In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects. You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the API section below.

Each client library has reference documentation the quickstart does not restate: JavaScript/TypeScript (npm install apify-client) and Python (pip install apify-client).

README

PDF to Markdown & Text Extractor

Convert a public PDF URL into page-aware Markdown or plain text for search indexing, RAG and AI-agent workflows. The Actor extracts text, links, metadata and tables, and can OCR scanned pages. One dataset item is written per input URL.

PDF to Markdown output: First UDHR PDF page beside the first 25 extracted Markdown lines.

Real output from a run on the English UN Universal Declaration of Human Rights PDF (2026-10-01): the first page beside the beginning of the extracted Markdown.

What you get

  • Markdown page markers, inferred headings, obvious lists, and preserved links
  • Tables extracted as Markdown: read from the PDF's own table tags where present, otherwise rebuilt from column alignment, each with a confidence score and its source
  • PDF metadata, page counts, hashes, and structured link records
  • Image-only page detection and optional OCR with per-page confidence
  • $0.003 per successfully processed PDF
  • $0.006 per OCR page with at least 20 recognized characters

OCR defaults to auto, which reads image-only pages. Use off for text-layer extraction only, or all to OCR every page when a PDF's text layer is unusable.

Quick start

{"pdfUrls":[{"url":"https://example.com/report.pdf"}],"outputFormat":"markdown","includeTables":true}

Try a ready-made example: PDF to text for RAG, extract tables from a PDF or OCR a scanned PDF to text. Open one, click Start, then copy it and swap in your own PDF URLs.

Example: a table extracted from a PDF

Real output from a run on the W3C WCAG example table PDF (2026-10-02). The table came from the PDF's own tags, so tableConfidence reports {"rows": 5, "columns": 6, "confidence": 1, "source": "tags"}:

| Disability Category | Participants | Ballots Completed | Ballots Incomplete/ Terminated | Results: Accuracy | Results: Time to complete |
| --- | --- | --- | --- | --- | --- |
| Blind | 5 | 1 | 4 | 34.5%, n=1 | 1199 sec, n=1 |
| Low Vision | 5 | 2 | 3 | 98.3% n=2 (97.7%, n=3) | 1716 sec, n=3 (1934 sec, n=2) |
| Dexterity | 5 | 4 | 1 | 98.3%, n=4 | 1672.1 sec, n=4 |
| Mobility | 3 | 3 | 0 | 95.4%, n=3 | 1416 sec, n=3 |

Untagged PDFs are rebuilt from column alignment instead, scored by how well the columns line up (tables with missing cells are capped at 0.6), so you can route uncertain tables to a manual check.

Which tables to trust

Every table gets a review flag, and every document gets a tableReport summary, so your pipeline can send clean tables straight to the index and hold the doubtful ones for a person:

  • needsReview is true when confidence is below 0.8, a row has missing cells, or the table has only one column.
  • raggedRows counts rows with fewer filled cells than the table's column count.
  • headerDetected is true when the PDF tags a header row, or the first row is text-only above a numeric column.

Real tableReport from the W3C table run above (2026-10-02):

{"tables": 1, "needsReview": 0, "raggedRows": 0, "lowConfidence": 0, "bySource": {"tags": 1, "layout": 0}}

Confidence measures how well the extraction is supported, not whether the numbers are right, so check values that matter against the PDF even when needsReview is false.

Use with AI agents / MCP

Call lintlab/pdf-to-markdown through the Apify API or Apify MCP server, then read markdown, perPage, metadata, and imageOnlyPages from the default dataset before sending text to an LLM or indexer.

Overview

Convert public PDF URLs into page-aware Markdown or plain text with an Apify Actor. It downloads each document safely, respects robots.txt, verifies that the response is a PDF, and writes one dataset item per URL.

What it does

  • Extracts text without a browser and runs OCR locally in the container.
  • Marks every Markdown page with <!-- page N -->.
  • Infers headings from relative font sizes and preserves obvious lists.
  • Extracts tables as Markdown: tagged PDFs from their Table/row/cell structure, untagged ones from consistent column alignment (rows with missing cells kept, cells left empty), with a per-table confidence score.
  • Converts link annotations into Markdown links and returns a structured link list.
  • Returns document title, author, creator, producer, creation date, and modification date when embedded in the PDF.
  • Reports pages with little or no extractable text as image-only and OCRs them in auto mode.
  • Uses SSRF-safe downloads with DNS pinning, public-unicast-only addresses, redirect checks, byte limits, timeouts, and fail-closed robots handling.

PDFs are processed two at a time in off mode and one at a time in auto or all mode. Password-protected or encrypted PDFs are reported as encrypted; the Actor never attempts to bypass encryption, DRM, paywalls, or logins.

Input

{
  "pdfUrls": [
    { "url": "https://example.com/report.pdf" }
  ],
  "maxPdfs": 50,
  "maxPagesPerPdf": 200,
  "maxPdfBytes": 26214400,
  "includeTables": true,
  "outputFormat": "markdown",
  "ocrMode": "auto",
  "ocrLanguages": "eng",
  "maxOcrPages": 20,
  "timeoutSecs": 30
}

outputFormat accepts markdown, text, or both. Duplicate URLs are removed. maxPdfs limits unique input URLs, and the run's maximum charge setting can lower the number attempted.

ocrMode accepts auto (image-only pages), off (no OCR), or all (OCR every page and use recognized text when available). ocrLanguages accepts eng, deu, fra, spa, por, ita, and nld; join codes with + for multilingual pages. maxOcrPages defaults to 20 and cannot exceed 200 per PDF.

Output

Each input produces one dataset item, including failures:

{
  "url": "https://example.com/report.pdf",
  "finalUrl": "https://cdn.example.com/report.pdf",
  "status": 200,
  "bytes": 48192,
  "sha256": "54eaf...",
  "pageCount": 2,
  "pagesProcessed": 2,
  "metadata": {
    "title": "Quarterly report",
    "author": "Example organization",
    "creator": null,
    "producer": null,
    "creationDate": null,
    "modDate": null
  },
  "markdown": "<!-- page 1 -->\n\n# Quarterly report\n\n...",
  "perPage": [
    { "page": 1, "chars": 831, "imageOnly": false, "ocr": false, "ocrConfidence": null, "ocrChars": 0 },
    { "page": 2, "chars": 286, "imageOnly": true, "ocr": true, "ocrConfidence": 91, "ocrChars": 286 }
  ],
  "imageOnlyPages": [2],
  "ocrPages": [2],
  "ocrEngine": "Tesseract.js 6.0.1 / tesseract.js-core 6.1.2",
  "links": [
    { "page": 1, "text": "Source", "url": "https://example.com/source" }
  ],
  "tables": 1,
  "tableConfidence": [
    { "page": 1, "rows": 3, "columns": 2, "confidence": 1, "source": "layout", "raggedRows": 0, "headerDetected": true, "needsReview": false }
  ],
  "tableReport": { "tables": 1, "needsReview": 0, "raggedRows": 0, "lowConfidence": 0, "bySource": { "tags": 0, "layout": 1 } },
  "warnings": [],
  "error": null
}

The default key-value store also contains SUMMARY, a JSON run summary with processed, failed, charged, page, byte, and error counts.

Pricing

The Actor costs $0.003 per successfully processed PDF, or $3 per 1,000 PDFs, plus $0.006 per OCR page when OCR returns at least 20 characters. Failed or skipped OCR pages are free. The dataset item is saved before charging. Failed downloads, robots-blocked URLs, non-PDF responses, oversized responses, parse failures, and encrypted PDFs are free. A clean image-only report is successful and chargeable because it identifies affected pages without failing.

Limits and responsible use

OCR uses bundled language data without run-time network access. It renders pages at up to 300 DPI with a 9-million-pixel limit and a 30-second timeout per page. Handwriting and low-quality scans may produce poor or empty text. OCR output is paragraphs; it does not infer tables. Heading, paragraph, list, link-label, and table reconstruction for text-layer pages are geometry-based heuristics; complex layouts, multi-column prose, forms, and nested tables may not reproduce perfectly. Tagged tables use PDF structure tree Table/TR/TH/TD nodes where available. In Markdown, spanned or missing cell positions are left empty; nested tables are not reconstructed. Untagged tables use row and column alignment, up to 10 columns. The tableConfidence source is tags or layout. Tagged confidence is below 1 when a tagged cell cannot be mapped to text. Layout confidence is below 1 when cell starts drift enough to lower the rounded alignment score; it is capped at 0.6 when a kept row has empty cells or nearby short text may have been omitted. Confidence measures extraction evidence, not semantic correctness. Right-to-left text-layer pages (Arabic, Hebrew) come out in logical reading order with searchable letters, but paired brackets and punctuation next to embedded left-to-right text can still land on the wrong side, for example )word( or (word,).

Only process documents you own or are authorized to access and transform. The Actor honors robots.txt, accepts only public HTTP(S) destinations, and does not authenticate or bypass access controls.

More lintlab tools


Built by lintlab — small, reliable data tools. Tested before release. Support: open an issue on this Actor's Issues tab here on Apify.

Actor input Schema

pdfUrls (type: array):

HTTP or HTTPS URLs of PDF documents. robots.txt is checked for every origin, including redirect targets.

maxPdfs (type: integer):

Maximum number of unique PDF URLs to attempt. The run charge limit may reduce this value.

maxPagesPerPdf (type: integer):

Extract only the first N pages of each document. The total page count is still reported.

maxPdfBytes (type: integer):

Hard streaming download and decompression limit per PDF. Default is 25 MiB (26,214,400 bytes).

includeTables (type: boolean):

Convert tagged tables and consistently aligned untagged rows into Markdown tables, with source and extraction confidence. This is best effort.

outputFormat (type: string):

Return Markdown, plain text, or both. Markdown includes explicit page boundary comments.

ocrMode (type: string):

Auto reads only image-only pages with OCR; off skips OCR; all attempts OCR on every page and uses recognized text when available.

ocrLanguages (type: string):

Tesseract language codes, separated by + for multilingual pages: eng, deu, fra, spa, por, ita, or nld.

maxOcrPages (type: integer):

OCR at most this many pages per PDF; additional eligible pages are reported in warnings.

timeoutSecs (type: integer):

Timeout applied independently to robots.txt and PDF HTTP requests.

Actor input object example

{
  "pdfUrls": [
    {
      "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    }
  ],
  "maxPdfs": 50,
  "maxPagesPerPdf": 200,
  "maxPdfBytes": 26214400,
  "includeTables": true,
  "outputFormat": "markdown",
  "ocrMode": "auto",
  "ocrLanguages": "eng",
  "maxOcrPages": 20,
  "timeoutSecs": 30
}

Actor output Schema

dataset (type: string):

No description

summary (type: string):

No description

API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

JavaScript example

import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pdfUrls": [
        {
            "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("lintlab/pdf-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

Python example

from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "pdfUrls": [{ "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf" }] }

# Run the Actor and wait for it to finish
run = client.actor("lintlab/pdf-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

CLI example

echo '{
  "pdfUrls": [
    {
      "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    }
  ]
}' |
apify call lintlab/pdf-to-markdown --silent --output-dataset

MCP server setup

{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,lintlab/pdf-to-markdown"
        }
    }
}

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an Authorization: Bearer <APIFY_API_TOKEN> header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Ov5GMUqzEh2ZUEmck/builds/TalUdWge5DbykDuGC/openapi.json