lintlab PDF to Markdown
PDF a Markdown por lintlab: convierte URLs públicas de PDF en Markdown o texto plano con conciencia de página, incluyendo enlaces, metadatos y tablas simples para RAG e indexación de búsqueda; marca páginas solo de imagen (sin OCR). Un Actor de Apify de pago por evento, utilizado como herramienta MCP a través del servidor MCP alojado de Apify; requiere un token de Apify.
Servidor MCP alojado
npx add-mcp 'https://mcp.apify.com?tools=lintlab/pdf-to-markdown'Se instala en Claude Code, Codex, Cursor y más
Documentación
PDF to Markdown & Text Extractor (lintlab/pdf-to-markdown) Actor
PDF to Markdown converter and PDF text extractor for RAG and AI agents, including scanned PDFs via OCR (7 Latin-script languages). Convert a public PDF to text or page-aware Markdown with metadata, links and simple tables; OCR reads image-only pages instead of only flagging them.
- URL: https://apify.com/lintlab/pdf-to-markdown.md
- Developed by: lintlab (community)
- Categories: Developer tools, AI
- Stats: 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- User rating: No ratings yet
Pricing
from $2.40 / 1,000 pdf processeds
This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events. Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.
Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event
What's an Apify Actor?
An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes. In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours, and optionally produces a well-defined JSON output, datasets with results, or files in key-value store. In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.
How to integrate an Actor?
If asked about integration, you help developers integrate Actors into their projects. You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.
For examples already wired to this Actor's own input schema, see the API section below.
Each client library has reference documentation the quickstart does not restate: JavaScript/TypeScript (npm install apify-client) and Python (pip install apify-client).
README
PDF to Markdown & Text Extractor
Convert a public PDF URL into page-aware Markdown or plain text for search indexing, RAG and AI-agent workflows. The Actor extracts text, links, metadata and tables, and can OCR scanned pages. One dataset item is written per input URL.
Real output from a run on the English UN Universal Declaration of Human Rights PDF (2026-10-01): the first page beside the beginning of the extracted Markdown.
What you get
- Markdown page markers, inferred headings, obvious lists, and preserved links
- Tables extracted as Markdown: read from the PDF's own table tags where present, otherwise rebuilt from column alignment, each with a confidence score and its source
- PDF metadata, page counts, hashes, and structured link records
- Image-only page detection and optional OCR with per-page confidence
- $0.003 per successfully processed PDF
- $0.006 per OCR page with at least 20 recognized characters
OCR defaults to auto, which reads image-only pages. Use off for text-layer extraction only, or all to OCR every page when a PDF's text layer is unusable.
Quick start
{"pdfUrls":[{"url":"https://example.com/report.pdf"}],"outputFormat":"markdown","includeTables":true}
Try a ready-made example: PDF to text for RAG, extract tables from a PDF or OCR a scanned PDF to text. Open one, click Start, then copy it and swap in your own PDF URLs.
Example: a table extracted from a PDF
Real output from a run on the W3C WCAG example table PDF (2026-10-02). The table came from the PDF's own tags, so tableConfidence reports {"rows": 5, "columns": 6, "confidence": 1, "source": "tags"}:
| Disability Category | Participants | Ballots Completed | Ballots Incomplete/ Terminated | Results: Accuracy | Results: Time to complete |
| --- | --- | --- | --- | --- | --- |
| Blind | 5 | 1 | 4 | 34.5%, n=1 | 1199 sec, n=1 |
| Low Vision | 5 | 2 | 3 | 98.3% n=2 (97.7%, n=3) | 1716 sec, n=3 (1934 sec, n=2) |
| Dexterity | 5 | 4 | 1 | 98.3%, n=4 | 1672.1 sec, n=4 |
| Mobility | 3 | 3 | 0 | 95.4%, n=3 | 1416 sec, n=3 |
Untagged PDFs are rebuilt from column alignment instead, scored by how well the columns line up (tables with missing cells are capped at 0.6), so you can route uncertain tables to a manual check.
Which tables to trust
Every table gets a review flag, and every document gets a tableReport summary, so your pipeline can send clean tables straight to the index and hold the doubtful ones for a person:
needsReviewis true when confidence is below 0.8, a row has missing cells, or the table has only one column.raggedRowscounts rows with fewer filled cells than the table's column count.headerDetectedis true when the PDF tags a header row, or the first row is text-only above a numeric column.
Real tableReport from the W3C table run above (2026-10-02):
{"tables": 1, "needsReview": 0, "raggedRows": 0, "lowConfidence": 0, "bySource": {"tags": 1, "layout": 0}}
Confidence measures how well the extraction is supported, not whether the numbers are right, so check values that matter against the PDF even when needsReview is false.
Use with AI agents / MCP
Call lintlab/pdf-to-markdown through the Apify API or Apify MCP server, then read markdown, perPage, metadata, and imageOnlyPages from the default dataset before sending text to an LLM or indexer.
Overview
Convert public PDF URLs into page-aware Markdown or plain text with an Apify Actor. It downloads each document safely, respects robots.txt, verifies that the response is a PDF, and writes one dataset item per URL.
What it does
- Extracts text without a browser and runs OCR locally in the container.
- Marks every Markdown page with
<!-- page N -->. - Infers headings from relative font sizes and preserves obvious lists.
- Extracts tables as Markdown: tagged PDFs from their Table/row/cell structure, untagged ones from consistent column alignment (rows with missing cells kept, cells left empty), with a per-table confidence score.
- Converts link annotations into Markdown links and returns a structured link list.
- Returns document title, author, creator, producer, creation date, and modification date when embedded in the PDF.
- Reports pages with little or no extractable text as image-only and OCRs them in
automode. - Uses SSRF-safe downloads with DNS pinning, public-unicast-only addresses, redirect checks, byte limits, timeouts, and fail-closed robots handling.
PDFs are processed two at a time in off mode and one at a time in auto or all mode. Password-protected or encrypted PDFs are reported as encrypted; the Actor never attempts to bypass encryption, DRM, paywalls, or logins.
Input
{
"pdfUrls": [
{ "url": "https://example.com/report.pdf" }
],
"maxPdfs": 50,
"maxPagesPerPdf": 200,
"maxPdfBytes": 26214400,
"includeTables": true,
"outputFormat": "markdown",
"ocrMode": "auto",
"ocrLanguages": "eng",
"maxOcrPages": 20,
"timeoutSecs": 30
}
outputFormat accepts markdown, text, or both. Duplicate URLs are removed. maxPdfs limits unique input URLs, and the run's maximum charge setting can lower the number attempted.
ocrMode accepts auto (image-only pages), off (no OCR), or all (OCR every page and use recognized text when available). ocrLanguages accepts eng, deu, fra, spa, por, ita, and nld; join codes with + for multilingual pages. maxOcrPages defaults to 20 and cannot exceed 200 per PDF.
Output
Each input produces one dataset item, including failures:
{
"url": "https://example.com/report.pdf",
"finalUrl": "https://cdn.example.com/report.pdf",
"status": 200,
"bytes": 48192,
"sha256": "54eaf...",
"pageCount": 2,
"pagesProcessed": 2,
"metadata": {
"title": "Quarterly report",
"author": "Example organization",
"creator": null,
"producer": null,
"creationDate": null,
"modDate": null
},
"markdown": "<!-- page 1 -->\n\n# Quarterly report\n\n...",
"perPage": [
{ "page": 1, "chars": 831, "imageOnly": false, "ocr": false, "ocrConfidence": null, "ocrChars": 0 },
{ "page": 2, "chars": 286, "imageOnly": true, "ocr": true, "ocrConfidence": 91, "ocrChars": 286 }
],
"imageOnlyPages": [2],
"ocrPages": [2],
"ocrEngine": "Tesseract.js 6.0.1 / tesseract.js-core 6.1.2",
"links": [
{ "page": 1, "text": "Source", "url": "https://example.com/source" }
],
"tables": 1,
"tableConfidence": [
{ "page": 1, "rows": 3, "columns": 2, "confidence": 1, "source": "layout", "raggedRows": 0, "headerDetected": true, "needsReview": false }
],
"tableReport": { "tables": 1, "needsReview": 0, "raggedRows": 0, "lowConfidence": 0, "bySource": { "tags": 0, "layout": 1 } },
"warnings": [],
"error": null
}
The default key-value store also contains SUMMARY, a JSON run summary with processed, failed, charged, page, byte, and error counts.
Pricing
The Actor costs $0.003 per successfully processed PDF, or $3 per 1,000 PDFs, plus $0.006 per OCR page when OCR returns at least 20 characters. Failed or skipped OCR pages are free. The dataset item is saved before charging. Failed downloads, robots-blocked URLs, non-PDF responses, oversized responses, parse failures, and encrypted PDFs are free. A clean image-only report is successful and chargeable because it identifies affected pages without failing.
Limits and responsible use
OCR uses bundled language data without run-time network access. It renders pages at up to 300 DPI with a 9-million-pixel limit and a 30-second timeout per page. Handwriting and low-quality scans may produce poor or empty text. OCR output is paragraphs; it does not infer tables. Heading, paragraph, list, link-label, and table reconstruction for text-layer pages are geometry-based heuristics; complex layouts, multi-column prose, forms, and nested tables may not reproduce perfectly. Tagged tables use PDF structure tree Table/TR/TH/TD nodes where available. In Markdown, spanned or missing cell positions are left empty; nested tables are not reconstructed. Untagged tables use row and column alignment, up to 10 columns. The tableConfidence source is tags or layout. Tagged confidence is below 1 when a tagged cell cannot be mapped to text. Layout confidence is below 1 when cell starts drift enough to lower the rounded alignment score; it is capped at 0.6 when a kept row has empty cells or nearby short text may have been omitted. Confidence measures extraction evidence, not semantic correctness. Right-to-left text-layer pages (Arabic, Hebrew) come out in logical reading order with searchable letters, but paired brackets and punctuation next to embedded left-to-right text can still land on the wrong side, for example )word( or (word,).
Only process documents you own or are authorized to access and transform. The Actor honors robots.txt, accepts only public HTTP(S) destinations, and does not authenticate or bypass access controls.
More lintlab tools
- Website Screenshot & Visual Regression Diff: full-page screenshots with pixel diffs
- Broken Link Checker & Technical SEO Audit: page-level SEO findings and broken internal links
- XML Sitemap Checker & URL Extractor: validate sitemaps and export every URL
Built by lintlab — small, reliable data tools. Tested before release. Support: open an issue on this Actor's Issues tab here on Apify.
Actor input Schema
pdfUrls (type: array):
HTTP or HTTPS URLs of PDF documents. robots.txt is checked for every origin, including redirect targets.
maxPdfs (type: integer):
Maximum number of unique PDF URLs to attempt. The run charge limit may reduce this value.
maxPagesPerPdf (type: integer):
Extract only the first N pages of each document. The total page count is still reported.
maxPdfBytes (type: integer):
Hard streaming download and decompression limit per PDF. Default is 25 MiB (26,214,400 bytes).
includeTables (type: boolean):
Convert tagged tables and consistently aligned untagged rows into Markdown tables, with source and extraction confidence. This is best effort.
outputFormat (type: string):
Return Markdown, plain text, or both. Markdown includes explicit page boundary comments.
ocrMode (type: string):
Auto reads only image-only pages with OCR; off skips OCR; all attempts OCR on every page and uses recognized text when available.
ocrLanguages (type: string):
Tesseract language codes, separated by + for multilingual pages: eng, deu, fra, spa, por, ita, or nld.
maxOcrPages (type: integer):
OCR at most this many pages per PDF; additional eligible pages are reported in warnings.
timeoutSecs (type: integer):
Timeout applied independently to robots.txt and PDF HTTP requests.
Actor input object example
{
"pdfUrls": [
{
"url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
}
],
"maxPdfs": 50,
"maxPagesPerPdf": 200,
"maxPdfBytes": 26214400,
"includeTables": true,
"outputFormat": "markdown",
"ocrMode": "auto",
"ocrLanguages": "eng",
"maxOcrPages": 20,
"timeoutSecs": 30
}
Actor output Schema
dataset (type: string):
No description
summary (type: string):
No description
API
You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.
JavaScript example
import { ApifyClient } from 'apify-client';
// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
token: '<YOUR_API_TOKEN>',
});
// Prepare Actor input
const input = {
"pdfUrls": [
{
"url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
}
]
};
// Run the Actor and wait for it to finish
const run = await client.actor("lintlab/pdf-to-markdown").call(input);
// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
console.dir(item);
});
// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs
Python example
from apify_client import ApifyClient
# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")
# Prepare the Actor input
run_input = { "pdfUrls": [{ "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf" }] }
# Run the Actor and wait for it to finish
run = client.actor("lintlab/pdf-to-markdown").call(run_input=run_input)
# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
print(item)
# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start
CLI example
echo '{
"pdfUrls": [
{
"url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
}
]
}' |
apify call lintlab/pdf-to-markdown --silent --output-dataset
MCP server setup
{
"mcpServers": {
"apify": {
"type": "http",
"url": "https://mcp.apify.com/?tools=fetch-actor-details,lintlab/pdf-to-markdown"
}
}
}
The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an Authorization: Bearer <APIFY_API_TOKEN> header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).
OpenAPI specification
Download the OpenAPI definition: https://api.apify.com/v2/actors/Ov5GMUqzEh2ZUEmck/builds/TalUdWge5DbykDuGC/openapi.json