PyreCrawl
13 MCP tools for AI agents: scrape, extract, crawl, map, search, academic papers (arXiv + Crossref), batch scrape, deep research, monitor page changes, persistent browser session, cache control. 3-tier auto-fallback ladder with free Cloudflare bypass, no API keys. Self-hosted Firecrawl alternative. `uvx pyrecrawl serve`
Documentation
π₯ PyreCrawl β Web Browsing Superpowers for Your AI Agent
One command gives any AI agent the whole web. Scrape, extract, crawl, map, and search β self-hosted, no API keys, no rate limits, no subscription.
PyreCrawl speaks MCP (Model Context Protocol), the standard tool interface for Claude, Cursor, VS Code, Codex, OpenCode, Hermes, and any MCP-compatible agent.
A smart auto-fallback ladder always picks the cheapest method that succeeds:
fast HTTP
β (403/503/Cloudflare challenge or empty body)
βΌ
stealth browser (real Chromium + Cloudflare solver)
β (still blocked, or the page needs full JS rendering)
βΌ
deep processing (LLM-ready markdown, citations, structured extraction)
β‘ Tools exposed
| Tool | What it does |
|---|---|
scrape(url, prefer="auto") | Single URL β LLM-ready markdown |
extract(url, schema) | Scrape + structured extraction (JsonCss schema) |
map_site(root, include_pattern=None, limit=200) | Enumerate all internal URLs |
crawl(root, max_pages=5, prefer="auto", include_paths=None, exclude_paths=None, max_depth=0) | Multi-page crawl with path filters + true BFS depth |
document(url) | PDF/DOCX/PPTX β markdown (no browser, optional [docs] extras) |
search(query, limit=10) | Web search via DuckDuckGo HTML (no API key) |
search_papers(query, limit=8, source="arxiv", category=None) | Academic search via arXiv + Crossref (no API key) β feed pdf_url into document |
batch_scrape(urls[], ...) | Many URLs in ONE call β parallel, deduped, cache-aware |
deep_research(query, limit=5, scrape_top=3) | Search β evidence pack with [n] citations (no LLM synthesis β your agent does that) |
monitor(url, action, css_selector=None) | Change detection with persisted snapshots + unified diff |
session(session, action, ...) | Persistent browser session (cookies kept) β login walls, multi-step flows, screenshots |
cache(action) | Inspect/clear/enable/disable the HTTP response cache |
health() | Versions + import sanity check |
MCP Resources (read-only state without a tool call):
pyrecrawl://cache/stats Β· pyrecrawl://sessions Β· pyrecrawl://monitors
MCP Prompts (ready-made playbooks): research(topic) Β· rag_ingest(site) Β· watch_page(url)
Env flags
| Variable | Default | Effect |
|---|---|---|
PYRECRAWL_CACHE | off | 1 = in-memory LRU (128 pages), or a directory path (reserved for disk mode) |
PYRECRAWL_CACHE_TTL | 900 | Cache entry lifetime in seconds |
PYRECRAWL_MONITOR_DIR | ~/.pyrecrawl/monitors | Where monitor snapshots persist |
prefer options: "auto" (default ladder) Β· "fast" (HTTP only) Β· "stealth" (CF bypass) Β· "llm" (deep processing).
π Install & Use (one-liner)
1. Install
UV (recommended β one command, zero Python setup)
UV is a fast Python package manager that handles Python itself β no need to install Python separately. Get it once:
# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
Then run PyreCrawl directly β no venv, no pip install, no Python download:
uvx pyrecrawl@latest
Or via uv tool install (persistent, recommended for regular use)
uv tool install pyrecrawl
Or via pipx (alternative)
pipx install pyrecrawl
Or via pip into a venv
pip install pyrecrawl
2. One-time browser engines
pyrecrawl setup
This installs Chromium + stealth browser engines (~2 min, one-time).
3. Register with your AI agent
# Auto-detect installed agents and write their MCP configs
pyrecrawl install
# Or target specific agents
pyrecrawl install claude-desktop cursor
# Dry-run to preview what would change
pyrecrawl install --dry-run
Supported agents: claude-desktop, claude-code, cursor, vscode, codex, opencode, hermes.
4. Start chatting
After installing + registering, restart your agent (or start a new session). Then ask:
"Scrape https://example.com and summarize it."
The tools appear as mcp_pyrecrawl_scrape, mcp_pyrecrawl_extract, mcp_pyrecrawl_map_site, mcp_pyrecrawl_crawl, mcp_pyrecrawl_search, mcp_pyrecrawl_health.
π Manual config (if pyrecrawl install doesn't match your setup)
Claude Desktop
Config file
- Linux:
~/.config/Claude/claude_desktop_config.json - macOS:
~/Library/Application Support/Claude/claude_desktop_config.json - Windows:
%AppData%\Claude\claude_desktop_config.json
{
"mcpServers": {
"pyrecrawl": {
"command": "uvx",
"args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
}
}
}
Claude Code
Config file: project-scoped .mcp.json
{
"mcpServers": {
"pyrecrawl": {
"command": "uvx",
"args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
}
}
}
Cursor
Config file: ~/.cursor/mcp.json
{
"mcpServers": {
"pyrecrawl": {
"command": "uvx",
"args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
}
}
}
VS Code / Copilot
Config file: .vscode/mcp.json (project-scoped)
{
"servers": {
"pyrecrawl": {
"command": "uvx",
"args": ["--from", "pyrecrawl", "pyrecrawl", "serve"],
"type": "stdio"
}
}
}
Codex CLI
Config file: ~/.codex/config.toml
[mcp_servers.pyrecrawl]
command = "uvx"
args = ["--from", "pyrecrawl", "pyrecrawl", "serve"]
OpenCode
Config file: ~/.config/opencode/opencode.json
{
"mcp": {
"pyrecrawl": {
"type": "local",
"command": ["uvx", "--from", "pyrecrawl", "pyrecrawl", "serve"],
"enabled": true
}
}
}
Hermes
Config file
- Linux/macOS:
~/.hermes/config.yaml - Windows:
%LocalAppData%\hermes\config.yaml
mcp_servers:
pyrecrawl:
command: uvx
args:
- --from
- pyrecrawl
- pyrecrawl
- serve
enabled: true
Windows note:
uvxmust be on PATH. If not, use the full path touvx.exe(e.g.C:\Users\<you>\AppData\Local\hermes\bin\uvx.exe).
π§ How the ladder chooses
PyreCrawl runs each request through three tiers, stopping at the first one that returns a complete, LLM-ready result:
| Concern | Fast tier | Stealth tier | Deep tier |
|---|---|---|---|
| Static HTML page | β ~200ms | β | β |
| Cloudflare-protected | β | β Turnstile solver | β |
| JS-heavy SPA | β | β real Chromium | β |
Live DOM data (input .value, JS state) | β | β
js param | β |
| LLM-ready markdown + citations | β | β | β BM25, fit-markdown |
| Structured extraction (CSS schema) | β | β | β |
| Deep crawl (BFS/DFS/BestFirst) | β | β | β adaptive |
The agent never has to pick. prefer="auto" does it every call.
Live DOM data with js and wait_for
Some sites keep the data you want in a DOM property (e.g. an <input>'s .value)
that JS writes after an XHR β it never appears in the serialized HTML. The
scrape tool accepts two stealth-tier params for exactly this:
{
"url": "https://temp-mail.org/id",
"prefer": "stealth",
"wait_for": "document.getElementById('mail').value.includes('@')",
"js": "document.getElementById('mail').value"
}
wait_forβ a JS predicate expression polled until truthy (bounded bytimeout). Use it instead of guessing a sleep for anything that arrives asynchronously.jsβ a JS expression evaluated once the page settles; the value comes back inmeta.js_result. Errors are captured inmeta.js_error(the page result is still returned, never a crash).
π Compared to Firecrawl (hosted)
| Firecrawl | PyreCrawl | |
|---|---|---|
| Cost | Free 1k/mo, then $16β333/mo | Free, self-hosted |
| Local LLM support | β | β Ollama / any LLM |
| Cloudflare bypass | β (Fire-Engine, paid) | β (free, built-in) |
| Markdown + BM25 | β | β |
| Self-host | β | β |
| Academic paper search | β | β
arXiv + Crossref (search_papers) |
| Hosted search API | β /search | β οΈ DuckDuckGo HTML + arXiv/Crossref (no key) |
π§ Development
git clone https://github.com/SanggonBoy/PyreCrawl.git
cd PyreCrawl
uv venv --python 3.12 .venv
source .venv/Scripts/activate # Windows; or .venv/bin/activate on macOS/Linux
uv pip install -e ".[dev]"
python -m playwright install chromium
scrapling install
Run tests
python scripts/selfcheck.py # real-network smoke test
python scripts/probe_stdio.py # stdio JSON-RPC probe
π¦ Publish
Maintainers only:
git tag v0.8.0
git push origin v0.8.0
GitHub Actions builds + uploads to PyPI via trusted publishing.
π Stay up to date
PyreCrawl checks PyPI on every startup and reports the latest version β your
MCP agent sees this automatically via the health() tool response and can
notify you inline.
To check manually:
pyrecrawl version
To upgrade:
pyrecrawl update # runs: uv tool upgrade pyrecrawl
Get notified of new releases: click Watch β Releases only at the GitHub repo to receive email notifications when a new version is published.
π Uninstall
# Remove from all agent configs
pyrecrawl uninstall
# Remove the package
uv tool uninstall pyrecrawl
π‘οΈ License
MIT β see LICENSE.
π Credits
Built on the shoulders of Scrapling and Crawl4AI β both MIT, both excellent.