Web-curl
ดึงข้อมูล แยกเนื้อหา และประมวลผลเนื้อหาจากเว็บและ API รองรับการบล็อกทรัพยากร การยืนยันตัวตน และการค้นหาแบบกำหนดเองของ Google
เอกสาร
Web-curl

Developed by Rayss
🚀 Open Source Project
🛠️ Built with Node.js & TypeScript (Node.js v18+ required)
🎬 Demo Video
Click here to watch the demo video directly in your browser.
If your platform supports it, you can also download and play demo/demo.mp4 directly.
📚 Table of Contents
- Changelog / Update History
- Overview
- Features
- Architecture
- Installation
- Usage
- CLI Usage
- MCP Server Usage
- Configuration
- Examples
- Troubleshooting
- Tips & Best Practices
- Contributing & Issues
- Academic Publication
- License & Attribution
📝 Changelog / Update History
See CHANGELOG.md for a complete history of updates and new features.
📝 Overview
Web-curl is a powerful tool for fetching and extracting text content from web pages and APIs. Use it as a standalone CLI or as an MCP (Model Context Protocol) server. Web-curl leverages Puppeteer for robust web scraping and supports advanced features such as resource blocking, custom headers, authentication, and External Search API integration.
✨ Features
🚀 Deep Research & Automation (v1.4.2)
- Advanced Browser Automation: Full control over Chromium via Puppeteer (click, type, scroll, hover, key presses).
- Always-On Session Persistence: Browser profiles are now always persistent. Login sessions, cookies, and cache are automatically saved in a local
user_data/directory. - Token-Efficient Snapshots (available via the hidden
browser_snapshothandler):- Accessibility Tree: Clean, structured snapshots instead of messy HTML.
- HTML Slice Mode: Raw HTML with
startIndex/endIndexfor safe chunking when needed.
- Chrome DevTools Integration (implemented, but hidden from
list_tools):- Network Monitoring (
browser_network_requests) - Console Logs (
browser_console_messages)
- Network Monitoring (
- External Search API:
multi_search: Run multiple queries in parallel using one configured external search API.
- Intelligent Resource Management:
- Idle Auto-Close: Browser automatically shuts down after 15 minutes of inactivity to save RAM/CPU.
- Tab Rotation: Automatically replaces the oldest tab when the 10-tab limit is reached.
- Media & Documents:
- Full-Page Screenshots: Capture high-quality screenshots with a 5-day auto-cleanup lifecycle and custom destination support.
- Document Parsing: Extract text from PDF and DOCX files directly from URLs.
Storage & Download Details
- 🗂️ Error log rotation:
logs/error-log.txtis rotated when it exceeds ~1MB (renamed toerror-log.txt.bak) to prevent unbounded growth. - 🧹 Logs & temp cleanup: old temporary files in the
logs/directory are cleaned up at startup. - 🛑 Browser lifecycle: Puppeteer browser instances are closed in finally blocks to avoid Chromium temp file leaks.
- 🔎 Content extraction:
- Returns raw text, HTML, and Readability "main article" when available. Readability attempts to extract the primary content of a webpage, removing headers, footers, sidebars, and other non-essential elements, providing a cleaner, more focused text.
- Readability output is subject to
startIndex/maxLength/chunkSizeslicing when requested.
- ⏱️ Timeout control: navigation and API request timeouts are configurable via tool arguments.
- 💾 Output: results can be printed to stdout or written to a file via CLI options.
- ⬇️ Download behavior (
download_file):destinationFolderaccepts relative paths (resolved against the project root) or absolute paths.- The server creates
destinationFolderif it does not exist. - Downloads are streamed using Node streams +
pipelineto minimize memory use and ensure robust writes. - Filenames are derived from the URL path (e.g.,
https://.../path/file.jpg->file.jpg). If no filename is present, the fallback name isdownloaded_file. - Overwrite semantics: by default the implementation will overwrite an existing file with the same name.
- 🖥️ Usage modes: CLI and MCP server (stdin/stdout transport).
- 🌐 REST client:
fetch_apitries a direct GET first and uses the configured External API web fetch when the request fails or returns a non-2xx response. Non-GET methods are sent directly. - 🔍 Search uses Google Custom Search first, then falls back to the External Search API when Google reports a quota or rate limit.
- 🤖 Smart command:
- Auto language detection (franc-min) and optional translation (dynamic
translateimport). - Query enrichment is heuristic-based; results depend on the detected intent.
- Auto language detection (franc-min) and optional translation (dynamic
🏗️ Architecture
This section outlines the high-level architecture of Web-curl.
graph TD
A[User/MCP Host] --> B(CLI / MCP Server)
B --> C{Tool Handlers}
C -- extract/crawl/agent --> D["Puppeteer (Web Scraping)"]
C -- fetch_api --> E["REST Client"]
C -- multi_search --> F["External Search API"]
C -- parse_document --> G["Document Parser (PDF/DOCX)"]
C -- download_file --> H["File System (Downloads)"]
D --> I["Web Content"]
E --> J["External APIs"]
F --> K["External Search Results"]
H --> L["Local Storage"]
- CLI & MCP Server:
src/index.tsImplements both the CLI entry point and the MCP server. - Web Scraping: Uses Puppeteer for headless browsing and content extraction.
- REST Client:
src/rest-client.tsProvides a flexible HTTP client for API requests.
⚙️ MCP Server Configuration Example
To integrate web-curl as an MCP server, add the following configuration to your mcp_settings.json:
{
"mcpServers": {
"web-curl": {
"command": "node",
"args": [
"build/index.js"
],
"disabled": false,
"alwaysAllow": [
"extract",
"crawl",
"agent",
"research",
"browser_configure",
"browser_close",
"multi_search",
"fetch_api",
"download_file",
"parse_document"
],
"env": {
"SEARCH_BASE_URL": "https://example.com/v1/search",
"SEARCH_MODEL": "search-combo",
"SEARCH_API_KEY": "YOUR_EXTERNAL_SEARCH_API_KEY",
"APIKEY_GOOGLE_SEARCH": "YOUR_GOOGLE_API_KEY",
"CX_GOOGLE_SEARCH": "YOUR_CX_ID"
}
}
}
}
🔑 Configure the External Search API
Search always tries Google Custom Search first, using APIKEY_GOOGLE_SEARCH and CX_GOOGLE_SEARCH. If Google reports a quota or rate limit (429, or a quota/rate-limit 403), search automatically falls back to the External Search API. Configure SEARCH_BASE_URL (for example, https://example.com/v1/search), SEARCH_MODEL (for example, search-combo), and SEARCH_API_KEY to enable the fallback. Other Google errors are returned without fallback.
🛠️ Installation
# Clone the repository
git clone https://github.com/rayss868/MCP-Web-Curl
cd web-curl
# Install dependencies
npm install
# Build the project
npm run build
- Prerequisites: Ensure you have Node.js (v18+) and Git installed on your system.
Puppeteer installation notes
-
Windows: Just run
npm install. -
Linux / Ubuntu Server: You must install extra dependencies for Chromium to handle rendering and screenshots in a headless environment. Run:
sudo apt-get update && sudo apt-get install -y \ fonts-liberation \ libasound2 \ libatk-bridge2.0-0 \ libatk1.0-0 \ libc6 \ libcairo2 \ libcups2 \ libdbus-1-3 \ libexpat1 \ libfontconfig1 \ libgbm1 \ libgcc1 \ libglib2.0-0 \ libgtk-3-0 \ libnspr4 \ libnss3 \ libpango-1-0-0 \ libpangocairo-1.0-0 \ libstdc++6 \ libx11-6 \ libx11-xcb1 \ libxcb1 \ libxcomposite1 \ libxcursor1 \ libxdamage1 \ libxext6 \ libxfixes3 \ libxi6 \ libxrandr2 \ libxrender1 \ libxss1 \ libxtst6 \ lsb-release \ wget \ xdg-utils
For more details, see the Puppeteer troubleshooting guide.
🚀 Usage
CLI Usage
The CLI supports fetching and extracting text content from web pages.
# Basic usage
node build/index.js https://example.com
# With options
node build/index.js --timeout 30000 https://example.com
# Save output to a file
node build/index.js -o result.json https://example.com
Command Line Options
--timeout <ms>: Set navigation timeout (default: 60000)-o <file>: Output result to specified file
MCP Server Usage
Web-curl can be run as an MCP server for integration with Roo Context or other MCP-compatible environments.
Exposed Tools (v1.4.2)
Only the tools below are exposed via list_tools to reduce tool-chaining in agent clients.
- browser_configure: Set proxy/user-agent/viewport (session persistence is always on via
user_data/). - browser_close: Close browser and tabs (also auto-closes after 15 minutes of inactivity).
- multi_search: Run multiple searches in parallel using the selected search backend.
- fetch_api: REST API request with response truncation (
limit). - download_file: Download a file from a URL.
- parse_document: Extract text from PDF/DOCX URLs.
- research: Decompose a question into sub-queries, search in parallel, and return a cited markdown report.
- extract: Pull structured fields from a page via CSS selectors, tables, meta tags, JSON-LD, or Readability. Non-HTML responses (plain text, JSON, source files) come back as raw text in
mainContentwithisHtml: false. - crawl: Traverse a site (BFS/DFS/sitemap/link-map) with include/exclude filters and a politeness delay.
- agent: Collect flat records from many pages using a field schema (CSS selectors and/or JSON-LD paths).
Lower-level browser tools still have handlers in CallToolRequestSchema but are intentionally not exposed.
Running as MCP Server
npm run start
The server will communicate via stdin/stdout and expose the tools as defined in src/index.ts.
🚦 Keeping Responses Small (Recommended for Large Pages)
Use extract with maxTextChars to cap the main-content text, and turn off the extra
extractors you do not need. This keeps large pages from flooding the context.
Client request for a trimmed extraction:
{
"name": "extract",
"arguments": {
"url": "https://example.com/article",
"includeTables": false,
"includeJsonLd": false,
"includeMeta": false,
"maxTextChars": 20000
}
}
Response (example):
{
"url": "https://example.com/article",
"title": "Example Article",
"mainContent": "The first 20000 characters of readable text...",
"truncated": true
}
🧩 Configuration
- Session Persistence: Always enabled. Logins and cookies are automatically reused across restarts.
- Timeout: Set navigation and API request timeouts.
- Environment Variables: Configure Google Custom Search and the External Search API fallback. For direct GET fallback, configure
EXTERNAL_API_URL,EXTERNAL_API_KEY, andEXTERNAL_API_MODEL.
💡 Examples {#examples}
Make a REST API Request
{
"name": "fetch_api",
"arguments": {
"url": "https://api.github.com/repos/nodejs/node",
"method": "GET",
"headers": {
"Accept": "application/vnd.github.v3+json"
},
"limit": 10000
}
}
Download File
{
"name": "download_file",
"arguments": {
"url": "https://example.com/image.jpg",
"destinationFolder": "downloads"
}
}
Note: destinationFolder can be either a relative path (resolved against the project root) or an absolute path. The server will create the destination folder if it does not exist.
Configure Browser
{
"name": "browser_configure",
"arguments": {
"proxy": "http://proxy.example.com:8080",
"viewport": { "width": 1920, "height": 1080 }
}
}
Note: Session persistence is always enabled. Cookies and login sessions are automatically stored in the user_data/ directory.
🛠️ Troubleshooting {#troubleshooting}
- Timeout Errors: Increase the
timeoutparameter if requests are timing out. - External Search Fails: Ensure
SEARCH_PROVIDER=externalandSEARCH_BASE_URL,SEARCH_MODEL, andSEARCH_API_KEYmatch your provider. For Google Custom Search, setSEARCH_PROVIDER=google,APIKEY_GOOGLE_SEARCH, andCX_GOOGLE_SEARCH. - Error Logs: Check the
logs/error-log.txtfile for detailed error messages.
🧠 Tips & Best Practices {#tips--best-practices}
Click for advanced tips
- For large pages, use
maxLengthandstartIndexto fetch content in slices. - Always validate your tool arguments to avoid errors.
- Secure your API keys and sensitive data using environment variables.
- Review the MCP tool schemas in
src/index.tsfor all available options.
🤝 Contributing & Issues {#contributing--issues}
Contributions are welcome! If you want to contribute, fork this repository and submit a pull request.
If you find any issues or have suggestions, please open an issue on the repository page.
📄 Academic Publication
This project is the subject of a peer-reviewed journal article:
Saleh, R. Z., & Lubis, M. (2026). Design and Implementation of MCP-Web-Curl: A Model Context Protocol Server for Web and API Access in Agentic Coding Assistants. JURNAL TEKNIK INFORMATIKA, 19(1), 122–134.
- DOI: 10.15408/jti.v19i1.49625
- Article: journal.uinjkt.ac.id/index.php/ti/article/view/49625
- PDF: Download
- License: CC BY-SA 4.0
📄 License & Attribution {#license--attribution}
This project was developed by Rayss.
For questions, improvements, or contributions, please contact the author or open an issue in the repository.
Note: Search availability, quotas, and pricing depend on the selected backend, either your External Search API provider or Google Custom Search.