Web-curl

Tìm nạp, trích xuất và xử lý nội dung web và API. Hỗ trợ chặn tài nguyên, xác thực và Tìm kiếm tùy chỉnh của Google.

GitHub
10
Dùng thử MCP nàyĐược tài trợ

Tài liệu

Web-curl

Web-curl Logo

Developed by Rayss

🚀 Open Source Project
🛠️ Built with Node.js & TypeScript (Node.js v18+ required)


Node.js License Status



🎬 Demo Video

Watch the demo

Click here to watch the demo video directly in your browser.

If your platform supports it, you can also download and play demo/demo.mp4 directly.

Your browser does not support the video tag.

📚 Table of Contents


📝 Changelog / Update History

See CHANGELOG.md for a complete history of updates and new features.

📝 Overview

Web-curl is a powerful tool for fetching and extracting text content from web pages and APIs. Use it as a standalone CLI or as an MCP (Model Context Protocol) server. Web-curl leverages Puppeteer for robust web scraping and supports advanced features such as resource blocking, custom headers, authentication, and External Search API integration.


✨ Features

🚀 Deep Research & Automation (v1.4.2)

  • Advanced Browser Automation: Full control over Chromium via Puppeteer (click, type, scroll, hover, key presses).
  • Always-On Session Persistence: Browser profiles are now always persistent. Login sessions, cookies, and cache are automatically saved in a local user_data/ directory.
  • Token-Efficient Snapshots (available via the hidden browser_snapshot handler):
    • Accessibility Tree: Clean, structured snapshots instead of messy HTML.
    • HTML Slice Mode: Raw HTML with startIndex/endIndex for safe chunking when needed.
  • Chrome DevTools Integration (implemented, but hidden from list_tools):
    • Network Monitoring (browser_network_requests)
    • Console Logs (browser_console_messages)
  • External Search API:
    • multi_search: Run multiple queries in parallel using one configured external search API.
  • Intelligent Resource Management:
    • Idle Auto-Close: Browser automatically shuts down after 15 minutes of inactivity to save RAM/CPU.
    • Tab Rotation: Automatically replaces the oldest tab when the 10-tab limit is reached.
  • Media & Documents:
    • Full-Page Screenshots: Capture high-quality screenshots with a 5-day auto-cleanup lifecycle and custom destination support.
    • Document Parsing: Extract text from PDF and DOCX files directly from URLs.

Storage & Download Details

  • 🗂️ Error log rotation: logs/error-log.txt is rotated when it exceeds ~1MB (renamed to error-log.txt.bak) to prevent unbounded growth.
  • 🧹 Logs & temp cleanup: old temporary files in the logs/ directory are cleaned up at startup.
  • 🛑 Browser lifecycle: Puppeteer browser instances are closed in finally blocks to avoid Chromium temp file leaks.
  • 🔎 Content extraction:
    • Returns raw text, HTML, and Readability "main article" when available. Readability attempts to extract the primary content of a webpage, removing headers, footers, sidebars, and other non-essential elements, providing a cleaner, more focused text.
    • Readability output is subject to startIndex/maxLength/chunkSize slicing when requested.
  • ⏱️ Timeout control: navigation and API request timeouts are configurable via tool arguments.
  • 💾 Output: results can be printed to stdout or written to a file via CLI options.
  • ⬇️ Download behavior (download_file):
    • destinationFolder accepts relative paths (resolved against the project root) or absolute paths.
    • The server creates destinationFolder if it does not exist.
    • Downloads are streamed using Node streams + pipeline to minimize memory use and ensure robust writes.
    • Filenames are derived from the URL path (e.g., https://.../path/file.jpg -> file.jpg). If no filename is present, the fallback name is downloaded_file.
    • Overwrite semantics: by default the implementation will overwrite an existing file with the same name.
  • 🖥️ Usage modes: CLI and MCP server (stdin/stdout transport).
  • 🌐 REST client: fetch_api tries a direct GET first and uses the configured External API web fetch when the request fails or returns a non-2xx response. Non-GET methods are sent directly.
  • 🔍 Search uses Google Custom Search first, then falls back to the External Search API when Google reports a quota or rate limit.
  • 🤖 Smart command:
    • Auto language detection (franc-min) and optional translation (dynamic translate import).
    • Query enrichment is heuristic-based; results depend on the detected intent.

🏗️ Architecture

This section outlines the high-level architecture of Web-curl.

graph TD
    A[User/MCP Host] --> B(CLI / MCP Server)
    B --> C{Tool Handlers}
    C -- extract/crawl/agent --> D["Puppeteer (Web Scraping)"]
    C -- fetch_api --> E["REST Client"]
    C -- multi_search --> F["External Search API"]
    C -- parse_document --> G["Document Parser (PDF/DOCX)"]
    C -- download_file --> H["File System (Downloads)"]
    D --> I["Web Content"]
    E --> J["External APIs"]
    F --> K["External Search Results"]
    H --> L["Local Storage"]
  • CLI & MCP Server: src/index.ts Implements both the CLI entry point and the MCP server.
  • Web Scraping: Uses Puppeteer for headless browsing and content extraction.
  • REST Client: src/rest-client.ts Provides a flexible HTTP client for API requests.

⚙️ MCP Server Configuration Example

To integrate web-curl as an MCP server, add the following configuration to your mcp_settings.json:

{
  "mcpServers": {
    "web-curl": {
      "command": "node",
      "args": [
        "build/index.js"
      ],
      "disabled": false,
      "alwaysAllow": [
        "extract",
        "crawl",
        "agent",
        "research",
        "browser_configure",
        "browser_close",
        "multi_search",
        "fetch_api",
        "download_file",
        "parse_document"
      ],
      "env": {
        "SEARCH_BASE_URL": "https://example.com/v1/search",
        "SEARCH_MODEL": "search-combo",
        "SEARCH_API_KEY": "YOUR_EXTERNAL_SEARCH_API_KEY",
        "APIKEY_GOOGLE_SEARCH": "YOUR_GOOGLE_API_KEY",
        "CX_GOOGLE_SEARCH": "YOUR_CX_ID"
      }
    }
  }
}

🔑 Configure the External Search API

Search always tries Google Custom Search first, using APIKEY_GOOGLE_SEARCH and CX_GOOGLE_SEARCH. If Google reports a quota or rate limit (429, or a quota/rate-limit 403), search automatically falls back to the External Search API. Configure SEARCH_BASE_URL (for example, https://example.com/v1/search), SEARCH_MODEL (for example, search-combo), and SEARCH_API_KEY to enable the fallback. Other Google errors are returned without fallback.


🛠️ Installation

# Clone the repository
git clone https://github.com/rayss868/MCP-Web-Curl
cd web-curl

# Install dependencies
npm install

# Build the project
npm run build
  • Prerequisites: Ensure you have Node.js (v18+) and Git installed on your system.

Puppeteer installation notes

  • Windows: Just run npm install.

  • Linux / Ubuntu Server: You must install extra dependencies for Chromium to handle rendering and screenshots in a headless environment. Run:

    sudo apt-get update && sudo apt-get install -y \
      fonts-liberation \
      libasound2 \
      libatk-bridge2.0-0 \
      libatk1.0-0 \
      libc6 \
      libcairo2 \
      libcups2 \
      libdbus-1-3 \
      libexpat1 \
      libfontconfig1 \
      libgbm1 \
      libgcc1 \
      libglib2.0-0 \
      libgtk-3-0 \
      libnspr4 \
      libnss3 \
      libpango-1-0-0 \
      libpangocairo-1.0-0 \
      libstdc++6 \
      libx11-6 \
      libx11-xcb1 \
      libxcb1 \
      libxcomposite1 \
      libxcursor1 \
      libxdamage1 \
      libxext6 \
      libxfixes3 \
      libxi6 \
      libxrandr2 \
      libxrender1 \
      libxss1 \
      libxtst6 \
      lsb-release \
      wget \
      xdg-utils
    

For more details, see the Puppeteer troubleshooting guide.


🚀 Usage

CLI Usage

The CLI supports fetching and extracting text content from web pages.

# Basic usage
node build/index.js https://example.com

# With options
node build/index.js --timeout 30000 https://example.com

# Save output to a file
node build/index.js -o result.json https://example.com

Command Line Options

  • --timeout <ms>: Set navigation timeout (default: 60000)
  • -o <file>: Output result to specified file

MCP Server Usage

Web-curl can be run as an MCP server for integration with Roo Context or other MCP-compatible environments.

Exposed Tools (v1.4.2)

Only the tools below are exposed via list_tools to reduce tool-chaining in agent clients.

  • browser_configure: Set proxy/user-agent/viewport (session persistence is always on via user_data/).
  • browser_close: Close browser and tabs (also auto-closes after 15 minutes of inactivity).
  • multi_search: Run multiple searches in parallel using the selected search backend.
  • fetch_api: REST API request with response truncation (limit).
  • download_file: Download a file from a URL.
  • parse_document: Extract text from PDF/DOCX URLs.
  • research: Decompose a question into sub-queries, search in parallel, and return a cited markdown report.
  • extract: Pull structured fields from a page via CSS selectors, tables, meta tags, JSON-LD, or Readability. Non-HTML responses (plain text, JSON, source files) come back as raw text in mainContent with isHtml: false.
  • crawl: Traverse a site (BFS/DFS/sitemap/link-map) with include/exclude filters and a politeness delay.
  • agent: Collect flat records from many pages using a field schema (CSS selectors and/or JSON-LD paths).

Lower-level browser tools still have handlers in CallToolRequestSchema but are intentionally not exposed.

Running as MCP Server

npm run start

The server will communicate via stdin/stdout and expose the tools as defined in src/index.ts.


🚦 Keeping Responses Small (Recommended for Large Pages)

Use extract with maxTextChars to cap the main-content text, and turn off the extra extractors you do not need. This keeps large pages from flooding the context.

Client request for a trimmed extraction:

{
  "name": "extract",
  "arguments": {
    "url": "https://example.com/article",
    "includeTables": false,
    "includeJsonLd": false,
    "includeMeta": false,
    "maxTextChars": 20000
  }
}

Response (example):

{
  "url": "https://example.com/article",
  "title": "Example Article",
  "mainContent": "The first 20000 characters of readable text...",
  "truncated": true
}

🧩 Configuration

  • Session Persistence: Always enabled. Logins and cookies are automatically reused across restarts.
  • Timeout: Set navigation and API request timeouts.
  • Environment Variables: Configure Google Custom Search and the External Search API fallback. For direct GET fallback, configure EXTERNAL_API_URL, EXTERNAL_API_KEY, and EXTERNAL_API_MODEL.

💡 Examples {#examples}

Make a REST API Request
{
  "name": "fetch_api",
  "arguments": {
    "url": "https://api.github.com/repos/nodejs/node",
    "method": "GET",
    "headers": {
      "Accept": "application/vnd.github.v3+json"
    },
    "limit": 10000
  }
}
Download File
{
  "name": "download_file",
  "arguments": {
    "url": "https://example.com/image.jpg",
    "destinationFolder": "downloads"
  }
}

Note: destinationFolder can be either a relative path (resolved against the project root) or an absolute path. The server will create the destination folder if it does not exist.

Configure Browser
{
  "name": "browser_configure",
  "arguments": {
    "proxy": "http://proxy.example.com:8080",
    "viewport": { "width": 1920, "height": 1080 }
  }
}

Note: Session persistence is always enabled. Cookies and login sessions are automatically stored in the user_data/ directory.


🛠️ Troubleshooting {#troubleshooting}

  • Timeout Errors: Increase the timeout parameter if requests are timing out.
  • External Search Fails: Ensure SEARCH_PROVIDER=external and SEARCH_BASE_URL, SEARCH_MODEL, and SEARCH_API_KEY match your provider. For Google Custom Search, set SEARCH_PROVIDER=google, APIKEY_GOOGLE_SEARCH, and CX_GOOGLE_SEARCH.
  • Error Logs: Check the logs/error-log.txt file for detailed error messages.

🧠 Tips & Best Practices {#tips--best-practices}

Click for advanced tips
  • For large pages, use maxLength and startIndex to fetch content in slices.
  • Always validate your tool arguments to avoid errors.
  • Secure your API keys and sensitive data using environment variables.
  • Review the MCP tool schemas in src/index.ts for all available options.

🤝 Contributing & Issues {#contributing--issues}

Contributions are welcome! If you want to contribute, fork this repository and submit a pull request.
If you find any issues or have suggestions, please open an issue on the repository page.


📄 Academic Publication

This project is the subject of a peer-reviewed journal article:

Saleh, R. Z., & Lubis, M. (2026). Design and Implementation of MCP-Web-Curl: A Model Context Protocol Server for Web and API Access in Agentic Coding Assistants. JURNAL TEKNIK INFORMATIKA, 19(1), 122–134.


📄 License & Attribution {#license--attribution}

This project was developed by Rayss.
For questions, improvements, or contributions, please contact the author or open an issue in the repository.


Note: Search availability, quotas, and pricing depend on the selected backend, either your External Search API provider or Google Custom Search.