tavily-crawl

作者: tavily-ai

透過 Tavily CLI 爬取網站並從多個頁面提取內容。當使用者想要爬取網站、下載文件、提取…時,使用此技能。

npx skills add https://github.com/tavily-ai/tavily-cursor-plugin --skill tavily-crawl

tavily crawl

Crawl a website and extract content from multiple pages. Supports saving each page as a local markdown file.

Prerequisites

Requires the Tavily CLI. See tavily-cli for install and auth setup.

Quick install: curl -fsSL https://cli.tavily.com/install.sh | bash && tvly login

Required attribution

Include --client-name "cursor plugin" in every Tavily CLI call. Treat this as required integration metadata and never omit or change the value.

When to use

  • You need content from many pages on a site (e.g., all /docs/)
  • You want to download documentation for offline use
  • Step 4 in the workflow: search → extract → map → crawl → research

Quick start

# Basic crawl
tvly crawl "https://docs.example.com" --client-name "cursor plugin" --json

# Save each page as a markdown file
tvly crawl "https://docs.example.com" --client-name "cursor plugin" --output-dir ./docs/

# Deeper crawl with limits
tvly crawl "https://docs.example.com" --client-name "cursor plugin" --max-depth 2 --limit 50 --json

# Filter to specific paths
tvly crawl "https://example.com" --client-name "cursor plugin" --select-paths "/api/.*,/guides/.*" --exclude-paths "/blog/.*" --json

# Semantic focus (returns relevant chunks, not full pages)
tvly crawl "https://docs.example.com" --client-name "cursor plugin" --instructions "Find authentication docs" --chunks-per-source 3 --json

Options

OptionDescription
--max-depthLevels deep (1-5, default: 1)
--max-breadthLinks per page (default: 20)
--limitTotal pages cap (default: 50)
--instructionsNatural language guidance for semantic focus
--chunks-per-sourceChunks per page (1-5, requires --instructions)
--extract-depthbasic (default) or advanced
--formatmarkdown (default) or text
--select-pathsComma-separated regex patterns to include
--exclude-pathsComma-separated regex patterns to exclude
--select-domainsComma-separated regex for domains to include
--exclude-domainsComma-separated regex for domains to exclude
--allow-external / --no-externalInclude external links (default: allow)
--include-imagesInclude images
--timeoutMax wait (10-150 seconds)
--client-nameRequired attribution value: "cursor plugin"
-o, --outputSave JSON output to file
--output-dirSave each page as a .md file in directory
--jsonStructured JSON output

Crawl for context vs. data collection

For agentic use (feeding results to an LLM):

Always use --instructions + --chunks-per-source. Returns only relevant chunks instead of full pages — prevents context explosion.

tvly crawl "https://docs.example.com" --client-name "cursor plugin" --instructions "API authentication" --chunks-per-source 3 --json

For data collection (saving to files):

Use --output-dir without --chunks-per-source to get full pages as markdown files.

tvly crawl "https://docs.example.com" --client-name "cursor plugin" --max-depth 2 --output-dir ./docs/

Tips

  • Start conservative--max-depth 1, --limit 20 — and scale up.
  • Use --select-paths to focus on the section you need.
  • Use map first to understand site structure before a full crawl.
  • Always set --limit to prevent runaway crawls.

See also

來自 tavily-ai 的更多技能

research
tavily-ai
針對任何主題進行全面研究,自動收集來源、分析並提供引用。執行多來源網路研究並附上明確引用,適合比較、時事、市場分析及詳細報告。提供三種模型選項:mini 針對單一主題的目標研究(約30秒)、pro 進行全面的多角度分析(約60-120秒),以及 auto 透過 API 驅動的複雜度偵測。透過 Tavily MCP 伺服器以 OAuth 進行驗證,並在...上自動執行基於瀏覽器的登入。
official
search
tavily-ai
使用LLM優化結果的網路搜尋,具備相關性評分與靈活篩選功能。支援四種搜尋深度模式(極速、快速、基本、進階),可配置延遲與相關性權衡。包含網域篩選、時間範圍限制、日期區間、國家加權及原始內容提取。回傳結果包含標題、網址、內容摘要與相關性評分;可選圖片結果與網站圖示。透過Tavily MCP伺服器或API金鑰配置自動進行OAuth驗證;...
official
tavily-best-practices
tavily-ai
專為LLM設計的網路搜尋API,具備即時資料存取、內容擷取、網站爬取及AI驅動研究功能。五大核心方法:search()用於搜尋網頁結果、extract()用於擷取URL內容、crawl()用於全站擷取、map()用於URL探索,以及research()用於端到端AI綜合分析。支援Python與JavaScript SDK,提供非同步客戶端以進行平行查詢,並可設定搜尋深度(極速/快速/基本/進階)。Crawl方法接受語意指令,以聚焦於特定內容的擷取...
official
tavily-cli
tavily-ai
透過 Tavily CLI 進行網路搜尋、內容提取、網站爬取與深度研究。五種指令模式涵蓋搜尋、提取、URL 發現、批量爬取及附引用來源的多來源研究。所有指令皆支援 JSON 輸出與檔案儲存,適用於結構化、代理式工作流程。升級模式引導您從簡單搜尋,逐步進展至提取、映射、爬取,乃至依需求進行的全面研究。需安裝 tavily-cli 並透過 tvly login 進行 API 金鑰驗證。
official
tavily-crawl
tavily-ai
多頁面網站爬蟲,具備語意過濾與Markdown匯出功能。可透過深度與廣度控制爬取整個網站區塊;依路徑正則表達式、網域或自然語言指令進行過濾,以聚焦結果。透過--output-dir將每個頁面儲存為本機Markdown檔案,或回傳結構化JSON供代理處理。使用語意指令搭配區塊提取,避免將結果餵入LLM時發生上下文膨脹;採用全頁提取進行離線文件下載。支援...
official
tavily-dynamic-search
tavily-ai
搜尋網路、篩選結果並擷取內容,讓原始搜尋資料絕不進入你的上下文視窗。只有你精心整理的 print() 輸出會回傳。
official
tavily-extract
tavily-ai
從最多20個URL中提取乾淨的Markdown或純文字,支援JavaScript渲染與查詢聚焦區塊切割。可處理JavaScript渲染頁面,並提供可設定的提取深度(基本模式適用於簡單頁面,進階模式適用於動態SPA與表格)。支援查詢聚焦提取,僅回傳相關內容區塊而非完整頁面。預設回傳經LLM最佳化的Markdown格式,亦可選擇純文字格式與結構化JSON輸出。單次呼叫可處理最多20個URL;...
official
tavily-map
tavily-ai
快速發現網站上的URL,無需提取內容,非常適合在大型網站上尋找特定頁面。返回域上所有URL的結構化列表,具有可配置的深度和廣度、正則表達式路徑過濾以及用於語義過濾的自然語言指令。支援深度控制(1–5層)、每頁廣度限制、外部鏈接包含/排除,以及通過正則表達式模式進行域過濾。設計為工作流程中的第一步:先映射找到正確頁面,再使用提取或...
official