tavily-extract

作者: tavily-ai

从最多20个URL中提取干净的Markdown或文本,支持JavaScript渲染和查询聚焦分块。可处理JavaScript渲染页面,提取深度可配置(简单页面使用基础模式,动态SPA和表格使用高级模式)。支持查询聚焦提取,仅返回相关的内容块而非完整页面。默认返回经LLM优化的Markdown格式,也可选择纯文本格式和结构化JSON输出。单次调用最多处理20个URL;...

npx skills add https://github.com/tavily-ai/skills --skill tavily-extract

tavily extract

Extract clean markdown or text content from one or more URLs.

Before running

Run extract directly when tvly is available. Extract supports capped keyless access, so do not look for an API key or authenticate before the first request.

If tvly is missing, follow the tavily-cli setup before retrying. If the keyless cap is reached in an interactive session, run tvly login to open browser OAuth, then retry the original extraction once. In an unattended environment, report the cap and authentication options instead of starting an interactive flow. Do not start a second login immediately after guided setup has completed.

When to use

  • You have a specific URL and want its content
  • You need text from JavaScript-rendered pages
  • Step 2 in the workflow: search → extract → map → crawl → research

Quick start

# Single URL
tvly extract "https://example.com/article" --json

# Multiple URLs
tvly extract "https://example.com/page1" "https://example.com/page2" --json

# Query-focused extraction (returns relevant chunks only)
tvly extract "https://example.com/docs" --query "authentication API" --chunks-per-source 3 --json

# JS-heavy pages
tvly extract "https://app.example.com" --extract-depth advanced --json

# Save to file
tvly extract "https://example.com/article" -o article.json

Options

OptionDescription
--queryRerank chunks by relevance to this query
--chunks-per-sourceChunks per URL (1-5, requires --query)
--extract-depthbasic (default) or advanced (for JS pages)
--formatmarkdown (default) or text
--include-imagesInclude image URLs
--timeoutMax wait time (1-60 seconds)
-o, --outputSave the JSON response to a file
--jsonStructured JSON output

Extract depth

DepthWhen to use
basicSimple pages, fast — try this first
advancedJS-rendered SPAs, dynamic content, tables

Tips

  • Max 20 URLs per request — batch larger lists into multiple calls.
  • Use --query + --chunks-per-source to get only relevant content instead of full pages.
  • Try basic first, fall back to advanced if content is missing.
  • Set --timeout for slow pages (up to 60s).
  • Inspect failed_results even after exit code 0. A successful request can still return no extracted pages. Retry the affected URL with advanced when appropriate, otherwise report the per-URL failure instead of treating the request as complete.
  • If search results already contain the content you need (via --include-raw-content), skip the extract step.

See also

来自 tavily-ai 的更多技能

research
tavily-ai
针对任意主题进行综合研究,自动收集来源、分析并生成引用。通过多源网络研究并附带明确引用,适用于对比分析、时事追踪、市场调研及详细报告。提供三种模型选项:mini用于针对性单主题研究(约30秒),pro用于全面多角度分析(约60-120秒),auto通过API自动检测复杂度。通过Tavily MCP服务器的OAuth进行身份验证,支持基于浏览器的自动登录...
search
tavily-ai
基于LLM优化的网页搜索,具备相关性评分与灵活筛选功能。支持四种搜索深度模式(极速、快速、基础、高级),可配置延迟与相关性权衡。包含域名过滤、时间范围限制、日期区间、国家加权及原始内容提取。返回结果含标题、URL、内容摘要及相关性评分;可选图片结果与网站图标。通过Tavily MCP服务器或API密钥配置实现自动OAuth认证;...
tavily-best-practices
tavily-ai
面向LLM的网页搜索API,支持实时数据访问、内容提取、站点爬取及AI驱动研究。五大核心方法:search()获取网页结果,extract()提取URL内容,crawl()进行全站提取,map()发现URL,research()实现端到端AI综合。提供Python和JavaScript SDK,支持异步客户端并行查询及可配置搜索深度(极速/快速/基础/高级)。crawl方法接受语义指令以聚焦提取内容...
tavily-cli
tavily-ai
通过Tavily CLI实现网页搜索、内容提取、站点爬取及深度研究。五种命令模式涵盖搜索、提取、URL发现、批量爬取及带引用的多源研究。所有命令支持JSON输出及文件保存,适用于结构化、智能体工作流。升级模式根据需求引导您从简单搜索逐步过渡到提取、映射、爬取乃至全面研究。需安装tavily-cli并通过tvly login进行API密钥认证。
tavily-crawl
tavily-ai
多页网站爬虫,具备语义过滤和Markdown导出功能。可控制深度和广度爬取整个网站部分;通过路径正则表达式、域名或自然语言指令过滤结果,聚焦所需内容。使用--output-dir参数将每个页面保存为本地Markdown文件,或返回结构化JSON供智能体处理。采用语义指令与分块提取,防止向LLM输入结果时上下文膨胀;支持全页提取用于离线文档下载。支持...
tavily-dynamic-search
tavily-ai
搜索网络、过滤结果并提取内容,使原始搜索数据永远不会进入您的上下文窗口。只有您精心整理的print()输出会返回。
tavily-research
tavily-ai
基于AI的综合研究,支持多源整合与引用。生成基于网络来源的结构化报告,耗时30-120秒,具体取决于模型选择(mini适用于定向查询,pro适用于复杂对比)。支持多种输出格式:Markdown报告、自定义模式的JSON以及可配置的引用样式(编号、MLA、APA、Chicago)。包含异步工作流,通过--no-wait、status和poll命令支持长时间运行的研究,并具备实时...
tavily-search
tavily-ai
基于LLM优化的网页搜索结果,包含内容片段和相关性评分。支持四种搜索深度(极速、快速、基础、高级),可配置最多20条结果,并支持域名过滤和时间范围限制。返回结构化JSON输出,包含内容片段、相关性评分和针对LLM优化的元数据。提供新闻和金融主题的专用搜索模式,可选AI生成答案和完整页面内容提取。集成至...