tavily-dynamic-search

作者: tavily-ai

搜索网络、过滤结果并提取内容,使原始搜索数据永远不会进入您的上下文窗口。只有您精心整理的print()输出会返回。

npx skills add https://github.com/tavily-ai/skills --skill tavily-dynamic-search

Tavily Dynamic Search

Keep large raw web payloads on disk and return only the evidence needed for the task. This is useful when using --include-raw-content, combining several queries, or extracting multiple long pages. Do not use this workflow for a simple lookup that a normal tvly search --json can answer directly.

Before running

Search and extract support capped keyless access. Run them directly when tvly is available. If tvly is missing, follow the tavily-cli setup. Do not look for an API key or authenticate before the first request. If the keyless cap is reached in an interactive session, run tvly login to open browser OAuth, then retry the blocked request once. In an unattended environment, report the cap and authentication options instead of starting an interactive flow.

Workflow

  1. Search broadly without raw content and inspect titles, URLs, scores, and snippets.
  2. Fetch full content only for the best sources.
  3. When raw output could be large, save it with -o and filter the file before printing anything to the model context.
  4. Preserve source URLs beside every extracted fact.

When the user restricts evidence to official or named domains, validate the hostname of every selected URL during local filtering. --include-domains narrows the search but is not proof that every returned result belongs to an allowed host. If full-page extraction is unavailable, label conclusions as search-snippet evidence instead of implying that the page body was verified.

Keep the process in one turn when the relevant sources and filters are already known. Use another turn only when the first search changes what should be extracted.

Create a unique temporary task directory before saving evidence so concurrent agents do not overwrite one another. Python's tempfile.mkdtemp() is available when mktemp is not permitted. Reuse that directory for all raw and filtered artifacts from the task.

Small result: filter a direct JSON response

For a small search response, a pipe is enough:

tvly search "query" --max-results 5 --json | python3 -c '
import json, sys
data = json.load(sys.stdin)
for result in data.get("results", []):
    score = result.get("score") or 0
    title = result.get("title") or ""
    print(f"[{score:.2f}] {title}")
    print(result.get("url", ""))
    print(result.get("content", "")[:300])
'

Do not discard stderr. Authentication failures, keyless-cap messages, and API errors are actionable and must remain visible.

Large result: save first, then filter

Use the CLI's file output so raw page content does not pass through the tool response:

tvly search "query" \
  --include-raw-content markdown \
  --max-results 8 \
  --json \
  -o /tmp/tavily-search-results.json

Then print only bounded evidence:

python3 -c '
import json
from pathlib import Path

data = json.loads(Path("/tmp/tavily-search-results.json").read_text())
for result in data.get("results", []):
    title = result.get("title") or ""
    url = result.get("url") or ""
    print(f"## {title}")
    print(f"URL: {url}")
    print((result.get("raw_content") or result.get("content") or "")[:1200])
    print()
'

Adjust the filtering logic to the question. Prefer relevant paragraphs or fields over fixed character slices when the target information is known. Aim for roughly 150-600 tokens per source unless a table or code block genuinely requires more.

Targeted extraction

When search identifies the right URLs, extract only those pages:

tvly extract "https://example.com/article" \
  --json \
  -o /tmp/tavily-extract-results.json

For topic-focused pages, let Tavily reduce the response before local filtering:

tvly extract "https://example.com/docs" \
  --query "authentication API" \
  --chunks-per-source 3 \
  --json \
  -o /tmp/tavily-extract-results.json

Multiple queries

For multi-angle research, run a small set of focused searches, deduplicate by URL, and rank before extracting. Use subprocess.run(..., capture_output=True, text=True) when orchestrating commands in Python. Check returncode; if a command fails, surface its stderr and stop or retry deliberately. Never use a blanket except Exception: continue that hides missing evidence.

Response shapes

tvly search --json returns query, optional answer, results, and response_time. Each result commonly contains url, title, content, score, and optional raw_content.

tvly extract --json returns results, failed_results, and response_time. Each successful result commonly contains url, raw_content, and optional images.

Treat fields as optional and use .get() while filtering. Inspect failed_results instead of assuming every requested URL succeeded.

Useful options

OptionPurpose
--max-resultsBound the search result count; default 5, maximum 20
--depthChoose ultra-fast, fast, basic, or advanced
--time-rangeRestrict results to day, week, month, or year
--include-domainsRestrict results to a comma-separated list of trusted domains
--exclude-domainsExclude a comma-separated list of domains
--include-raw-contentInclude full content as markdown or text
-o, --outputSave the complete response to a file

Use jq only for short filters when Python is unavailable:

tvly search "query" --json | jq '[.results[] | {title, url, score, content}]'

来自 tavily-ai 的更多技能

research
tavily-ai
针对任意主题进行综合研究,自动收集来源、分析并生成引用。通过多源网络研究并附带明确引用,适用于对比分析、时事追踪、市场调研及详细报告。提供三种模型选项:mini用于针对性单主题研究(约30秒),pro用于全面多角度分析(约60-120秒),auto通过API自动检测复杂度。通过Tavily MCP服务器的OAuth进行身份验证,支持基于浏览器的自动登录...
search
tavily-ai
基于LLM优化的网页搜索,具备相关性评分与灵活筛选功能。支持四种搜索深度模式(极速、快速、基础、高级),可配置延迟与相关性权衡。包含域名过滤、时间范围限制、日期区间、国家加权及原始内容提取。返回结果含标题、URL、内容摘要及相关性评分;可选图片结果与网站图标。通过Tavily MCP服务器或API密钥配置实现自动OAuth认证;...
tavily-best-practices
tavily-ai
面向LLM的网页搜索API,支持实时数据访问、内容提取、站点爬取及AI驱动研究。五大核心方法:search()获取网页结果,extract()提取URL内容,crawl()进行全站提取,map()发现URL,research()实现端到端AI综合。提供Python和JavaScript SDK,支持异步客户端并行查询及可配置搜索深度(极速/快速/基础/高级)。crawl方法接受语义指令以聚焦提取内容...
tavily-cli
tavily-ai
通过Tavily CLI实现网页搜索、内容提取、站点爬取及深度研究。五种命令模式涵盖搜索、提取、URL发现、批量爬取及带引用的多源研究。所有命令支持JSON输出及文件保存,适用于结构化、智能体工作流。升级模式根据需求引导您从简单搜索逐步过渡到提取、映射、爬取乃至全面研究。需安装tavily-cli并通过tvly login进行API密钥认证。
tavily-crawl
tavily-ai
多页网站爬虫,具备语义过滤和Markdown导出功能。可控制深度和广度爬取整个网站部分;通过路径正则表达式、域名或自然语言指令过滤结果,聚焦所需内容。使用--output-dir参数将每个页面保存为本地Markdown文件,或返回结构化JSON供智能体处理。采用语义指令与分块提取,防止向LLM输入结果时上下文膨胀;支持全页提取用于离线文档下载。支持...
tavily-extract
tavily-ai
从最多20个URL中提取干净的Markdown或文本,支持JavaScript渲染和查询聚焦分块。可处理JavaScript渲染页面,提取深度可配置(简单页面使用基础模式,动态SPA和表格使用高级模式)。支持查询聚焦提取,仅返回相关的内容块而非完整页面。默认返回经LLM优化的Markdown格式,也可选择纯文本格式和结构化JSON输出。单次调用最多处理20个URL;...
tavily-research
tavily-ai
基于AI的综合研究,支持多源整合与引用。生成基于网络来源的结构化报告,耗时30-120秒,具体取决于模型选择(mini适用于定向查询,pro适用于复杂对比)。支持多种输出格式:Markdown报告、自定义模式的JSON以及可配置的引用样式(编号、MLA、APA、Chicago)。包含异步工作流,通过--no-wait、status和poll命令支持长时间运行的研究,并具备实时...
tavily-search
tavily-ai
基于LLM优化的网页搜索结果,包含内容片段和相关性评分。支持四种搜索深度(极速、快速、基础、高级),可配置最多20条结果,并支持域名过滤和时间范围限制。返回结构化JSON输出,包含内容片段、相关性评分和针对LLM优化的元数据。提供新闻和金融主题的专用搜索模式,可选AI生成答案和完整页面内容提取。集成至...