tavily-dynamic-search

작성자: tavily-ai

웹을 검색하고, 결과를 필터링하며, 콘텐츠를 추출하여 원시 검색 데이터가 컨텍스트 창에 들어오지 않도록 합니다. 선별된 print() 출력만 반환됩니다.

npx skills add https://github.com/tavily-ai/skills --skill tavily-dynamic-search

Tavily Dynamic Search

Keep large raw web payloads on disk and return only the evidence needed for the task. This is useful when using --include-raw-content, combining several queries, or extracting multiple long pages. Do not use this workflow for a simple lookup that a normal tvly search --json can answer directly.

Before running

Search and extract support capped keyless access. Run them directly when tvly is available. If tvly is missing, follow the tavily-cli setup. Do not look for an API key or authenticate before the first request. If the keyless cap is reached in an interactive session, run tvly login to open browser OAuth, then retry the blocked request once. In an unattended environment, report the cap and authentication options instead of starting an interactive flow.

Workflow

  1. Search broadly without raw content and inspect titles, URLs, scores, and snippets.
  2. Fetch full content only for the best sources.
  3. When raw output could be large, save it with -o and filter the file before printing anything to the model context.
  4. Preserve source URLs beside every extracted fact.

When the user restricts evidence to official or named domains, validate the hostname of every selected URL during local filtering. --include-domains narrows the search but is not proof that every returned result belongs to an allowed host. If full-page extraction is unavailable, label conclusions as search-snippet evidence instead of implying that the page body was verified.

Keep the process in one turn when the relevant sources and filters are already known. Use another turn only when the first search changes what should be extracted.

Create a unique temporary task directory before saving evidence so concurrent agents do not overwrite one another. Python's tempfile.mkdtemp() is available when mktemp is not permitted. Reuse that directory for all raw and filtered artifacts from the task.

Small result: filter a direct JSON response

For a small search response, a pipe is enough:

tvly search "query" --max-results 5 --json | python3 -c '
import json, sys
data = json.load(sys.stdin)
for result in data.get("results", []):
    score = result.get("score") or 0
    title = result.get("title") or ""
    print(f"[{score:.2f}] {title}")
    print(result.get("url", ""))
    print(result.get("content", "")[:300])
'

Do not discard stderr. Authentication failures, keyless-cap messages, and API errors are actionable and must remain visible.

Large result: save first, then filter

Use the CLI's file output so raw page content does not pass through the tool response:

tvly search "query" \
  --include-raw-content markdown \
  --max-results 8 \
  --json \
  -o /tmp/tavily-search-results.json

Then print only bounded evidence:

python3 -c '
import json
from pathlib import Path

data = json.loads(Path("/tmp/tavily-search-results.json").read_text())
for result in data.get("results", []):
    title = result.get("title") or ""
    url = result.get("url") or ""
    print(f"## {title}")
    print(f"URL: {url}")
    print((result.get("raw_content") or result.get("content") or "")[:1200])
    print()
'

Adjust the filtering logic to the question. Prefer relevant paragraphs or fields over fixed character slices when the target information is known. Aim for roughly 150-600 tokens per source unless a table or code block genuinely requires more.

Targeted extraction

When search identifies the right URLs, extract only those pages:

tvly extract "https://example.com/article" \
  --json \
  -o /tmp/tavily-extract-results.json

For topic-focused pages, let Tavily reduce the response before local filtering:

tvly extract "https://example.com/docs" \
  --query "authentication API" \
  --chunks-per-source 3 \
  --json \
  -o /tmp/tavily-extract-results.json

Multiple queries

For multi-angle research, run a small set of focused searches, deduplicate by URL, and rank before extracting. Use subprocess.run(..., capture_output=True, text=True) when orchestrating commands in Python. Check returncode; if a command fails, surface its stderr and stop or retry deliberately. Never use a blanket except Exception: continue that hides missing evidence.

Response shapes

tvly search --json returns query, optional answer, results, and response_time. Each result commonly contains url, title, content, score, and optional raw_content.

tvly extract --json returns results, failed_results, and response_time. Each successful result commonly contains url, raw_content, and optional images.

Treat fields as optional and use .get() while filtering. Inspect failed_results instead of assuming every requested URL succeeded.

Useful options

OptionPurpose
--max-resultsBound the search result count; default 5, maximum 20
--depthChoose ultra-fast, fast, basic, or advanced
--time-rangeRestrict results to day, week, month, or year
--include-domainsRestrict results to a comma-separated list of trusted domains
--exclude-domainsExclude a comma-separated list of domains
--include-raw-contentInclude full content as markdown or text
-o, --outputSave the complete response to a file

Use jq only for short filters when Python is unavailable:

tvly search "query" --json | jq '[.results[] | {title, url, score, content}]'

tavily-ai의 다른 스킬

research
tavily-ai
모든 주제에 대해 자동 소스 수집, 분석 및 인용을 포함한 포괄적 연구를 수행합니다. 명확한 인용과 함께 다중 소스 웹 연구를 진행하며, 비교 분석, 최신 이슈, 시장 분석 및 상세 보고서에 적합합니다. 세 가지 모델 옵션을 제공합니다: 미니(단일 주제 집중 연구, 약 30초), 프로(포괄적 다각도 분석, 약 60-120초), 오토(API 기반 복잡도 자동 감지). Tavily MCP 서버를 통해 OAuth 인증을 하며, 자동 브라우저 기반 로그인을 지원합니다.
search
tavily-ai
LLM 최적화 결과, 관련성 점수, 유연한 필터링을 갖춘 웹 검색. 네 가지 검색 심도 모드(초고속, 빠름, 기본, 고급)를 지원하며 지연 시간과 관련성 간의 균형을 설정 가능. 도메인 필터링, 시간 범위 제약, 날짜 범위, 국가 가중치 부여, 원본 콘텐츠 추출 포함. 제목, URL, 콘텐츠 스니펫, 관련성 점수와 함께 결과 반환; 선택적 이미지 결과 및 파비콘. Tavily MCP 서버 또는 API 키 구성을 통한 자동 OAuth 인증;...
tavily-best-practices
tavily-ai
LLM을 위한 웹 검색 API로, 실시간 데이터 접근, 콘텐츠 추출, 사이트 크롤링, AI 기반 리서치를 제공합니다. 다섯 가지 핵심 메서드: 웹 결과 검색을 위한 search(), URL 콘텐츠 추출을 위한 extract(), 사이트 전체 추출을 위한 crawl(), URL 발견을 위한 map(), 종단 간 AI 합성을 위한 research()를 지원합니다. Python 및 JavaScript SDK를 제공하며, 병렬 쿼리와 설정 가능한 검색 심도(초고속/고속/기본/고급)를 위한 비동기 클라이언트를 포함합니다. Crawl 메서드는 추출 대상을 집중시키기 위해 의미론적 지시를 받아들입니다...
tavily-cli
tavily-ai
Tavily CLI를 통한 웹 검색, 콘텐츠 추출, 사이트 크롤링 및 심층 리서치. 검색, 추출, URL 발견, 대량 크롤링, 인용 포함 다중 소스 리서치를 아우르는 다섯 가지 명령 모드. 모든 명령은 JSON 출력 및 파일 저장을 지원하여 구조화된 에이전트 워크플로우에 적합. 에스컬레이션 패턴은 단순 검색에서 추출, 매핑, 크롤링, 필요에 따른 종합 리서치까지 안내. tavily-cli 설치 및 tvly login을 통한 API 키 인증 필요.
tavily-crawl
tavily-ai
다중 페이지 웹사이트 크롤러로, 의미론적 필터링과 마크다운 내보내기 기능을 제공합니다. 깊이와 범위를 제어하여 사이트 전체 섹션을 크롤링하고, 경로 정규식, 도메인 또는 자연어 명령어로 필터링하여 결과를 집중시킬 수 있습니다. --output-dir 옵션을 통해 각 페이지를 로컬 마크다운 파일로 저장하거나, 구조화된 JSON을 반환하여 에이전트 처리에 활용할 수 있습니다. 결과를 LLM에 전달할 때 컨텍스트 팽창을 방지하기 위해 청크 추출과 함께 의미론적 명령어를 사용하고, 오프라인 문서 다운로드를 위해 전체 페이지 추출을 지원합니다.
tavily-extract
tavily-ai
최대 20개의 URL에서 깨끗한 마크다운 또는 텍스트를 추출하며, JavaScript 렌더링 및 쿼리 중심 청킹을 지원합니다. JavaScript로 렌더링된 페이지를 처리하며, 추출 깊이를 구성할 수 있습니다(기본 페이지는 기본, 동적 SPA 및 테이블은 고급). 쿼리 중심 추출을 지원하여 전체 페이지 대신 관련 콘텐츠 청크만 반환합니다. 기본적으로 LLM에 최적화된 마크다운을 반환하며, 일반 텍스트 형식 및 구조화된 JSON 출력 옵션을 제공합니다. 단일 호출에서 최대 20개의 URL을 처리합니다.
tavily-research
tavily-ai
다중 소스 종합 및 인용을 포함한 포괄적인 AI 기반 리서치. 웹 출처에 기반한 구조화된 보고서를 생성하며, 모델 선택에 따라 30~120초 소요됨(미니는 특정 질의, 프로는 복잡한 비교에 적합). 여러 출력 형식 지원: 마크다운 보고서, 사용자 정의 스키마를 포함한 JSON, 설정 가능한 인용 스타일(번호형, MLA, APA, 시카고). --no-wait, status, poll 명령어를 통한 장기 실행 리서치용 비동기 워크플로우 및 실시간...
tavily-search
tavily-ai
LLM에 최적화된 결과, 콘텐츠 스니펫 및 관련성 점수를 제공하는 웹 검색. 4가지 검색 깊이(초고속, 빠름, 기본, 고급)를 지원하며, 최대 20개까지 구성 가능한 결과 수, 도메인 필터링 및 시간 범위 제약 조건을 포함합니다. 콘텐츠 스니펫, 관련성 점수 및 LLM 소비에 최적화된 메타데이터가 포함된 구조화된 JSON 출력을 반환합니다. 뉴스 및 금융 주제를 위한 특수 검색 모드, 선택적 AI 생성 답변 및 전체 페이지 콘텐츠 추출을 포함합니다.