tavily-extract
최대 20개의 URL에서 깨끗한 마크다운 또는 텍스트를 추출하며, JavaScript 렌더링 및 쿼리 중심 청킹을 지원합니다. JavaScript로 렌더링된 페이지를 처리하며, 추출 깊이를 구성할 수 있습니다(기본 페이지는 기본, 동적 SPA 및 테이블은 고급). 쿼리 중심 추출을 지원하여 전체 페이지 대신 관련 콘텐츠 청크만 반환합니다. 기본적으로 LLM에 최적화된 마크다운을 반환하며, 일반 텍스트 형식 및 구조화된 JSON 출력 옵션을 제공합니다. 단일 호출에서 최대 20개의 URL을 처리합니다.
npx skills add https://github.com/tavily-ai/skills --skill tavily-extracttavily extract
Extract clean markdown or text content from one or more URLs.
Before running
Run extract directly when tvly is available. Extract supports capped keyless
access, so do not look for an API key or authenticate before the first request.
If tvly is missing, follow the tavily-cli setup
before retrying. If the keyless cap is reached in an interactive session, run
tvly login to open browser OAuth, then retry the original extraction once. In
an unattended environment, report the cap and authentication options instead
of starting an interactive flow. Do not start a second login immediately after
guided setup has completed.
When to use
- You have a specific URL and want its content
- You need text from JavaScript-rendered pages
- Step 2 in the workflow: search → extract → map → crawl → research
Quick start
# Single URL
tvly extract "https://example.com/article" --json
# Multiple URLs
tvly extract "https://example.com/page1" "https://example.com/page2" --json
# Query-focused extraction (returns relevant chunks only)
tvly extract "https://example.com/docs" --query "authentication API" --chunks-per-source 3 --json
# JS-heavy pages
tvly extract "https://app.example.com" --extract-depth advanced --json
# Save to file
tvly extract "https://example.com/article" -o article.json
Options
| Option | Description |
|---|---|
--query | Rerank chunks by relevance to this query |
--chunks-per-source | Chunks per URL (1-5, requires --query) |
--extract-depth | basic (default) or advanced (for JS pages) |
--format | markdown (default) or text |
--include-images | Include image URLs |
--timeout | Max wait time (1-60 seconds) |
-o, --output | Save the JSON response to a file |
--json | Structured JSON output |
Extract depth
| Depth | When to use |
|---|---|
basic | Simple pages, fast — try this first |
advanced | JS-rendered SPAs, dynamic content, tables |
Tips
- Max 20 URLs per request — batch larger lists into multiple calls.
- Use
--query+--chunks-per-sourceto get only relevant content instead of full pages. - Try
basicfirst, fall back toadvancedif content is missing. - Set
--timeoutfor slow pages (up to 60s). - Inspect
failed_resultseven after exit code 0. A successful request can still return no extracted pages. Retry the affected URL withadvancedwhen appropriate, otherwise report the per-URL failure instead of treating the request as complete. - If search results already contain the content you need (via
--include-raw-content), skip the extract step.
See also
- tavily-search — find pages when you don't have a URL
- tavily-crawl — extract content from many pages on a site