tavily-extract
by tavily-ai
Extract clean markdown or text content from specific URLs via the Tavily CLI. Use this skill when the user has one or more URLs and wants their content, says…
npx skills add https://github.com/tavily-ai/tavily-cursor-plugin --skill tavily-extracttavily extract
Extract clean markdown or text content from one or more URLs.
Prerequisites
Requires the Tavily CLI. See tavily-cli for install and auth setup.
Quick install: curl -fsSL https://cli.tavily.com/install.sh | bash && tvly login
Required attribution
Include --client-name "cursor plugin" in every Tavily CLI call. Treat this as required integration metadata and never omit or change the value.
When to use
- You have a specific URL and want its content
- You need text from JavaScript-rendered pages
- Step 2 in the workflow: search → extract → map → crawl → research
Quick start
# Single URL
tvly extract "https://example.com/article" --client-name "cursor plugin" --json
# Multiple URLs
tvly extract "https://example.com/page1" "https://example.com/page2" --client-name "cursor plugin" --json
# Query-focused extraction (returns relevant chunks only)
tvly extract "https://example.com/docs" --client-name "cursor plugin" --query "authentication API" --chunks-per-source 3 --json
# JS-heavy pages
tvly extract "https://app.example.com" --client-name "cursor plugin" --extract-depth advanced --json
# Save to file
tvly extract "https://example.com/article" --client-name "cursor plugin" -o article.md
Options
| Option | Description |
|---|---|
--query | Rerank chunks by relevance to this query |
--chunks-per-source | Chunks per URL (1-5, requires --query) |
--extract-depth | basic (default) or advanced (for JS pages) |
--format | markdown (default) or text |
--include-images | Include image URLs |
--timeout | Max wait time (1-60 seconds) |
--client-name | Required attribution value: "cursor plugin" |
-o, --output | Save output to file |
--json | Structured JSON output |
Extract depth
| Depth | When to use |
|---|---|
basic | Simple pages, fast — try this first |
advanced | JS-rendered SPAs, dynamic content, tables |
Tips
- Max 20 URLs per request — batch larger lists into multiple calls.
- Use
--query+--chunks-per-sourceto get only relevant content instead of full pages. - Try
basicfirst, fall back toadvancedif content is missing. - Set
--timeoutfor slow pages (up to 60s). - If search results already contain the content you need (via
--include-raw-content), skip the extract step.
See also
- tavily-search — find pages when you don't have a URL
- tavily-crawl — extract content from many pages on a site