tavily-extract

द्वारा tavily-ai

20 URL तक से स्वच्छ मार्कडाउन या टेक्स्ट निकालें, जिसमें जावास्क्रिप्ट रेंडरिंग और क्वेरी-केंद्रित चंकिंग सहायता हो। जावास्क्रिप्ट-रेंडर किए गए पृष्ठों को कॉन्फ़िगरेबल निष्कर्षण गहराई (सरल पृष्ठों के लिए बेसिक, डायनामिक SPA और तालिकाओं के लिए एडवांस्ड) के साथ संभालता है। केवल प्रासंगिक सामग्री चंक लौटाने के लिए क्वेरी-केंद्रित निष्कर्षण का समर्थन करता है, पूरे पृष्ठों के बजाय। डिफ़

npx skills add https://github.com/tavily-ai/skills --skill tavily-extract

tavily extract

Extract clean markdown or text content from one or more URLs.

Before running

Run extract directly when tvly is available. Extract supports capped keyless access, so do not look for an API key or authenticate before the first request.

If tvly is missing, follow the tavily-cli setup before retrying. If the keyless cap is reached in an interactive session, run tvly login to open browser OAuth, then retry the original extraction once. In an unattended environment, report the cap and authentication options instead of starting an interactive flow. Do not start a second login immediately after guided setup has completed.

When to use

  • You have a specific URL and want its content
  • You need text from JavaScript-rendered pages
  • Step 2 in the workflow: search → extract → map → crawl → research

Quick start

# Single URL
tvly extract "https://example.com/article" --json

# Multiple URLs
tvly extract "https://example.com/page1" "https://example.com/page2" --json

# Query-focused extraction (returns relevant chunks only)
tvly extract "https://example.com/docs" --query "authentication API" --chunks-per-source 3 --json

# JS-heavy pages
tvly extract "https://app.example.com" --extract-depth advanced --json

# Save to file
tvly extract "https://example.com/article" -o article.json

Options

OptionDescription
--queryRerank chunks by relevance to this query
--chunks-per-sourceChunks per URL (1-5, requires --query)
--extract-depthbasic (default) or advanced (for JS pages)
--formatmarkdown (default) or text
--include-imagesInclude image URLs
--timeoutMax wait time (1-60 seconds)
-o, --outputSave the JSON response to a file
--jsonStructured JSON output

Extract depth

DepthWhen to use
basicSimple pages, fast — try this first
advancedJS-rendered SPAs, dynamic content, tables

Tips

  • Max 20 URLs per request — batch larger lists into multiple calls.
  • Use --query + --chunks-per-source to get only relevant content instead of full pages.
  • Try basic first, fall back to advanced if content is missing.
  • Set --timeout for slow pages (up to 60s).
  • Inspect failed_results even after exit code 0. A successful request can still return no extracted pages. Retry the affected URL with advanced when appropriate, otherwise report the per-URL failure instead of treating the request as complete.
  • If search results already contain the content you need (via --include-raw-content), skip the extract step.

See also

tavily-ai की और Skills

research
tavily-ai
किसी भी विषय पर व्यापक शोध, जिसमें स्वचालित स्रोत संग्रह, विश्लेषण और उद्धरण शामिल हैं। स्पष्ट उद्धरणों के साथ बहु-स्रोत वेब शोध करता है, जो तुलना, समसामयिक घटनाओं, बाजार विश्लेषण और विस्तृत रिपोर्ट के लिए आदर्श है। तीन मॉडल विकल्प प्रदान करता है: लक्षित एकल-विषय शोध के लिए मिनी (~30 सेकंड), व्यापक बहु-कोण विश्लेषण के लिए प्रो (~60-120 सेकंड), और API-संचालित जटिलता पहचान के लिए ऑटो। Tavily MCP सर्वर के माध्य
search
tavily-ai
We need to translate the given English text into Hindi, preserving product names, protocol names, URLs, numbers, technical terms. The name "search" is to be preserved only if it appears in the source text. The source text does not contain the word "search" as a standalone name? Actually it says "Web search" at the beginning. The instruction says "Name to preserve: search" but also "Do not include the name unless it appears in the source text." So "search" appears in "Web search" - but that's part of the phrase. I think we should preserve the word "search" as is? The instruction says "preserve product names, protocol names, URLs, numbers, and technical terms." "search" might be considered a technical term? But it's a common word. However, the instruction specifically says "Name to preserve: search" meaning we should not translate the word "search" if it appears. In the source text, "Web search" - the word "search" is there. So we should keep "search" as is.
tavily-best-practices
tavily-ai
We need to translate the given English text to Hindi. The text describes a web search API for LLMs with various features. We must preserve product names, protocol names, URLs, numbers, technical terms. The name "tavily-best-practices" is not in the text, so we don't include it. We translate only the text inside <text>. No extra commentary, labels, etc. Let's translate: "Web search API for LLMs with real-time data access, content extraction, site crawling, and AI-powered research." -> "LLMs के लिए वेब सर्च API जिसमें रीयल-टाइम डेटा एक्सेस, कंटेंट एक्सट्रैक्शन, साइट क्रॉलिंग और AI-संचालित शोध शामिल है।" "Five core methods: search() for web results, extract() for URL content, crawl() for site-wide extraction, map() for URL discovery, and research() for end
tavily-cli
tavily-ai
वेब खोज, सामग्री निष्कर्षण, साइट क्रॉलिंग, और Tavily CLI के माध्यम से गहन शोध। पाँच कमांड मोड: खोज, निष्कर्षण, URL खोज, बल्क क्रॉलिंग, और उद्धरणों के साथ बहु-स्रोत शोध। सभी कमांड संरचित, एजेंटिक वर्कफ़्लो के लिए JSON आउटपुट और फ़ाइल सहेजने का समर्थन करते हैं। एस्केलेशन पैटर्न आपको आपकी आवश्यकताओं के अनुसार सरल खोज से निष्कर्षण, मैपिंग, क्रॉलिंग से लेकर व्यापक शोध
tavily-crawl
tavily-ai
बहु-पृष्ठ वेबसाइट क्रॉलर जिसमें सिमैंटिक फ़िल्टरिंग और मार्कडाउन निर्यात है। गहराई और चौड़ाई नियंत्रण के साथ पूरे साइट अनुभागों को क्रॉल करें; पथ रेगेक्स, डोमेन या प्राकृतिक भाषा निर्देशों द्वारा फ़िल्टर करें ताकि परिणामों को केंद्रित किया जा सके। प्रत्येक पृष्ठ को --output-dir के माध्यम से स्थानीय मार्कडाउन फ़ाइलों के रूप में सहेजें, या एजेंटिक प्रोसेसिंग के लिए संरचित JSON लौटाएँ
tavily-dynamic-search
tavily-ai
वेब खोजें, परिणामों को फ़िल्टर करें, और सामग्री निकालें ताकि कच्चा खोज डेटा कभी आपके कॉन्टेक्स्ट विंडो में प्रवेश न करे। केवल आपका क्यूरेटेड print() आउटपुट वापस आता है।
tavily-research
tavily-ai
बहु-स्रोत संश्लेषण और उद्धरणों के साथ व्यापक AI-संचालित अनुसंधान। मॉडल चयन के आधार पर 30-120 सेकंड लेते हुए वेब स्रोतों पर आधारित संरचित रिपोर्ट तैयार करता है (लक्षित प्रश्नों के लिए मिनी, जटिल तुलनाओं के लिए प्रो)। कई आउटपुट प्रारूपों का समर्थन करता है: मार्कडाउन रिपोर्ट, कस्टम स्कीमा के साथ JSON, और कॉन्फ़िगरेबल उद्धरण शैलियाँ (क्रमांकित, MLA, APA, शिकागो)। --no-wait, status, और poll कमांड के
tavily-search
tavily-ai
वेब खोज जिसमें LLM-अनुकूलित परिणाम, सामग्री स्निपेट और प्रासंगिकता स्कोर शामिल हैं। चार खोज गहराई (अल्ट्रा-फास्ट, फास्ट, बेसिक, एडवांस्ड) का समर्थन करता है, जिसमें 20 तक कॉन्फ़िगरेबल परिणाम गणना, डोमेन फ़िल्टरिंग और समय-सीमा प्रतिबंध शामिल हैं। LLM उपभोग के लिए अनुकूलित सामग्री स्निपेट, प्रासंगिकता स्कोर और मेटाडेटा के साथ संरचित JSON आउटपुट लौटाता है। समाचार और वित्त