tavily-ai/skills

tavily-extract

Extract clean markdown or text content from specific URLs via the Tavily CLI.

查看源码
仓库原始内容

按源仓库内容呈现,保留标题、案例、代码、表格、链接以及原文引用的演示图片。

tavily extract

Extract clean markdown or text content from one or more URLs.

Before running

Run extract directly when tvly is available. Extract supports capped keyless access, so do not look for an API key or authenticate before the first request.

If tvly is missing, follow the tavily-cli setup before retrying. If the keyless cap is reached in an interactive session, run tvly login to open browser OAuth, then retry the original extraction once. In an unattended environment, report the cap and authentication options instead of starting an interactive flow. Do not start a second login immediately after guided setup has completed.

When to use

  • You have a specific URL and want its content
  • You need text from JavaScript-rendered pages
  • Step 2 in the workflow: search → extract → map → crawl → research

Quick start

bash
# Single URL
tvly extract "https://example.com/article" --json

# Multiple URLs
tvly extract "https://example.com/page1" "https://example.com/page2" --json

# Query-focused extraction (returns relevant chunks only)
tvly extract "https://example.com/docs" --query "authentication API" --chunks-per-source 3 --json

# JS-heavy pages
tvly extract "https://app.example.com" --extract-depth advanced --json

# Save to file
tvly extract "https://example.com/article" -o article.json

Options

OptionDescription
--queryRerank chunks by relevance to this query
--chunks-per-sourceChunks per URL (1-5, requires --query)
--extract-depthbasic (default) or advanced (for JS pages)
--formatmarkdown (default) or text
--include-imagesInclude image URLs
--timeoutMax wait time (1-60 seconds)
-o, --outputSave the JSON response to a file
--jsonStructured JSON output

Extract depth

DepthWhen to use
basicSimple pages, fast — try this first
advancedJS-rendered SPAs, dynamic content, tables

Tips

  • Max 20 URLs per request — batch larger lists into multiple calls.
  • Use `--query` + `--chunks-per-source` to get only relevant content instead of full pages.
  • Try `basic` first, fall back to advanced if content is missing.
  • Set `--timeout` for slow pages (up to 60s).
  • Inspect `failed_results` even after exit code 0. A successful request can

still return no extracted pages. Retry the affected URL with advanced when appropriate, otherwise report the per-URL failure instead of treating the request as complete.

  • If search results already contain the content you need (via --include-raw-content), skip the extract step.

See also

来自同一仓库

更多 Skills

全部 Skills
tavily-ai
官方

tavily-search

Search the web with LLM-optimized results via the Tavily CLI. Use this skill when the user wants to search the web, find articles, look up information, get recent news, discover sources, or says "search for", "find me", "look up", "what's the latest on", "find articles about", or needs current information from the internet. Returns relevant results with content snippets, relevance scores, and metadata — optimized for LLM consumption. Supports domain filtering, time ranges, and multiple search depths.

安装量
9
GitHub Stars
476
最近更新
9月4日
tavily-ai
官方

tavily-best-practices

Build production-ready Tavily integrations with best practices baked in. Reference documentation for developers using coding assistants (Claude Code, Cursor, etc.) to implement web search, content extraction, crawling, and research in agentic workflows, RAG systems, or autonomous agents.

安装量
8
GitHub Stars
476
最近更新
9月4日
tavily-ai
官方

tavily-cli

Set up, authenticate, update, troubleshoot, or choose between Tavily CLI web commands. Use when the user asks about the Tavily CLI, installing Tavily skills, first-time setup, authentication, keyless limits, CLI updates, or which Tavily command to use. For an ordinary web task, use the specific search, extract, map, crawl, research, or dynamic-search skill instead.

安装量
8
GitHub Stars
476
最近更新
9月4日
tavily-ai
官方

tavily-crawl

Crawl websites and extract content from multiple pages via the Tavily CLI. Use this skill when the user wants to crawl a site, download documentation, extract an entire docs section, bulk-extract pages, save a site as local markdown files, or says "crawl", "get all the pages", "download the docs", "extract everything under /docs", "bulk extract", or needs content from many pages on the same domain. Supports depth/breadth control, path filtering, semantic instructions, and saving each page as a local markdown file.

安装量
8
GitHub Stars
476
最近更新
9月4日