exa-labs/agent-skills

exa-contents

Call Exa Contents directly with cURL or raw HTTP.

Quelltext ansehen
Originales Skill-Dokument

Aus dem Quell-Repository gerendert; Überschriften, Beispiele, Code, Tabellen, Links und Bilder bleiben erhalten.

Exa Contents

Requires API key: Get one at https://dashboard.exa.ai/api-keys Header: x-api-key: $EXA_API_KEY

Use POST https://api.exa.ai/contents when the agent already knows the URLs and needs clean, LLM-ready extraction without running a new search. Start with one content mode: highlights for compact agent context, text for broad page context, or summary for Exa-side compression.

Quick Start (cURL)

Basic text extraction

bash
curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://example.com"],
    "text": true
  }'

Highlights with freshness control

bash
curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://arxiv.org/abs/2307.06435"],
    "highlights": {
      "query": "methodology and results"
    },
    "maxAgeHours": 24,
    "livecrawlTimeout": 12000
  }'

Endpoint

text
POST https://api.exa.ai/contents

Authentication: x-api-key: <API_KEY> header. Exa also accepts Authorization: Bearer <API_KEY>, but prefer x-api-key in cURL examples for consistency.

Use this endpoint for known-URL extraction. If the agent needs discovery or ranking first, use POST /search.

Parameters

Core request parameters

ParameterTypeRequiredDefaultDescription
urlsstring[]Yes-URLs to extract content from. Use this for known URLs.
textboolean or objectNo-Return full page text as markdown. Object form supports maxCharacters, includeHtmlTags, verbosity, includeSections, and excludeSections.
highlightsboolean or objectNo-Return key excerpts. Prefer true for agent workflows unless a custom focus is needed.
summaryboolean or objectNo-Return per-page LLM summaries. Use when the caller wants Exa-side compression or structured extraction.
maxAgeHoursintegerNo-Freshness control. 0 always live crawls; -1 uses cache only; omit for default cache-first behavior with crawl fallback.
livecrawlTimeoutintegerNo10000Timeout for live crawling in milliseconds. Use 10000 to 15000 for most freshness-sensitive calls.
subpagesintegerNo0Number of linked subpages to crawl from each URL.
subpageTargetstring or string[]No-Terms used to prioritize which subpages matter, such as ["api", "reference", "pricing"].
extras.linksintegerNo0Number of links to extract from each page.
extras.imageLinksintegerNo0Number of image URLs to extract from each page.
compliancestringNo-Enterprise-only compliance mode, such as hipaa, when enabled for the account.

Text object options

ParameterTypeDefaultDescription
maxCharactersinteger-Character limit for returned text. Use this instead of tokensNum.
includeHtmlTagsbooleanfalsePreserve HTML tags in output.
verbositystringcompactcompact, standard, or full. Pair fresh section-aware extraction with maxAgeHours: 0.
includeSectionsstring[]-Only include selected sections: header, navigation, banner, body, sidebar, footer, metadata.
excludeSectionsstring[]-Exclude selected sections from the same section list.

Highlights object options

Prefer highlights: true for the highest-quality default. Only use object form when the agent needs a custom focus or budget.

ParameterTypeDefaultDescription
querystring-Custom query guiding which excerpts are returned.
maxCharactersinteger-Cap highlight characters per URL. Omit unless the caller has a strict budget.

Summary object options

ParameterTypeDefaultDescription
querystring-Custom query for the summary.
schemaobject-JSON Schema for structured per-page summaries.

Content Modes

On /contents, text, highlights, and summary are top-level request fields.

ModeBest forNotes
textDeep analysis and broad page contextUse maxCharacters to keep payloads bounded.
highlightsAgent workflows and factual lookupsMost token-efficient default. Excerpts are grounded in the source page.
summaryCompression or structured per-page extractionAdds Exa-side synthesis per page.

Avoid requesting multiple modes unless the caller truly needs multiple views of the same page.

Freshness and Crawling

Use maxAgeHours as the normative freshness control.

ValueBehavior
omittedUse default cache-first behavior with crawl fallback when needed.
positive integerUse cache if it is less than N hours old, otherwise live crawl.
0Always live crawl. Highest freshness, higher latency.
-1Cache only. Fastest, but fails if no cached content exists.

Set livecrawlTimeout when live crawling should not block past a fixed budget.

Subpages and extras

Use subpages and subpageTarget when linked pages matter.

bash
curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://docs.example.com"],
    "text": {
      "maxCharacters": 5000
    },
    "subpages": 10,
    "subpageTarget": ["api", "reference", "guide"],
    "extras": {
      "links": 10,
      "imageLinks": 5
    }
  }'

Start with subpages around 5 to 10, then increase only when the caller needs broader site coverage.

Response Fields and Statuses

Inspect statuses even when the HTTP status is 200. The endpoint can succeed for one URL and fail for another in the same request.

bash
curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://example.com", "https://failed-url.example.com"],
    "highlights": true
  }' | jq '{results, statuses}'
FieldTypeDescription
requestIdstringUnique request identifier.
resultsarrayExtracted content result objects.
results[].titlestringPage title.
results[].urlstringPage URL.
results[].publishedDatestring or nullEstimated publication date when available.
results[].authorstring or nullAuthor when available.
results[].textstringReturned when text is requested.
results[].highlightsstring[]Returned when highlights is requested.
results[].highlightScoresnumber[]Similarity scores for highlights.
results[].summarystringReturned when summary is requested.
results[].subpagesarrayNested result objects from subpage crawling.
results[].extras.linksstring[]Extracted links when requested.
statusesarrayPer-URL success or error states. Always inspect this field.
statuses[].idstringRequested URL.
statuses[].statusstringsuccess or error.
statuses[].error.tagstringError type for failed URLs.
statuses[].error.httpStatusCodeinteger or nullHTTP code associated with a per-URL failure.
costDollars.totalnumberTotal request cost when returned.

Common per-URL error tags include CRAWL_NOT_FOUND, CRAWL_TIMEOUT, CRAWL_LIVECRAWL_TIMEOUT, SOURCE_NOT_AVAILABLE, UNSUPPORTED_URL, and CRAWL_UNKNOWN_ERROR.

Critical Pitfalls

  • Keep text, highlights, and summary at the top level on /contents.
  • Do not wrap extraction options in a contents object; that nesting belongs to /search.
  • Do not assume HTTP 200 means every URL succeeded; inspect statuses.
  • Do not send stream: true; /contents is not a streaming endpoint.
  • Do not send tokensNum; use text.maxCharacters to cap extracted text.
  • Do not use useAutoprompt, numSentences, highlightsPerUrl, or older livecrawl string values in new requests.
  • Prefer maxAgeHours for freshness and pair it with livecrawlTimeout when crawl latency matters.
  • Use subpageTarget with subpages; otherwise subpage selection is best effort.
  • Pick one of highlights, text, or summary by default. Stack modes only when the caller truly needs multiple views of each page.
aus demselben Repository

Weitere Skills

Alle Skills
exa-labs
Community

build-with-exa

Build applications and agents with Exa's API: search, contents extraction, answer, Agent API, monitors, websets, OpenAI-compatible endpoints, and exa-py/exa-js SDKs. Use when choosing Exa endpoints, writing Exa API calls, integrating semantic web search or research into products, or debugging Exa request shapes.

Installationen
1
GitHub Stars
45
Aktualisiert
3. Sept.
exa-labs
Community

company-research

Company research using Exa. Finds company info, competitors, news, financials, LinkedIn profiles, builds company lists. Use when researching companies, doing competitor analysis, market research, or building company lists.

Installationen
1
GitHub Stars
45
Aktualisiert
3. Sept.
exa-labs
Community

exa-search

Call Exa Search directly with cURL or raw HTTP. Use when an agent needs Exa semantic web retrieval from POST /search without an SDK, including ranked results, domain or category filters, freshness-aware result content, highlights or text extraction, structured output, or streaming search responses.

Installationen
1
GitHub Stars
45
Aktualisiert
3. Sept.
exa-labs
Community

lead-generation

Generate enriched lead lists using Exa Agent. Finds companies matching an ICP, enriches with signals/news/scores, and outputs CSV. Use when generating leads, building prospect lists, finding companies to sell to, doing outbound research, or ICP-based company discovery. Triggers on "leads", "lead gen", "prospect list", "find companies", "ICP", "outbound list".

Installationen
1
GitHub Stars
45
Aktualisiert
3. Sept.