nvidia logo
nvidia

NVIDIA: Nemotron 3 Ultra (free)

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). This model is free to use. 1,000,000 token context window, maximum output of 65,536 tokens. Higher uptime with 5 providers. Includes independent benchmarks from Artificial Analysis.

Überblick

Modellspezifikationen

Kontext
1.000.000 tokens
Maximale Ausgabe
65.536 tokens
Architektur
text->text
Tokenizer
Other
Wissensstand
Nicht angegeben
Moderiert
Nein
OPENROUTER

Vollständige Preise

Mit OpenRouter synchronisierte Preise; Tokenpreise gelten pro eine Million Token.

Eingabe
Kostenlos
/M tokens
Ausgabe
Kostenlos
/M tokens
API

Modellkonfiguration

Standardparameter

temperature
1
top_p
0.95

Schlussfolgern

Moderiert
Nein
Standardparameter
high
Fähigkeiten und Modalitäten
high, medium
API

Schnellstart

Dieses Modell über die OpenAI-kompatible API von OpenRouter aufrufen.

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nvidia/nemotron-3-ultra-550b-a55b:free","messages":[{"role":"user","content":"Hello!"}]}'
Fähigkeiten und Modalitäten

EingabeAusgabe

Eingabe
text
Ausgabe
text
Schlussfolgern

Ja · high · medium

API

Unterstützte API-Parameter

include_reasoningmax_tokensreasoningreasoning_effortseedtemperaturetool_choicetoolstop_p
Verfügbare Anbieter

3 Anbieter

Live-Anbieter bei OpenRouter

Verfügbarkeit, Latenz, Durchsatz und Routing ändern sich laufend. Aktuelle Betriebsdaten stehen auf der Quellseite.

OpenRouter

DeepInfra

fp4
Kontext
262K
Maximale Ausgabe
16K
Eingabe
$0.500
Ausgabe
$2.20
Cache-Lesen
$0.100
Cache-Schreiben
Nicht angegeben

BaseTen

fp4
Kontext
203K
Maximale Ausgabe
183K
Eingabe
$0.600
Ausgabe
$2.40
Cache-Lesen
$0.120
Cache-Schreiben
Nicht angegeben

Venice

fp8
Kontext
256K
Maximale Ausgabe
33K
Eingabe
$0.625
Ausgabe
$3.13
Cache-Lesen
$0.188
Cache-Schreiben
Nicht angegeben
nvidia

Modelle

Alle Modelle
nvidia logonvidia

NVIDIA: Nemotron 3.5 ASR Streaming Multilingual 0.6B

Nemotron 3.5 ASR Streaming Multilingual 0.6B is a speech recognition model from NVIDIA. Its prompt-conditioned, cache-aware FastConformer-RNNT design targets low-latency transcription across more than 40 languages for real-time captioning, voice...

Kontext
Nicht angegeben
Eingabe
audio
Ausgabe
transcription
Eingabe: $3.33Ausgabe: Kostenlospro 1 Mio. Token
Modelldetails ansehen
nvidia logonvidia

NVIDIA: Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Kontext
262K
Eingabe
text
Ausgabe
text
Eingabe: $0.080Ausgabe: $0.200pro 1 Mio. Token
Modelldetails ansehen
nvidia logonvidia

NVIDIA: Nemotron 3.5 Lightning (free)

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Kontext
1M
Eingabe
text
Ausgabe
text
Eingabe: KostenlosAusgabe: Kostenlospro 1 Mio. Token
Modelldetails ansehen
nvidia logonvidia

NVIDIA: Nemotron 3 Embed 1B (free)

NVIDIA Nemotron 3 Embed 1B is an open text embedding model from NVIDIA, optimized for high-throughput, low-latency retrieval. It is suited for enterprise search, RAG, code retrieval, and agentic retrieval...

Kontext
33K
Eingabe
text
Ausgabe
embeddings
Eingabe: KostenlosAusgabe: Kostenlospro 1 Mio. Token
Modelldetails ansehen