nvidia logo
nvidia

NVIDIA: Nemotron 3 Ultra (free)

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). This model is free to use. 1,000,000 token context window, maximum output of 65,536 tokens. Higher uptime with 5 providers. Includes independent benchmarks from Artificial Analysis.

Présentation

Caractéristiques du modèle

Contexte
1 000 000 tokens
Sortie maximale
65 536 tokens
Architecture
text->text
Tokenizer
Other
Limite des connaissances
Non indiqué
Modéré
Non
OPENROUTER

Tarification complète

Tarifs synchronisés depuis OpenRouter, par million de tokens.

Entrée
Gratuit
/M tokens
Sortie
Gratuit
/M tokens
API

Configuration du modèle

Paramètres par défaut

temperature
1
top_p
0.95

Raisonnement

Modéré
Non
Paramètres par défaut
high
Capacités et modalités
high, medium
API

Démarrage rapide

Appelez ce modèle via l’API compatible OpenAI d’OpenRouter.

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nvidia/nemotron-3-ultra-550b-a55b:free","messages":[{"role":"user","content":"Hello!"}]}'
Capacités et modalités

EntréeSortie

Entrée
text
Sortie
text
Raisonnement

Oui · high · medium

API

Paramètres API pris en charge

include_reasoningmax_tokensreasoningreasoning_effortseedtemperaturetool_choicetoolstop_p
Fournisseurs disponibles

3 fournisseurs

Fournisseurs en direct sur OpenRouter

Disponibilité, latence, débit et routage évoluent en continu. Consultez la source pour les données en temps réel.

OpenRouter

DeepInfra

fp4
Contexte
262K
Sortie maximale
16K
Entrée
$0.500
Sortie
$2.20
Lecture du cache
$0.100
Écriture du cache
Non indiqué

BaseTen

fp4
Contexte
203K
Sortie maximale
183K
Entrée
$0.600
Sortie
$2.40
Lecture du cache
$0.120
Écriture du cache
Non indiqué

Venice

fp8
Contexte
256K
Sortie maximale
33K
Entrée
$0.625
Sortie
$3.13
Lecture du cache
$0.188
Écriture du cache
Non indiqué
nvidia

modèles

Tous les modèles
nvidia logonvidia

NVIDIA: Nemotron 3.5 ASR Streaming Multilingual 0.6B

Nemotron 3.5 ASR Streaming Multilingual 0.6B is a speech recognition model from NVIDIA. Its prompt-conditioned, cache-aware FastConformer-RNNT design targets low-latency transcription across more than 40 languages for real-time captioning, voice...

Contexte
Non indiqué
Entrée
audio
Sortie
transcription
Entrée: $3.33Sortie: Gratuitpar million de tokens
Voir la fiche du modèle
nvidia logonvidia

NVIDIA: Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Contexte
262K
Entrée
text
Sortie
text
Entrée: $0.080Sortie: $0.200par million de tokens
Voir la fiche du modèle
nvidia logonvidia

NVIDIA: Nemotron 3.5 Lightning (free)

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Contexte
1M
Entrée
text
Sortie
text
Entrée: GratuitSortie: Gratuitpar million de tokens
Voir la fiche du modèle
nvidia logonvidia

NVIDIA: Nemotron 3 Embed 1B (free)

NVIDIA Nemotron 3 Embed 1B is an open text embedding model from NVIDIA, optimized for high-throughput, low-latency retrieval. It is suited for enterprise search, RAG, code retrieval, and agentic retrieval...

Contexte
33K
Entrée
text
Sortie
embeddings
Entrée: GratuitSortie: Gratuitpar million de tokens
Voir la fiche du modèle