nvidia logo
nvidia

NVIDIA: Nemotron 3 Ultra (free)

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). This model is free to use. 1,000,000 token context window, maximum output of 65,536 tokens. Higher uptime with 5 providers. Includes independent benchmarks from Artificial Analysis.

Resumen

Especificaciones del modelo

Contexto
1.000.000 tokens
Salida máxima
65.536 tokens
Arquitectura
text->text
Tokenizador
Other
Corte de conocimiento
No indicado
Moderado
No
OPENROUTER

Precios completos

Tarifas sincronizadas desde OpenRouter, por millón de tokens.

Entrada
Gratis
/M tokens
Salida
Gratis
/M tokens
API

Configuración del modelo

Parámetros predeterminados

temperature
1
top_p
0.95

Razonamiento

Moderado
No
Parámetros predeterminados
high
Capacidades y modalidades
high, medium
API

Inicio rápido

Usa este modelo mediante la API compatible con OpenAI de OpenRouter.

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nvidia/nemotron-3-ultra-550b-a55b:free","messages":[{"role":"user","content":"Hello!"}]}'
Capacidades y modalidades

EntradaSalida

Entrada
text
Salida
text
Razonamiento

· high · medium

API

Parámetros API compatibles

include_reasoningmax_tokensreasoningreasoning_effortseedtemperaturetool_choicetoolstop_p
Proveedores disponibles

3 proveedores

Proveedores en vivo en OpenRouter

Disponibilidad, latencia, rendimiento y enrutamiento cambian continuamente. Consulta la fuente para datos actuales.

OpenRouter

DeepInfra

fp4
Contexto
262K
Salida máxima
16K
Entrada
$0.500
Salida
$2.20
Lectura de caché
$0.100
Escritura de caché
No indicado

BaseTen

fp4
Contexto
203K
Salida máxima
183K
Entrada
$0.600
Salida
$2.40
Lectura de caché
$0.120
Escritura de caché
No indicado

Venice

fp8
Contexto
256K
Salida máxima
33K
Entrada
$0.625
Salida
$3.13
Lectura de caché
$0.188
Escritura de caché
No indicado
nvidia

modelos

Todos los modelos
nvidia logonvidia

NVIDIA: Nemotron 3.5 ASR Streaming Multilingual 0.6B

Nemotron 3.5 ASR Streaming Multilingual 0.6B is a speech recognition model from NVIDIA. Its prompt-conditioned, cache-aware FastConformer-RNNT design targets low-latency transcription across more than 40 languages for real-time captioning, voice...

Contexto
No indicado
Entrada
audio
Salida
transcription
Entrada: $3.33Salida: Gratispor millón de tokens
Ver detalles del modelo
nvidia logonvidia

NVIDIA: Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Contexto
262K
Entrada
text
Salida
text
Entrada: $0.080Salida: $0.200por millón de tokens
Ver detalles del modelo
nvidia logonvidia

NVIDIA: Nemotron 3.5 Lightning (free)

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Contexto
1M
Entrada
text
Salida
text
Entrada: GratisSalida: Gratispor millón de tokens
Ver detalles del modelo
nvidia logonvidia

NVIDIA: Nemotron 3 Embed 1B (free)

NVIDIA Nemotron 3 Embed 1B is an open text embedding model from NVIDIA, optimized for high-throughput, low-latency retrieval. It is suited for enterprise search, RAG, code retrieval, and agentic retrieval...

Contexto
33K
Entrada
text
Salida
embeddings
Entrada: GratisSalida: Gratispor millón de tokens
Ver detalles del modelo