nvidia logo
nvidia

NVIDIA: Nemotron 3.5 Lightning

Description de la source (anglais)

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Présentation

Caractéristiques du modèle

Contexte
262 144 tokens
Sortie maximale
131 072 tokens
Architecture
text->text
Tokenizer
Other
Limite des connaissances
Non indiqué
Modéré
Non
OPENROUTER

Tarification complète

Tarifs synchronisés depuis OpenRouter, par million de tokens.

Entrée
$0.08
par million de tokens
Sortie
$0.2
par million de tokens
Lecture du cache
$0.04
par million de tokens
API

Configuration du modèle

Raisonnement

Raisonnement obligatoire
Non
Paramètres par défaut
Non
API

Démarrage rapide

Définissez OPENROUTER_API_KEY localement. Python nécessite requests ; JavaScript s’exécute dans Node.js. Gardez la clé côté serveur.

Documentation de l’API

curl --fail-with-body https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nvidia/nemotron-3.5-lightning","messages":[{"role":"user","content":"Hello!"}]}'
Capacités et modalités

EntréeSortie

Entrée
text
Sortie
text
Raisonnement

Non

API

Paramètres API pris en charge

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Fournisseurs disponibles

4 fournisseurs

Vérifié: 23 septembre 2026

Fournisseurs en direct sur OpenRouter

Disponibilité, latence, débit et routage évoluent en continu. Consultez la source pour les données en temps réel.

OpenRouter

Darkbloom

int4
Contexte
262K
Sortie maximale
33K
Entrée
$0.065
Sortie
$0.18
Lecture du cache
Non indiqué
Écriture du cache
Non indiqué

Phala

Non indiqué
Contexte
262K
Sortie maximale
236K
Entrée
$0.07
Sortie
$0.2
Lecture du cache
$0.04
Écriture du cache
Non indiqué

CoreWeave

bf16
Contexte
262K
Sortie maximale
236K
Entrée
$0.07
Sortie
$0.2
Lecture du cache
$0.04
Écriture du cache
Non indiqué

DeepInfra

bf16
Contexte
262K
Sortie maximale
131K
Entrée
$0.08
Sortie
$0.2
Lecture du cache
$0.04
Écriture du cache
Non indiqué
nvidia

4 modèles

Tous les modèles
nvidia logonvidia

NVIDIA: Nemotron 3.5 ASR Streaming Multilingual 0.6B

Nemotron 3.5 ASR Streaming Multilingual 0.6B is a speech recognition model from NVIDIA. Its prompt-conditioned, cache-aware FastConformer-RNNT design targets low-latency transcription across more than 40 languages for real-time captioning, voice...

Contexte
Non indiqué
Entrée
audio
Sortie
transcription
Durée audio: $0.000003 par seconde
Voir la fiche du modèle
nvidia logonvidia

NVIDIA: Nemotron 3.5 Lightning (free)

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Contexte
1M
Entrée
text
Sortie
text
Entrée: GratuitSortie: Gratuitpar million de tokens
Voir la fiche du modèle
nvidia logonvidia

NVIDIA: Nemotron 3 Embed 1B (free)

NVIDIA Nemotron 3 Embed 1B is an open text embedding model from NVIDIA, optimized for high-throughput, low-latency retrieval. It is suited for enterprise search, RAG, code retrieval, and agentic retrieval...

Contexte
33K
Entrée
text
Sortie
embeddings
Entrée: Gratuitpar million de tokens
Voir la fiche du modèle
nvidia logonvidia

NVIDIA: Llama Nemotron Rerank VL 1B V2 (free)

Llama Nemotron Rerank VL 1B V2 is a 1.7B multimodal reranking model from NVIDIA. It evaluates the relevance of document images and text against user queries, designed for vision RAG...

Contexte
10K
Entrée
text · image
Sortie
rerank
Entrée: Gratuit par million de tokensSortie: Gratuit par million de tokens
Voir la fiche du modèle