N
nvidia

NVIDIA: Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Overview

Model specifications

Context
262,144 tokens
Maximum output
131,072 tokens
Architecture
text->text
Tokenizer
Other
Knowledge cutoff
Not provided
Moderated
No
Capabilities

InputOutput

Input
text
Output
text
Reasoning

No

API

Supported API parameters

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Available providers

2 providers

DeepInfra

bf16
Context
262K
Maximum output
131K
Input
$0.080
Output
$0.200
Cache read
$0.040
Cache write
Not provided

CoreWeave

bf16
Context
262K
Maximum output
236K
Input
$0.100
Output
$0.250
Cache read
$0.050
Cache write
Not provided
nvidia

models

All models