N
nvidia

NVIDIA: Nemotron 3.5 Lightning (free)

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Overview

Model specifications

Context
1,000,000 tokens
Maximum output
65,536 tokens
Architecture
text->text
Tokenizer
Other
Knowledge cutoff
Not provided
Moderated
No
Capabilities

InputOutput

Input
text
Output
text
Reasoning

No

API

Supported API parameters

include_reasoningmax_tokensreasoningseedtemperaturetool_choicetoolstop_p
Available providers

2 providers

DeepInfra

bf16
Context
262K
Maximum output
131K
Input
$0.080
Output
$0.200
Cache read
$0.040
Cache write
Not provided

CoreWeave

bf16
Context
262K
Maximum output
236K
Input
$0.100
Output
$0.250
Cache read
$0.050
Cache write
Not provided
nvidia

models

All models