z-ai logo
z-ai

Z.ai: GLM 4.6

Descrizione della fonte (inglese)

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

Panoramica

Specifiche del modello

Contesto
204.800 tokens
Output massimo
16.384 tokens
Architettura
text->text
Tokenizer
Other
Limite di conoscenza
2025-03-31
Moderato
No
OPENROUTER

Prezzi completi

Tariffe sincronizzate da OpenRouter, per milione di token.

Input
$0.43
per milione di token
Output
$1.75
per milione di token
Lettura cache
$0.08
per milione di token
API

Configurazione del modello

Parametri predefiniti

temperature
0.6

Ragionamento

Ragionamento obbligatorio
No
Parametri predefiniti
No
API

Avvio rapido

Imposta OPENROUTER_API_KEY localmente. Python richiede requests; JavaScript viene eseguito in Node.js. Conserva la chiave sul server.

Documentazione API

curl --fail-with-body https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"z-ai/glm-4.6","messages":[{"role":"user","content":"Hello!"}]}'
Capacità e modalità

InputOutput

Input
text
Output
text
Ragionamento

No

API

Parametri API supportati

frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
Provider disponibili

4 provider

Verificato: 20 settembre 2026

Provider live su OpenRouter

Disponibilità, latenza, throughput e routing cambiano continuamente. Consulta la fonte per i dati correnti.

OpenRouter

Venice

fp4
Contesto
198K
Output massimo
16K
Input
$0.43
Output
$1.75
Lettura cache
$0.08
Scrittura cache
Non indicato

DeepInfra

fp4
Contesto
203K
Output massimo
131K
Input
$0.5
Output
$2
Lettura cache
$0.1
Scrittura cache
Non indicato

Novita

bf16
Contesto
205K
Output massimo
131K
Input
$0.55
Output
$2.2
Lettura cache
$0.11
Scrittura cache
Non indicato

Z.AI

fp4
Contesto
203K
Output massimo
131K
Input
$0.6
Output
$2.2
Lettura cache
$0.11
Scrittura cache
Non indicato
z-ai

4 modelli

Tutti i modelli
z-ai logoz-ai

Z.ai: GLM 5.3 FlashX

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Contesto
1.0M
Input
text · image · video
Output
text
Input: $0.37Output: $1.25per milione di token
Vedi i dettagli del modello
z-ai logoz-ai

Z.ai: GLM 5.3 Flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Contesto
1.3M
Input
text · image · video
Output
text
Input: $0.15Output: $0.5per milione di token
Vedi i dettagli del modello
z-ai logoz-ai

Z.ai: GLM 5.3 Flash (batch)

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Contesto
1.0M
Input
text · image · video
Output
text
Input: $0.06Output: $0.2per milione di token
Vedi i dettagli del modello
z-ai logoz-ai

Z.ai: GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

Contesto
1.3M
Input
text
Output
text
Input: $0.5614Output: $1.7644per milione di token
Vedi i dettagli del modello