ZA
z-ai

Z.ai: GLM 5.3 Flash (batch)

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Présentation

Caractéristiques du modèle

Context
1 048 575 tokens
Sortie maximale
943 717 tokens
Architecture
text+image+video->text
Tokenizer
Other
Limite des connaissances
Not provided
Modéré
Non
Capacités et modalités

InputOutput

Input
textimagevideo
Output
text
Raisonnement

Oui · max · high · low

API

Paramètres API pris en charge

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Fournisseurs disponibles

21 fournisseurs

Relace

fp4
Context
1.0M
Sortie maximale
131K
Input
$0.050
Output
$0.167
Lecture du cache
$0.010
Écriture du cache
Not provided

Z.AI

fp8
Context
1.0M
Sortie maximale
131K
Input
$0.075
Output
$0.250
Lecture du cache
$0.015
Écriture du cache
Not provided

Novita

fp8
Context
1.0M
Sortie maximale
131K
Input
$0.075
Output
$0.250
Lecture du cache
$0.015
Écriture du cache
Not provided

DeepInfra

fp8
Context
1.0M
Sortie maximale
944K
Input
$0.075
Output
$0.250
Lecture du cache
$0.015
Écriture du cache
Not provided

GMICloud

fp8
Context
1.0M
Sortie maximale
944K
Input
$0.075
Output
$0.250
Lecture du cache
$0.015
Écriture du cache
Not provided

Makora

unknown
Context
1.0M
Sortie maximale
944K
Input
$0.140
Output
$0.470
Lecture du cache
$0.024
Écriture du cache
Not provided

Modal

fp8
Context
1.0M
Sortie maximale
944K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

Morph

fp8
Context
1.0M
Sortie maximale
944K
Input
$0.150
Output
$0.420
Lecture du cache
$0.010
Écriture du cache
Not provided

Parasail

fp8
Context
1.0M
Sortie maximale
944K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

Reka

fp8
Context
262K
Sortie maximale
236K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

Together

unknown
Context
1.0M
Sortie maximale
944K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

Wafer

unknown
Context
1.0M
Sortie maximale
944K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

DigitalOcean

unknown
Context
1.0M
Sortie maximale
944K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

SiliconFlow

fp8
Context
1.0M
Sortie maximale
262K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

Friendli

unknown
Context
1.0M
Sortie maximale
944K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

Phala

fp8
Context
1.0M
Sortie maximale
131K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

Fireworks

unknown
Context
1.0M
Sortie maximale
944K
Input
$0.150
Output
$0.500
Lecture du cache
$0.029
Écriture du cache
Not provided

Cloudflare

unknown
Context
1.3M
Sortie maximale
1.2M
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

Io Net

fp8
Context
262K
Sortie maximale
131K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

BaseTen

fp8
Context
1.0M
Sortie maximale
131K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided

Venice

unknown
Context
1.0M
Sortie maximale
131K
Input
$0.150
Output
$0.500
Lecture du cache
$0.030
Écriture du cache
Not provided