ZA
z-ai

Z.ai: GLM 5.3 Flash (batch)

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Overview

Model specifications

Context
1,048,575 tokens
Maximum output
943,717 tokens
Architecture
text+image+video->text
Tokenizer
Other
Knowledge cutoff
Not provided
Moderated
No
Capabilities

InputOutput

Input
textimagevideo
Output
text
Reasoning

Yes · max · high · low

API

Supported API parameters

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Available providers

21 providers

Relace

fp4
Context
1.0M
Maximum output
131K
Input
$0.050
Output
$0.167
Cache read
$0.010
Cache write
Not provided

Z.AI

fp8
Context
1.0M
Maximum output
131K
Input
$0.075
Output
$0.250
Cache read
$0.015
Cache write
Not provided

Novita

fp8
Context
1.0M
Maximum output
131K
Input
$0.075
Output
$0.250
Cache read
$0.015
Cache write
Not provided

DeepInfra

fp8
Context
1.0M
Maximum output
944K
Input
$0.075
Output
$0.250
Cache read
$0.015
Cache write
Not provided

GMICloud

fp8
Context
1.0M
Maximum output
944K
Input
$0.075
Output
$0.250
Cache read
$0.015
Cache write
Not provided

Makora

unknown
Context
1.0M
Maximum output
944K
Input
$0.140
Output
$0.470
Cache read
$0.024
Cache write
Not provided

Modal

fp8
Context
1.0M
Maximum output
944K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

Morph

fp8
Context
1.0M
Maximum output
944K
Input
$0.150
Output
$0.420
Cache read
$0.010
Cache write
Not provided

Parasail

fp8
Context
1.0M
Maximum output
944K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

Reka

fp8
Context
262K
Maximum output
236K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

Together

unknown
Context
1.0M
Maximum output
944K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

Wafer

unknown
Context
1.0M
Maximum output
944K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

DigitalOcean

unknown
Context
1.0M
Maximum output
944K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

SiliconFlow

fp8
Context
1.0M
Maximum output
262K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

Friendli

unknown
Context
1.0M
Maximum output
944K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

Phala

fp8
Context
1.0M
Maximum output
131K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

Fireworks

unknown
Context
1.0M
Maximum output
944K
Input
$0.150
Output
$0.500
Cache read
$0.029
Cache write
Not provided

Cloudflare

unknown
Context
1.3M
Maximum output
1.2M
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

Io Net

fp8
Context
262K
Maximum output
131K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

BaseTen

fp8
Context
1.0M
Maximum output
131K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided

Venice

unknown
Context
1.0M
Maximum output
131K
Input
$0.150
Output
$0.500
Cache read
$0.030
Cache write
Not provided