google logo
google

Google: Gemma 4 31B

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. $0.09 per million input tokens, $0.34 per million output tokens. 262,144 token context window, maximum output of 16,384 tokens. Higher uptime with 15 providers. Includes independent benchmarks from Artificial Analysis.

Überblick

Modellspezifikationen

Kontext
262.144 tokens
Maximale Ausgabe
16.384 tokens
Architektur
text+image+video->text
Tokenizer
Gemma
Wissensstand
Nicht angegeben
Moderiert
Nein
OPENROUTER

Vollständige Preise

Mit OpenRouter synchronisierte Preise; Tokenpreise gelten pro eine Million Token.

Eingabe
$0.09
/M tokens
Ausgabe
$0.34
/M tokens
Cache-Lesen
$0.05
/M tokens
API

Modellkonfiguration

Standardparameter

temperature
1
top_p
0.95
top_k
64

Schlussfolgern

Moderiert
Nein
Standardparameter
Nein
API

Schnellstart

Dieses Modell über die OpenAI-kompatible API von OpenRouter aufrufen.

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"google/gemma-4-31b-it","messages":[{"role":"user","content":"Hello!"}]}'
Fähigkeiten und Modalitäten

EingabeAusgabe

Eingabe
imagetextvideo
Ausgabe
text
Schlussfolgern

Nein

API

Unterstützte API-Parameter

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Verfügbare Anbieter

16 Anbieter

Live-Anbieter bei OpenRouter

Verfügbarkeit, Latenz, Durchsatz und Routing ändern sich laufend. Aktuelle Betriebsdaten stehen auf der Quellseite.

OpenRouter

DeepInfra

fp4
Kontext
262K
Maximale Ausgabe
16K
Eingabe
$0.090
Ausgabe
$0.340
Cache-Lesen
$0.050
Cache-Schreiben
Nicht angegeben

CoreWeave

fp4
Kontext
262K
Maximale Ausgabe
236K
Eingabe
$0.100
Ausgabe
$0.340
Cache-Lesen
$0.100
Cache-Schreiben
Nicht angegeben

Venice

bf16
Kontext
256K
Maximale Ausgabe
8K
Eingabe
$0.120
Ausgabe
$0.360
Cache-Lesen
$0.090
Cache-Schreiben
Nicht angegeben

Chutes

fp4
Kontext
131K
Maximale Ausgabe
66K
Eingabe
$0.120
Ausgabe
$0.370
Cache-Lesen
$0.012
Cache-Schreiben
Nicht angegeben

DeepInfra

fp8
Kontext
262K
Maximale Ausgabe
16K
Eingabe
$0.130
Ausgabe
$0.380
Cache-Lesen
Nicht angegeben
Cache-Schreiben
Nicht angegeben

SiliconFlow

fp8
Kontext
262K
Maximale Ausgabe
236K
Eingabe
$0.130
Ausgabe
$0.400
Cache-Lesen
Nicht angegeben
Cache-Schreiben
Nicht angegeben

Crusoe

unknown
Kontext
262K
Maximale Ausgabe
262K
Eingabe
$0.140
Ausgabe
$0.400
Cache-Lesen
$0.140
Cache-Schreiben
Nicht angegeben

Friendli

unknown
Kontext
262K
Maximale Ausgabe
8K
Eingabe
$0.140
Ausgabe
$0.400
Cache-Lesen
Nicht angegeben
Cache-Schreiben
Nicht angegeben

Novita

bf16
Kontext
262K
Maximale Ausgabe
131K
Eingabe
$0.140
Ausgabe
$0.400
Cache-Lesen
Nicht angegeben
Cache-Schreiben
Nicht angegeben

Parasail

fp8
Kontext
262K
Maximale Ausgabe
236K
Eingabe
$0.150
Ausgabe
$0.400
Cache-Lesen
$0.060
Cache-Schreiben
Nicht angegeben

Phala

unknown
Kontext
262K
Maximale Ausgabe
236K
Eingabe
$0.150
Ausgabe
$0.460
Cache-Lesen
$0.075
Cache-Schreiben
Nicht angegeben

DeepInfra

fp8
Kontext
131K
Maximale Ausgabe
8K
Eingabe
$0.270
Ausgabe
$0.760
Cache-Lesen
Nicht angegeben
Cache-Schreiben
Nicht angegeben

SambaNova

unknown
Kontext
131K
Maximale Ausgabe
118K
Eingabe
$0.380
Ausgabe
$1.15
Cache-Lesen
Nicht angegeben
Cache-Schreiben
Nicht angegeben

Together

unknown
Kontext
262K
Maximale Ausgabe
236K
Eingabe
$0.390
Ausgabe
$0.970
Cache-Lesen
Nicht angegeben
Cache-Schreiben
Nicht angegeben

ModelRun

fp4
Kontext
262K
Maximale Ausgabe
236K
Eingabe
$0.750
Ausgabe
$1.00
Cache-Lesen
$0.750
Cache-Schreiben
Nicht angegeben

Cerebras

fp16
Kontext
131K
Maximale Ausgabe
41K
Eingabe
$0.990
Ausgabe
$1.49
Cache-Lesen
$0.990
Cache-Schreiben
Nicht angegeben
google

Modelle

Alle Modelle
google logogoogle

Google: Gemini 3.7 Flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Kontext
1.0M
Eingabe
text · image · video · file · audio
Ausgabe
text
Eingabe: $0.750Ausgabe: $3.75pro 1 Mio. Token
Modelldetails ansehen
google logogoogle

Google: Gemini 3.7 Flash (batch)

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Kontext
1.0M
Eingabe
text · image · video · file · audio
Ausgabe
text
Eingabe: $0.188Ausgabe: $0.938pro 1 Mio. Token
Modelldetails ansehen
google logogoogle

Google: Gemini 3.6 Flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Kontext
1.0M
Eingabe
text · image · video · file · audio
Ausgabe
text
Eingabe: $0.750Ausgabe: $3.75pro 1 Mio. Token
Modelldetails ansehen
google logogoogle

Google: Gemini 3.6 Flash (batch)

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Kontext
1.0M
Eingabe
text · image · video · file · audio
Ausgabe
text
Eingabe: $0.375Ausgabe: $1.88pro 1 Mio. Token
Modelldetails ansehen