google logo
google

Google: Gemma 4 31B (free)

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. This model is free to use. 262,144 token context window, maximum output of 32,768 tokens. Higher uptime with 15 providers. Includes independent benchmarks from Artificial Analysis.

Overview

Model specifications

Context
262,144 tokens
Maximum output
32,768 tokens
Architecture
text+image+video->text
Tokenizer
Gemma
Knowledge cutoff
Not provided
Moderated
No
OPENROUTER

Complete pricing

Synchronized OpenRouter rates. Token prices are shown per one million tokens.

Input
Free
/M tokens
Output
Free
/M tokens
API

Model configuration

Default parameters

temperature
1
top_p
0.95
top_k
64

Reasoning

Moderated
No
Default parameters
No
API

Quick start

Call this exact model ID through OpenRouter’s OpenAI-compatible API.

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"google/gemma-4-31b-it:free","messages":[{"role":"user","content":"Hello!"}]}'
Capabilities

InputOutput

Input
imagetextvideo
Output
text
Reasoning

No

API

Supported API parameters

include_reasoningmax_tokensreasoningresponse_formatseedtemperaturetool_choicetoolstop_p
Available providers

16 providers

Live providers on OpenRouter

Provider availability, latency, throughput and routing can change continuously. Open the source page for current operational data.

OpenRouter

DeepInfra

fp4
Context
262K
Maximum output
16K
Input
$0.090
Output
$0.340
Cache read
$0.050
Cache write
Not provided

CoreWeave

fp4
Context
262K
Maximum output
236K
Input
$0.100
Output
$0.340
Cache read
$0.100
Cache write
Not provided

Venice

bf16
Context
256K
Maximum output
8K
Input
$0.120
Output
$0.360
Cache read
$0.090
Cache write
Not provided

Chutes

fp4
Context
131K
Maximum output
66K
Input
$0.120
Output
$0.370
Cache read
$0.012
Cache write
Not provided

DeepInfra

fp8
Context
262K
Maximum output
16K
Input
$0.130
Output
$0.380
Cache read
Not provided
Cache write
Not provided

SiliconFlow

fp8
Context
262K
Maximum output
236K
Input
$0.130
Output
$0.400
Cache read
Not provided
Cache write
Not provided

Crusoe

unknown
Context
262K
Maximum output
262K
Input
$0.140
Output
$0.400
Cache read
$0.140
Cache write
Not provided

Friendli

unknown
Context
262K
Maximum output
8K
Input
$0.140
Output
$0.400
Cache read
Not provided
Cache write
Not provided

Novita

bf16
Context
262K
Maximum output
131K
Input
$0.140
Output
$0.400
Cache read
Not provided
Cache write
Not provided

Parasail

fp8
Context
262K
Maximum output
236K
Input
$0.150
Output
$0.400
Cache read
$0.060
Cache write
Not provided

Phala

unknown
Context
262K
Maximum output
236K
Input
$0.150
Output
$0.460
Cache read
$0.075
Cache write
Not provided

DeepInfra

fp8
Context
131K
Maximum output
8K
Input
$0.270
Output
$0.760
Cache read
Not provided
Cache write
Not provided

SambaNova

unknown
Context
131K
Maximum output
118K
Input
$0.380
Output
$1.15
Cache read
Not provided
Cache write
Not provided

Together

unknown
Context
262K
Maximum output
236K
Input
$0.390
Output
$0.970
Cache read
Not provided
Cache write
Not provided

ModelRun

fp4
Context
262K
Maximum output
236K
Input
$0.750
Output
$1.00
Cache read
$0.750
Cache write
Not provided

Cerebras

fp16
Context
131K
Maximum output
41K
Input
$0.990
Output
$1.49
Cache read
$0.990
Cache write
Not provided
google

models

All models
google logogoogle

Google: Gemini 3.7 Flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Context
1.0M
Input
text · image · video · file · audio
Output
text
Input: $0.750Output: $3.75per 1M tokens
View model details
google logogoogle

Google: Gemini 3.7 Flash (batch)

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Context
1.0M
Input
text · image · video · file · audio
Output
text
Input: $0.188Output: $0.938per 1M tokens
View model details
google logogoogle

Google: Gemini 3.6 Flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Context
1.0M
Input
text · image · video · file · audio
Output
text
Input: $0.750Output: $3.75per 1M tokens
View model details
google logogoogle

Google: Gemini 3.6 Flash (batch)

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Context
1.0M
Input
text · image · video · file · audio
Output
text
Input: $0.375Output: $1.88per 1M tokens
View model details