モデルレーダー · OPENROUTER

AIモデルを、同じ基準で比較。

OpenRouterで現在公開されているすべてのモデルを検索し、モダリティ、コンテキスト、料金、設定、プロバイダー情報を比較できます。

カタログ範囲
公開モデルすべて
モデル
631
プロバイダー
86
最終同期
2026/09/29
631 モデル
stepfun logostepfun

StepFun: Step 3.7 Flash

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...

コンテキスト
262K
入力
text · image · video
出力
text
入力: $0.2出力: $1.15100万トークンあたり
モデル詳細を見る →
anthropic logoanthropic

Anthropic: Claude Opus 4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

コンテキスト
1M
入力
text · image · file
出力
text
入力: $5出力: $25100万トークンあたり
モデル詳細を見る →
anthropic logoanthropic

Anthropic: Claude Opus 4.8 (batch)

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

コンテキスト
1M
入力
text · image · file
出力
text
入力: $2.5出力: $12.5100万トークンあたり
モデル詳細を見る →
nvidia logonvidia

NVIDIA: Parakeet TDT 0.6B v3

Parakeet TDT 0.6B v3 is NVIDIA's 600M-parameter multilingual speech-to-text model built on the FastConformer-TDT architecture. Trained on the Granary dataset (670,000+ hours of audio), it supports automatic language detection across...

コンテキスト
情報なし
入力
audio
出力
transcription
音声の長さ: $0.000025 1秒あたり
モデル詳細を見る →
qwen logoqwen

Qwen: Qwen3.7 Max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

コンテキスト
1M
入力
text
出力
text
入力: $1.475出力: $4.425100万トークンあたり
モデル詳細を見る →
x-ai logox-ai

SpaceXAI: Grok Build 0.1

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

コンテキスト
256K
入力
text · image · file
出力
text
入力: $1出力: $2100万トークンあたり
モデル詳細を見る →
google logogoogle

Google: Gemini Embedding 2

Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports...

コンテキスト
8K
入力
text · image · file · audio · video
出力
embeddings
入力: $0.2100万トークンあたり
モデル詳細を見る →
google logogoogle

Google: Gemini Embedding 2 (batch)

Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports...

コンテキスト
8K
入力
text · image · file · audio · video
出力
embeddings
入力: $0.1100万トークンあたり
モデル詳細を見る →
google logogoogle

Google: Gemini 3.5 Flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

コンテキスト
1.0M
入力
text · image · video · file · audio
出力
text
入力: $1.5出力: $9100万トークンあたり
モデル詳細を見る →
google logogoogle

Google: Gemini 3.5 Flash (batch)

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

コンテキスト
1.0M
入力
text · image · video · file · audio
出力
text
入力: $0.75出力: $4.5100万トークンあたり
モデル詳細を見る →
x-ai logox-ai

SpaceXAI: Grok Imagine Video

Grok Imagine Video is SpaceXAI's fast, text-, image-, and reference-conditioned video generation model. It produces short videos (1–15 seconds, 24 fps) at 480p or 720p across seven aspect ratios -...

コンテキスト
情報なし
入力
text · image
出力
video
画像入力: $0.002 画像1枚あたり動画出力: $0.05 1秒あたり
モデル詳細を見る →
x-ai logox-ai

SpaceXAI: Grok Imagine Image Quality

Grok Imagine Image Quality is SpaceXAI's fast, high-fidelity image generation and editing model. It accepts text prompts and optional reference images, producing photorealistic outputs at 1K or 2K across a...

コンテキスト
66K
入力
text · image
出力
image
画像入力: $0.01 画像1枚あたり画像出力: $0.05 画像1枚あたり
モデル詳細を見る →
mistralai logomistralai

Mistral: Voxtral Mini Transcribe

Voxtral Mini Transcribe is Mistral's speech-to-text model, derived from the Voxtral Mini family. It accepts audio input and returns transcribed text via the standard transcription API. Suited for transcribing meetings,...

コンテキスト
16K
入力
audio
出力
transcription
音声の長さ: $0.00005 1秒あたり
モデル詳細を見る →
x-ai logox-ai

SpaceXAI: Grok Voice TTS 1.0

Grok Voice TTS 1.0 is a text-to-speech model from SpaceXAI. It converts text into spoken audio across 20+ languages with automatic language detection, and offers five built-in voices (Eve, Ara,...

コンテキスト
15K
入力
text
出力
speech
文字: $15 100万文字あたり
モデル詳細を見る →
qwen logoqwen

Qwen: Qwen3 ASR Flash

Qwen3-ASR-Flash is Alibaba's automatic speech recognition service, built on the Qwen3-Omni foundation and trained on tens of millions of hours of multimodal speech data. The model handles 11 languages —...

コンテキスト
情報なし
入力
audio
出力
transcription
音声の長さ: $0.000035 1秒あたり
モデル詳細を見る →
recraft logorecraft

Recraft: Recraft V4.1 Pro Vector

Recraft V4.1 Pro Vector is the vector (SVG) variant of Recraft V4.1 Pro, tuned for high aesthetics. It supports text and image inputs and produces higher-resolution SVG image output across...

コンテキスト
66K
入力
text · image
出力
image
画像出力: $0.3 画像1枚あたり
モデル詳細を見る →
recraft logorecraft

Recraft: Recraft V4.1 Vector

Recraft V4.1 Vector is the vector (SVG) variant of Recraft V4.1, tuned for high aesthetics. It supports text and image inputs and produces SVG image output across multiple aspect ratios,...

コンテキスト
66K
入力
text · image
出力
image
画像出力: $0.08 画像1枚あたり
モデル詳細を見る →
recraft logorecraft

Recraft: Recraft V4.1 Utility Pro

Recraft V4.1 Utility Pro is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at 2K resolution across multiple aspect ratios — double...

コンテキスト
66K
入力
text · image
出力
image
画像出力: $0.21 画像1枚あたり
モデル詳細を見る →
recraft logorecraft

Recraft: Recraft V4.1 Utility

Recraft V4.1 Utility is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at 1K resolution across multiple aspect ratios, with typical generation...

コンテキスト
66K
入力
text · image
出力
image
画像出力: $0.035 画像1枚あたり
モデル詳細を見る →
recraft logorecraft

Recraft: Recraft V4.1 Pro

Recraft V4.1 Pro is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at 2K resolution across multiple aspect ratios...

コンテキスト
66K
入力
text · image
出力
image
画像出力: $0.21 画像1枚あたり
モデル詳細を見る →
recraft logorecraft

Recraft: Recraft V4.1

Recraft V4.1 is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at 1K resolution across multiple aspect ratios, with...

コンテキスト
66K
入力
text · image
出力
image
画像出力: $0.035 画像1枚あたり
モデル詳細を見る →
recraft logorecraft

Recraft: Recraft V4 Pro Vector

Recraft V4 Pro Vector is the vector (SVG) variant of Recraft V4 Pro. It supports text and image inputs and produces vector image output across multiple aspect ratios at the...

コンテキスト
66K
入力
text · image
出力
image
画像出力: $0.3 画像1枚あたり
モデル詳細を見る →
recraft logorecraft

Recraft: Recraft V4 Vector

Recraft V4 Vector is the vector (SVG) variant of Recraft V4. It supports text and image inputs and produces vector image output across multiple aspect ratios. Compared to the raster...

コンテキスト
66K
入力
text · image
出力
image
画像出力: $0.08 画像1枚あたり
モデル詳細を見る →
perceptron logoperceptron

Perceptron: Perceptron Mk1

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning. It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...

コンテキスト
33K
入力
text · image · video
出力
text
入力: $0.15出力: $1.5100万トークンあたり
モデル詳細を見る →