模型雷达 · OPENROUTER
把 AI 模型放在一起比较。
收录 OpenRouter 当前公开的全部模型,统一呈现模态、上下文窗口、价格、配置与提供方信息,并持续增量更新。
- 目录范围
- 全部公开
- 个模型
- 534
- 个提供方
- 75
- 最近同步
- 2026年8月31日
Anthropic: Claude Fable 5 (batch)
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
- 上下文
- 1M
- 输入
- text · image · file
- 输出
- text
Nex AGI: Nex-N2-Pro
Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...
- 上下文
- 262K
- 输入
- text · image
- 输出
- text
Sourceful: Riverflow V2.5 Pro
Riverflow V2.5 Pro is the most powerful variant of Sourceful's Riverflow 2.5 lineup, best for top-tier control and quality-sensitive outputs. The Riverflow 2.5 series is a unified text-to-image and image-to-image...
- 上下文
- 33K
- 输入
- text · image
- 输出
- image
Sourceful: Riverflow V2.5 Fast
Riverflow V2.5 Fast is the speed-optimized variant of Sourceful's Riverflow 2.5 lineup, best for production deployments and latency-critical workflows. The Riverflow 2.5 series is a unified text-to-image and image-to-image family...
- 上下文
- 33K
- 输入
- text · image
- 输出
- image
NVIDIA: Nemotron 3.5 Content Safety (free)
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...
- 上下文
- 128K
- 输入
- text · image
- 输出
- text
NVIDIA: Nemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
- 上下文
- 262K
- 输入
- text
- 输出
- text
NVIDIA: Nemotron 3 Ultra (batch)
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
- 上下文
- 512K
- 输入
- text
- 输出
- text
NVIDIA: Nemotron 3 Ultra (free)
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
- 上下文
- 1M
- 输入
- text
- 输出
- text
Qwen: Qwen3.7 Plus
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
- 上下文
- 1M
- 输入
- text · image
- 输出
- text
Microsoft: MAI-Voice-2
MAI-Voice-2 is an expressive text-to-speech model from Microsoft. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18 locales,...
- 上下文
- 暂未提供
- 输入
- text
- 输出
- speech
Microsoft: MAI-Transcribe 1.5
MAI-Transcribe 1.5 is a multilingual speech-to-text model from Microsoft AI. It is suited for captions, call transcription, subtitling, accessibility, and other voice-enabled applications, with reliable transcription across 43 languages, diverse...
- 上下文
- 暂未提供
- 输入
- audio
- 输出
- transcription
Microsoft: MAI-Image-2.5
Microsoft's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.
- 上下文
- 4K
- 输入
- text · image
- 输出
- image
MiniMax: MiniMax M3
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
- 上下文
- 1.0M
- 输入
- text · image · video
- 输出
- text
MiniMax: MiniMax M3 (batch)
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
- 上下文
- 524K
- 输入
- text · image · video
- 输出
- text
MiniMax: MiniMax M3 (free)
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
- 上下文
- 1.0M
- 输入
- text · image · video
- 输出
- text
StepFun: Step 3.7 Flash
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
- 上下文
- 262K
- 输入
- text · image · video
- 输出
- text
Anthropic: Claude Opus 4.8 (Fast)
Fast-mode variant of Opus 4.8 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode
- 上下文
- 1M
- 输入
- text · image · file
- 输出
- text
Anthropic: Claude Opus 4.8
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
- 上下文
- 1M
- 输入
- text · image · file
- 输出
- text
Anthropic: Claude Opus 4.8 (batch)
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
- 上下文
- 1M
- 输入
- text · image · file
- 输出
- text
NVIDIA: Parakeet TDT 0.6B v3
Parakeet TDT 0.6B v3 is NVIDIA's 600M-parameter multilingual speech-to-text model built on the FastConformer-TDT architecture. Trained on the Granary dataset (670,000+ hours of audio), it supports automatic language detection across...
- 上下文
- 暂未提供
- 输入
- audio
- 输出
- transcription
Qwen: Qwen3.7 Max
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
- 上下文
- 1M
- 输入
- text
- 输出
- text
SpaceXAI: Grok Build 0.1
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...
- 上下文
- 256K
- 输入
- text · image · file
- 输出
- text
Google: Gemini Embedding 2
Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports...
- 上下文
- 8K
- 输入
- text · image · file · audio · video
- 输出
- embeddings
Google: Gemini Embedding 2 (batch)
Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports...
- 上下文
- 8K
- 输入
- text · image · file · audio · video
- 输出
- embeddings