模型雷达 · OPENROUTER
把 AI 模型放在一起比较。
收录 OpenRouter 当前公开的全部模型,统一呈现模态、上下文窗口、价格、配置与提供方信息,并持续增量更新。
- 目录范围
- 全部公开
- 个模型
- 631
- 个提供方
- 86
- 最近同步
- 2026年9月29日
Thinking Machines: Inkling Small (free)
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
- 上下文
- 1.0M
- 输入
- text · image · audio
- 输出
- text
MiniMax: H3
MiniMax H3 is a lightweight, open-weights video generation model from MiniMax. It is designed for precise multimodal editing and controlled content generation, including instruction-guided edits, text and brand rendering, and...
- 上下文
- 暂未提供
- 输入
- text · image · video · audio
- 输出
- video
Fish Audio: Transcribe 1
Transcribe 1 is a speech-to-text model from Fish Audio. It is suited for audio transcription with automatic language detection and can return timestamped word-level segments when alignment details are requested.
- 上下文
- 暂未提供
- 输入
- audio
- 输出
- transcription
Fish Audio: S1
S1 is a multilingual text-to-speech model from Fish Audio. It is suited for voice applications that need broad emotional expression, using parenthetical controls to guide speaking style across its supported...
- 上下文
- 暂未提供
- 输入
- text
- 输出
- speech
Fish Audio: S2 Pro
S2 Pro is a multilingual text-to-speech model from Fish Audio. It is suited for expressive narration and multi-speaker dialogue, with natural-language controls for speaking style and emotion.
- 上下文
- 暂未提供
- 输入
- text
- 输出
- speech
Fish Audio: S2.1 Pro Free (free)
S2.1 Pro Free is the no-cost variant of Fish Audio S2.1 Pro, intended for testing, prototyping, and low-volume applications. It provides the same synthesis capabilities without production latency or availability...
- 上下文
- 暂未提供
- 输入
- text
- 输出
- speech
Fish Audio: S2.1 Pro
S2.1 Pro is a production-oriented text-to-speech model from Fish Audio. It is suited for multilingual voice applications, expressive narration, and dialogue synthesis, with open-ended natural-language controls for speaking style and...
- 上下文
- 暂未提供
- 输入
- text
- 输出
- speech
Runway: Aleph 2.0
Runway Aleph 2.0 is an in-context video editing model from Runway. It applies text instructions and keyframe-guided edits across existing footage while preserving details that are not meant to change....
- 上下文
- 暂未提供
- 输入
- text · image · video
- 输出
- video
Runway: Gen-4.5
Runway Gen-4.5 is a video generation model from Runway for text-to-video and image-to-video workflows. It is designed for cinematic scene creation with strong motion quality, visual fidelity, and prompt adherence....
- 上下文
- 暂未提供
- 输入
- text · image
- 输出
- video
Qwen: Qwen3.7 Flash
Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...
- 上下文
- 1M
- 输入
- text · image · video
- 输出
- text
VoyageAI by MongoDB: rerank-2.5-lite
rerank-2.5-lite is a reranker optimized for both latency and quality, delivering a 7.16% improvement in retrieval accuracy over Cohere Rerank v3.5 across 93 datasets. It also outperformed Cohere Rerank v3.5...
- 上下文
- 32K
- 输入
- text
- 输出
- rerank
VoyageAI by MongoDB: rerank-2.5
rerank-2.5 is a cutting-edge reranker optimized for quality, delivering a 7.94% improvement in retrieval accuracy over Cohere Rerank v3.5 across 93 datasets. It also outperformed Cohere Rerank v3.5 by 12.70%...
- 上下文
- 32K
- 输入
- text
- 输出
- rerank
VoyageAI by MongoDB: voyage-multimodal-3.5
voyage-multimodal-3.5 is a state-of-the-art multimodal embedding model capable of vectorizing not only text, images, and video individually, but also content that interleaves all three modalities. It delivers excellent performance for...
- 上下文
- 32K
- 输入
- text · image
- 输出
- embeddings
VoyageAI by MongoDB: voyage-4-lite
voyage-4-lite is a lightweight, general-purpose embedding model optimized for low latency and cost. Enabled by Matryoshka learning and quantization-aware training, voyage-4-lite supports embeddings in 2048, 1024, 512, and 256 dimensions,...
- 上下文
- 32K
- 输入
- text
- 输出
- embeddings
VoyageAI by MongoDB: voyage-4
voyage-4 is a general-purpose (including multilingual) embedding model optimized for retrieval/search and AI applications. voyage-4 supports embeddings in 2048, 1024, 512, and 256 dimensions, with multiple quantization options. Learn more...
- 上下文
- 32K
- 输入
- text
- 输出
- embeddings
VoyageAI by MongoDB: voyage-4-large
voyage-4-large is a state-of-the-art general-purpose and multilingual embedding optimized for retrieval quality. Enabled by Matryoshka learning and quantization-aware training, voyage-4-large supports embeddings in 2048, 1024, 512, and 256 dimensions, with...
- 上下文
- 32K
- 输入
- text
- 输出
- embeddings
Anthropic: Claude Opus 5
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
- 上下文
- 1M
- 输入
- text · image · file
- 输出
- text
Anthropic: Claude Opus 5 (batch)
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
- 上下文
- 1M
- 输入
- text · image · file
- 输出
- text
Microsoft AI: MAI-Image-2.5 Pro
Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.
- 上下文
- 4K
- 输入
- text · image
- 输出
- image
Microsoft AI: MAI-Voice-2-Flash
MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft AI for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15...
- 上下文
- 暂未提供
- 输入
- text
- 输出
- speech
inclusionAI: Ling 3.0 Flash
Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers...
- 上下文
- 262K
- 输入
- text
- 输出
- text
Qwen: Qwen-Audio-3.0-TTS Flash
Qwen-Audio-3.0-TTS Flash is Alibaba's fast, cost-efficient text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.
- 上下文
- 暂未提供
- 输入
- text
- 输出
- speech
Qwen: Qwen-Audio-3.0-TTS Plus
Qwen-Audio-3.0-TTS Plus is Alibaba's higher-quality text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.
- 上下文
- 暂未提供
- 输入
- text
- 输出
- speech
SpaceXAI: Grok STT 1.0
Grok STT is SpaceXAI's speech-to-text model, available via the REST /v1/stt endpoint. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio.
- 上下文
- 暂未提供
- 输入
- audio
- 输出
- transcription