MODEL RADAR · OPENROUTER
KI-Modelle auf einen Blick vergleichen.
Alle derzeit bei OpenRouter gelisteten Modelle mit einheitlichen Angaben zu Modalitäten, Kontext, Preisen, Konfiguration und Anbietern.
- KATALOGUMFANG
- ALLE ÖFFENTLICHEN
- Modelle
- 534
- Anbieter
- 75
- zuletzt synchronisiert
- 31.08.2026
OpenAI: GPT-5.5 Pro
GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...
- Kontext
- 1.1M
- Eingabe
- file · image · text
- Ausgabe
- text
OpenAI: GPT-5.5
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...
- Kontext
- 1.1M
- Eingabe
- file · image · text
- Ausgabe
- text
DeepSeek: DeepSeek V4 Pro 0423
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...
- Kontext
- 1.0M
- Eingabe
- text
- Ausgabe
- text
DeepSeek: DeepSeek V4 Flash 0423
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
- Kontext
- 1.0M
- Eingabe
- text
- Ausgabe
- text
Google: Gemini 3.1 Flash TTS Preview
Gemini 3.1 Flash TTS Preview is a text-to-speech model from Google, and a substantial generational step up from Gemini 2.5 Flash TTS. It takes text input and produces audio output...
- Kontext
- 33K
- Eingabe
- text
- Ausgabe
- speech
Google: Veo 3.1 Fast
Google's mid-tier video generation model balancing speed and quality. Veo 3.1 Fast generates high-quality video from text or image prompts with native synchronized audio, offering faster turnaround than Veo 3.1...
- Kontext
- Nicht angegeben
- Eingabe
- text · image
- Ausgabe
- video
Canopy Labs: Orpheus 3B
Orpheus 3B is an English text-to-speech model from Canopy Labs, fine-tuned for natural prosody and expressive delivery. It offers 7 preset voices and is suited for narration, voice assistants, and...
- Kontext
- 4K
- Eingabe
- text
- Ausgabe
- speech
Sesame: CSM 1B
CSM 1B is a conversational speech model from Sesame. It accepts text input and produces English speech output, with voice options spanning conversational and read-speech styles. At 1B parameters, it...
- Kontext
- 4K
- Eingabe
- text
- Ausgabe
- speech
hexgrad: Kokoro 82M
Kokoro 82M is a lightweight, open-weight text-to-speech model from hexgrad. It converts text to speech across 8 languages (American and British English, Spanish, French, Hindi, Italian, Japanese, Portuguese, and Chinese)...
- Kontext
- 4K
- Eingabe
- text
- Ausgabe
- speech
Google: Veo 3.1 Lite
Google's most cost-effective video generation model, designed for high-volume applications and rapid iteration. Veo 3.1 Lite generates 720p and 1080p video from text or image prompts with native synchronized audio...
- Kontext
- Nicht angegeben
- Eingabe
- text · image
- Ausgabe
- video
Tencent: Hy3 preview
Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...
- Kontext
- 262K
- Eingabe
- text
- Ausgabe
- text
Xiaomi: MiMo-V2.5-Pro
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....
- Kontext
- 1.1M
- Eingabe
- text
- Ausgabe
- text
Xiaomi: MiMo-V2.5
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
- Kontext
- 1.1M
- Eingabe
- text · audio · image · video
- Ausgabe
- text
OpenAI: GPT-5.4 Image 2
GPT-5.4 Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...
- Kontext
- 272K
- Eingabe
- image · text · file
- Ausgabe
- image · text
Anthropic: Claude Opus Latest
This model always redirects to the latest model in the Claude Opus family.
- Kontext
- 1M
- Eingabe
- text · image · file
- Ausgabe
- text
Pareto Code Router
The Pareto Router maintains a tiered shortlist of strong coding models, ranked by Artificial Analysis coding percentiles. Set mincodingscore between 0 and 1 on the pareto-router plugin to control how...
- Kontext
- 2M
- Eingabe
- text
- Ausgabe
- text
Kling: Video O1
Kling Video O1 is a video generation model from Kuaishou. It supports text and image inputs with video output, enabling text-to-video and image-to-video workflows. It is suited for cinematic content...
- Kontext
- Nicht angegeben
- Eingabe
- text · image
- Ausgabe
- video
MiniMax: Hailuo 2.3
Hailuo 2.3 is a video generation model from MiniMax. It accepts text prompts and reference images as input and generates video output, supporting both text-to-video and image-to-video workflows. It is...
- Kontext
- Nicht angegeben
- Eingabe
- text · image
- Ausgabe
- video
MoonshotAI: Kimi K2.6
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...
- Kontext
- 262K
- Eingabe
- text · image
- Ausgabe
- text
Mistral: Voxtral Mini TTS
Voxtral Mini TTS is Mistral's text-to-speech model featuring zero-shot voice cloning and multilingual support. It converts text input into natural-sounding audio output.
- Kontext
- 4K
- Eingabe
- text
- Ausgabe
- speech
Google: Gemini Embedding 2 Preview
Gemini Embedding 2 Preview is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It...
- Kontext
- 8K
- Eingabe
- text · image · file · audio · video
- Ausgabe
- embeddings
Anthropic: Claude Opus 4.7
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...
- Kontext
- 1M
- Eingabe
- text · image · file
- Ausgabe
- text
Anthropic: Claude Opus 4.7 (batch)
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...
- Kontext
- 1M
- Eingabe
- text · image · file
- Ausgabe
- text
Alibaba: Wan 2.7
Wan 2.7 is a video generation model from Alibaba. It supports text-to-video, image-to-video with first and last frame control, and reference-to-video, where multiple reference images guide the style and content...
- Kontext
- Nicht angegeben
- Eingabe
- text · image
- Ausgabe
- video