MODEL RADAR · OPENROUTER
Porównuj modele AI według tych samych kryteriów.
Wszystkie modele obecnie dostępne w OpenRouter wraz ze spójnymi danymi o modalnościach, kontekście, cenach, konfiguracji i dostawcach.
- ZAKRES KATALOGU
- WSZYSTKIE PUBLICZNE
- modele
- 631
- dostawców
- 86
- ostatnia synchronizacja
- 29 wrz 2026
StepFun: Step 3.7 Flash
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
- Kontekst
- 262K
- Wejście
- text · image · video
- Wyjście
- text
Anthropic: Claude Opus 4.8
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
- Kontekst
- 1M
- Wejście
- text · image · file
- Wyjście
- text
Anthropic: Claude Opus 4.8 (batch)
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
- Kontekst
- 1M
- Wejście
- text · image · file
- Wyjście
- text
NVIDIA: Parakeet TDT 0.6B v3
Parakeet TDT 0.6B v3 is NVIDIA's 600M-parameter multilingual speech-to-text model built on the FastConformer-TDT architecture. Trained on the Granary dataset (670,000+ hours of audio), it supports automatic language detection across...
- Kontekst
- Brak danych
- Wejście
- audio
- Wyjście
- transcription
Qwen: Qwen3.7 Max
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
- Kontekst
- 1M
- Wejście
- text
- Wyjście
- text
SpaceXAI: Grok Build 0.1
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...
- Kontekst
- 256K
- Wejście
- text · image · file
- Wyjście
- text
Google: Gemini Embedding 2
Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports...
- Kontekst
- 8K
- Wejście
- text · image · file · audio · video
- Wyjście
- embeddings
Google: Gemini Embedding 2 (batch)
Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports...
- Kontekst
- 8K
- Wejście
- text · image · file · audio · video
- Wyjście
- embeddings
Google: Gemini 3.5 Flash
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
- Kontekst
- 1.0M
- Wejście
- text · image · video · file · audio
- Wyjście
- text
Google: Gemini 3.5 Flash (batch)
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
- Kontekst
- 1.0M
- Wejście
- text · image · video · file · audio
- Wyjście
- text
SpaceXAI: Grok Imagine Video
Grok Imagine Video is SpaceXAI's fast, text-, image-, and reference-conditioned video generation model. It produces short videos (1–15 seconds, 24 fps) at 480p or 720p across seven aspect ratios -...
- Kontekst
- Brak danych
- Wejście
- text · image
- Wyjście
- video
SpaceXAI: Grok Imagine Image Quality
Grok Imagine Image Quality is SpaceXAI's fast, high-fidelity image generation and editing model. It accepts text prompts and optional reference images, producing photorealistic outputs at 1K or 2K across a...
- Kontekst
- 66K
- Wejście
- text · image
- Wyjście
- image
Mistral: Voxtral Mini Transcribe
Voxtral Mini Transcribe is Mistral's speech-to-text model, derived from the Voxtral Mini family. It accepts audio input and returns transcribed text via the standard transcription API. Suited for transcribing meetings,...
- Kontekst
- 16K
- Wejście
- audio
- Wyjście
- transcription
SpaceXAI: Grok Voice TTS 1.0
Grok Voice TTS 1.0 is a text-to-speech model from SpaceXAI. It converts text into spoken audio across 20+ languages with automatic language detection, and offers five built-in voices (Eve, Ara,...
- Kontekst
- 15K
- Wejście
- text
- Wyjście
- speech
Qwen: Qwen3 ASR Flash
Qwen3-ASR-Flash is Alibaba's automatic speech recognition service, built on the Qwen3-Omni foundation and trained on tens of millions of hours of multimodal speech data. The model handles 11 languages —...
- Kontekst
- Brak danych
- Wejście
- audio
- Wyjście
- transcription
Recraft: Recraft V4.1 Pro Vector
Recraft V4.1 Pro Vector is the vector (SVG) variant of Recraft V4.1 Pro, tuned for high aesthetics. It supports text and image inputs and produces higher-resolution SVG image output across...
- Kontekst
- 66K
- Wejście
- text · image
- Wyjście
- image
Recraft: Recraft V4.1 Vector
Recraft V4.1 Vector is the vector (SVG) variant of Recraft V4.1, tuned for high aesthetics. It supports text and image inputs and produces SVG image output across multiple aspect ratios,...
- Kontekst
- 66K
- Wejście
- text · image
- Wyjście
- image
Recraft: Recraft V4.1 Utility Pro
Recraft V4.1 Utility Pro is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at 2K resolution across multiple aspect ratios — double...
- Kontekst
- 66K
- Wejście
- text · image
- Wyjście
- image
Recraft: Recraft V4.1 Utility
Recraft V4.1 Utility is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at 1K resolution across multiple aspect ratios, with typical generation...
- Kontekst
- 66K
- Wejście
- text · image
- Wyjście
- image
Recraft: Recraft V4.1 Pro
Recraft V4.1 Pro is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at 2K resolution across multiple aspect ratios...
- Kontekst
- 66K
- Wejście
- text · image
- Wyjście
- image
Recraft: Recraft V4.1
Recraft V4.1 is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at 1K resolution across multiple aspect ratios, with...
- Kontekst
- 66K
- Wejście
- text · image
- Wyjście
- image
Recraft: Recraft V4 Pro Vector
Recraft V4 Pro Vector is the vector (SVG) variant of Recraft V4 Pro. It supports text and image inputs and produces vector image output across multiple aspect ratios at the...
- Kontekst
- 66K
- Wejście
- text · image
- Wyjście
- image
Recraft: Recraft V4 Vector
Recraft V4 Vector is the vector (SVG) variant of Recraft V4. It supports text and image inputs and produces vector image output across multiple aspect ratios. Compared to the raster...
- Kontekst
- 66K
- Wejście
- text · image
- Wyjście
- image
Perceptron: Perceptron Mk1
Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning. It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...
- Kontekst
- 33K
- Wejście
- text · image · video
- Wyjście
- text