模型雷达 · OPENROUTER
只看近期发布的 AI 模型。
这里展示最近 60 天发布的模型,并统一呈现输入输出模态、上下文窗口与 Token 价格,内容会持续滚动更新。
- 滚动时间窗
- 60 DAYS
- 个模型
- 119
- 个提供方
- 36
- 最近同步
- 2026年8月28日
Tencent: Hy4 preview
Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that...
- 上下文
- 1.0M
- 输入
- text
- 输出
- text
Alibaba: Wan 3.0 Prime
Wan 3.0 Prime is a fast-mode variant of Wan 3.0 from Alibaba. It supports text-to-video and first-frame image-to-video generation.
- 上下文
- 暂未提供
- 输入
- text · image
- 输出
- video
Ling 3.0 Flash Fin (free)
Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...
- 上下文
- 262K
- 输入
- text
- 输出
- text
Qwen: Qwen3.8 Flash
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.
- 上下文
- 1M
- 输入
- text · image · video
- 输出
- text
Meta: Muse Image
Muse Image is an agentic image generation model from Meta that generates and edits images from text and reference images. Unlike single-pass image models, it reasons before it renders, breaking...
- 上下文
- 66K
- 输入
- text · image
- 输出
- image
Recraft: Recraft V4 Styles Pro
Recraft V4 Styles Pro is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's...
- 上下文
- 66K
- 输入
- text · image
- 输出
- image
Recraft: Recraft V4 Styles Vector
Recraft V4 Styles Vector is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's...
- 上下文
- 66K
- 输入
- text · image
- 输出
- image
Recraft: Recraft V4 Styles Pro Vector
Recraft V4 Styles Pro Vector is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the...
- 上下文
- 66K
- 输入
- text · image
- 输出
- image
Recraft: Recraft V4 Styles
Recraft V4 Styles is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's rendering...
- 上下文
- 66K
- 输入
- text · image
- 输出
- image
Alibaba: Wan 3.0
Wan 3.0 is a video generation model from Alibaba for text-to-video, image-to-video, and reference-guided video generation. It produces 480p, 720p, or 1080p video with durations from 2 to 30 seconds.
- 上下文
- 暂未提供
- 输入
- text · image
- 输出
- video
HeyGen: Avatar IV
HeyGen: Avatar IV is an image-to-video model that animates a single photo into an expressive, lip-synced talking-head video. Rather than only matching mouth shapes to words, it interprets the vocal...
- 上下文
- 暂未提供
- 输入
- text · image · audio
- 输出
- video
Meta: Muse Spark 1.2 Contributor
Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...
- 上下文
- 1.0M
- 输入
- text · image · video · file · audio
- 输出
- text
DeepSeek: DeepSeek V4 Flash Vision Exp
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...
- 上下文
- 1.0M
- 输入
- text · image
- 输出
- text
Tencent: Hy-MT2-1.8B
Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided...
- 上下文
- 8K
- 输入
- text
- 输出
- text
Tencent: Hy-MT2-30B-A3B
Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and...
- 上下文
- 8K
- 输入
- text
- 输出
- text
Black Forest Labs: FLUX Video Upscale
FLUX Video Upscale is a video upscaling model from Black Forest Labs. It enlarges a single source video by 1.5× to 3× while preserving its duration, with an optional prompt...
- 上下文
- 暂未提供
- 输入
- text · video
- 输出
- video
Tencent: Hy-MT2-7B
Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided translation.
- 上下文
- 8K
- 输入
- text
- 输出
- text
LiquidAI: LFM2.5-Embedding-350M (free)
LFM2.5-Embedding-350M is a text embedding model from Liquid AI. It produces 1,024-dimensional embeddings for retrieval and semantic search. Successful OpenRouter requests and embeddings may be retained and used to train...
- 上下文
- 1K
- 输入
- text
- 输出
- embeddings
Qwen: Qwen3.8 27B
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...
- 上下文
- 1M
- 输入
- text · image · video
- 输出
- text
NVIDIA: Nemotron 3.5 ASR Streaming Multilingual 0.6B
Nemotron 3.5 ASR Streaming Multilingual 0.6B is a speech recognition model from NVIDIA. Its prompt-conditioned, cache-aware FastConformer-RNNT design targets low-latency transcription across more than 40 languages for real-time captioning, voice...
- 上下文
- 暂未提供
- 输入
- audio
- 输出
- transcription
Mistral: Voxtral Small 24B 2507 STT
Voxtral Small 24B 2507 STT is a speech transcription model from Mistral AI. It is suited for transcription, translation, and audio understanding workloads that benefit from its larger model capacity.
- 上下文
- 暂未提供
- 输入
- audio
- 输出
- transcription
Mistral: Voxtral Mini 3B 2507
Voxtral Mini 3B 2507 is a speech and audio understanding model from Mistral AI. It is suited for transcription, translation, and compact audio processing workloads.
- 上下文
- 暂未提供
- 输入
- audio
- 输出
- transcription
ByteDance Seed: Seedream 5.0 Lite
Seedream 5.0 Lite is an image generation model from ByteDance Seed. It is suited for professional visual creation that benefits from web-connected retrieval, complex-prompt comprehension, visual references, and broad knowledge...
- 上下文
- 暂未提供
- 输入
- text · image
- 输出
- image
Google: Gemini 3.7 Flash
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
- 上下文
- 1.0M
- 输入
- text · image · video · file · audio
- 输出
- text
Google: Gemini 3.7 Flash (batch)
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
- 上下文
- 1.0M
- 输入
- text · image · video · file · audio
- 输出
- text
VoyageAI by MongoDB: voyage-code-4
voyage-code-4 is a code embedding model from Voyage AI, a MongoDB company. It is designed for coding agents and code retrieval, with Matryoshka embeddings at 2048, 1024, 512, and 256...
- 上下文
- 32K
- 输入
- text
- 输出
- embeddings
Qwen3 Reranker 8B
Qwen3 Reranker 8B is a text reranking model from Alibaba Cloud built on the Qwen3 architecture. It evaluates query-document pairs to produce relevance scores for use in retrieval and RAG...
- 上下文
- 41K
- 输入
- text
- 输出
- rerank
Qwen: Qwen3 ASR 1.7B
Qwen3 ASR 1.7B is an automatic speech recognition model from Qwen. It supports multilingual language identification and transcription across 30 languages and 22 Chinese dialects, with streaming and offline inference...
- 上下文
- 暂未提供
- 输入
- audio
- 输出
- transcription
Qwen: Qwen3 ASR 0.6B
Qwen3 ASR 0.6B is a compact automatic speech recognition model from Qwen. It supports multilingual language identification and transcription across 30 languages and 22 Chinese dialects, with streaming and offline...
- 上下文
- 暂未提供
- 输入
- audio
- 输出
- transcription
ByteDance Seed: Seedream 5.0 Pro
Seedream 5.0 Pro is an image generation and editing model from ByteDance Seed. It is suited for commercial visual-production workflows that require precise editing control, lifelike scenes, and natural rendering.
- 上下文
- 暂未提供
- 输入
- text · image
- 输出
- image
Deepgram: Flux TTS (free)
Flux TTS is a text-to-speech model from Deepgram. It is suited for natural, expressive English speech synthesis across Deepgram's Flux voice catalog.
- 上下文
- 暂未提供
- 输入
- text
- 输出
- speech
ByteDance: Seedance 2.0 Mini
Seedance 2.0 Mini is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video with image, video, and audio inputs. It...
- 上下文
- 暂未提供
- 输入
- text · image · video · audio
- 输出
- video
ByteDance Seed: Seed 2.1 Turbo
Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...
- 上下文
- 262K
- 输入
- text · image · video
- 输出
- text
Qwen: Qwen3.8 2.4T A95B
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of Qwen3.8 Max, with 95 billion active parameters out of 2.4 trillion total. It is...
- 上下文
- 1.0M
- 输入
- text
- 输出
- text
ByteDance Seed: Seed-2.0-Code
Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-agent workflows in tools such as Claude...
- 上下文
- 262K
- 输入
- text · image · video
- 输出
- text
DeepSeek: DeepSeek V4 Pro 0813
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
- 上下文
- 1.0M
- 输入
- text
- 输出
- text
SpaceXAI: Grok 4.6
Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
- 上下文
- 500K
- 输入
- text · image · file
- 输出
- text
xAI: Grok Imagine Image 2.0
Grok Imagine Image 2.0 is an image generation and editing model from xAI. It is suited for creating images from text prompts and editing images from references, with low and...
- 上下文
- 66K
- 输入
- text · image
- 输出
- image
LiquidAI: LFM2.5-2.6B (free)
LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises against using it for agentic coding or...
- 上下文
- 66K
- 输入
- text
- 输出
- text
NVIDIA: Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
- 上下文
- 262K
- 输入
- text
- 输出
- text
NVIDIA: Nemotron 3.5 Lightning (free)
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
- 上下文
- 1M
- 输入
- text
- 输出
- text
Sakana: Sakana Namazu
Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts. It is suited for Japanese instruction following,...
- 上下文
- 262K
- 输入
- text · image · file
- 输出
- text
Upstage: Solar Pro 4
Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...
- 上下文
- 524K
- 输入
- text
- 输出
- text
Meta: Muse Glimmer 30B
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon...
- 上下文
- 131K
- 输入
- text · image
- 输出
- text
ByteDance: Seedance 2.5
Seedance 2.5 is a video generation model from ByteDance. It is suited for long-form storytelling, multimodal reference-based generation, video editing, and video extension. It supports first-frame and first-and-last-frame control, up...
- 上下文
- 暂未提供
- 输入
- text · image · video · audio
- 输出
- video
OpenAI: GPT Transcribe
GPT Transcribe is a high-accuracy speech-to-text model from OpenAI. It is suited for recorded audio, streamed file transcription, and committed Realtime turns, with free-form context, keyword hints, and multiple language...
- 上下文
- 暂未提供
- 输入
- audio
- 输出
- transcription
Meta: Muse Spark 1.2
Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...
- 上下文
- 1.0M
- 输入
- text · image · video · file · audio
- 输出
- text
Qwen: Qwen Image 3
Qwen Image 3 is a unified image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with a richer world...
- 上下文
- 66K
- 输入
- text · image
- 输出
- image