Processing...Please wait while we secure this action
/ models

70 models,one endpoint.

Curated model records are wired into this public marketplace with price, context, capability, routing status, and source labels. Shortlist the workload first, then run candidates through one OpenAI-compatible endpoint.

routing candidates70/70
Claude Fable 5$1.45/1M
Claude Haiku 4.5$0.145/1M
Claude Opus 4.5$0.723/1M
Claude Opus 4.6$0.723/1M
10providers1sources1.1Mmax context$0.0043lowest input
Reset

70 of 70 models

Model cards with source labels and copyable OpenAI-compatible calls.

What can you compare on this page?

The NextModel marketplace compares model provider, input price, output price, context length, latency estimate, capabilities, use cases, availability, routing status, and source label so teams can shortlist candidates before sending production traffic.

AnthropicCatalog

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

$1.45 / 1M tokensInput$7.23 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

$0.145 / 1M tokensInput$0.723 / 1M tokensOutput200kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

$0.723 / 1M tokensInput$3.62 / 1M tokensOutput200kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

$0.723 / 1M tokensInput$3.62 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

$0.723 / 1M tokensInput$3.62 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

$0.723 / 1M tokensInput$3.62 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

$0.723 / 1M tokensInput$3.62 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

$0.434 / 1M tokensInput$2.17 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

$0.434 / 1M tokensInput$2.17 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

$0.289 / 1M tokensInput$1.45 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
DeepSeekCatalog

Available through the NextModel gateway via DeepSeek.

$0.042 / 1M tokensInput$0.064 / 1M tokensOutputContext
Best forChinese Q&A, general chat
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
DeepSeekCatalog

Available through the NextModel gateway via DeepSeek.

$0.02 / 1M tokensInput$0.041 / 1M tokensOutputContext
Best forChinese Q&A, general chat
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
VolcengineCatalog

Available through the NextModel gateway via Volcengine.

$0.129 / 1M tokensInput$0.639 / 1M tokensOutputContext
Best forChinese Q&A, general chat
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
VolcengineCatalog

Available through the NextModel gateway via Volcengine.

$0.065 / 1M tokensInput$0.32 / 1M tokensOutputContext
Best forChinese Q&A, general chat
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
GoogleCatalog

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

$0.043 / 1M tokensInput$0.362 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
GoogleCatalog

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

$0.072 / 1M tokensInput$0.434 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
GoogleCatalog

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

$0.036 / 1M tokensInput$0.217 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
GoogleCatalog

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

$0.217 / 1M tokensInput$1.30 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
GoogleCatalog

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

$0.043 / 1M tokensInput$0.362 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
MiniMaxCatalog

Available through the NextModel gateway via MiniMax.

$0.045 / 1M tokensInput$0.177 / 1M tokensOutputContext
Best forChinese Q&A, general chat
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
MoonshotaiCatalog

Available through the NextModel gateway via Moonshotai.

$0.084 / 1M tokensInput$0.437 / 1M tokensOutputContext
Best forGeneral chat via Moonshotai, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
MoonshotaiCatalog

Available through the NextModel gateway via Moonshotai.

$0.13 / 1M tokensInput$0.538 / 1M tokensOutputContext
Best forGeneral chat via Moonshotai, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
MoonshotaiCatalog

Available through the NextModel gateway via Moonshotai.

$0.13 / 1M tokensInput$0.538 / 1M tokensOutputContext
Best forGeneral chat via Moonshotai, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
MoonshotaiCatalog

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

$0.427 / 1M tokensInput$2.13 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...

$0.289 / 1M tokensInput$1.16 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...

$0.058 / 1M tokensInput$0.231 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

$0.014 / 1M tokensInput$0.058 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...

$0.362 / 1M tokensInput$1.45 / 1M tokensOutput128kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...

$0.022 / 1M tokensInput$0.087 / 1M tokensOutput128kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

$0.181 / 1M tokensInput$1.45 / 1M tokensOutput400kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

$0.036 / 1M tokensInput$0.289 / 1M tokensOutput400kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...

$0.0072 / 1M tokensInput$0.058 / 1M tokensOutput400kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...

$0.181 / 1M tokensInput$1.45 / 1M tokensOutput400kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...

$0.253 / 1M tokensInput$2.03 / 1M tokensOutput400kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

Starting at $0.362 / 1M tokensInputStarting at $2.17 / 1M tokensOutput1.1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

$0.109 / 1M tokensInput$0.651 / 1M tokensOutput400kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

$0.029 / 1M tokensInput$0.181 / 1M tokensOutput400kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

Starting at $0.723 / 1M tokensInputStarting at $4.34 / 1M tokensOutput1.1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

Starting at $0.029 / 1M tokensInputStarting at $0.174 / 1M tokensOutput1.1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

Starting at $0.723 / 1M tokensInputStarting at $4.34 / 1M tokensOutput1.1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

Starting at $0.289 / 1M tokensInputStarting at $1.74 / 1M tokensOutput1.1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates...

$0.723 / 1M tokensInput$4.34 / 1M tokensOutput400kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....

$0.289 / 1M tokensInput$1.16 / 1M tokensOutput200kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
OpenAICatalog

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...

$0.159 / 1M tokensInput$0.637 / 1M tokensOutput200kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.174 / 1M tokensInputStarting at $0.868 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.052 / 1M tokensInputStarting at $0.208 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.0043 / 1M tokensInputStarting at $0.032 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.0072 / 1M tokensInputStarting at $0.058 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.029 / 1M tokensInputStarting at $0.231 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.022 / 1M tokensInputStarting at $0.208 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.0043 / 1M tokensInputStarting at $0.042 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

$0.014 / 1M tokensInput$0.058 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.017 / 1M tokensInputStarting at $0.1 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.058 / 1M tokensInputStarting at $0.347 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.025 / 1M tokensInputStarting at $0.143 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.036 / 1M tokensInputStarting at $0.217 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.179 / 1M tokensInputStarting at $1.07 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.188 / 1M tokensInputStarting at $1.13 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.041 / 1M tokensInputStarting at $0.24 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

Starting at $0.072 / 1M tokensInputStarting at $0.434 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

$0.239 / 1M tokensInput$0.718 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
QwenCatalog

Available through the NextModel gateway via Qwen.

$0.362 / 1M tokensInput$1.09 / 1M tokensOutputContext
Best forGeneral chat via Qwen, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
X AiCatalog

Available through the NextModel gateway via X Ai.

$0.029 / 1M tokensInput$0.072 / 1M tokensOutputContext
Best forGeneral chat via X Ai, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
X AiCatalog

Available through the NextModel gateway via X Ai.

$0.029 / 1M tokensInput$0.072 / 1M tokensOutputContext
Best forGeneral chat via X Ai, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
X AiCatalog

Available through the NextModel gateway via X Ai.

$0.181 / 1M tokensInput$0.362 / 1M tokensOutputContext
Best forGeneral chat via X Ai, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
X AiCatalog

Available through the NextModel gateway via X Ai.

$0.181 / 1M tokensInput$0.362 / 1M tokensOutputContext
Best forGeneral chat via X Ai, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
X AiCatalog

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

$0.181 / 1M tokensInput$0.362 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
Z AiCatalog

Available through the NextModel gateway via Z Ai.

Starting at $0.084 / 1M tokensInputStarting at $0.373 / 1M tokensOutputContext
Best forGeneral chat via Z Ai, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
Z AiCatalog

Available through the NextModel gateway via Z Ai.

Starting at $0.12 / 1M tokensInputStarting at $0.479 / 1M tokensOutputContext
Best forGeneral chat via Z Ai, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details
Z AiCatalog

Available through the NextModel gateway via Z Ai.

$0.171 / 1M tokensInput$0.596 / 1M tokensOutputContext
Best forGeneral chat via Z Ai, API workloads
RoutingConfigured
Streaming
Platform curatedNextModel gateway catalog (Go origin)
View details

0 video models

Video generation models

Decision table

Compare price, context, capabilities, status, and source in one scan.

Use the table when you are narrowing a shortlist for production tests, cost estimates, or provider policy decisions.

ModelProviderInputOutputContextCapabilitiesBest forLatencyStatusSource
Anthropic: Claude Fable 5anthropic/claude-fable-5Anthropic$1.45 / 1M tokens$7.23 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5Anthropic$0.145 / 1M tokens$0.723 / 1M tokens200k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5Anthropic$0.723 / 1M tokens$3.62 / 1M tokens200k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Opus 4.6anthropic/claude-opus-4.6Anthropic$0.723 / 1M tokens$3.62 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Opus 4.7anthropic/claude-opus-4.7Anthropic$0.723 / 1M tokens$3.62 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Opus 4.8anthropic/claude-opus-4.8Anthropic$0.723 / 1M tokens$3.62 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Claude Opus 5anthropic/claude-opus-5Anthropic$0.723 / 1M tokens$3.62 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5Anthropic$0.434 / 1M tokens$2.17 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6Anthropic$0.434 / 1M tokens$2.17 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Sonnet 5anthropic/claude-sonnet-5Anthropic$0.289 / 1M tokens$1.45 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Deepseek V3.2 CNdeepseek/deepseek-v3.2-cnDeepSeek$0.042 / 1M tokens$0.064 / 1M tokens
Streaming
Chinese Q&A, general chat1000-3000msCatalogPlatform curated
Deepseek V4 Flash CNdeepseek/deepseek-v4-flash-cnDeepSeek$0.02 / 1M tokens$0.041 / 1M tokens
Streaming
Chinese Q&A, general chat1000-3000msCatalogPlatform curated
Doubao Seed 2 1 PROdoubao-seed-2-1-proVolcengine$0.129 / 1M tokens$0.639 / 1M tokens
Streaming
Chinese Q&A, general chat1000-3000msCatalogPlatform curated
Doubao Seed 2 1 Turbodoubao-seed-2-1-turboVolcengine$0.065 / 1M tokens$0.32 / 1M tokens
Streaming
Chinese Q&A, general chat1000-3000msCatalogPlatform curated
Google: Gemini 2.5 Flashgoogle/gemini-2.5-flashGoogle$0.043 / 1M tokens$0.362 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Google: Gemini 3 Flash Previewgoogle/gemini-3-flash-previewGoogle$0.072 / 1M tokens$0.434 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Google: Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-liteGoogle$0.036 / 1M tokens$0.217 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Google: Gemini 3.5 Flashgoogle/gemini-3.5-flashGoogle$0.217 / 1M tokens$1.30 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Google: Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-liteGoogle$0.043 / 1M tokens$0.362 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Minimax M2.5 CNminimax/minimax-m2.5-cnMiniMax$0.045 / 1M tokens$0.177 / 1M tokens
Streaming
Chinese Q&A, general chat1000-3000msCatalogPlatform curated
Kimi K2.5 CNmoonshotai/kimi-k2.5-cnMoonshotai$0.084 / 1M tokens$0.437 / 1M tokens
Streaming
General chat via Moonshotai, API workloads1000-3000msCatalogPlatform curated
Kimi K2.6 CNmoonshotai/kimi-k2.6-cnMoonshotai$0.13 / 1M tokens$0.538 / 1M tokens
Streaming
General chat via Moonshotai, API workloads1000-3000msCatalogPlatform curated
Kimi K2.7 Code CNmoonshotai/kimi-k2.7-code-cnMoonshotai$0.13 / 1M tokens$0.538 / 1M tokens
Streaming
General chat via Moonshotai, API workloads1000-3000msCatalogPlatform curated
MoonshotAI: Kimi K3moonshotai/kimi-k3Moonshotai$0.427 / 1M tokens$2.13 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-4.1openai/gpt-4.1OpenAI$0.289 / 1M tokens$1.16 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-4.1 Miniopenai/gpt-4.1-miniOpenAI$0.058 / 1M tokens$0.231 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-4.1 Nanoopenai/gpt-4.1-nanoOpenAI$0.014 / 1M tokens$0.058 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-4oopenai/gpt-4oOpenAI$0.362 / 1M tokens$1.45 / 1M tokens128k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-4o-miniopenai/gpt-4o-miniOpenAI$0.022 / 1M tokens$0.087 / 1M tokens128k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5openai/gpt-5OpenAI$0.181 / 1M tokens$1.45 / 1M tokens400k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5 Miniopenai/gpt-5-miniOpenAI$0.036 / 1M tokens$0.289 / 1M tokens400k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5 Nanoopenai/gpt-5-nanoOpenAI$0.0072 / 1M tokens$0.058 / 1M tokens400k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5.1openai/gpt-5.1OpenAI$0.181 / 1M tokens$1.45 / 1M tokens400k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5.2openai/gpt-5.2OpenAI$0.253 / 1M tokens$2.03 / 1M tokens400k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5.4openai/gpt-5.4OpenAI$0.362 / 1M tokens$2.17 / 1M tokens1.1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5.4 Miniopenai/gpt-5.4-miniOpenAI$0.109 / 1M tokens$0.651 / 1M tokens400k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5.4 Nanoopenai/gpt-5.4-nanoOpenAI$0.029 / 1M tokens$0.181 / 1M tokens400k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5.5openai/gpt-5.5OpenAI$0.723 / 1M tokens$4.34 / 1M tokens1.1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5.6 Lunaopenai/gpt-5.6-lunaOpenAI$0.029 / 1M tokens$0.174 / 1M tokens1.1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5.6 Solopenai/gpt-5.6-solOpenAI$0.723 / 1M tokens$4.34 / 1M tokens1.1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT-5.6 Terraopenai/gpt-5.6-terraOpenAI$0.289 / 1M tokens$1.74 / 1M tokens1.1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: GPT Chat Latestopenai/gpt-chat-latestOpenAI$0.723 / 1M tokens$4.34 / 1M tokens400k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: o3openai/o3OpenAI$0.289 / 1M tokens$1.16 / 1M tokens200k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
OpenAI: o4 Miniopenai/o4-miniOpenAI$0.159 / 1M tokens$0.637 / 1M tokens200k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Qwen3 MAX 2026 01 23 GLBqwen/qwen3-max-2026-01-23-glbQwen$0.174 / 1M tokens$0.868 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3 MAX CNqwen/qwen3-max-cnQwen$0.052 / 1M tokens$0.208 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3 VL Flash CNqwen/qwen3-vl-flash-cnQwen$0.0043 / 1M tokens$0.032 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3 VL Flash GLBqwen/qwen3-vl-flash-glbQwen$0.0072 / 1M tokens$0.058 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3 VL Plus 2025 12 19 GLBqwen/qwen3-vl-plus-2025-12-19-glbQwen$0.029 / 1M tokens$0.231 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3 VL Plus CNqwen/qwen3-vl-plus-cnQwen$0.022 / 1M tokens$0.208 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.5 Flash CNqwen/qwen3.5-flash-cnQwen$0.0043 / 1M tokens$0.042 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.5 Flash GLBqwen/qwen3.5-flash-glbQwen$0.014 / 1M tokens$0.058 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.5 Plus CNqwen/qwen3.5-plus-cnQwen$0.017 / 1M tokens$0.1 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.5 Plus GLBqwen/qwen3.5-plus-glbQwen$0.058 / 1M tokens$0.347 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.6 Flash CNqwen/qwen3.6-flash-cnQwen$0.025 / 1M tokens$0.143 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.6 Flash GLBqwen/qwen3.6-flash-glbQwen$0.036 / 1M tokens$0.217 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.6 MAX Preview CNqwen/qwen3.6-max-preview-cnQwen$0.179 / 1M tokens$1.07 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.6 MAX Preview GLBqwen/qwen3.6-max-preview-glbQwen$0.188 / 1M tokens$1.13 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.6 Plus CNqwen/qwen3.6-plus-cnQwen$0.041 / 1M tokens$0.24 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.6 Plus GLBqwen/qwen3.6-plus-glbQwen$0.072 / 1M tokens$0.434 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.7 MAX CNqwen/qwen3.7-max-cnQwen$0.239 / 1M tokens$0.718 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Qwen3.7 MAX GLBqwen/qwen3.7-max-glbQwen$0.362 / 1M tokens$1.09 / 1M tokens
Streaming
General chat via Qwen, API workloads1000-3000msCatalogPlatform curated
Grok 4.1 Fast NON Reasoningx-ai/grok-4.1-fast-non-reasoningX Ai$0.029 / 1M tokens$0.072 / 1M tokens
Streaming
General chat via X Ai, API workloads1000-3000msCatalogPlatform curated
Grok 4.1 Fast Reasoningx-ai/grok-4.1-fast-reasoningX Ai$0.029 / 1M tokens$0.072 / 1M tokens
Streaming
General chat via X Ai, API workloads1000-3000msCatalogPlatform curated
Grok 4.20 NON Reasoningx-ai/grok-4.20-non-reasoningX Ai$0.181 / 1M tokens$0.362 / 1M tokens
Streaming
General chat via X Ai, API workloads1000-3000msCatalogPlatform curated
Grok 4.20 Reasoningx-ai/grok-4.20-reasoningX Ai$0.181 / 1M tokens$0.362 / 1M tokens
Streaming
General chat via X Ai, API workloads1000-3000msCatalogPlatform curated
SpaceXAI: Grok 4.3x-ai/grok-4.3X Ai$0.181 / 1M tokens$0.362 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
GLM 5 CNz-ai/glm-5-cnZ Ai$0.084 / 1M tokens$0.373 / 1M tokens
Streaming
General chat via Z Ai, API workloads1000-3000msCatalogPlatform curated
GLM 5.1 CNz-ai/glm-5.1-cnZ Ai$0.12 / 1M tokens$0.479 / 1M tokens
Streaming
General chat via Z Ai, API workloads1000-3000msCatalogPlatform curated
GLM 5.2 CNz-ai/glm-5.2-cnZ Ai$0.171 / 1M tokens$0.596 / 1M tokens
Streaming
General chat via Z Ai, API workloads1000-3000msCatalogPlatform curated