Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
70 models,one endpoint.
Curated model records are wired into this public marketplace with price, context, capability, routing status, and source labels. Shortlist the workload first, then run candidates through one OpenAI-compatible endpoint.
70 of 70 models
Model cards with source labels and copyable OpenAI-compatible calls.
What can you compare on this page?
The NextModel marketplace compares model provider, input price, output price, context length, latency estimate, capabilities, use cases, availability, routing status, and source label so teams can shortlist candidates before sending production traffic.
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...
Available through the NextModel gateway via DeepSeek.
Available through the NextModel gateway via DeepSeek.
Available through the NextModel gateway via Volcengine.
Available through the NextModel gateway via Volcengine.
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Available through the NextModel gateway via MiniMax.
Available through the NextModel gateway via Moonshotai.
Available through the NextModel gateway via Moonshotai.
Available through the NextModel gateway via Moonshotai.
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates...
o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via X Ai.
Available through the NextModel gateway via X Ai.
Available through the NextModel gateway via X Ai.
Available through the NextModel gateway via X Ai.
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
GLM 5 CN
80Available through the NextModel gateway via Z Ai.
Available through the NextModel gateway via Z Ai.
Available through the NextModel gateway via Z Ai.
0 video models
Video generation models
Decision table
Compare price, context, capabilities, status, and source in one scan.
Use the table when you are narrowing a shortlist for production tests, cost estimates, or provider policy decisions.
| Model | Provider | Input | Output | Context | Capabilities | Best for | Latency | Status | Source |
|---|---|---|---|---|---|---|---|---|---|
| Anthropic: Claude Fable 5anthropic/claude-fable-5 | Anthropic | $1.45 / 1M tokens | $7.23 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | Anthropic | $0.145 / 1M tokens | $0.723 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5 | Anthropic | $0.723 / 1M tokens | $3.62 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | $0.723 / 1M tokens | $3.62 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.7anthropic/claude-opus-4.7 | Anthropic | $0.723 / 1M tokens | $3.62 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.8anthropic/claude-opus-4.8 | Anthropic | $0.723 / 1M tokens | $3.62 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Claude Opus 5anthropic/claude-opus-5 | Anthropic | $0.723 / 1M tokens | $3.62 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Anthropic | $0.434 / 1M tokens | $2.17 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | $0.434 / 1M tokens | $2.17 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Sonnet 5anthropic/claude-sonnet-5 | Anthropic | $0.289 / 1M tokens | $1.45 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Deepseek V3.2 CNdeepseek/deepseek-v3.2-cn | DeepSeek | $0.042 / 1M tokens | $0.064 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 1000-3000ms | Catalog | Platform curated |
| Deepseek V4 Flash CNdeepseek/deepseek-v4-flash-cn | DeepSeek | $0.02 / 1M tokens | $0.041 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 1000-3000ms | Catalog | Platform curated |
| Doubao Seed 2 1 PROdoubao-seed-2-1-pro | Volcengine | $0.129 / 1M tokens | $0.639 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 1000-3000ms | Catalog | Platform curated |
| Doubao Seed 2 1 Turbodoubao-seed-2-1-turbo | Volcengine | $0.065 / 1M tokens | $0.32 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 1000-3000ms | Catalog | Platform curated |
| Google: Gemini 2.5 Flashgoogle/gemini-2.5-flash | $0.043 / 1M tokens | $0.362 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated | |
| Google: Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | $0.072 / 1M tokens | $0.434 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated | |
| Google: Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite | $0.036 / 1M tokens | $0.217 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated | |
| Google: Gemini 3.5 Flashgoogle/gemini-3.5-flash | $0.217 / 1M tokens | $1.30 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated | |
| Google: Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | $0.043 / 1M tokens | $0.362 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated | |
| Minimax M2.5 CNminimax/minimax-m2.5-cn | MiniMax | $0.045 / 1M tokens | $0.177 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 1000-3000ms | Catalog | Platform curated |
| Kimi K2.5 CNmoonshotai/kimi-k2.5-cn | Moonshotai | $0.084 / 1M tokens | $0.437 / 1M tokens | — | Streaming | General chat via Moonshotai, API workloads | 1000-3000ms | Catalog | Platform curated |
| Kimi K2.6 CNmoonshotai/kimi-k2.6-cn | Moonshotai | $0.13 / 1M tokens | $0.538 / 1M tokens | — | Streaming | General chat via Moonshotai, API workloads | 1000-3000ms | Catalog | Platform curated |
| Kimi K2.7 Code CNmoonshotai/kimi-k2.7-code-cn | Moonshotai | $0.13 / 1M tokens | $0.538 / 1M tokens | — | Streaming | General chat via Moonshotai, API workloads | 1000-3000ms | Catalog | Platform curated |
| MoonshotAI: Kimi K3moonshotai/kimi-k3 | Moonshotai | $0.427 / 1M tokens | $2.13 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-4.1openai/gpt-4.1 | OpenAI | $0.289 / 1M tokens | $1.16 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-4.1 Miniopenai/gpt-4.1-mini | OpenAI | $0.058 / 1M tokens | $0.231 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-4.1 Nanoopenai/gpt-4.1-nano | OpenAI | $0.014 / 1M tokens | $0.058 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-4oopenai/gpt-4o | OpenAI | $0.362 / 1M tokens | $1.45 / 1M tokens | 128k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-4o-miniopenai/gpt-4o-mini | OpenAI | $0.022 / 1M tokens | $0.087 / 1M tokens | 128k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5openai/gpt-5 | OpenAI | $0.181 / 1M tokens | $1.45 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5 Miniopenai/gpt-5-mini | OpenAI | $0.036 / 1M tokens | $0.289 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5 Nanoopenai/gpt-5-nano | OpenAI | $0.0072 / 1M tokens | $0.058 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5.1openai/gpt-5.1 | OpenAI | $0.181 / 1M tokens | $1.45 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5.2openai/gpt-5.2 | OpenAI | $0.253 / 1M tokens | $2.03 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5.4openai/gpt-5.4 | OpenAI | $0.362 / 1M tokens | $2.17 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5.4 Miniopenai/gpt-5.4-mini | OpenAI | $0.109 / 1M tokens | $0.651 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5.4 Nanoopenai/gpt-5.4-nano | OpenAI | $0.029 / 1M tokens | $0.181 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5.5openai/gpt-5.5 | OpenAI | $0.723 / 1M tokens | $4.34 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5.6 Lunaopenai/gpt-5.6-luna | OpenAI | $0.029 / 1M tokens | $0.174 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5.6 Solopenai/gpt-5.6-sol | OpenAI | $0.723 / 1M tokens | $4.34 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT-5.6 Terraopenai/gpt-5.6-terra | OpenAI | $0.289 / 1M tokens | $1.74 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: GPT Chat Latestopenai/gpt-chat-latest | OpenAI | $0.723 / 1M tokens | $4.34 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: o3openai/o3 | OpenAI | $0.289 / 1M tokens | $1.16 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI: o4 Miniopenai/o4-mini | OpenAI | $0.159 / 1M tokens | $0.637 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Qwen3 MAX 2026 01 23 GLBqwen/qwen3-max-2026-01-23-glb | Qwen | $0.174 / 1M tokens | $0.868 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3 MAX CNqwen/qwen3-max-cn | Qwen | $0.052 / 1M tokens | $0.208 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3 VL Flash CNqwen/qwen3-vl-flash-cn | Qwen | $0.0043 / 1M tokens | $0.032 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3 VL Flash GLBqwen/qwen3-vl-flash-glb | Qwen | $0.0072 / 1M tokens | $0.058 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3 VL Plus 2025 12 19 GLBqwen/qwen3-vl-plus-2025-12-19-glb | Qwen | $0.029 / 1M tokens | $0.231 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3 VL Plus CNqwen/qwen3-vl-plus-cn | Qwen | $0.022 / 1M tokens | $0.208 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.5 Flash CNqwen/qwen3.5-flash-cn | Qwen | $0.0043 / 1M tokens | $0.042 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.5 Flash GLBqwen/qwen3.5-flash-glb | Qwen | $0.014 / 1M tokens | $0.058 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.5 Plus CNqwen/qwen3.5-plus-cn | Qwen | $0.017 / 1M tokens | $0.1 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.5 Plus GLBqwen/qwen3.5-plus-glb | Qwen | $0.058 / 1M tokens | $0.347 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.6 Flash CNqwen/qwen3.6-flash-cn | Qwen | $0.025 / 1M tokens | $0.143 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.6 Flash GLBqwen/qwen3.6-flash-glb | Qwen | $0.036 / 1M tokens | $0.217 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.6 MAX Preview CNqwen/qwen3.6-max-preview-cn | Qwen | $0.179 / 1M tokens | $1.07 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.6 MAX Preview GLBqwen/qwen3.6-max-preview-glb | Qwen | $0.188 / 1M tokens | $1.13 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.6 Plus CNqwen/qwen3.6-plus-cn | Qwen | $0.041 / 1M tokens | $0.24 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.6 Plus GLBqwen/qwen3.6-plus-glb | Qwen | $0.072 / 1M tokens | $0.434 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.7 MAX CNqwen/qwen3.7-max-cn | Qwen | $0.239 / 1M tokens | $0.718 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Qwen3.7 MAX GLBqwen/qwen3.7-max-glb | Qwen | $0.362 / 1M tokens | $1.09 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1000-3000ms | Catalog | Platform curated |
| Grok 4.1 Fast NON Reasoningx-ai/grok-4.1-fast-non-reasoning | X Ai | $0.029 / 1M tokens | $0.072 / 1M tokens | — | Streaming | General chat via X Ai, API workloads | 1000-3000ms | Catalog | Platform curated |
| Grok 4.1 Fast Reasoningx-ai/grok-4.1-fast-reasoning | X Ai | $0.029 / 1M tokens | $0.072 / 1M tokens | — | Streaming | General chat via X Ai, API workloads | 1000-3000ms | Catalog | Platform curated |
| Grok 4.20 NON Reasoningx-ai/grok-4.20-non-reasoning | X Ai | $0.181 / 1M tokens | $0.362 / 1M tokens | — | Streaming | General chat via X Ai, API workloads | 1000-3000ms | Catalog | Platform curated |
| Grok 4.20 Reasoningx-ai/grok-4.20-reasoning | X Ai | $0.181 / 1M tokens | $0.362 / 1M tokens | — | Streaming | General chat via X Ai, API workloads | 1000-3000ms | Catalog | Platform curated |
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | X Ai | $0.181 / 1M tokens | $0.362 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| GLM 5 CNz-ai/glm-5-cn | Z Ai | $0.084 / 1M tokens | $0.373 / 1M tokens | — | Streaming | General chat via Z Ai, API workloads | 1000-3000ms | Catalog | Platform curated |
| GLM 5.1 CNz-ai/glm-5.1-cn | Z Ai | $0.12 / 1M tokens | $0.479 / 1M tokens | — | Streaming | General chat via Z Ai, API workloads | 1000-3000ms | Catalog | Platform curated |
| GLM 5.2 CNz-ai/glm-5.2-cn | Z Ai | $0.171 / 1M tokens | $0.596 / 1M tokens | — | Streaming | General chat via Z Ai, API workloads | 1000-3000ms | Catalog | Platform curated |