GLM 5.3 CNNew
Available through the NextModel gateway via Z Ai.
Curated model records are wired into this public marketplace with price, context, capability, routing status, and source labels. Shortlist the workload first, then run candidates through one OpenAI-compatible endpoint.
124 of 124 models
The NextModel marketplace compares model provider, input price, output price, context length, latency estimate, capabilities, use cases, availability, routing status, and source label so teams can shortlist candidates before sending production traffic.
Available through the NextModel gateway via Z Ai.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Black Forest Labs.
Available through the NextModel gateway via Black Forest Labs.
Available through the NextModel gateway via OpenAI.
Available through the NextModel gateway via OpenAI.
Available through the NextModel gateway via OpenAI.
Available through the NextModel gateway via OpenAI.
Available through the NextModel gateway via OpenAI.
Available through the NextModel gateway via OpenAI.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Volcengine.
Available through the NextModel gateway via Volcengine.
Available through the NextModel gateway via Volcengine.
Available through the NextModel gateway via Volcengine.
Available through the NextModel gateway via OpenAI.
Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,...
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Volcengine.
Available through the NextModel gateway via Volcengine.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via X Ai.
Available through the NextModel gateway via X Ai.
Available through the NextModel gateway via X Ai.
Available through the NextModel gateway via X Ai.
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...
GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates...
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...
Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...
GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...
Available through the NextModel gateway via DeepSeek.
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K...
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...
GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results...
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...
GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...
Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed...
GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....
GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based on an updated version of the 5.1 reasoning stack and trained on agentic...
GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex
GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...
GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...
GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...
o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
Available through the NextModel gateway via Doubao.
Available through the NextModel gateway via Doubao.
Available through the NextModel gateway via Doubao.
Available through the NextModel gateway via Dreamina.
Available through the NextModel gateway via Dreamina.
Available through the NextModel gateway via Dreamina.
Available through the NextModel gateway via Gemini.
Available through the NextModel gateway via Gemini.
Available through the NextModel gateway via Gemini.
Available through the NextModel gateway via Kling.
Available through the NextModel gateway via Kling.
Available through the NextModel gateway via Kling.
Available through the NextModel gateway via Kling.
Available through the NextModel gateway via OpenAI.
Available through the NextModel gateway via OpenAI.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Qwen.
Available through the NextModel gateway via Vidu.
Available through the NextModel gateway via Vidu.
18 video models
Decision table
Use the table when you are narrowing a shortlist for production tests, cost estimates, or provider policy decisions.
| Model | Provider | Input | Output | Context | Capabilities | Best for | Latency | Status | Source |
|---|---|---|---|---|---|---|---|---|---|
| GLM 5.3 CNz-ai/glm-5.3-cn | Z Ai | $1.40 / 1M tokens | $4.40 / 1M tokens | — | Streaming | General chat via Z Ai, API workloads | 0-0ms | Catalog | Platform curated |
| Qwen3.8 MAX GLBqwen/qwen3.8-max-glb | Qwen | $2 / 1M tokens | $6 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 0-0ms | Catalog | Platform curated |
| FLUX.2 Flexblack-forest-labs/FLUX.2-flex | Black Forest Labs | $0.2 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| FLUX.2 PROblack-forest-labs/FLUX.2-pro | Black Forest Labs | $0.075 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| GPT Image 1openai/gpt-image-1 | OpenAI | $0.26 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| GPT Image 1 Miniopenai/gpt-image-1-mini | OpenAI | $0.0539 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| GPT Image 1.5openai/gpt-image-1.5 | OpenAI | $0.218 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| GPT Image 2openai/gpt-image-2 | OpenAI | $0.221 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Text Embedding 3 Smallopenai/text-embedding-3-small | OpenAI | $0.02 / 1M tokens | $0 / 1M tokens | — | Streaming | General chat via OpenAI, API workloads | 0-0ms | Catalog | Platform curated |
| Text Embedding 3 Largeopenai/text-embedding-3-large | OpenAI | $0.13 / 1M tokens | $0 / 1M tokens | — | Streaming | General chat via OpenAI, API workloads | 0-0ms | Catalog | Platform curated |
| Qwen3.5 Flashqwen/qwen3.5-flash | Qwen | Starting at $0.03 / 1M tokens | $0.29 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 0-0ms | Catalog | Platform curated |
| Qwen3.5 Plusqwen/qwen3.5-plus | Qwen | Starting at $0.12 / 1M tokens | $0.69 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 0-0ms | Catalog | Platform curated |
| DeepSeek: DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | DeepSeek | $1.33 / 1M tokens | $3.98 / 1M tokens | 1M | StreamingTool callingJSON modeLong context | Chinese Q&A, general chat | 0-0ms | Catalog | Platform curated |
| Qwen3 VL Flash GLBqwen/qwen3-vl-flash-glb | Qwen | Starting at $0.05 / 1M tokens | $0.4 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 485-1081ms | Catalog | Platform curated |
| Qwen3 VL Plus CNqwen/qwen3-vl-plus-cn | Qwen | Starting at $0.15 / 1M tokens | $1.44 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 976-1270ms | Catalog | Platform curated |
| Qwen3 VL Plus 2025 12 19 GLBqwen/qwen3-vl-plus-2025-12-19-glb | Qwen | Starting at $0.2 / 1M tokens | $1.60 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 539-984ms | Catalog | Platform curated |
| Qwen3 MAX 2026 01 23 GLBqwen/qwen3-max-2026-01-23-glb | Qwen | Starting at $1.20 / 1M tokens | $6 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1115-1620ms | Catalog | Platform curated |
| Qwen3 VL Flash CNqwen/qwen3-vl-flash-cn | Qwen | Starting at $0.03 / 1M tokens | $0.22 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 738-968ms | Catalog | Platform curated |
| Qwen3.6 Flash GLBqwen/qwen3.6-flash-glb | Qwen | Starting at $0.25 / 1M tokens | $1.50 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 3725-5091ms | Catalog | Platform curated |
| Qwen3 MAX CNqwen/qwen3-max-cn | Qwen | Starting at $0.36 / 1M tokens | $1.44 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 1252-1644ms | Catalog | Platform curated |
| Qwen3.5 Plus GLBqwen/qwen3.5-plus-glb | Qwen | Starting at $0.4 / 1M tokens | $2.40 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 4273-23436ms | Catalog | Platform curated |
| Qwen3.6 Plus GLBqwen/qwen3.6-plus-glb | Qwen | Starting at $0.5 / 1M tokens | $3 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 5504-5937ms | Catalog | Platform curated |
| Qwen3.6 MAX Preview GLBqwen/qwen3.6-max-preview-glb | Qwen | Starting at $1.30 / 1M tokens | $7.80 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 7036-14296ms | Catalog | Platform curated |
| Doubao Seed 2 0 Litedoubao-seed-2-0-lite | Volcengine | Starting at $0.09 / 1M tokens | $0.53 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 18958-26364ms | Catalog | Platform curated |
| Doubao Seed 2 0 Codedoubao-seed-2-0-code | Volcengine | Starting at $0.48 / 1M tokens | $2.36 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 3987-7615ms | Catalog | Platform curated |
| Doubao Seed 2 0 PROdoubao-seed-2-0-pro | Volcengine | Starting at $0.48 / 1M tokens | $2.36 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 6712-7583ms | Catalog | Platform curated |
| Doubao Seed 2 0 Minidoubao-seed-2-0-mini | Volcengine | Starting at $0.03 / 1M tokens | $0.3 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 2007-2427ms | Catalog | Platform curated |
| GPT 5 Codexopenai/gpt-5-codex | OpenAI | $1.25 / 1M tokens | $10 / 1M tokens | — | Streaming | General chat via OpenAI, API workloads | 1202-5837ms | Catalog | Platform curated |
| Qwen: Qwen3.8 Maxqwen/qwen3.8-max | Qwen | $1.77 / 1M tokens | $5.30 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Qwen3.5 Flash GLBqwen/qwen3.5-flash-glb | Qwen | $0.1 / 1M tokens | $0.4 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 2165-3952ms | Catalog | Platform curated |
| Doubao Seed 2 1 Turbodoubao-seed-2-1-turbo | Volcengine | $0.45 / 1M tokens | $2.21 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 7716-9899ms | Catalog | Platform curated |
| Doubao Seed 2 1 PROdoubao-seed-2-1-pro | Volcengine | $0.89 / 1M tokens | $4.42 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 12630-17025ms | Catalog | Platform curated |
| Qwen3.7 MAX GLBqwen/qwen3.7-max-glb | Qwen | $2.50 / 1M tokens | $7.50 / 1M tokens | — | Streaming | General chat via Qwen, API workloads | 6984-13532ms | Catalog | Platform curated |
| Grok 4.1 Fast NON Reasoningx-ai/grok-4.1-fast-non-reasoning | X Ai | $0.2 / 1M tokens | $0.5 / 1M tokens | — | Streaming | General chat via X Ai, API workloads | 596-17561ms | Catalog | Platform curated |
| Grok 4.1 Fast Reasoningx-ai/grok-4.1-fast-reasoning | X Ai | $0.2 / 1M tokens | $0.5 / 1M tokens | — | Streaming | General chat via X Ai, API workloads | 1412-3439ms | Catalog | Platform curated |
| Grok 4.20 NON Reasoningx-ai/grok-4.20-non-reasoning | X Ai | $1.25 / 1M tokens | $2.50 / 1M tokens | — | Streaming | General chat via X Ai, API workloads | 825-2803ms | Catalog | Platform curated |
| Grok 4.20 Reasoningx-ai/grok-4.20-reasoning | X Ai | $1.25 / 1M tokens | $2.50 / 1M tokens | — | Streaming | General chat via X Ai, API workloads | 2477-3116ms | Catalog | Platform curated |
| DeepSeek: DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | DeepSeek | $0.45 / 1M tokens | $1.33 / 1M tokens | 1.3M | StreamingTool callingJSON modeLong context | Chinese Q&A, general chat | 0-0ms | Catalog | Platform curated |
| Qwen: Qwen3.7 Flashqwen/qwen3.7-flash | Qwen | Starting at $0.03 / 1M tokens | $0.12 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Google: Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | $0.3 / 1M tokens | $2.50 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 973-1444ms | Catalog | Platform curated | |
| OpenAI: GPT-5.6 Solopenai/gpt-5.6-sol | OpenAI | Starting at $5 / 1M tokens | $30 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 2004-3263ms | Catalog | Platform curated |
| OpenAI: GPT-5.6 Lunaopenai/gpt-5.6-luna | OpenAI | Starting at $0.2 / 1M tokens | $1.20 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 2007-3004ms | Catalog | Platform curated |
| OpenAI: GPT-5.6 Terraopenai/gpt-5.6-terra | OpenAI | Starting at $2 / 1M tokens | $12 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1819-3372ms | Catalog | Platform curated |
| Z.ai: GLM 5.2z-ai/glm-5.2 | Z Ai | $1.18 / 1M tokens | $4.12 / 1M tokens | 1M | StreamingTool callingJSON modeLong context | General chat via Z Ai, API workloads | 0-0ms | Catalog | Platform curated |
| MoonshotAI: Kimi K2.7 Codemoonshotai/kimi-k2.7-code | Moonshotai | $0.9 / 1M tokens | $3.72 / 1M tokens | 262.1k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | Qwen | Starting at $0.3 / 1M tokens | $1.18 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1151-1755ms | Catalog | Platform curated |
| Qwen: Qwen3.7 Maxqwen/qwen3.7-max | Qwen | $1.65 / 1M tokens | $4.96 / 1M tokens | 1M | StreamingTool callingJSON modeLong context | General chat via Qwen, API workloads | 0-0ms | Catalog | Platform curated |
| Google: Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite | $0.25 / 1M tokens | $1.50 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1001-1682ms | Catalog | Platform curated | |
| OpenAI: GPT Chat Latestopenai/gpt-chat-latest | OpenAI | $5 / 1M tokens | $30 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 2075-4931ms | Catalog | Platform curated |
| SpaceXAI: Grok 4.3x-ai/grok-4.3 | X Ai | $1.25 / 1M tokens | $2.50 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 3013-4858ms | Catalog | Platform curated |
| Qwen: Qwen3.6 Flashqwen/qwen3.6-flash | Qwen | Starting at $0.17 / 1M tokens | $0.99 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 299-578ms | Catalog | Platform curated |
| Qwen: Qwen3.6 Max Previewqwen/qwen3.6-max-preview | Qwen | Starting at $1.24 / 1M tokens | $7.43 / 1M tokens | 262.1k | StreamingTool callingJSON modeLong context | General chat via Qwen, API workloads | 0-0ms | Catalog | Platform curated |
| DeepSeek: DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro | DeepSeek | $1.65 / 1M tokens | $3.31 / 1M tokens | 1M | StreamingTool callingJSON modeLong context | Chinese Q&A, general chat | 1579-3500ms | Catalog | Platform curated |
| DeepSeek: DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | DeepSeek | $0.14 / 1M tokens | $0.28 / 1M tokens | 1M | StreamingTool callingJSON modeLong context | Chinese Q&A, general chat | 1282-3456ms | Catalog | Platform curated |
| OpenAI: GPT-5.5openai/gpt-5.5 | OpenAI | Starting at $5 / 1M tokens | $30 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1703-3835ms | Catalog | Platform curated |
| MoonshotAI: Kimi K2.6moonshotai/kimi-k2.6 | Moonshotai | $0.9 / 1M tokens | $3.72 / 1M tokens | 262.1k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 746-1563ms | Catalog | Platform curated |
| OpenAI: GPT-4o-miniopenai/gpt-4o-mini | OpenAI | $0.15 / 1M tokens | $0.6 / 1M tokens | 128k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1601-2808ms | Catalog | Platform curated |
| Z.ai: GLM 5.1z-ai/glm-5.1 | Z Ai | Starting at $0.83 / 1M tokens | $3.31 / 1M tokens | 204.8k | StreamingTool callingJSON modeLong context | General chat via Z Ai, API workloads | 0-0ms | Catalog | Platform curated |
| Qwen: Qwen3.6 Plusqwen/qwen3.6-plus | Qwen | Starting at $0.28 / 1M tokens | $1.66 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 747-1108ms | Catalog | Platform curated |
| OpenAI: GPT-5.4 Nanoopenai/gpt-5.4-nano | OpenAI | $0.2 / 1M tokens | $1.25 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1803-3982ms | Catalog | Platform curated |
| OpenAI: GPT-5.4 Miniopenai/gpt-5.4-mini | OpenAI | $0.75 / 1M tokens | $4.50 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1574-3485ms | Catalog | Platform curated |
| Deepseek V3.2 CNdeepseek/deepseek-v3.2-cn | DeepSeek | $0.29 / 1M tokens | $0.44 / 1M tokens | — | Streaming | Chinese Q&A, general chat | 899-2063ms | Catalog | Platform curated |
| OpenAI: GPT-5.4 Proopenai/gpt-5.4-pro | OpenAI | Starting at $30 / 1M tokens | $180 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 10078-21029ms | Catalog | Platform curated |
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | Anthropic | $1 / 1M tokens | $5 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1510-2510ms | Catalog | Platform curated |
| OpenAI: GPT-5.4openai/gpt-5.4 | OpenAI | Starting at $2.50 / 1M tokens | $15 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1450-3122ms | Catalog | Platform curated |
| OpenAI: GPT-5.3-Codexopenai/gpt-5.3-codex | OpenAI | $1.75 / 1M tokens | $14 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 2440-8057ms | Catalog | Platform curated |
| Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1861-2676ms | Catalog | Platform curated |
| Z.ai: GLM 5z-ai/glm-5 | Z Ai | Starting at $0.58 / 1M tokens | $2.58 / 1M tokens | 204.8k | StreamingTool callingJSON modeLong context | General chat via Z Ai, API workloads | 1417-2913ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 2663-3214ms | Catalog | Platform curated |
| MiniMax: MiniMax M2.5minimax/minimax-m2.5 | MiniMax | $0.31 / 1M tokens | $1.22 / 1M tokens | 204.8k | StreamingTool callingJSON modeLong context | Chinese Q&A, general chat | 1216-2730ms | Catalog | Platform curated |
| MoonshotAI: Kimi K2.5moonshotai/kimi-k2.5 | Moonshotai | $0.58 / 1M tokens | $3.02 / 1M tokens | 262.1k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 2698-6761ms | Catalog | Platform curated |
| OpenAI: GPT-5.2-Codexopenai/gpt-5.2-codex | OpenAI | $1.75 / 1M tokens | $14 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 2127-13503ms | Catalog | Platform curated |
| OpenAI: GPT-5.1-Codex-Maxopenai/gpt-5.1-codex-max | OpenAI | $1.25 / 1M tokens | $10 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1232-3138ms | Catalog | Platform curated |
| OpenAI: GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini | OpenAI | $0.25 / 1M tokens | $2 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1813-3242ms | Catalog | Platform curated |
| OpenAI: GPT-5.2openai/gpt-5.2 | OpenAI | $1.75 / 1M tokens | $14 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1941-3076ms | Catalog | Platform curated |
| OpenAI: GPT-5.1-Codexopenai/gpt-5.1-codex | OpenAI | $1.25 / 1M tokens | $10 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1850-2885ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1867-2824ms | Catalog | Platform curated |
| OpenAI: GPT-5 Proopenai/gpt-5-pro | OpenAI | $15 / 1M tokens | $120 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 14129-21589ms | Catalog | Platform curated |
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 2218-6376ms | Catalog | Platform curated |
| OpenAI: GPT-5.1openai/gpt-5.1 | OpenAI | $1.25 / 1M tokens | $10 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1762-2702ms | Catalog | Platform curated |
| OpenAI: GPT-4.1openai/gpt-4.1 | OpenAI | $2 / 1M tokens | $8 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1364-2415ms | Catalog | Platform curated |
| OpenAI: o4 Miniopenai/o4-mini | OpenAI | $1.10 / 1M tokens | $4.40 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1510-2671ms | Catalog | Platform curated |
| OpenAI: GPT-5openai/gpt-5 | OpenAI | $1.25 / 1M tokens | $10 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 2058-8840ms | Catalog | Platform curated |
| OpenAI: GPT-5 Nanoopenai/gpt-5-nano | OpenAI | $0.05 / 1M tokens | $0.4 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1681-3917ms | Catalog | Platform curated |
| OpenAI: o3openai/o3 | OpenAI | $2 / 1M tokens | $8 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1650-3772ms | Catalog | Platform curated |
| OpenAI: GPT-4.1 Miniopenai/gpt-4.1-mini | OpenAI | $0.4 / 1M tokens | $1.60 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1530-2650ms | Catalog | Platform curated |
| OpenAI: GPT-4oopenai/gpt-4o | OpenAI | $2.50 / 1M tokens | $10 / 1M tokens | 128k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1427-2784ms | Catalog | Platform curated |
| OpenAI: GPT-4.1 Nanoopenai/gpt-4.1-nano | OpenAI | $0.1 / 1M tokens | $0.4 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1692-2754ms | Catalog | Platform curated |
| Google: Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | $0.5 / 1M tokens | $3 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated | |
| OpenAI: GPT-5 Miniopenai/gpt-5-mini | OpenAI | $0.25 / 1M tokens | $2 / 1M tokens | 400k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1898-3460ms | Catalog | Platform curated |
| Google: Gemini 2.5 Flashgoogle/gemini-2.5-flash | $0.3 / 1M tokens | $2.50 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated | |
| Doubao Seedance 2 0 260128doubao/doubao-seedance-2-0-260128 | Doubao | $0.148 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Doubao Seedance 2 0 Fast 260128doubao/doubao-seedance-2-0-fast-260128 | Doubao | $0.119 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Doubao Seedance 2 0 Mini 260615doubao/doubao-seedance-2-0-mini-260615 | Doubao | $0.0738 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Dreamina Seedance 2 0 260128dreamina/dreamina-seedance-2-0-260128 | Dreamina | $0.153 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Dreamina Seedance 2 0 Fast 260128dreamina/dreamina-seedance-2-0-fast-260128 | Dreamina | $0.122 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Dreamina Seedance 2 0 Mini 260615dreamina/dreamina-seedance-2-0-mini-260615 | Dreamina | $0.0764 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Gemini 3 PRO Imagegemini/gemini-3-pro-image | Gemini | $0.314 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Gemini 3.1 Flash Imagegemini/gemini-3.1-flash-image | Gemini | $0.151 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Gemini 3.1 Flash Lite Imagegemini/gemini-3.1-flash-lite-image | Gemini | $0.0336 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Kling V1 6 CNkling/kling-v1-6-cn | Kling | $0.525 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Kling V2 5 Turbo CNkling/kling-v2-5-turbo-cn | Kling | $0.075 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Kling V2 6 CNkling/kling-v2-6-cn | Kling | $0.18 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Kling V3 CNkling/kling-v3-cn | Kling | $0.45 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| TTS 1openai/tts-1 | OpenAI | $0 / 1M tokens | $0 / 1M tokens | — | General chat via OpenAI, API workloads | 0-0ms | Catalog | Platform curated | |
| TTS 1 HDopenai/tts-1-hd | OpenAI | $0 / 1M tokens | $0 / 1M tokens | — | General chat via OpenAI, API workloads | 0-0ms | Catalog | Platform curated | |
| Qwen Image 2.0 CNqwen/qwen-image-2.0-cn | Qwen | $0.0287 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Qwen Image 2.0 GLBqwen/qwen-image-2.0-glb | Qwen | $0.035 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Qwen Image 2.0 PRO CNqwen/qwen-image-2.0-pro-cn | Qwen | $0.0717 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Qwen Image 2.0 PRO GLBqwen/qwen-image-2.0-pro-glb | Qwen | $0.075 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.6 T2I CNqwen/wan2.6-t2i-cn | Qwen | $0.0287 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.6 T2I GLBqwen/wan2.6-t2i-glb | Qwen | $0.03 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.7 I2V CNqwen/wan2.7-i2v-cn | Qwen | $0.138 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.7 I2V GLBqwen/wan2.7-i2v-glb | Qwen | $0.15 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.7 Image CNqwen/wan2.7-image-cn | Qwen | $0.0275 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.7 Image GLBqwen/wan2.7-image-glb | Qwen | $0.03 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.7 Image PRO CNqwen/wan2.7-image-pro-cn | Qwen | $0.0688 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.7 Image PRO GLBqwen/wan2.7-image-pro-glb | Qwen | $0.075 / image | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.7 R2V CNqwen/wan2.7-r2v-cn | Qwen | $0.138 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.7 R2V GLBqwen/wan2.7-r2v-glb | Qwen | $0.15 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.7 T2V CNqwen/wan2.7-t2v-cn | Qwen | $0.138 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Wan2.7 T2V GLBqwen/wan2.7-t2v-glb | Qwen | $0.15 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Viduq3 PRO CNvidu/viduq3-pro-cn | Vidu | $0.11 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |
| Viduq3 Turbo CNvidu/viduq3-turbo-cn | Vidu | $0.0598 / s | — | — | StreamingVision | image understanding, multimodal chat | 0-0ms | Catalog | Platform curated |