/ models

129 models,one endpoint.

Pick by workload first: every card states who supplies the model, what it costs in and out, and how much context you get. Then run your candidates through one OpenAI-compatible endpoint.

routing candidates129/129
Claude Fable 5.1$10/1M
Claude Haiku 4.5$1/1M
Claude Opus 4.5$5/1M
Claude Opus 4.6$5/1M
16providers1sources2Mmax context$0lowest input
Reset

129 of 129 models

Model cards with source labels and copyable OpenAI-compatible calls.

What matters when you pick a model?

Three things decide most of your cost: who supplies the model, what it charges in and out, and whether the context window fits your task — every card states all three. Latency, routing status, and the rest are filters.

OpenAICatalogHealthytypical

GPT-6 Astra via NextModel: 1.1M-token context; Starting at $10 / 1M tokens in / $50 / 1M tokens out per 1M tokens. Capabilities: streaming, tool, json.

Starting at $10 / 1M tokensInput$50 / 1M tokensOutput1.1MContext
Best forimage understanding, multimodal chat
StreamingTool callingJSON modeVisionLong context
Platform curatedCurated and maintained in the NextModel catalog
View details
QwenCatalogHealthytypical

Qwen3.8-Max via NextModel: 1M-token context; $1.77 / 1M tokens in / $5.30 / 1M tokens out per 1M tokens. Capabilities: streaming.

$1.77 / 1M tokensInput$5.30 / 1M tokensOutput1MContext
Best forGeneral chat via Qwen, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
AnthropicCatalogHealthyslow

Claude Fable 5.1 via NextModel: 1M-token context; $10 / 1M tokens in / $50 / 1M tokens out per 1M tokens. Capabilities: streaming, tool, json.

$10 / 1M tokensInput$50 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
StreamingTool callingJSON modeVisionLong context
Platform curatedCurated and maintained in the NextModel catalog
View details
QwenCatalogHealthytypical

Qwen3.8-Max-GLB via NextModel: 1M-token context; $2 / 1M tokens in / $6 / 1M tokens out per 1M tokens. Capabilities: streaming.

$2 / 1M tokensInput$6 / 1M tokensOutput1MContext
Best forGeneral chat via Qwen, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
Black Forest LabsCatalog

FLUX.2 Flex via NextModel: context unpublished; $0.2 / image in / — out per 1M tokens. Capabilities: streaming, vision.

$0.2 / imagePer imageContext
Best forimage understanding, multimodal chat
StreamingVision
Platform curatedCurated and maintained in the NextModel catalog
View details
Black Forest LabsCatalog

FLUX.2 PRO via NextModel: context unpublished; $0.075 / image in / — out per 1M tokens. Capabilities: streaming, vision.

$0.075 / imagePer imageContext
Best forimage understanding, multimodal chat
StreamingVision
Platform curatedCurated and maintained in the NextModel catalog
View details
OpenAICatalog

GPT Image 1 via NextModel: context unpublished; $0.26 / image in / — out per 1M tokens. Capabilities: streaming, vision.

$0.26 / imagePer imageContext
Best forimage understanding, multimodal chat
StreamingVision
Platform curatedCurated and maintained in the NextModel catalog
View details
OpenAICatalog

GPT Image 1 Mini via NextModel: context unpublished; $0.0539 / image in / — out per 1M tokens. Capabilities: streaming, vision.

$0.0539 / imagePer imageContext
Best forimage understanding, multimodal chat
StreamingVision
Platform curatedCurated and maintained in the NextModel catalog
View details
OpenAICatalog

GPT Image 1.5 via NextModel: context unpublished; $0.218 / image in / — out per 1M tokens. Capabilities: streaming, vision.

$0.218 / imagePer imageContext
Best forimage understanding, multimodal chat
StreamingVision
Platform curatedCurated and maintained in the NextModel catalog
View details
OpenAICatalog

GPT Image 2.0 via NextModel: context unpublished; $0.221 / image in / — out per 1M tokens. Capabilities: streaming, vision.

$0.221 / imagePer imageContext
Best forimage understanding, multimodal chat
StreamingVision
Platform curatedCurated and maintained in the NextModel catalog
View details
OpenAICatalog

Text Embedding 3 Small via NextModel: context unpublished; $0.02 / 1M tokens in / $0 / 1M tokens out per 1M tokens. Capabilities: streaming.

$0.02 / 1M tokensInputOutputContext
Best forGeneral chat via OpenAI, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
OpenAICatalog

Text Embedding 3 Large via NextModel: context unpublished; $0.13 / 1M tokens in / $0 / 1M tokens out per 1M tokens. Capabilities: streaming.

$0.13 / 1M tokensInputOutputContext
Best forGeneral chat via OpenAI, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
Z AiCatalogHealthyslow

GLM 5.3 via NextModel: 1M-token context; $1.40 / 1M tokens in / $4.40 / 1M tokens out per 1M tokens. Capabilities: streaming, tool, json.

$1.40 / 1M tokensInput$4.40 / 1M tokensOutput1MContext
Best forGeneral chat via Z Ai, API workloads
StreamingTool callingJSON modeLong context
Platform curatedCurated and maintained in the NextModel catalog
View details
DeepSeekCatalogHealthytypical

Deepseek-V4-Pro-0813 via NextModel: 1M-token context; $1.33 / 1M tokens in / $3.98 / 1M tokens out per 1M tokens. Capabilities: streaming, tool, json.

$1.33 / 1M tokensInput$3.98 / 1M tokensOutput1MContext
Best forChinese Q&A, general chat
StreamingTool callingJSON modeLong context
Platform curatedCurated and maintained in the NextModel catalog
View details
VolcengineCatalogHealthyslow

Doubao Seed 2.0 Lite via NextModel: 256k-token context; Starting at $0.09 / 1M tokens in / $0.53 / 1M tokens out per 1M tokens. Capabilities: streaming.

Starting at $0.09 / 1M tokensInput$0.53 / 1M tokensOutput256kContext
Best forChinese Q&A, general chat
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
VolcengineCatalogHealthyslow

Doubao Seed 2.0 Pro via NextModel: 256k-token context; Starting at $0.48 / 1M tokens in / $2.36 / 1M tokens out per 1M tokens. Capabilities: streaming.

Starting at $0.48 / 1M tokensInput$2.36 / 1M tokensOutput256kContext
Best forChinese Q&A, general chat
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
QwenCatalogHealthyfast

Qwen3-VL-Flash CN via NextModel: 256k-token context; Starting at $0.03 / 1M tokens in / $0.22 / 1M tokens out per 1M tokens. Capabilities: streaming.

Starting at $0.03 / 1M tokensInput$0.22 / 1M tokensOutput256kContext
Best forGeneral chat via Qwen, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
QwenCatalogHealthyfast

Qwen3.5 Flash via NextModel: 1M-token context; Starting at $0.03 / 1M tokens in / $0.29 / 1M tokens out per 1M tokens. Capabilities: streaming.

Starting at $0.03 / 1M tokensInput$0.29 / 1M tokensOutput1MContext
Best forGeneral chat via Qwen, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
QwenCatalogHealthytypical

Qwen3.5 Plus via NextModel: 1M-token context; Starting at $0.12 / 1M tokens in / $0.69 / 1M tokens out per 1M tokens. Capabilities: streaming.

Starting at $0.12 / 1M tokensInput$0.69 / 1M tokensOutput1MContext
Best forGeneral chat via Qwen, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
QwenCatalogHealthyfast

Qwen3-VL-Plus CN via NextModel: 256k-token context; Starting at $0.15 / 1M tokens in / $1.44 / 1M tokens out per 1M tokens. Capabilities: streaming.

Starting at $0.15 / 1M tokensInput$1.44 / 1M tokensOutput256kContext
Best forGeneral chat via Qwen, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
QwenCatalogHealthytypical

Qwen3.6 Flash GLB via NextModel: 1M-token context; Starting at $0.25 / 1M tokens in / $1.50 / 1M tokens out per 1M tokens. Capabilities: streaming.

Starting at $0.25 / 1M tokensInput$1.50 / 1M tokensOutput1MContext
Best forGeneral chat via Qwen, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
QwenCatalogHealthyfast

Qwen3-Max CN via NextModel: 256k-token context; Starting at $0.36 / 1M tokens in / $1.44 / 1M tokens out per 1M tokens. Capabilities: streaming.

Starting at $0.36 / 1M tokensInput$1.44 / 1M tokensOutput256kContext
Best forGeneral chat via Qwen, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
QwenCatalogHealthyfast

Qwen3-VL-Plus GLB via NextModel: 256k-token context; Starting at $0.2 / 1M tokens in / $1.60 / 1M tokens out per 1M tokens. Capabilities: streaming.

Starting at $0.2 / 1M tokensInput$1.60 / 1M tokensOutput256kContext
Best forGeneral chat via Qwen, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details
QwenCatalogHealthyslow

Qwen3.6 Plus GLB via NextModel: 1M-token context; Starting at $0.5 / 1M tokens in / $3 / 1M tokens out per 1M tokens. Capabilities: streaming.

Starting at $0.5 / 1M tokensInput$3 / 1M tokensOutput1MContext
Best forGeneral chat via Qwen, API workloads
Streaming
Platform curatedCurated and maintained in the NextModel catalog
View details

18 video models

Video generation models

vapeurcertified

kling/kling-v1-6-cn

$0.525 / soutput video
T2VI2V first frameI2V first + lastR2V references
Durations5s
Resolutions720p
References1-4 reference images
vapeurcertified

kling/kling-v2-6-cn

$0.18 / soutput video
T2VI2V first frameI2V first + last
Durations5s
Resolutions720p
vapeurcertified

kling/kling-v3-cn

$0.45 / soutput video
T2VI2V first frameI2V first + last
Durations5s
Resolutions720p
vapeurcertified

qwen/wan2.7-i2v-cn

$0.138 / soutput video
I2V first frameI2V first + last
Durations5s
Resolutions720p, 1080p
vapeurcertified

qwen/wan2.7-i2v-glb

$0.15 / soutput video
I2V first frameI2V first + last
Durations5s
Resolutions720p, 1080p
vapeurcertified

qwen/wan2.7-r2v-cn

$0.138 / soutput video
R2V references
Durations5s
Resolutions720p, 1080p
References1-4 reference images
vapeurcertified

qwen/wan2.7-r2v-glb

$0.15 / soutput video
R2V references
Durations5s
Resolutions720p, 1080p
References1-4 reference images
vapeurcertified

qwen/wan2.7-t2v-cn

$0.138 / soutput video
T2V
Durations5s
Resolutions720p, 1080p
vapeurcertified

vidu/viduq3-pro-cn

$0.11 / soutput video
T2VI2V first frameI2V first + last
Durations5s
Resolutions540p, 720p, 1080p
vapeurcertified

vidu/viduq3-turbo-cn

$0.0598 / soutput video
T2VI2V first frameI2V first + last
Durations5s
Resolutions540p, 720p, 1080p
ModelProviderInputOutputContextCapabilitiesBest forLatencyStatusSource
GPT-6 Astraopenai/gpt-6-astraOpenAIStarting at $10 / 1M tokens$50 / 1M tokens1.1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat2766-6283msCatalogPlatform curated
Qwen3.8-Maxqwen/qwen3.8-maxQwen$1.77 / 1M tokens$5.30 / 1M tokens1M
Streaming
General chat via Qwen, API workloads1672-5802msCatalogPlatform curated
Claude Fable 5.1anthropic/claude-fable-5.1Anthropic$10 / 1M tokens$50 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat4209-4963msCatalogPlatform curated
Qwen3.8-Max-GLBqwen/qwen3.8-max-glbQwen$2 / 1M tokens$6 / 1M tokens1M
Streaming
General chat via Qwen, API workloads1957-5865msCatalogPlatform curated
FLUX.2 Flexblack-forest-labs/FLUX.2-flexBlack Forest Labs$0.2 / image
StreamingVision
image understanding, multimodal chat0-0msCatalogPlatform curated
FLUX.2 PROblack-forest-labs/FLUX.2-proBlack Forest Labs$0.075 / image
StreamingVision
image understanding, multimodal chat0-0msCatalogPlatform curated
GPT Image 1openai/gpt-image-1OpenAI$0.26 / image
StreamingVision
image understanding, multimodal chat0-0msCatalogPlatform curated
GPT Image 1 Miniopenai/gpt-image-1-miniOpenAI$0.0539 / image
StreamingVision
image understanding, multimodal chat0-0msCatalogPlatform curated
GPT Image 1.5openai/gpt-image-1.5OpenAI$0.218 / image
StreamingVision
image understanding, multimodal chat0-0msCatalogPlatform curated
GPT Image 2.0openai/gpt-image-2OpenAI$0.221 / image
StreamingVision
image understanding, multimodal chat0-0msCatalogPlatform curated
Text Embedding 3 Smallopenai/text-embedding-3-smallOpenAI$0.02 / 1M tokens$0 / 1M tokens
Streaming
General chat via OpenAI, API workloads0-0msCatalogPlatform curated
Text Embedding 3 Largeopenai/text-embedding-3-largeOpenAI$0.13 / 1M tokens$0 / 1M tokens
Streaming
General chat via OpenAI, API workloads0-0msCatalogPlatform curated
GLM 5.3z-ai/glm-5.3Z Ai$1.40 / 1M tokens$4.40 / 1M tokens1M
StreamingTool callingJSON modeLong context
General chat via Z Ai, API workloads5137-7205msCatalogPlatform curated
Deepseek-V4-Pro-0813deepseek/deepseek-v4-pro-0813DeepSeek$1.33 / 1M tokens$3.98 / 1M tokens1M
StreamingTool callingJSON modeLong context
Chinese Q&A, general chat2075-14930msCatalogPlatform curated
Doubao Seed 2.0 Litedoubao-seed-2-0-liteVolcengineStarting at $0.09 / 1M tokens$0.53 / 1M tokens256k
Streaming
Chinese Q&A, general chat19256-25059msCatalogPlatform curated
Doubao Seed 2.0 Prodoubao-seed-2-0-proVolcengineStarting at $0.48 / 1M tokens$2.36 / 1M tokens256k
Streaming
Chinese Q&A, general chat5387-7239msCatalogPlatform curated
Qwen3-VL-Flash CNqwen/qwen3-vl-flash-cnQwenStarting at $0.03 / 1M tokens$0.22 / 1M tokens256k
Streaming
General chat via Qwen, API workloads835-1062msCatalogPlatform curated
Qwen3.5 Flashqwen/qwen3.5-flashQwenStarting at $0.03 / 1M tokens$0.29 / 1M tokens1M
Streaming
General chat via Qwen, API workloads1085-1339msCatalogPlatform curated
Qwen3.5 Plusqwen/qwen3.5-plusQwenStarting at $0.12 / 1M tokens$0.69 / 1M tokens1M
Streaming
General chat via Qwen, API workloads2698-2964msCatalogPlatform curated
Qwen3-VL-Plus CNqwen/qwen3-vl-plus-cnQwenStarting at $0.15 / 1M tokens$1.44 / 1M tokens256k
Streaming
General chat via Qwen, API workloads1076-1257msCatalogPlatform curated
Qwen3.6 Flash GLBqwen/qwen3.6-flash-glbQwenStarting at $0.25 / 1M tokens$1.50 / 1M tokens1M
Streaming
General chat via Qwen, API workloads3643-4402msCatalogPlatform curated
Qwen3-Max CNqwen/qwen3-max-cnQwenStarting at $0.36 / 1M tokens$1.44 / 1M tokens256k
Streaming
General chat via Qwen, API workloads1152-1427msCatalogPlatform curated
Qwen3-VL-Plus GLBqwen/qwen3-vl-plus-2025-12-19-glbQwenStarting at $0.2 / 1M tokens$1.60 / 1M tokens256k
Streaming
General chat via Qwen, API workloads606-670msCatalogPlatform curated
Qwen3.6 Plus GLBqwen/qwen3.6-plus-glbQwenStarting at $0.5 / 1M tokens$3 / 1M tokens1M
Streaming
General chat via Qwen, API workloads5518-6156msCatalogPlatform curated