处理中…正在保护此关键操作,请稍候
模型短名单

适合图像理解的视觉模型 API

比较截图、文档、商品图、客服附件等场景的视觉模型 API,并标价格与能力。

这份候选名单适合什么用途: 视觉模型

视觉 API 适合截图、收据、商品图、工单附件。关键是模型能不能吃你的图、要不要 JSON 输出、在你自己的样本上错得有多离谱。价格放第二。这页只放一小撮候选,方便你 A/B,而不是扫整本目录。

来源依据: NextModel 的能力映射,以及可用时的 OpenRouter 输入模态元数据。 · 更新日期 2026-07-01

步骤

如何使用这份名单(视觉模型)

  1. 对照任务选模型 先看「视觉模型」短名单是否覆盖你的真实任务,不要只比标价。
  2. 用同一批提示词试跑 挑 2–3 个候选,用业务提示词对比质量与输出长度。
  3. 估算月费 用价格页或成本计算器,按预计 token 量估算月度花费。
  4. 定兜底与预算 定主模型、兜底模型,设项目预算后再接生产流量。

匹配分

推荐候选 视觉模型

先从候选名单开始,再用真实提示词测试,并在接入生产路由前比较月度成本。

AnthropicCatalog

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

$1 / 1M tokensInput$5 / 1M tokensOutput200kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

$5 / 1M tokensInput$25 / 1M tokensOutput200kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

$5 / 1M tokensInput$25 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

$3 / 1M tokensInput$15 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details

比较表

按价格、提供方、上下文、能力和来源比较这份候选名单。

缩小生产候选、建立兜底策略或比较模型经济性时用。

ModelProviderInputOutputContextCapabilitiesBest forLatencyStatusSource
Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5Anthropic$1 / 1M tokens$5 / 1M tokens200k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5Anthropic$5 / 1M tokens$25 / 1M tokens200k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Opus 4.6anthropic/claude-opus-4.6Anthropic$5 / 1M tokens$25 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5Anthropic$3 / 1M tokens$15 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6Anthropic$3 / 1M tokens$15 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
FLUX.2 Flexblack-forest-labs/FLUX.2-flexBlack Forest Labs$0.2 / image
StreamingVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
FLUX.2 PROblack-forest-labs/FLUX.2-proBlack Forest Labs$0.075 / image
StreamingVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Doubao Seedance 2 0 260128doubao/doubao-seedance-2-0-260128Doubao$0.148 / s
StreamingVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated

常见问题

视觉模型 常见问题

选视觉模型 API 前应该比较什么?

图像输入支持、是否需要 JSON、延迟、输出成本,以及真实图片上的质量。合成基准经常骗人。

低成本模型可以处理视觉任务吗?

轻量任务可以。密文档和高精度抽取通常需要更强模型和真实评测集。

相关

相关排行榜