处理中…正在保护此关键操作,请稍候
模型短名单

适合工具调用工作流的 Agent 模型 API

比较需要工具、JSON、长上下文和可承受预算的 Agent 模型 API。

这份候选名单适合什么用途: Agent 模型

Agent 跑起来输出 token 多,工具循环一旦跑偏就烧钱。接线前先看工具调用、JSON 稳定性、上下文、延迟、输出价。再设预算,避免坏循环一夜掏空账户。

来源依据: NextModel 的能力映射,以及可用时的支持参数元数据。 · 更新日期 2026-07-01

步骤

如何使用这份名单(Agent 模型)

  1. 对照任务选模型 先看「Agent 模型」短名单是否覆盖你的真实任务,不要只比标价。
  2. 用同一批提示词试跑 挑 2–3 个候选,用业务提示词对比质量与输出长度。
  3. 估算月费 用价格页或成本计算器,按预计 token 量估算月度花费。
  4. 定兜底与预算 定主模型、兜底模型,设项目预算后再接生产流量。

匹配分

推荐候选 Agent 模型

先从候选名单开始,再用真实提示词测试,并在接入生产路由前比较月度成本。

AnthropicCatalog

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

$1 / 1M tokensInput$5 / 1M tokensOutput200kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

$5 / 1M tokensInput$25 / 1M tokensOutput200kContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

$5 / 1M tokensInput$25 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details
AnthropicCatalog

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

$3 / 1M tokensInput$15 / 1M tokensOutput1MContext
Best forimage understanding, multimodal chat
RoutingConfigured
StreamingTool callingJSON modeVisionLong context
Platform curatedNextModel gateway catalog (Go origin)
View details

比较表

按价格、提供方、上下文、能力和来源比较这份候选名单。

缩小生产候选、建立兜底策略或比较模型经济性时用。

ModelProviderInputOutputContextCapabilitiesBest forLatencyStatusSource
Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5Anthropic$1 / 1M tokens$5 / 1M tokens200k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5Anthropic$5 / 1M tokens$25 / 1M tokens200k
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Opus 4.6anthropic/claude-opus-4.6Anthropic$5 / 1M tokens$25 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5Anthropic$3 / 1M tokens$15 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6Anthropic$3 / 1M tokens$15 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Google: Gemini 2.5 Flashgoogle/gemini-2.5-flashGoogle$0.3 / 1M tokens$2.50 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Google: Gemini 3 Flash Previewgoogle/gemini-3-flash-previewGoogle$0.5 / 1M tokens$3 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated
Google: Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-liteGoogle$0.25 / 1M tokens$1.50 / 1M tokens1M
StreamingTool callingJSON modeVision
image understanding, multimodal chat1000-3000msCatalogPlatform curated

常见问题

Agent 模型 常见问题

Agent 模型最重要的能力是什么?

工具调用、结构化 JSON、任务够用的上下文、指令跟得住。其余都是次要。

为什么 Agent 工作流特别烧钱?

规划文本、工具结果回灌上下文、重试都会拉长轨迹。限制步数,并按次记录 token。

Agent 是否总该用最强模型?

不必。简单规划或工具选择可走便宜模型;难步骤再升级。