Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
这份候选名单适合什么用途: Agent 模型
Agent 跑起来输出 token 多,工具循环一旦跑偏就烧钱。接线前先看工具调用、JSON 稳定性、上下文、延迟、输出价。再设预算,避免坏循环一夜掏空账户。
来源依据: NextModel 的能力映射,以及可用时的支持参数元数据。 · 更新日期 2026-07-01
步骤
如何使用这份名单(Agent 模型)
- 对照任务选模型。 先看「Agent 模型」短名单是否覆盖你的真实任务,不要只比标价。
- 用同一批提示词试跑。 挑 2–3 个候选,用业务提示词对比质量与输出长度。
- 估算月费。 用价格页或成本计算器,按预计 token 量估算月度花费。
- 定兜底与预算。 定主模型、兜底模型,设项目预算后再接生产流量。
匹配分
推荐候选 Agent 模型
先从候选名单开始,再用真实提示词测试,并在接入生产路由前比较月度成本。
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
比较表
按价格、提供方、上下文、能力和来源比较这份候选名单。
缩小生产候选、建立兜底策略或比较模型经济性时用。
| Model | Provider | Input | Output | Context | Capabilities | Best for | Latency | Status | Source |
|---|---|---|---|---|---|---|---|---|---|
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | Anthropic | $1 / 1M tokens | $5 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Anthropic | $3 / 1M tokens | $15 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Google: Gemini 2.5 Flashgoogle/gemini-2.5-flash | $0.3 / 1M tokens | $2.50 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated | |
| Google: Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | $0.5 / 1M tokens | $3 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated | |
| Google: Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite | $0.25 / 1M tokens | $1.50 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
常见问题
Agent 模型 常见问题
Agent 模型最重要的能力是什么?
工具调用、结构化 JSON、任务够用的上下文、指令跟得住。其余都是次要。
为什么 Agent 工作流特别烧钱?
规划文本、工具结果回灌上下文、重试都会拉长轨迹。限制步数,并按次记录 token。
Agent 是否总该用最强模型?
不必。简单规划或工具选择可走便宜模型;难步骤再升级。
指南