This model always redirects to the latest model in the Claude Fable family.
這份候選名單適合什麼用途?: LLM gateway
An OpenAI-compatible LLM gateway means you keep the OpenAI SDK, point base_url at the gateway, and route across providers from there. You can compare unit cost, add cache or budgets, and fail over without rewriting app code. This ranking is for teams picking default models for that setup: start with solid general models, define a fallback, then estimate monthly spend. Migration steps live at /docs/openai-compatible. Cost math is at /tools/ai-api-cost-calculator. If you are comparing multi-model marketplaces specifically, use /best/openrouter-alternatives.
來源依據: NextModel catalog taxonomy, OpenAI-compatible gateway positioning, and provider public pricing when available. · 更新日期 2026-08-05
如何使用这份名单
如何使用这份名单 (LLM gateway)
- 对照任务选模型. 先看「LLM gateway」短名单是否覆盖你的真实任务,不要只比标价。
- 用同一批提示词试跑. 挑 2–3 个候选,用业务提示词对比质量与输出长度。
- 估算月费. 用价格页或成本计算器,按预计 token 量估算月度花费。
- 定兜底与预算. 定主模型、兜底模型,设项目预算后再接生产流量。
匹配分
推薦候選 llm gateway
先從候選名單開始,再以真實提示詞測試,並在接入生產路由前比較月度成本。
This model always redirects to the latest model in the Anthropic Claude Haiku family.
This model always redirects to the latest model in the Claude Opus family.
This model always redirects to the latest model in the Anthropic Claude Sonnet family.
比較表
按價格、提供方、上下文、能力與來源比較這份候選名單。
當你在縮小正式環境候選名單、建立兜底策略或比較模型經濟性時,可使用此視圖。
| Model | Provider | Input | Output | Context | Capabilities | Best for | Latency | Status | Source |
|---|---|---|---|---|---|---|---|---|---|
| Anthropic: Claude Fable Latest~anthropic/claude-fable-latest | OpenRouter | $1.88 / 1M tokens | $9.40 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic Claude Haiku Latest~anthropic/claude-haiku-latest | OpenRouter | $0.188 / 1M tokens | $0.94 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Opus Latest~anthropic/claude-opus-latest | OpenRouter | $0.94 / 1M tokens | $4.70 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic Claude Sonnet Latest~anthropic/claude-sonnet-latest | OpenRouter | $0.564 / 1M tokens | $2.82 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Google Gemini Flash Latest~google/gemini-flash-latest | OpenRouter | $0.282 / 1M tokens | $1.69 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Google Gemini Pro Latest~google/gemini-pro-latest | OpenRouter | $0.376 / 1M tokens | $2.26 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| MoonshotAI Kimi Latest~moonshotai/kimi-latest | OpenRouter | $0.124 / 1M tokens | $0.658 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| OpenAI GPT Latest~openai/gpt-latest | OpenRouter | $0.94 / 1M tokens | $5.64 / 1M tokens | 1.1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
常見問題
LLM gateway 常見問題
What is an OpenAI-compatible LLM gateway?
A hosted API that speaks the OpenAI request shape (chat completions, models, keys) and forwards to one or more providers behind a single base URL.
Why use a gateway instead of calling each provider directly?
One integration surface, easier fallbacks, and one place for usage, budgets, and receipts before traffic gets loud.
How do I switch an existing OpenAI app to a gateway?
Change base_url and the API key in most cases. Then verify streaming, tools, JSON mode, and vision against the model IDs you will actually use.
How should teams control cost on a multi-model gateway?
Send cheap work to cheap models, keep stronger models for hard cases, cache only when it is safe, and set project budgets. Estimate monthly cost before you scale.