Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Best vision model APIs for image understanding
Compare vision model APIs for screenshots, documents, product images, and support tickets, with price and capability labels.
این فهرست کوتاه برای چیست؟: Vision models
Vision APIs are useful for screenshots, receipts, product photos, and ticket attachments. What matters is whether the model accepts the image type you have, whether you need JSON out, and how wrong answers look on your own samples. Price second. This page is a small set of vision-capable candidates so you can A/B a few instead of the whole catalog.
مبنای منبع: NextModel capability mapping and OpenRouter input-modality metadata when available. · بهروزرسانی شد 2026-07-01
Fit score
گزینههای پیشنهادی vision models
از فهرست کوتاه شروع کنید، پرامپتهای واقعی را آزمایش کنید و پیش از مسیردهی در پروداکشن هزینه ماهانه را مقایسه کنید.
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...
جدول مقایسه
فهرست را بر اساس قیمت، ارائهدهنده، زمینه، قابلیتها و منبع مقایسه کنید.
از این نما وقتی استفاده کنید که فهرست پروداکشن را محدود میکنید، سیاست پشتیبان میسازید یا اقتصاد مدلها را مقایسه میکنید.
| Model | Provider | Input | Output | Context | Capabilities | Best for | Latency | Status | Source |
|---|---|---|---|---|---|---|---|---|---|
| Anthropic: Claude Fable 5anthropic/claude-fable-5 | Anthropic | $1.45 / 1M tokens | $7.23 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | Anthropic | $0.145 / 1M tokens | $0.723 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.5anthropic/claude-opus-4.5 | Anthropic | $0.723 / 1M tokens | $3.62 / 1M tokens | 200k | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.6anthropic/claude-opus-4.6 | Anthropic | $0.723 / 1M tokens | $3.62 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.7anthropic/claude-opus-4.7 | Anthropic | $0.723 / 1M tokens | $3.62 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Opus 4.8anthropic/claude-opus-4.8 | Anthropic | $0.723 / 1M tokens | $3.62 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Claude Opus 5anthropic/claude-opus-5 | Anthropic | $0.723 / 1M tokens | $3.62 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Anthropic | $0.434 / 1M tokens | $2.17 / 1M tokens | 1M | StreamingTool callingJSON modeVision | image understanding, multimodal chat | 1000-3000ms | Catalog | Platform curated |
FAQ
Vision models FAQ
What should I compare before choosing a vision model API?
Image input support, JSON mode if you need it, latency, output cost, and quality on your real images. Synthetic benchmarks lie more often than people admit.
Can low-cost models handle vision tasks?
Sometimes, for light work. Dense documents and high-accuracy extraction usually need a stronger model and a real eval set.
رتبهبندیها