نُشر في 2026-08-16 · 林启明
الجواب المباشر
First-party CacheSafety numbers for openai/gpt-5-mini: 50 prompts, 80% Safe Hit Rate, 0% Bad Hit Rate, and what we did not measure. هذا الدليل موجه لفرق المنتج والمنصة التي تقارن جودة النماذج، والتكلفة، وسياسة التوجيه، ومخاطر الإطلاق.
What We Ran
On 2026-08-16, NextModel evaluated `openai/gpt-5-mini` with 50 prompts, 50 valid responses, and 0 errors. The run used cache-safety semantics v1 and promptset v3. All 10 exact repeats hit the exact cache, and 8 of 10 replayed the expected keyword answer. Trap-pair second calls and fresh questions served fresh responses. Cost saved per 1k requests was not recorded.
What The Numbers Mean
Safe Hit Rate is 0.8. That means 80% of exact repeated prompts returned the expected cached answer. Bad Hit Rate is 0, so no wrong cached answer was served in this run. Semantic Trap Failure Rate is 0, meaning no trap-pair second call incorrectly reused a cached response. This was the highest Safe Hit Rate in the wave. The numbers measure cache behavior, not general intelligence. Exact-cache semantics only cover exact repeats, not semantic similarity.
What We Did Not Measure
Release date is unpublished in the catalog snapshot. Independent third-party quality scores are not in these notes. Real latency was not measured; the catalog lists a latency band of 1000 to 3000ms. Tool, JSON, vision, long-context, and agent capabilities are listed but were not validated in this run. The catalog lists a 400000 token context window. No independent verification of that window appears in the notes.
When This Model Is The Wrong Pick
Pick a different model if you need proven reasoning quality, because independent quality scores are unmeasured. Pick a different model if agentic reliability is your main requirement, because this run did not test agents. If your prompts are mostly fresh, cache benefits will not matter. If you need every exact repeat to replay the expected answer, note that only 8 of 10 exact repeats did so. The low price of $0.036 per 1M input tokens and $0.289 per 1M output tokens is attractive, but price only helps when the cache behavior is useful to you.