ModelsHardware & Inference 🇨🇳 03.08.2026 11:03

DeepSeek-V4-Flash Official Benchmarks: About 98% Cache Hit for Ultimate Cost-Effectiveness

DeepSeekDeepSeek
DeepSeek has released official benchmarks for its V4-Flash model, highlighting an approximately 98% cache hit rate. This is said to deliver extreme cost-performance, as cached tokens are significantly cheaper. The model has been presented as a competitive option in the AI inference market.
DeepSeek has officially unveiled the benchmark results for its V4-Flash model, showcasing an impressive cache hit rate of about 98%. A high cache hit rate means that a large portion of requests can be served from a cache, reducing computational costs and latency. This efficiency translates into a notably lower price per token, positioning V4-Flash as a highly cost-effective solution for developers and businesses. The model competes in the same space as other fast inference models, but its cache optimization sets it apart in terms of economy. The announcement was made via Tencent News.
Source: DeepSeek (GNews) — original
Our earlier posts on this topic ↓
Fresh news