Models 🇺🇸 27.07.2026 02:02

Kimi K3: 2.8 Trillion Parameters and a New Look at the 'Pelican' Benchmark

Moonshot AIMoonshot AI DeepSeekDeepSeek AnthropicAnthropic OpenAIOpenAI Google/DeepMindGoogle/DeepMind
Chinese lab Moonshot AI has released the Kimi K3 model with 2.8 trillion parameters, calling it the first "open 3T-class" model. K3 surpasses Claude Opus 4.8 max and GPT-5.5 high, but falls behind Claude Fable 5 and GPT-5.6 Sol, with a cost of $3 per million input tokens and $15 per million output tokens, making it the most expensive among Chinese models. The author also notes that the 'pelican' benchmark is losing correlation with model quality but remains useful for quick evaluation. The model is already available through the website and API, with open weights promised by July 27.
Chinese lab Moonshot AI has announced the Kimi K3 model, which it describes as "the most powerful to date" with 2.8 trillion parameters. The company positions it as the first "open 3T-class" model (rounding 2.8 up to 3 trillion), surpassing DeepSeek with its 1.6T v4 Pro. According to its own benchmarks, K3 generally outperforms Claude Opus 4.8 max and GPT-5.5 high, but falls behind Claude Fable 5 and GPT-5.6 Sol. Per the Artificial Analysis report, on their private evaluation of "long-term knowledge," K3 scored an Elo of 1547, which is 732 points higher than Kimi K2.6 and second only to Claude Fable 5. The cost per task ($0.94) is close to GPT-5.6 Sol ($1.04), roughly half the price of Opus 4.8 ($1.80), but higher than open-source alternatives. Token usage for K3 has decreased by 21% compared to K2.6. The model also leads the Frontend Code arena on Arena.ai, even surpassing Claude Fable 5. API pricing: $3 per million input tokens and $15 per million output tokens—on par with Anthropic's Claude Sonnet series, making K3 the most expensive model among Chinese AI labs. This is a significant increase from Kimi K2.6 ($0.95/$4). With 2.8 trillion parameters, it more than doubles the previous 1T model. The author tested K3 via OpenRouter, generating an SVG of a pelican on a bicycle. The request consumed 95 input tokens and 16,658 output tokens (of which 13,241 were reasoning tokens), costing 25 cents. When re-processing the generated SVG, the model correctly described the image (for 0.6 cents). The author notes that over 21 months, the "pelican" test has lost correlation with overall model quality—currently, the best results come from GLM-5.2, although it does not reach the Fable class. Nevertheless, the test remains useful as a "hello world": it encourages trying the model, provides a rough estimate of cost and reasoning time, checks basic geometric understanding and the ability to generate valid SVG. For K3, the test revealed that the model currently uses only one level of reasoning effort ("max"), consumes many reasoning tokens, and also hinted at a possible hidden system prompt (85 tokens). The visual description from the model is quite good.
Source: Simon Willison — original
Our earlier posts on this topic ↓
Fresh news