Grok 4.6 released: benchmarks and the price of long tasks
SpaceXAI released Grok 4.6 with 500K context tokens, four reasoning modes, and a focus on long agentic work. In early tests it performs alongside top competitors, leading CursorBench at an average cost of $2.81 per task. However, the price doubles for contexts over 200K tokens.
On August 12, SpaceXAI released Grok 4.6, which features 500 thousand context tokens, four reasoning modes, and an explicit focus on extended agentic work. In initial benchmarks, the model performs on par with the strongest competitors and comes out ahead in CursorBench, with an average cost of 2.81 dollars per task. That number looks appealing, but the long context acts as a hidden extra: after 200 thousand tokens, the rate doubles. The article examines where the model genuinely improves, how much a long run costs, and who should try it now.
Source: Habr — хаб ИИ —
original
