⚡ BREAKING
ModelsAgents 🇺🇸 26.07.2026 16:05

Anthropic Unveils Claude Opus 5: Agentic Coding and Computer Tasks at Same Opus Price

AnthropicAnthropic
Anthropic has released Claude Opus 5, a new flagship model with improved agentic capabilities, unchanged pricing ($5/$25 per million tokens), and a 1 million token context window. The model shows significant gains in coding, cybersecurity, and autonomous task benchmarks, with an updated thinking mechanism and revised safety policies.
Anthropic has released Claude Opus 5, replacing Opus 4.8 as the flagship Opus-level model. Prices remain the same: $5 per million input tokens and $25 per million output tokens. The model approaches Claude Fable 5 in intelligence at half the price and is now the default model on Claude Max and the most powerful on Claude Pro. On the API level, three key changes have occurred: thinking is enabled by default (the effort parameter controls depth); disabling thinking (setting type: disabled) with effort xhigh or max now returns a 400 error; developers are advised to remove prompts like "enable final check" as the model already self-checks. The model ID is claude-opus-5, context window is 1M tokens (maximum and default), maximum output is 128k tokens via synchronous Messages API and up to 300k via Message Batches API. In benchmarks, the model showed significant growth: on FrontierBench v0.1 — 43.3% (Opus 4.8 — 18.7%), on SWE-bench Verified — 96.0%, OSWorld 2.0 — 70.57%, Zapier AutomationBench — 26.0%. Agentic results are the strongest: the model solved 42 out of 42 IMO 2026 problems (gold medal). On ARC-AGI-3 with high effort — 30.16% (4 times better than previous record). In multimodal tasks, tools (containers, image cropping) are significantly more effective than "thinking": on Chartography with tools — 83.0% vs. 29.6%. In cybersecurity, capabilities increased as a side effect: on ExploitBench the model got 10.14 flags and 99 full exploits (Mythos 5 — 132). Anthropic only relaxed restrictions on source code vulnerability hunting; binary scanning, penetration testing, and exploit generation remain blocked. Classifiers will trigger approximately 85% less often than on Fable 5. On UK AISI tests, the model solved the "The Last Ones" task in 8 out of 10 attempts. According to the RSP program, the model has CB-1 capabilities, but not CB-2. Outstanding result in prompt injection resistance: browser attacks dropped from 31.5% to 3.70% without protection and to 0% with auto mode. However, the company also publishes downsides, including a slightly higher rate of factual hallucinations compared to Opus 4.8.
Source: MarkTechPost — original
Our earlier posts on this topic ↓
Fresh news