Anthropic introduces Claude Opus 4.6, with improvements in reasoning, coding and long-context handling. The model leads several benchmarks against GPT-5.2 and Gemini 3 Pro.
Claude Opus 4.6 is the new version of Anthropic's most capable model, available on claude.ai, the API and all major cloud platforms, at the same price as its predecessor: $5 and $25 per million input and output tokens respectively.
The most significant technical addition is a one-million-token context window in beta, arriving for the first time in the Opus model family, enabling the model to process entire codebases, lengthy contracts or large volumes of documents in a single query. On the MRCR v2 benchmark, which evaluates information retrieval in very long texts, Opus 4.6 reaches 76% compared to 18.5% for Sonnet 4.5.
Comparative benchmarks against other models show strong results across several categories. In agentic terminal coding (Terminal-Bench 2.0), it scores 65.4%, ahead of GPT-5.2 at 64.7% and Gemini 3 Pro at 56.2%. In agentic search (BrowseComp), it reaches 84%, compared to 77.9% for GPT-5.2 and 59.2% for Gemini 3 Pro. In economically valuable office tasks (GDPVal-AA), it achieves 1,606 Elo points versus 1,462 for GPT-5.2 and 1,195 for Gemini 3 Pro, which translates to outperforming the second-best model on the market roughly 70% of the time on this evaluation.
In coding, Opus 4.6 reaches 80.8% on SWE-bench Verified, plans more carefully, sustains agentic tasks for longer and catches its own mistakes more reliably. Claude Code now supports agent teams that work in parallel on different parts of the same project.
On safety, the model maintains a low rate of undesired behaviors comparable to Opus 4.5, until now the most aligned model in the company, and records the lowest rate of incorrect refusals on legitimate queries among recent Claude models.
Opus 4.6 also extends output capacity to 128,000 tokens and comes alongside improvements to Claude in Excel and the research preview launch of Claude in PowerPoint, available for Max, Team and Enterprise plans.
For developers, the API introduces four effort control levels, adaptive thinking and context compaction in beta.
Anthropic develops reliable and interpretable artificial intelligence systems through a scientific approach to safety. The company integrates advanced research and multidisciplinary collaboration to ...
Claude is a conversational AI system from Anthropic designed to process natural language and images, providing analysis, logical reasoning, code generation, and multilingual communication under ...
24/07/2026
Anthropic unveils Claude Opus 5, its new Opus-tier model, which nears the intelligence of Fable 5 in coding and knowledge work at half the price, ...
21/07/2026
Hugging Face detected an intrusion it attributed to an autonomous AI agent. Days later, OpenAI confirmed its own models had accessed that ...
16/07/2026
SpaceXAI has introduced Grok 4.5, its most advanced model to date, trained alongside Cursor and built for coding, agentic tasks and knowledge work, ...
16/07/2026
Moonshot AI has introduced Kimi K3, an open-source model with 2.8 trillion parameters that, according to its own data, outperforms several commercial ...