Anthropic launches Claude Opus 4.6, the company’s most advanced model

05/02/2026

Anthropic introduces Claude Opus 4.6, with improvements in reasoning, coding and long-context handling. The model leads several benchmarks against GPT-5.2 and Gemini 3 Pro.

Anthropic launches Claude Opus 4.6, the company’s most advanced model

Claude Opus 4.6 is the new version of Anthropic's most capable model, available on claude.ai, the API and all major cloud platforms, at the same price as its predecessor: $5 and $25 per million input and output tokens respectively.

The most significant technical addition is a one-million-token context window in beta, arriving for the first time in the Opus model family, enabling the model to process entire codebases, lengthy contracts or large volumes of documents in a single query. On the MRCR v2 benchmark, which evaluates information retrieval in very long texts, Opus 4.6 reaches 76% compared to 18.5% for Sonnet 4.5.

Comparative benchmarks against other models show strong results across several categories. In agentic terminal coding (Terminal-Bench 2.0), it scores 65.4%, ahead of GPT-5.2 at 64.7% and Gemini 3 Pro at 56.2%. In agentic search (BrowseComp), it reaches 84%, compared to 77.9% for GPT-5.2 and 59.2% for Gemini 3 Pro. In economically valuable office tasks (GDPVal-AA), it achieves 1,606 Elo points versus 1,462 for GPT-5.2 and 1,195 for Gemini 3 Pro, which translates to outperforming the second-best model on the market roughly 70% of the time on this evaluation.

In coding, Opus 4.6 reaches 80.8% on SWE-bench Verified, plans more carefully, sustains agentic tasks for longer and catches its own mistakes more reliably. Claude Code now supports agent teams that work in parallel on different parts of the same project.

On safety, the model maintains a low rate of undesired behaviors comparable to Opus 4.5, until now the most aligned model in the company, and records the lowest rate of incorrect refusals on legitimate queries among recent Claude models.

Opus 4.6 also extends output capacity to 128,000 tokens and comes alongside improvements to Claude in Excel and the research preview launch of Claude in PowerPoint, available for Max, Team and Enterprise plans.

For developers, the API introduces four effort control levels, adaptive thinking and context compaction in beta.

Key points

  • Claude Opus 4.6 is Anthropic's most advanced model, priced the same as its predecessor.
  • It introduces a one-million-token context window in beta for the first time in the Opus family.
  • On the MRCR v2 long-context benchmark, it jumps from Sonnet 4.5's 18.5% to 76%.
  • It outperforms GPT-5.2 and Gemini 3 Pro on the main benchmarks for coding, agentic search and office tasks.
  • In coding (SWE-bench Verified), it reaches 80.8% and enables agent teams in Claude Code.
  • It maintains a safety profile comparable to Opus 4.5, the company's best-aligned model to date.
  • Output capacity is extended to 128,000 tokens.
  • Claude in PowerPoint launches in research preview and Claude in Excel receives improvements.
  • The API introduces adaptive thinking, four effort control levels and context compaction.

Videos

Related AI

Anthropic

AI systems you can rely on

Anthropic develops reliable and interpretable artificial intelligence systems through a scientific approach to safety. The company integrates advanced research and multidisciplinary collaboration to ...

Claude

Create with Claude

Claude is a conversational AI system from Anthropic designed to process natural language and images, providing analysis, logical reasoning, code generation, and multilingual communication under ...

Lastest news

★★★★★
Rate us on Google
This website uses technical, personalization and analysis cookies, both our own and from third parties, to facilitate anonymous browsing and analyze website usage statistics. We consider that if you continue browsing, you accept their use.