Claude Opus 4.5, Anthropic’s new model that dominates in software engineering

24/11/2025

Anthropic has introduced Claude Opus 4.5, an AI model that sets new highs in real-world software development tests. The system incorporates improvements in programming, agent usage and computer control, along with updates to products like Claude Code and Excel.

Claude Opus 4.5, Anthropic’s new model that dominates in software engineering

Anthropic has announced Claude Opus 4.5, available today on its API, applications and the three major cloud platforms. The system achieves 80.9% accuracy on SWE-bench Verified, the benchmark evaluation for software engineering under real conditions, surpassing models like Sonnet 4.5 (77.2%) and other industry competitors. In multilingual programming it leads in 7 out of 8 evaluated languages.

In addition to leading in agent programming, the model shows superior capabilities across multiple technical areas. In tool usage it achieves 98.2% in telecommunications scenarios and 88.9% in retail environments. In computer use tasks it registers 66.3%, and in visual reasoning 80.7%. The system also achieves 90.8% in multilingual responses and 87% in university-level reasoning.

A distinctive feature is the new effort parameter in the API, which allows developers to adjust the balance between capability and token consumption. With medium effort level, Opus 4.5 matches Sonnet 4.5's performance using 76% fewer output tokens. At its maximum level, it exceeds Sonnet 4.5 by 4.3 percentage points while consuming 48% fewer tokens.

Anthropic conducted internal testing where Claude Opus 4.5 completed a two-hour technical exam for performance engineering candidates, obtaining the highest score recorded among all evaluated human candidates. The company indicates that this result raises questions about how artificial intelligence will modify software development as a profession.

In terms of security, Claude Opus 4.5 presents greater resistance to prompt injection attacks than any other model on the market. In tests with a thousand queries, the model registers an attack success rate of 4.7%, compared to 7.3% for Sonnet 4.5, 12.5% for Gemini 3 Pro and 12.6% for GPT-5.1.

Anthropic has updated several products leveraging the model's capabilities. Claude Code incorporates a Plan mode that generates editable files before executing tasks. Conversations in applications no longer have length limits, as the system automatically summarizes previous context. Claude for Excel has expanded beta access to all Max, Team and Enterprise users.

Claude Opus 4.5 is available on the API with identifier claude-opus-4-5-20251101. Pricing is set at $5 per million input tokens and $25 per million output tokens.

Key points

  • Claude Opus 4.5 achieves 80.9% on SWE-bench Verified, surpassing Sonnet 4.5 and other competitors
  • Leads in 7 out of 8 programming languages in multilingual evaluations
  • Incorporates effort parameter allowing token consumption adjustment while maintaining or exceeding performance
  • With medium effort level matches Sonnet 4.5 using 76% fewer output tokens
  • Surpassed all human candidates in two-hour performance engineering technical exam
  • Records 4.7% vulnerability to prompt injection attacks, the lowest rate on the market
  • Claude Code incorporates Plan mode with editable files before executing tasks
  • Conversations in applications no longer have length limits

Videos

Related AI

Anthropic

AI systems you can rely on

Anthropic develops reliable and interpretable artificial intelligence systems through a scientific approach to safety. The company integrates advanced research and multidisciplinary collaboration to ...

Claude

Create with Claude

Claude is a conversational AI system from Anthropic designed to process natural language and images, providing analysis, logical reasoning, code generation, and multilingual communication under ...

Lastest news

★★★★★
Rate us on Google
This website uses technical, personalization and analysis cookies, both our own and from third parties, to facilitate anonymous browsing and analyze website usage statistics. We consider that if you continue browsing, you accept their use.