Anthropic launched the most recent Sonnet-class model, Claude Sonnet 5, which enhances performance compared to earlier models by improving coding capabilities, agentic performance, and token usage efficiency.

Anthropic highlighted Sonnet 5’s capability to independently perform complex tasks with minimal human intervention, including planning, utilizing tools like browsers and terminals, and operating at a level that previously necessitated larger and costlier models.

Sonnet 5 Uses Less Tokens Efficiently

Sonnet 5 offers better affordability and quality compared to 4.6, although Opus 4.8 remains more accurate. Anthropic suggests adjusting the effort level to optimize the balance between cost and performance when using Sonnet 5. Additionally, there is a special introductory pricing for Sonnet 5 until August 31st, with rates of $2 per MToK input and $10 per MToK output.

Sonnet 5 Performance Tests

Sonnet 5 surpasses Sonnet 4.6, GPT-5.5, and Gemini 3.5 in several comparisons.

The BrowseComp evaluates an AI agent’s ability to discover hard-to-find information on the internet.

BrowseComp results:

  • Claude Sonnet 5 discusses a single agent in line 84.7.
  • Claude’s Sonnet 4.6 can be found in line 76.2.
  • GPT-5.5: 84.4

Terminal-Bench 2.1 evaluates how well an AI model can perform coding tasks in a terminal and CLI environment.

Terminal-Bench 2.1 results:

  • Claude’s Sonnet 5 is found in line 80.4.
  • Claude’s Sonnet 4.6 is found in line 67 of the text.
  • GPT-5.5 scored 83.4 in Codex CLI.
  • Gemini 3.5 is compatible with Flash 76.2.

Sonnet 5 excelled compared to other similar LLMs in the software engineering benchmark known as SWE-bench Pro.

SWE-bench Pro results:

  • Claude’s Sonnet 5: 63.2
  • Claude’s Sonnet 4.6 is numbered as 58.1.
  • GPT-5.5: 58.6
  • Gemini 3.5 Flash recorded a speed of 55.1.

FrontierCode is a standard for independent coding in 150 tasks, with Sonnet 5 outperforming GPT-5.5 significantly in this standard.

The Claude Sonnet 5 System Card provides information.

The agent is provided with a repository and a description of a single issue for each task. It then independently operates in a containerized setting to create a final patch, without any human involvement or time constraints.

Patches are evaluated based on specific criteria such as unit tests and rubric criteria, including model-graded checks for test coverage and prohibited implementation patterns, ensuring fairness through manual verification.

The FrontierCode ratings are:

  • Claude’s Sonnet 5: 38.8
  • Claude’s Sonnet 4.6 can be found in line 15.1.
  • GPT-5.5: Twenty-five point five

Sonnet 5 is considered as “almost genius-level intelligence.”

Anthropic states that Sonnet 5 is not considered a groundbreaking model, but it is described as their most proficient in the Sonnet class. The system card indicates that it is not as advanced as Anthropic’s Opus and Mythos models. However, Anthropic asserts that Sonnet 5 offers intelligence close to that of Opus at a more affordable price for coding, agents, and professional tasks.

Visit Anthropic to read the complete announcement.

Image provided by Shutterstock/jackpress

LEAVE A REPLY

Please enter your comment!
Please enter your name here