Product产品

Introducing Claude Sonnet 5隆重推出 Claude Sonnet 5

2026年6月30日2026年6月30日
Introducing Claude Sonnet 5

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.Claude Sonnet 5 旨在成为迄今为止最具代理能力的 Sonnet 模型。它能够制定计划、使用浏览器和终端等工具,并以一种在几个月前还需要更大、更昂贵的模型才能达到的水平进行自主运行。

For many developers, the agentic AI era began with Sonnet-class models: Claude Sonnet 3.5, 3.6, and 3.7 were the first models that showed impressive skills in coding and tool use. More recently, though, the clearest gains in agentic capabilities have been in our Opus-class models.对于许多开发者而言,代理 AI 时代始于 Sonnet 级模型:Claude Sonnet 3.5、3.6 和 3.7 是首批在编码和工具使用方面展现出令人印象深刻技能的模型。然而,最近代理能力最显著的提升出现在我们的 Opus 级模型中。

Sonnet 5 narrows the gap: its performance is close to that of Opus 4.8, but at lower prices. It’s a substantial improvement over its predecessor, Sonnet 4.6, on important aspects of agentic performance like reasoning, tool use, coding, and knowledge work:Sonnet 5 缩小了这一差距:其性能接近 Opus 4.8,但价格更低。在推理、工具使用、编码和知识工作等重要的代理性能方面,它比其前身 Sonnet 4.6 有了实质性的改进:

Claude Sonnet 5 benchmark table
Scores for Sonnet 5 on a variety of evaluations compared to those of Sonnet 4.6 and Opus 4.8 (a more generally capable model, for reference). The Claude Sonnet 5 System Card reports a broader set of evaluations in detail.Sonnet 5 在各项评估中的得分与 Sonnet 4.6 和 Opus 4.8(作为参考,这是一款能力更全面的模型)的对比。Claude Sonnet 5 系统卡(System Card)中报告了更广泛的评估详情。

Our safety assessments found that Sonnet 5 shows an overall lower rate of undesirable behaviors than Sonnet 4.6, and is generally safer to use in agentic contexts. Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models.我们的安全评估发现,Sonnet 5 的不良行为发生率总体低于 Sonnet 4.6,在代理场景中使用通常更安全。评估还显示,它执行网络安全任务的能力远低于我们目前的 Opus 模型。

From today, Claude Sonnet 5 is available across all plans: it is the default model for Free and Pro plans, and is available to Max, Team, and Enterprise users. It’s also available in Claude Code and on the Claude Platform, where it launches with introductory pricing of $2 per million input tokens and $10 per million output tokens through August 31, 2026, after which it will be priced at $3 per million input tokens and $15 per million output tokens. Developers can use claude-sonnet-5 via the Claude API.从今天起,Claude Sonnet 5 已在所有计划中可用:它是 Free 和 Pro 计划的默认模型,并向 Max、Team 和 Enterprise 用户开放。它也可在 Claude Code 和 Claude Platform 上使用,发布时的入门价格为每百万输入 token 2 美元,每百万输出 token 10 美元(有效期至 2026 年 8 月 31 日),此后价格将调整为每百万输入 token 3 美元,每百万输出 token 15 美元。开发者可以通过 Claude API 使用 claude-sonnet-5。

Working with Claude Sonnet 5使用 Claude Sonnet 5

The charts below compare the performance of Sonnet 5 with Sonnet 4.6 and Opus 4.8 at different effort levels on the agentic search evaluation BrowseComp and the computer use evaluation OSWorld-Verified. Sonnet 5 (orange line) is a strict improvement over Sonnet 4.6 (gray line) and covers a much wider range of cost-performance options than Opus 4.8 (yellow line). It provides substantially improved cost efficiency at medium effort; its higher-effort performance can match Opus 4.8 on some tasks. Between Sonnet 5 and Opus 4.8, users can adjust the effort level to find the right balance of cost and performance.下表比较了 Sonnet 5 与 Sonnet 4.6 和 Opus 4.8 在代理搜索评估 BrowseComp 和计算机使用评估 OSWorld-Verified 中不同努力程度下的性能。Sonnet 5(橙色线)是对 Sonnet 4.6(灰色线)的严格改进,并且涵盖了比 Opus 4.8(黄色线)更广泛的性价比选择。它在中等努力程度下提供了显著提高的成本效率;其高努力程度下的性能在某些任务上可以媲美 Opus 4.8。在 Sonnet 5 和 Opus 4.8 之间,用户可以调整努力程度,以找到成本与性能之间的最佳平衡点。

Feedback from our early access partners has been consistent: Sonnet 5 is much more agentic than its predecessors. Testers described how it finishes complex tasks where previous Sonnet models would stop short, how it checks its own output without explicitly being asked, and how it does all this agentic work at an attractive price point:来自我们早期访问合作伙伴的反馈非常一致:Sonnet 5 比其前身更具代理能力。测试人员描述了它如何完成以前的 Sonnet 模型会中途停止的复杂任务,如何在未被明确要求的情况下检查自己的输出,以及它如何以极具吸引力的价格点完成所有这些代理工作:

Safety evaluations安全评估

Our pre-deployment safety evaluations found that Sonnet 5 was overall an improvement on Sonnet 4.6. On agentic safety, the model is better at refusing malicious requests and resisting hijack attempts in prompt injection attacks. The model shows lower rates of hallucination and sycophancy than Sonnet 4.6. On our automated behavioral audit, which tests a wide range of misaligned behaviors such as cooperation with misuse and deception, Sonnet 5 scored lower (that is, safer) overall. However, it did show somewhat higher rates of misaligned behavior on this assessment compared to the more capable Opus 4.8 and Claude Mythos Preview.我们的部署前安全评估发现,Sonnet 5 总体上是对 Sonnet 4.6 的改进。在代理安全方面,该模型在拒绝恶意请求和抵御提示词注入攻击中的劫持尝试方面表现更好。该模型表现出的幻觉和阿谀奉承率低于 Sonnet 4.6。在我们测试各种不一致行为(如配合滥用和欺骗)的自动化行为审计中,Sonnet 5 的总体得分更低(即更安全)。然而,与能力更强的 Opus 4.8 和 Claude Mythos Preview 相比,它在此评估中确实表现出稍高的不一致行为率。

Rates of misaligned behavior across Claude models
Rates of misaligned behavior on our automated behavioral audit, which tests for a very wide range of undesirable behaviors across many situations and contexts (see Section 6.4 of the Sonnet 5 System Card for a complete list and results for each specific behavior). Sonnet 5 shows an overall lower rate of misaligned behavior than Sonnet 4.6, though a higher rate than Mythos Preview and Opus 4.8.我们自动化行为审计中的不一致行为率,该审计测试了许多情况和背景下非常广泛的不良行为(完整列表及每种特定行为的结果请参阅 Sonnet 5 系统卡的第 6.4 节)。Sonnet 5 显示出比 Sonnet 4.6 更低的总体不一致行为率,尽管比 Mythos Preview 和 Opus 4.8 的比率要高。

We did not deliberately train Sonnet 5 on cybersecurity tasks. It can perform some routine, non-harmful cyber tasks, but on evaluations testing potentially dangerous cyber skills, such as developing software exploits, it shows substantially poorer performance than models such as Opus 4.8 and Mythos 5. Scores from one evaluation, which tested models’ ability to develop exploits for vulnerabilities in the Firefox browser, are shown in the chart below. Sonnet 5 was never able to develop a full working exploit, but it does show a slightly higher rate of partial success than Sonnet 4.6. This latter change is likely due to improvements in general intelligence rather than specific training.我们没有刻意在网络安全任务上训练 Sonnet 5。它可以执行一些常规的、无害的网络任务,但在测试潜在危险网络技能(如开发软件漏洞利用)的评估中,它的表现明显差于 Opus 4.8 和 Mythos 5 等模型。下表显示了其中一项评估的得分,该评估测试了模型开发 Firefox 浏览器漏洞利用的能力。Sonnet 5 从未能够开发出完整的可用漏洞利用,但它确实显示出比 Sonnet 4.6 略高的部分成功率。后者的变化很可能是由于通用智能的提升,而非特定训练的结果。

Scores measuring Claude models’ success at developing exploits for software vulnerabilities in Firefox 147
Scores measuring models’ success at developing exploits for software vulnerabilities in Firefox 147 (this evaluation was developed in collaboration with Mozilla; all vulnerabilities have been patched in Firefox 148). For each model, the left-hand bar shows how often the model (without safeguards) developed a working exploit; the right-hand bar shows how often the model had partial success. Neither of the Sonnet models could successfully develop a working exploit (both scored 0.0%); Sonnet 5 showed a slightly higher partial success rate than Sonnet 4.6. Both Sonnet models have substantially poorer cyber capabilities than Opus 4.8 and Mythos 5. For full details, see Section 3.2.4 of the Sonnet 5 System Card.衡量模型在 Firefox 147 中开发软件漏洞利用成功率的得分(此评估是与 Mozilla 合作开发的;所有漏洞已在 Firefox 148 中修补)。对于每个模型,左侧柱状图显示模型(无防护措施)开发出可用漏洞利用的频率;右侧柱状图显示模型获得部分成功的频率。两个 Sonnet 模型都无法成功开发出可用的漏洞利用(得分均为 0.0%);Sonnet 5 显示出比 Sonnet 4.6 略高的部分成功率。两个 Sonnet 模型的网络能力都明显弱于 Opus 4.8 和 Mythos 5。详情请参阅 Sonnet 5 系统卡的第 3.2.4 节。

Since Sonnet 5 is somewhat stronger than its predecessor on these tasks, we’ve launched it with cyber safeguards enabled by default. These safeguards—which detect and block dangerous cyber usage in real time—are the same as those present in Claude Opus 4.7 and 4.8 (because we judged that the overall level of cybersecurity risk from Sonnet 5 was low, the safeguards are less strict than those launched with Fable 5, which block a much wider range of cybersecurity tasks).1由于 Sonnet 5 在这些任务上比其前身稍强,我们默认启用了网络安全防护措施发布了它。这些防护措施——实时检测并阻止危险的网络使用——与 Claude Opus 4.7 和 4.8 中存在的相同(因为我们判断 Sonnet 5 带来的总体网络安全风险较低,所以防护措施不如 Fable 5 发布时那么严格,后者阻止了更广泛的网络安全任务)。1

Our full assessment of Sonnet 5 across many safety and capability evaluations is reported in the Claude Sonnet 5 System Card.我们对 Sonnet 5 在多项安全和能力评估中的全面评估报告在 Claude Sonnet 5 系统卡中。

Availability and pricing可用性和定价

Claude Sonnet 5 is available everywhere today at an introductory price of $2 per million input tokens and $10 per million output tokens through August 31, 2026. It then moves to standard pricing at $3 per million input tokens and $15 per million output tokens.2 We’ve increased rate limits across Chat, Cowork, Claude Code, and the Claude Platform3 to accommodate the higher token usage of higher effort levels; users can select whichever level makes sense for their particular project.Claude Sonnet 5 即日起在全球范围内可用,入门价格为每百万输入 token 2 美元,每百万输出 token 10 美元(有效期至 2026 年 8 月 31 日)。此后将采用标准定价,即每百万输入 token 3 美元,每百万输出 token 15 美元。2 我们提高了 Chat、Cowork、Claude Code 和 Claude Platform3 的速率限制,以适应更高努力程度下更高的 token 使用量;用户可以选择适合其特定项目的任何级别。

Changelog更新日志

Edit June 30, 2026: In the original version of this post, we included a cost-performance chart for the BrowseComp evaluation that was based on data from a simpler methodology that did not reflect the standard methodology we use for agentic search evaluations. This had the result of underestimating Sonnet 5's performance on the evaluation.2026 年 6 月 30 日编辑:在本文的原始版本中,我们包含了一张基于更简单方法论的 BrowseComp 评估性价比图表,该方法论并未反映我们用于代理搜索评估的标准方法论。这导致低估了 Sonnet 5 在该评估中的表现。

We have now updated the chart so that it matches the methodology that we used and discussed in the Sonnet 5 system card (which used a 10M token budget with compaction and programmatic tool calling). We have also updated the surrounding text.我们现在已经更新了图表,使其符合我们在 Sonnet 5 系统卡中使用和讨论的方法论(使用了带有压缩和程序化工具调用的 10M token 预算)。我们还更新了周围的文字。

Footnotes脚注

1 Sonnet 5 is part of our Cyber Verification Program, which is available today on the native Claude Platform, the Claude Platform on AWS, and Claude in Microsoft Foundry (hosted on Azure and Anthropic), and coming soon on Claude in Google Vertex. Organizations that are already enrolled in the Cyber Verification Program automatically have the same access on Sonnet 5, with no need to reapply. Overall, we recommend Claude Opus 4.8 for cybersecurity work that requires reduced guardrails.1 Sonnet 5 是我们网络验证计划(Cyber Verification Program)的一部分,该计划现已在原生 Claude Platform、AWS 上的 Claude Platform 以及 Microsoft Foundry 上的 Claude(托管在 Azure 和 Anthropic 上)中提供,并即将登陆 Google Vertex 上的 Claude。已经加入网络验证计划的组织在 Sonnet 5 上自动拥有相同的访问权限,无需重新申请。总体而言,我们建议将 Claude Opus 4.8 用于需要减少护栏的网络安全工作。

2 Sonnet 5 is an upgrade to Sonnet 4.6, but it uses an updated tokenizer that changes how the model processes text to improve performance (this is similar to the tokenizer change we introduced with Claude Opus 4.7). The tradeoff is that the same input can map to more tokens: roughly 1.0–1.35× depending on the content type. The introductory pricing is set so that the transition to Sonnet 5 is roughly cost-neutral.2 Sonnet 5 是 Sonnet 4.6 的升级版,但它使用了更新的分词器(tokenizer),改变了模型处理文本的方式以提高性能(这类似于我们在 Claude Opus 4.7 中引入的分词器更改)。代价是相同的输入可能会映射到更多的 token:根据内容类型,大约是 1.0–1.35 倍。入门定价的设定使得向 Sonnet 5 的过渡大致是成本中性的。

3 On April 26, 2026, we raised Sonnet and Haiku rate limits at every usage tier and simplified to three tiers (Start, Build, and Scale) on the native Claude Platform. You can view your tier and current limits in the Claude Console or read the documentation to learn more.3 2026 年 4 月 26 日,我们提高了所有使用层级的 Sonnet 和 Haiku 速率限制,并简化为原生 Claude Platform 上的三个层级(Start、Build 和 Scale)。您可以在 Claude 控制台中查看您的层级和当前限制,或阅读文档以了解更多信息。

  • Humanity’s Last Exam: We updated the grader model for Humanity’s Last Exam and have updated the Sonnet 4.6 score to 34.6% (no tools) and 46.8% (with tools). This is the reason the score differs from that reported in the Sonnet 4.6 launch blog.人类终极考试(Humanity’s Last Exam):我们更新了人类终极考试的评分模型,并将 Sonnet 4.6 的得分更新为 34.6%(无工具)和 46.8%(有工具)。这就是得分与 Sonnet 4.6 发布博客中报告的得分不同的原因。
  • OSWorld-Verified: We made changes to how we run the OSWorld-Verified evaluation to more accurately reflect the model’s performance in the real world, and have updated the Sonnet 4.6 score to 78.5%. This is the reason the score differs from that reported in the Sonnet 4.6 launch blog.OSWorld-Verified:我们更改了运行 OSWorld-Verified 评估的方式,以更准确地反映模型在现实世界中的表现,并将 Sonnet 4.6 的得分更新为 78.5%。这就是得分与 Sonnet 4.6 发布博客中报告的得分不同的原因。

Related content相关内容

Redeploying Fable 5重新部署 Fable 5

Fable 5 returns globally July 1. We're also proposing an industry-wide framework for scoring jailbreak severity, together with Amazon, Microsoft, Google, and other Glasswing partners. Fable 5 将于 7 月 1 日在全球回归。我们还与亚马逊、微软、谷歌和其他 Glasswing 合作伙伴一起,提议建立一个全行业的越狱严重程度评分框架。

Read more阅读更多

Claude Science, an AI workbench for scientists, is now availableClaude Science,一款面向科学家的 AI 工作台,现已推出

Claude Science is a customizable app that integrates the tools and packages researchers most often use, produces auditable artifacts, and provides flexible access to computing resources.Claude Science 是一款可定制的应用程序,它集成了研究人员最常使用的工具和包,生成可审计的工件,并提供对计算资源的灵活访问。

Read more

Introducing Claude Tag隆重推出 Claude Tag

Claude Tag is a new way for teams to work with Claude.Claude Tag 是团队使用 Claude 的一种新方式。

Read more