Introducing Claude Sonnet 5隆重推出 Claude Sonnet 5

Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.Claude Sonnet 5 旨在成为迄今为止最具代理能力的 Sonnet 模型。它能够制定计划、使用浏览器和终端等工具,并以一种在几个月前还需要更大、更昂贵的模型才能达到的水平进行自主运行。
For many developers, the agentic AI era began with Sonnet-class models: Claude Sonnet 3.5, 3.6, and 3.7 were the first models that showed impressive skills in coding and tool use. More recently, though, the clearest gains in agentic capabilities have been in our Opus-class models.对于许多开发者而言,代理 AI 时代始于 Sonnet 级模型:Claude Sonnet 3.5、3.6 和 3.7 是首批在编码和工具使用方面展现出令人印象深刻技能的模型。然而,最近代理能力最显著的提升出现在我们的 Opus 级模型中。
Sonnet 5 narrows the gap: its performance is close to that of Opus 4.8, but at lower prices. It’s a substantial improvement over its predecessor, Sonnet 4.6, on important aspects of agentic performance like reasoning, tool use, coding, and knowledge work:Sonnet 5 缩小了这一差距:其性能接近 Opus 4.8,但价格更低。在推理、工具使用、编码和知识工作等重要的代理性能方面,它比其前身 Sonnet 4.6 有了实质性的改进:

Our safety assessments found that Sonnet 5 shows an overall lower rate of undesirable behaviors than Sonnet 4.6, and is generally safer to use in agentic contexts. Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models.我们的安全评估发现,Sonnet 5 的不良行为发生率总体低于 Sonnet 4.6,在代理场景中使用通常更安全。评估还显示,它执行网络安全任务的能力远低于我们目前的 Opus 模型。
From today, Claude Sonnet 5 is available across all plans: it is the default model for Free and Pro plans, and is available to Max, Team, and Enterprise users. It’s also available in Claude Code and on the Claude Platform, where it launches with introductory pricing of $2 per million input tokens and $10 per million output tokens through August 31, 2026, after which it will be priced at $3 per million input tokens and $15 per million output tokens. Developers can use claude-sonnet-5 via the Claude API.从今天起,Claude Sonnet 5 已在所有计划中可用:它是 Free 和 Pro 计划的默认模型,并向 Max、Team 和 Enterprise 用户开放。它也可在 Claude Code 和 Claude Platform 上使用,发布时的入门价格为每百万输入 token 2 美元,每百万输出 token 10 美元(有效期至 2026 年 8 月 31 日),此后价格将调整为每百万输入 token 3 美元,每百万输出 token 15 美元。开发者可以通过 Claude API 使用 claude-sonnet-5。
Working with Claude Sonnet 5使用 Claude Sonnet 5
The charts below compare the performance of Sonnet 5 with Sonnet 4.6 and Opus 4.8 at different effort levels on the agentic search evaluation BrowseComp and the computer use evaluation OSWorld-Verified. Sonnet 5 (orange line) is a strict improvement over Sonnet 4.6 (gray line) and covers a much wider range of cost-performance options than Opus 4.8 (yellow line). It provides substantially improved cost efficiency at medium effort; its higher-effort performance can match Opus 4.8 on some tasks. Between Sonnet 5 and Opus 4.8, users can adjust the effort level to find the right balance of cost and performance.下表比较了 Sonnet 5 与 Sonnet 4.6 和 Opus 4.8 在代理搜索评估 BrowseComp 和计算机使用评估 OSWorld-Verified 中不同努力程度下的性能。Sonnet 5(橙色线)是对 Sonnet 4.6(灰色线)的严格改进,并且涵盖了比 Opus 4.8(黄色线)更广泛的性价比选择。它在中等努力程度下提供了显著提高的成本效率;其高努力程度下的性能在某些任务上可以媲美 Opus 4.8。在 Sonnet 5 和 Opus 4.8 之间,用户可以调整努力程度,以找到成本与性能之间的最佳平衡点。

Feedback from our early access partners has been consistent: Sonnet 5 is much more agentic than its predecessors. Testers described how it finishes complex tasks where previous Sonnet models would stop short, how it checks its own output without explicitly being asked, and how it does all this agentic work at an attractive price point:来自我们早期访问合作伙伴的反馈非常一致:Sonnet 5 比其前身更具代理能力。测试人员描述了它如何完成以前的 Sonnet 模型会中途停止的复杂任务,如何在未被明确要求的情况下检查自己的输出,以及它如何以极具吸引力的价格点完成所有这些代理工作:
Claude Sonnet 5 gives our agents a strong execution layer for multi-step software engineering work. It handles sustained coding, tool use, and debugging well across messy technical contexts, and has been especially useful for workflows where follow-through and technical grounding matter.Claude Sonnet 5 为我们的代理提供了强大的执行层,用于多步骤软件工程工作。它在混乱的技术环境中能很好地处理持续的编码、工具使用和调试,对于那些注重后续跟进和技术基础的工作流特别有用。Claude Sonnet 5 为我们的代理提供了强大的执行层,用于多步骤软件工程工作。它在混乱的技术环境中能很好地处理持续的编码、工具使用和调试,对于那些注重后续跟进和技术基础的工作流特别有用。Zimu Li,技术人员
We handed Claude Sonnet 5 a two-part job—update Salesforce account tiers, send a launch announcement to enterprise contacts—and it finished end to end. That used to stall halfway. For day-to-day automation, it’s a no-brainer.我们交给 Claude Sonnet 5 一项两部分的工作——更新 Salesforce 账户层级,向企业联系人发送发布公告——它从头到尾完成了。这以前总是会中途停滞。对于日常自动化来说,这是一个无需多想的选择。我们交给 Claude Sonnet 5 一项两部分的工作——更新 Salesforce 账户层级,向企业联系人发送发布公告——它从头到尾完成了。这以前总是会中途停滞。对于日常自动化来说,这是一个无需多想的选择。Daniel Shepard,高级工程师
Claude Sonnet 5 gets more done with less. Same output quality, fewer steps to get there. It refuses unsafe requests cleanly and consistently, too. At Lovable, we’re putting powerful tools in the hands of millions of builders. A model that knows when to say no is just as important as one that knows how to build.Claude Sonnet 5 用更少的资源完成了更多的工作。同样的输出质量,更少的步骤。它也能干净利落地拒绝不安全的请求。在 Lovable,我们正在将强大的工具交到数百万构建者手中。一个知道何时说“不”的模型,与一个知道如何构建的模型同样重要。Claude Sonnet 5 用更少的资源完成了更多的工作。同样的输出质量,更少的步骤。它也能干净利落地拒绝不安全的请求。在 Lovable,我们正在将强大的工具交到数百万构建者手中。一个知道何时说“不”的模型,与一个知道如何构建的模型同样重要。Fabian Hedin,联合创始人
We ran Claude Sonnet 5 against dozens of our most challenging real pull requests, and it carried each one through to a tested, verified result on its own — freeing our engineers to focus on the judgment, the decision, and the final sign-off.我们用 Claude Sonnet 5 测试了数十个最具挑战性的真实拉取请求(pull requests),它独自完成了每一个请求,并给出了经过测试和验证的结果——让我们的工程师能够专注于判断、决策和最终签字。我们用 Claude Sonnet 5 测试了数十个最具挑战性的真实拉取请求(pull requests),它独自完成了每一个请求,并给出了经过测试和验证的结果——让我们的工程师能够专注于判断、决策和最终签字。Yusuke Kaji,AI 商业总经理
I asked Claude Sonnet 5 to investigate a bug. Unprompted, it wrote a reproducing test, implemented the fix, then stashed it to confirm the bug came back without the change. All in a single pass.我让 Claude Sonnet 5 调查一个 bug。它在没有提示的情况下,编写了一个复现测试,实施了修复,然后将其暂存以确认在没有更改的情况下 bug 会复现。所有这些都在一次运行中完成。我让 Claude Sonnet 5 调查一个 bug。它在没有提示的情况下,编写了一个复现测试,实施了修复,然后将其暂存以确认在没有更改的情况下 bug 会复现。所有这些都在一次运行中完成。Neel Chotai,Rust 工程师和软件工程师
With Claude Sonnet 5, agents stay on plan, follow our conventions, and ship clean multi-step changes, all at an efficient cost.有了 Claude Sonnet 5,代理能够保持计划,遵循我们的约定,并以高效的成本交付干净的多步骤更改。有了 Claude Sonnet 5,代理能够保持计划,遵循我们的约定,并以高效的成本交付干净的多步骤更改。Sualeh Asif,联合创始人
Claude Sonnet 5 is at its best on brownfield code—race conditions, hidden tests, the parts nobody wants to touch. It traces a failure to its actual root cause and ships a durable fix instead of patching the symptom.Claude Sonnet 5 在处理遗留代码(brownfield code)方面表现最佳——竞态条件、隐藏测试、没人愿意碰的部分。它能追踪到故障的真正根源并提供持久的修复,而不是修补症状。Claude Sonnet 5 在处理遗留代码(brownfield code)方面表现最佳——竞态条件、隐藏测试、没人愿意碰的部分。它能追踪到故障的真正根源并提供持久的修复,而不是修补症状。Dominic Elm,创始工程师
Claude Sonnet 5 sits on the Pareto frontier for Eve’s plaintiff-law tasks. We see the clearest gains in legal research and analysis, at a price-to-performance ratio that made the choice to migrate easy.Claude Sonnet 5 处于 Eve 原告法律任务的帕累托前沿。我们在法律研究和分析方面看到了最明显的收益,其性价比使迁移的选择变得容易。Claude Sonnet 5 处于 Eve 原告法律任务的帕累托前沿。我们在法律研究和分析方面看到了最明显的收益,其性价比使迁移的选择变得容易。Mauricio Wulfovich,资深机器学习工程师
ClickHouse agents explore live data and produce insights on the fly, so time-to-insight matters when testing new models. Claude Sonnet 5 reasons in tighter steps and gets our users to answers noticeably faster. That speed is a difference our customers feel.ClickHouse 代理探索实时数据并即时产生见解,因此在测试新模型时,洞察时间(time-to-insight)至关重要。Claude Sonnet 5 以更紧凑的步骤进行推理,并让我们的用户明显更快地获得答案。这种速度是我们的客户能感受到的差异。ClickHouse 代理探索实时数据并即时产生见解,因此在测试新模型时,洞察时间(time-to-insight)至关重要。Claude Sonnet 5 以更紧凑的步骤进行推理,并让我们的用户明显更快地获得答案。这种速度是我们的客户能感受到的差异。Ryadh Dahimene,AI/ML 产品管理总监
At Pace, our computer-use agents run insurance workflows—submission intake, FNOL, loss runs—on the systems our operations teams already use. Claude Sonnet 5 consistently takes the right action and does it quickly, which is what real insurance work demands.在 Pace,我们的计算机使用代理在我们的运营团队已经使用的系统上运行保险工作流——提交接收、FNOL、损失运行。Claude Sonnet 5 始终能采取正确的行动并快速完成,这正是真正的保险工作所要求的。在 Pace,我们的计算机使用代理在我们的运营团队已经使用的系统上运行保险工作流——提交接收、FNOL、损失运行。Claude Sonnet 5 始终能采取正确的行动并快速完成,这正是真正的保险工作所要求的。Eric He,技术人员
Safety evaluations安全评估
Our pre-deployment safety evaluations found that Sonnet 5 was overall an improvement on Sonnet 4.6. On agentic safety, the model is better at refusing malicious requests and resisting hijack attempts in prompt injection attacks. The model shows lower rates of hallucination and sycophancy than Sonnet 4.6. On our automated behavioral audit, which tests a wide range of misaligned behaviors such as cooperation with misuse and deception, Sonnet 5 scored lower (that is, safer) overall. However, it did show somewhat higher rates of misaligned behavior on this assessment compared to the more capable Opus 4.8 and Claude Mythos Preview.我们的部署前安全评估发现,Sonnet 5 总体上是对 Sonnet 4.6 的改进。在代理安全方面,该模型在拒绝恶意请求和抵御提示词注入攻击中的劫持尝试方面表现更好。该模型表现出的幻觉和阿谀奉承率低于 Sonnet 4.6。在我们测试各种不一致行为(如配合滥用和欺骗)的自动化行为审计中,Sonnet 5 的总体得分更低(即更安全)。然而,与能力更强的 Opus 4.8 和 Claude Mythos Preview 相比,它在此评估中确实表现出稍高的不一致行为率。

We did not deliberately train Sonnet 5 on cybersecurity tasks. It can perform some routine, non-harmful cyber tasks, but on evaluations testing potentially dangerous cyber skills, such as developing software exploits, it shows substantially poorer performance than models such as Opus 4.8 and Mythos 5. Scores from one evaluation, which tested models’ ability to develop exploits for vulnerabilities in the Firefox browser, are shown in the chart below. Sonnet 5 was never able to develop a full working exploit, but it does show a slightly higher rate of partial success than Sonnet 4.6. This latter change is likely due to improvements in general intelligence rather than specific training.我们没有刻意在网络安全任务上训练 Sonnet 5。它可以执行一些常规的、无害的网络任务,但在测试潜在危险网络技能(如开发软件漏洞利用)的评估中,它的表现明显差于 Opus 4.8 和 Mythos 5 等模型。下表显示了其中一项评估的得分,该评估测试了模型开发 Firefox 浏览器漏洞利用的能力。Sonnet 5 从未能够开发出完整的可用漏洞利用,但它确实显示出比 Sonnet 4.6 略高的部分成功率。后者的变化很可能是由于通用智能的提升,而非特定训练的结果。

Since Sonnet 5 is somewhat stronger than its predecessor on these tasks, we’ve launched it with cyber safeguards enabled by default. These safeguards—which detect and block dangerous cyber usage in real time—are the same as those present in Claude Opus 4.7 and 4.8 (because we judged that the overall level of cybersecurity risk from Sonnet 5 was low, the safeguards are less strict than those launched with Fable 5, which block a much wider range of cybersecurity tasks).1由于 Sonnet 5 在这些任务上比其前身稍强,我们默认启用了网络安全防护措施发布了它。这些防护措施——实时检测并阻止危险的网络使用——与 Claude Opus 4.7 和 4.8 中存在的相同(因为我们判断 Sonnet 5 带来的总体网络安全风险较低,所以防护措施不如 Fable 5 发布时那么严格,后者阻止了更广泛的网络安全任务)。1
Our full assessment of Sonnet 5 across many safety and capability evaluations is reported in the Claude Sonnet 5 System Card.我们对 Sonnet 5 在多项安全和能力评估中的全面评估报告在 Claude Sonnet 5 系统卡中。
Availability and pricing可用性和定价
Claude Sonnet 5 is available everywhere today at an introductory price of $2 per million input tokens and $10 per million output tokens through August 31, 2026. It then moves to standard pricing at $3 per million input tokens and $15 per million output tokens.2 We’ve increased rate limits across Chat, Cowork, Claude Code, and the Claude Platform3 to accommodate the higher token usage of higher effort levels; users can select whichever level makes sense for their particular project.Claude Sonnet 5 即日起在全球范围内可用,入门价格为每百万输入 token 2 美元,每百万输出 token 10 美元(有效期至 2026 年 8 月 31 日)。此后将采用标准定价,即每百万输入 token 3 美元,每百万输出 token 15 美元。2 我们提高了 Chat、Cowork、Claude Code 和 Claude Platform3 的速率限制,以适应更高努力程度下更高的 token 使用量;用户可以选择适合其特定项目的任何级别。
Changelog更新日志
Edit June 30, 2026: In the original version of this post, we included a cost-performance chart for the BrowseComp evaluation that was based on data from a simpler methodology that did not reflect the standard methodology we use for agentic search evaluations. This had the result of underestimating Sonnet 5's performance on the evaluation.2026 年 6 月 30 日编辑:在本文的原始版本中,我们包含了一张基于更简单方法论的 BrowseComp 评估性价比图表,该方法论并未反映我们用于代理搜索评估的标准方法论。这导致低估了 Sonnet 5 在该评估中的表现。
We have now updated the chart so that it matches the methodology that we used and discussed in the Sonnet 5 system card (which used a 10M token budget with compaction and programmatic tool calling). We have also updated the surrounding text.我们现在已经更新了图表,使其符合我们在 Sonnet 5 系统卡中使用和讨论的方法论(使用了带有压缩和程序化工具调用的 10M token 预算)。我们还更新了周围的文字。
Footnotes脚注
1 Sonnet 5 is part of our Cyber Verification Program, which is available today on the native Claude Platform, the Claude Platform on AWS, and Claude in Microsoft Foundry (hosted on Azure and Anthropic), and coming soon on Claude in Google Vertex. Organizations that are already enrolled in the Cyber Verification Program automatically have the same access on Sonnet 5, with no need to reapply. Overall, we recommend Claude Opus 4.8 for cybersecurity work that requires reduced guardrails.1 Sonnet 5 是我们网络验证计划(Cyber Verification Program)的一部分,该计划现已在原生 Claude Platform、AWS 上的 Claude Platform 以及 Microsoft Foundry 上的 Claude(托管在 Azure 和 Anthropic 上)中提供,并即将登陆 Google Vertex 上的 Claude。已经加入网络验证计划的组织在 Sonnet 5 上自动拥有相同的访问权限,无需重新申请。总体而言,我们建议将 Claude Opus 4.8 用于需要减少护栏的网络安全工作。
2 Sonnet 5 is an upgrade to Sonnet 4.6, but it uses an updated tokenizer that changes how the model processes text to improve performance (this is similar to the tokenizer change we introduced with Claude Opus 4.7). The tradeoff is that the same input can map to more tokens: roughly 1.0–1.35× depending on the content type. The introductory pricing is set so that the transition to Sonnet 5 is roughly cost-neutral.2 Sonnet 5 是 Sonnet 4.6 的升级版,但它使用了更新的分词器(tokenizer),改变了模型处理文本的方式以提高性能(这类似于我们在 Claude Opus 4.7 中引入的分词器更改)。代价是相同的输入可能会映射到更多的 token:根据内容类型,大约是 1.0–1.35 倍。入门定价的设定使得向 Sonnet 5 的过渡大致是成本中性的。
3 On April 26, 2026, we raised Sonnet and Haiku rate limits at every usage tier and simplified to three tiers (Start, Build, and Scale) on the native Claude Platform. You can view your tier and current limits in the Claude Console or read the documentation to learn more.3 2026 年 4 月 26 日,我们提高了所有使用层级的 Sonnet 和 Haiku 速率限制,并简化为原生 Claude Platform 上的三个层级(Start、Build 和 Scale)。您可以在 Claude 控制台中查看您的层级和当前限制,或阅读文档以了解更多信息。
- Humanity’s Last Exam: We updated the grader model for Humanity’s Last Exam and have updated the Sonnet 4.6 score to 34.6% (no tools) and 46.8% (with tools). This is the reason the score differs from that reported in the Sonnet 4.6 launch blog.人类终极考试(Humanity’s Last Exam):我们更新了人类终极考试的评分模型,并将 Sonnet 4.6 的得分更新为 34.6%(无工具)和 46.8%(有工具)。这就是得分与 Sonnet 4.6 发布博客中报告的得分不同的原因。
- OSWorld-Verified: We made changes to how we run the OSWorld-Verified evaluation to more accurately reflect the model’s performance in the real world, and have updated the Sonnet 4.6 score to 78.5%. This is the reason the score differs from that reported in the Sonnet 4.6 launch blog.OSWorld-Verified:我们更改了运行 OSWorld-Verified 评估的方式,以更准确地反映模型在现实世界中的表现,并将 Sonnet 4.6 的得分更新为 78.5%。这就是得分与 Sonnet 4.6 发布博客中报告的得分不同的原因。
Related content相关内容
Redeploying Fable 5重新部署 Fable 5
Fable 5 returns globally July 1. We're also proposing an industry-wide framework for scoring jailbreak severity, together with Amazon, Microsoft, Google, and other Glasswing partners. Fable 5 将于 7 月 1 日在全球回归。我们还与亚马逊、微软、谷歌和其他 Glasswing 合作伙伴一起,提议建立一个全行业的越狱严重程度评分框架。
Read more阅读更多Claude Science, an AI workbench for scientists, is now availableClaude Science,一款面向科学家的 AI 工作台,现已推出
Claude Science is a customizable app that integrates the tools and packages researchers most often use, produces auditable artifacts, and provides flexible access to computing resources.Claude Science 是一款可定制的应用程序,它集成了研究人员最常使用的工具和包,生成可审计的工件,并提供对计算资源的灵活访问。
Read moreIntroducing Claude Tag隆重推出 Claude Tag
Claude Tag is a new way for teams to work with Claude.Claude Tag 是团队使用 Claude 的一种新方式。
Read more