Introducing Claude Opus 5隆重推出 Claude Opus 5

Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.Claude Opus 5 即日起正式发布。这是一款深思熟虑且具有主动性的模型,其智能水平已接近 Claude Fable 5 的前沿水准,而价格仅为其一半。
On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks.在 Frontier-Bench 和 GDPval-AA 等编程与知识工作评估中,Opus 5 树立了新的行业标杆,不过在网络安全任务方面仍落后于 Mythos 5。
Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro.Opus 5 专为日常使用而设计,运行效率高于其他模型。它是 Claude Max 的默认模型,也是 Claude Pro 中性能最强的模型。

Performance and cost-effectiveness性能与成本效益
Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. The charts in this section show how performance changes according to the model’s effort setting, which customers can use to optimize for intelligence or conserve tokens for faster and cheaper results.Claude Opus 5 在保持与前代产品 Opus 4.8 相同成本的同时,实现了性能的显著提升。本节图表展示了性能如何随模型的“努力程度”(effort)设置而变化,客户可以利用此设置来优化智能水平,或通过节省 Token 来获得更快、更经济的结果。
Opus 5 excels on valuable software engineering tasks. For example, on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench 3.2, at max effort, the model performs within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater performance at a given cost than all other models on high, xhigh, and max effort.Opus 5 在高价值软件工程任务中表现卓越。例如,在 Frontier-Bench v0.1 上,Opus 5 超越了所有其他模型,且以更低的单任务成本实现了两倍于 Opus 4.8 的性能。在 CursorBench 3.2 上,当设置为最大努力程度时,该模型的性能与 Fable 5 的峰值得分相差不到 0.5%,但单任务成本却降低了一半;此外,在“高”、“超高”和“最大”努力程度下,它在给定成本下的性能表现均优于所有其他模型。

We see similar results on knowledge work and problem-solving tasks. For example:我们在知识工作和问题解决任务中也看到了类似的结果。例如:
- On ARC-AGI 3, an evaluation where the model has to solve novel problems, Opus 5’s score is three times as high as the next-best model.在 ARC-AGI 3(一项要求模型解决新颖问题的评估)中,Opus 5 的得分是排名第二模型的 3 倍。
- On Zapier AutomationBench, which measures whether models can complete business tasks from start to finish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its lowest effort setting, Opus 5 passes more tasks than any other model.在衡量模型能否从头到尾完成业务任务的 Zapier AutomationBench 上,Opus 5 的任务通过率约为排名第二模型的 1.5 倍(成本相同)。即使在最低努力程度设置下,Opus 5 完成的任务量也超过了任何其他模型。
- On OSWorld 2.0, a computer use benchmark, Opus 5 outperforms every other model at any given cost, surpassing Fable 5’s best result at just over a third of the cost.在计算机操作基准测试 OSWorld 2.0 中,Opus 5 在任何成本下都优于其他所有模型,仅以三分之一多一点的成本就超越了 Fable 5 的最佳成绩。
It’s also our best and most cost-efficient model on several related evaluations:它也是我们在多项相关评估中性能最强且最具成本效益的模型:

Opus 5 is a meaningful improvement over Opus 4.8 for scientific research. It shows better performance than Opus 4.8 on every one of our life sciences evaluations, which cover topics including structural biology, organic chemistry, and bioinformatics. Its improvements are most notable on organic chemistry tasks, like inferring molecular structures from spectroscopy data (it scores 10.2 percentage points higher than Opus 4.8 on our internal benchmark), and on protein-related tasks like predicting how variations in a protein’s sequence affect how it functions (here, it scores 7.7 percentage points higher).对于科学研究而言,Opus 5 相比 Opus 4.8 有了显著进步。在我们涵盖结构生物学、有机化学和生物信息学等主题的所有生命科学评估中,它的表现均优于 Opus 4.8。其改进在有机化学任务中尤为突出,例如从光谱数据推断分子结构(在我们内部基准测试中得分比 Opus 4.8 高出 10.2 个百分点),以及蛋白质相关任务,如预测蛋白质序列变异如何影响其功能(得分高出 7.7 个百分点)。
Finally, Opus 5 is capable of producing much stronger visual outputs:最后,Opus 5 能够生成更出色的视觉输出:
Working with Claude Opus 5与 Claude Opus 5 协作
Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. In evaluations and early-access testing, we and our users found many examples of Opus 5’s agency and thoroughness:Claude Opus 5 在验证自身工作并仔细迭代直至成功方面能力更强。在评估和抢先体验测试中,我们和用户发现了许多体现 Opus 5 主动性和严谨性的案例:
- On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It succeeded in doing so repeatedly; no competing model with the same setup could solve it after five attempts.在一项 Frontier-Bench 任务中,Opus 5 获得了一张机械零件图纸,并被要求编写代码将其重建为 3D FreeCAD 模型。然而,在这个任务中,模型被刻意限制无法直接查看图纸。Opus 5 的应对方式是编写了自己的计算机视觉流水线,从原始像素中提取几何形状,然后重建了完整的机械零件。它多次成功完成了任务;而竞争模型在相同设置下尝试五次后均无法解决。
- Given a real bug in a popular open-source package manager, Opus 5 found the root cause and fixed an edge case that the community’s patch had missed. A competing model fixed only the surface symptom (not the underlying cause), then reported the bug resolved.针对一个流行开源包管理器中的真实漏洞,Opus 5 找到了根本原因并修复了一个社区补丁遗漏的边缘情况。而竞争模型仅修复了表面症状(而非根本原因),随后便报告漏洞已解决。
- An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly.一家贸易公司的工程师使用 Opus 5 在单次会话中构建了一个新交易所的市场数据馈送。之前的模型即便在获得工程师详细计划的情况下也无法完成此任务。由于没有实时馈送进行验证,Opus 5 甚至自行构建了一个测试工具来检查其代码是否正确解析了交易所的数据。
Below are further reports from our early-access customers on their experience of working with Opus 5:以下是我们抢先体验客户在使用 Opus 5 后的更多反馈:
On FrontierCode 1.1, Claude Opus 5 approaches Fable-level performance at half the cost. Within Devin, it also shows particular strength on difficult debugging and root-cause analysis tasks.在 FrontierCode 1.1 上,Claude Opus 5 以一半的成本接近了 Fable 级别的性能。在 Devin 中,它在困难的调试和根本原因分析任务中也表现出了特别的优势。在 FrontierCode 1.1 上,Claude Opus 5 以一半的成本接近了 Fable 级别的性能。在 Devin 中,它在困难的调试和根本原因分析任务中也表现出了特别的优势。——Scott Wu,首席执行官
Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it’s just under Fable 5 and has many of the same behaviors. We are excited to see how developers use it in Cursor.Claude Opus 5 以 Opus 的速度和成本提供了接近 Fable 5 的智能水平。在 CursorBench 上,它略逊于 Fable 5,并具有许多相同的行为特征。我们很高兴看到开发者如何在 Cursor 中使用它。Claude Opus 5 以 Opus 的速度和成本提供了接近 Fable 5 的智能水平。在 CursorBench 上,它略逊于 Fable 5,并具有许多相同的行为特征。我们很高兴看到开发者如何在 Cursor 中使用它。——Sualeh Asif,联合创始人
Claude Opus 5 topped Zapier’s AutomationBench leaderboard without spending more tokens than prior Claude models. It took a raw account-health workbook and ran a full churn-prevention sequence end to end: flagging at-risk accounts, alerting the right owner, and summarizing for retention ops. Previous models didn’t pass; Opus 5 hit 100%.Claude Opus 5 在没有消耗比以往 Claude 模型更多 Token 的情况下,登顶了 Zapier 的 AutomationBench 排行榜。它处理了一个原始的账户健康状况工作簿,并端到端地运行了完整的流失预防流程:标记高风险账户、通知相关负责人,并为留存运营进行总结。之前的模型未能通过测试,而 Opus 5 的准确率达到了 100%。Claude Opus 5 在没有消耗比以往 Claude 模型更多 Token 的情况下,登顶了 Zapier 的 AutomationBench 排行榜。它处理了一个原始的账户健康状况工作簿,并端到端地运行了完整的流失预防流程:标记高风险账户、通知相关负责人,并为留存运营进行总结。之前的模型未能通过测试,而 Opus 5 的准确率达到了 100%。——Wade Foster,首席执行官
On our genomics analysis work, Claude Opus 5 behaves more like a careful scientist than any model we’ve run. It reaches for the right statistical tests to rule out confounders, cross-checks its own results by independent methods, and stays on track through long multi-step analyses.在我们的基因组分析工作中,Claude Opus 5 的表现比我们运行过的任何模型都更像一位严谨的科学家。它会选择正确的统计检验来排除混杂因素,通过独立方法交叉验证自己的结果,并在漫长的多步骤分析中保持进度。在我们的基因组分析工作中,Claude Opus 5 的表现比我们运行过的任何模型都更像一位严谨的科学家。它会选择正确的统计检验来排除混杂因素,通过独立方法交叉验证自己的结果,并在漫长的多步骤分析中保持进度。——Alfredo Andere,首席执行官
Claude Opus 5 came out ahead of every model in its family on our internal evals. It isn’t just better on our hardest agentic coding tasks, up 22% over Opus 4.7, it’s steadier, with far less variance run to run. For the millions of builders on Lovable, that consistency is the whole game. Reliable results, build after build.Claude Opus 5 在我们的内部评估中领先于其系列中的所有模型。它不仅在我们最困难的智能体编程任务上表现更好(较 Opus 4.7 提升了 22%),而且更加稳定,运行间的差异极小。对于 Lovable 的数百万构建者来说,这种一致性至关重要。每一次构建,都能获得可靠的结果。Claude Opus 5 在我们的内部评估中领先于其系列中的所有模型。它不仅在我们最困难的智能体编程任务上表现更好(较 Opus 4.7 提升了 22%),而且更加稳定,运行间的差异极小。对于 Lovable 的数百万构建者来说,这种一致性至关重要。每一次构建,都能获得可靠的结果。——Fabian Hedin,联合创始人
Claude Opus 5 is the biggest leap in the Opus family since 4.5. On the same full-stack app builds, the front end shows it first: the best animations, games, and 3D work we have seen from an Opus model.Claude Opus 5 是 Opus 系列自 4.5 版本以来最大的飞跃。在相同的全栈应用构建中,前端效果最先显现:这是我们在 Opus 模型中见过的最好的动画、游戏和 3D 作品。Claude Opus 5 是 Opus 系列自 4.5 版本以来最大的飞跃。在相同的全栈应用构建中,前端效果最先显现:这是我们在 Opus 模型中见过的最好的动画、游戏和 3D 作品。——Madhav Jha,联合创始人兼首席技术官
We’re loving Claude Opus 5. For the kind of open-ended analytical work our agent handles, it’s a strict upgrade over Opus 4.8, and the gains are biggest exactly where it matters: the harder, vaguer tasks. Responses are clearer and more concise, and we see improved efficiency at higher effort levels too.我们非常喜欢 Claude Opus 5。对于我们的智能体处理的那种开放式分析工作,它比 Opus 4.8 有了严格的提升,而且在最重要的地方(更困难、更模糊的任务)增益最大。响应更清晰、更简洁,我们在更高的努力程度下也看到了效率的提升。我们非常喜欢 Claude Opus 5。对于我们的智能体处理的那种开放式分析工作,它比 Opus 4.8 有了严格的提升,而且在最重要的地方(更困难、更模糊的任务)增益最大。响应更清晰、更简洁,我们在更高的努力程度下也看到了效率的提升。——Izzy Miller,AI 研究主管
Claude Opus 5 is a striking improvement over Opus 4.8 for the financial research workflows our analysts run every day. It stands out on numerical reasoning, table work, and sharper critical thinking where precision matters.Claude Opus 5 对于我们的分析师每天进行的金融研究工作流而言,是相对于 Opus 4.8 的惊人改进。它在数值推理、表格处理和需要精度的敏锐批判性思维方面表现突出。Claude Opus 5 对于我们的分析师每天进行的金融研究工作流而言,是相对于 Opus 4.8 的惊人改进。它在数值推理、表格处理和需要精度的敏锐批判性思维方面表现突出。——Shirley Zhang,应用 AI 高级 AI 工程师
Claude Opus 5 delivers the industry intelligence and accuracy that is essential for the analysis of specialized enterprise content. Box found that Opus 5 outperforms Opus 4.8 by 8% and delivers notable performance gains in the data analysis (11% improvement) and due diligence (17% improvement) workflows that technology, healthcare, and public sector organizations rely on daily.Claude Opus 5 提供了分析专业企业内容所必需的行业智能和准确性。Box 发现,Opus 5 比 Opus 4.8 领先 8%,并在技术、医疗保健和公共部门组织每天依赖的数据分析(提升 11%)和尽职调查(提升 17%)工作流中提供了显著的性能提升。Claude Opus 5 提供了分析专业企业内容所必需的行业智能和准确性。Box 发现,Opus 5 比 Opus 4.8 领先 8%,并在技术、医疗保健和公共部门组织每天依赖的数据分析(提升 11%)和尽职调查(提升 17%)工作流中提供了显著的性能提升。——Ben Kus,首席技术官
Claude Opus 5 is a clear generational step up from Opus 4.8. Over one weekend I gave it a chief-of-staff role over my dev environments: it built its own monitor, drove each box, and pulled me in only for the judgment calls.Claude Opus 5 是相对于 Opus 4.8 的明确代际提升。在一个周末,我赋予了它对我的开发环境的参谋长角色:它构建了自己的监控器,驱动了每个盒子,并且只在需要判断时才叫我。Claude Opus 5 是相对于 Opus 4.8 的明确代际提升。在一个周末,我赋予了它对我的开发环境的参谋长角色:它构建了自己的监控器,驱动了每个盒子,并且只在需要判断时才叫我。——Cristian Rivera,资深软件工程师
Claude Opus 5 made large scale changes across our Fundamental Research Assistant codebase, adapting to feedback throughout an agentic workflow and explaining its reasoning more clearly than any model we’ve used. It handled work we would normally have broken into much smaller pieces.Claude Opus 5 在我们的基础研究助手代码库中进行了大规模更改,适应了整个智能体工作流中的反馈,并比我们使用过的任何模型都更清晰地解释了其推理过程。它处理了我们通常会拆分成更小部分的工作。Claude Opus 5 在我们的基础研究助手代码库中进行了大规模更改,适应了整个智能体工作流中的反馈,并比我们使用过的任何模型都更清晰地解释了其推理过程。它处理了我们通常会拆分成更小部分的工作。——Conor Kiernan,首席技术官
On some of our hardest financial-modeling tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both accuracy and efficiency. Its performance floor is materially higher, especially on deep finance domain logic. Across effort levels it averaged 9 percentage points higher accuracy with a third fewer turns and tool calls and 60% less time.在我们一些最困难的金融建模任务中,Claude Opus 5 在准确性和效率方面明显优于 Opus 4.8。它的性能下限显著提高,特别是在深层金融领域逻辑方面。在各种努力程度下,它的平均准确率提高了 9 个百分点,轮次和工具调用减少了三分之一,时间缩短了 60%。在我们一些最困难的金融建模任务中,Claude Opus 5 在准确性和效率方面明显优于 Opus 4.8。它的性能下限显著提高,特别是在深层金融领域逻辑方面。在各种努力程度下,它的平均准确率提高了 9 个百分点,轮次和工具调用减少了三分之一,时间缩短了 60%。——Richard Pham,评估与产品主管
Claude Opus 5 checks its own work the way a real frontend developer would. On our benchmark it opened its pages in a browser at desktop and phone widths, caught a product hidden below the mobile fold and an off-screen checkout button, and fixed both before handing the work back.Claude Opus 5 像真正的前端开发人员一样检查自己的工作。在我们的基准测试中,它在桌面和手机宽度下打开页面,捕获了隐藏在移动端折叠下方的产品和屏幕外的结账按钮,并在交还工作前修复了这两者。Claude Opus 5 像真正的前端开发人员一样检查自己的工作。在我们的基准测试中,它在桌面和手机宽度下打开页面,捕获了隐藏在移动端折叠下方的产品和屏幕外的结账按钮,并在交还工作前修复了这两者。——AJ Orbach,联合创始人兼首席执行官
Claude Opus 5 is a clear step up in performance on legal agent work compared to prior Opus models, and we saw the biggest gains in practice areas like corporate governance and arbitration. We were also impressed with Opus 5’s ability to maintain quality at lower reasoning levels, achieving similar performance while generating 26% fewer tokens on average compared to Opus 4.8 at max reasoning.与之前的 Opus 模型相比,Claude Opus 5 在法律智能体工作中的性能有明显提升,我们在公司治理和仲裁等实践领域看到了最大的收益。我们还对 Opus 5 在较低推理水平下保持质量的能力印象深刻,它在实现类似性能的同时,与 Opus 4.8 在最大推理水平下相比,平均生成的 Token 减少了 26%。与之前的 Opus 模型相比,Claude Opus 5 在法律智能体工作中的性能有明显提升,我们在公司治理和仲裁等实践领域看到了最大的收益。我们还对 Opus 5 在较低推理水平下保持质量的能力印象深刻,它在实现类似性能的同时,与 Opus 4.8 在最大推理水平下相比,平均生成的 Token 减少了 26%。——Niko Grupen,应用研究主管
Claude Opus 5’s biggest gains for us are on longer-horizon work: building a full deck, then revising it. Artifact quality is what decides which model we ship, and this is the clearest step up we’ve seen — better visual understanding, cleaner formatting, fewer slide issues.Claude Opus 5 对我们最大的收益在于长期的工作:构建完整的演示文稿,然后进行修订。Artifact 的质量决定了我们发布哪个模型,这是我们所见过的最明显的提升——更好的视觉理解、更整洁的格式、更少的幻灯片问题。Claude Opus 5 对我们最大的收益在于长期的工作:构建完整的演示文稿,然后进行修订。Artifact 的质量决定了我们发布哪个模型,这是我们所见过的最明显的提升——更好的视觉理解、更整洁的格式、更少的幻灯片问题。——Alex Wang,应用 AI
Claude Opus 5’s judgment is what stands out. Handing off a PR, it doesn’t rush to publish: it verifies the branches, checks the template, and thinks through test implications so the handoff is clean. The older models tended to jump ahead and get caught on our checks.Claude Opus 5 的判断力非常突出。在移交 PR 时,它不会急于发布:它会验证分支、检查模板,并思考测试的影响,从而使移交非常清晰。旧模型往往会抢先一步,并在我们的检查中被卡住。Claude Opus 5 的判断力非常突出。在移交 PR 时,它不会急于发布:它会验证分支、检查模板,并思考测试的影响,从而使移交非常清晰。旧模型往往会抢先一步,并在我们的检查中被卡住。——Zimu Li,技术人员
During a rearchitecting session, Claude Opus 5 pushed back on a design I proposed, and it didn’t fold when I insisted. Instead, it explained exactly what was valuable in my idea, narrowed its objection to a single design question, and proposed a compromise that kept the good part while fixing the flaw. That’s the kind of judgment that lets us trust it with less oversight.在一次重新架构会议中,Claude Opus 5 对我提出的设计提出了反对意见,并且在我坚持时没有妥协。相反,它准确地解释了我的想法中哪些是有价值的,将其反对意见缩小到一个单一的设计问题,并提出了一个折衷方案,既保留了好的部分又修复了缺陷。这种判断力让我们能够以更少的监督来信任它。在一次重新架构会议中,Claude Opus 5 对我提出的设计提出了反对意见,并且在我坚持时没有妥协。相反,它准确地解释了我的想法中哪些是有价值的,将其反对意见缩小到一个单一的设计问题,并提出了一个折衷方案,既保留了好的部分又修复了缺陷。这种判断力让我们能够以更少的监督来信任它。——Marquis Wang,首席 AI 工程师
On first-turn redlines, Claude Opus 5 scored the highest of any model we tested, nearly double Opus 4.8. Commenting is better too: on NDAs it gets to the redline in less time and with fewer passes, with accuracy maintained or better.在第一轮红线修改中,Claude Opus 5 的得分是我们测试过的所有模型中最高的,几乎是 Opus 4.8 的两倍。评论功能也更好了:在 NDA(保密协议)方面,它能以更少的时间和更少的轮次完成红线修改,且准确性保持不变或更好。在第一轮红线修改中,Claude Opus 5 的得分是我们测试过的所有模型中最高的,几乎是 Opus 4.8 的两倍。评论功能也更好了:在 NDA(保密协议)方面,它能以更少的时间和更少的轮次完成红线修改,且准确性保持不变或更好。——Ryan Tanenholz,技术人员
Claude Opus 5 writes clean, tight diffs with no dead code, and it’s the stronger hazard spotter on subtle, codebase-specific issues. We’re adopting it for production workloads.Claude Opus 5 编写的代码差异(diffs)干净、紧凑,且没有死代码,它是针对细微、代码库特定问题的更强风险发现者。我们正在将其用于生产工作负载。Claude Opus 5 编写的代码差异(diffs)干净、紧凑,且没有死代码,它是针对细微、代码库特定问题的更强风险发现者。我们正在将其用于生产工作负载。——Neeraj Deshmukh,工程总监
We will definitely migrate a number of use cases in Cosmos, our unified agent platform. We’re looking forward to increasingly using Claude Opus 5 for code review, and I am confident in saying we would rather people be using Opus 5 than Opus 4.8.我们肯定会迁移 Cosmos(我们的统一智能体平台)中的许多用例。我们期待越来越多地使用 Claude Opus 5 进行代码审查,我可以自信地说,我们更希望人们使用 Opus 5 而不是 Opus 4.8。我们肯定会迁移 Cosmos(我们的统一智能体平台)中的许多用例。我们期待越来越多地使用 Claude Opus 5 进行代码审查,我可以自信地说,我们更希望人们使用 Opus 5 而不是 Opus 4.8。——Igor Ostrovsky,联合创始人兼首席技术官
What stands out about Claude Opus 5 is judgment. It thinks harder before it writes a single line, catches its own logical faults during planning rather than after the fact, and reasons about why an answer is right, not just whether it works. It’s the clearest jump in problem-solving we’ve seen from one Claude model to the next, and we’re looking forward to seeing it adopted in JetBrains IDEs.Claude Opus 5 的突出之处在于判断力。它在写下一行代码之前会更深入地思考,在规划过程中而不是事后捕获自己的逻辑错误,并推理答案为什么是对的,而不仅仅是它是否有效。这是我们从一个 Claude 模型到下一个模型所见过的最清晰的问题解决能力的飞跃,我们期待看到它在 JetBrains IDE 中被采用。Claude Opus 5 的突出之处在于判断力。它在写下一行代码之前会更深入地思考,在规划过程中而不是事后捕获自己的逻辑错误,并推理答案为什么是对的,而不仅仅是它是否有效。这是我们从一个 Claude 模型到下一个模型所见过的最清晰的问题解决能力的飞跃,我们期待看到它在 JetBrains IDE 中被采用。——Denis Shiryaev,IDE AI 主管
Claude Opus 5 is the strongest Opus model we’ve tested on our trading benchmark, and it gets there using roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8. Better answers at a fraction of the compute.Claude Opus 5 是我们在交易基准测试中测试过的最强的 Opus 模型,它实现这一目标所使用的推理 Token 大约是 Opus 4.8 的七分之一,延迟不到其一半。以极少的计算量获得更好的答案。Claude Opus 5 是我们在交易基准测试中测试过的最强的 Opus 模型,它实现这一目标所使用的推理 Token 大约是 Opus 4.8 的七分之一,延迟不到其一半。以极少的计算量获得更好的答案。——Matt Nassr,全球数据工程与 AI 转型主管
Alignment and safety对齐与安全
Alignment. During pre-deployment testing, our automated behavioral audit found Opus 5 to be our most aligned model to date (as shown in the graph below). It adheres to Claude’s Constitution better than Opus 4.8, Sonnet 5, or Fable 5; exhibits the lowest rates of deceptive behavior; and is the least susceptible to being tricked into misuse. It’s also our safest model yet in terms of avoiding reckless actions that could have hard-to-reverse side effects.对齐。在部署前测试期间,我们的自动化行为审计发现 Opus 5 是我们迄今为止最对齐的模型(如下图所示)。它比 Opus 4.8、Sonnet 5 或 Fable 5 更能遵守 Claude 的宪法;表现出最低的欺骗性行为发生率;并且最不容易被诱导进行滥用。就避免可能产生难以逆转的副作用的鲁莽行为而言,它也是我们迄今为止最安全的模型。

Safety. Opus 5 does not advance the frontier in risky, dual-use capabilities. In rigorous evaluations conducted alongside private-sector and government partners, we found it remains behind Mythos 5 in both biology research and offensive cybersecurity. More information about these evaluations can be found in our System Card.安全。Opus 5 没有在风险性、双重用途能力方面推进前沿。在与私营部门和政府合作伙伴共同进行的严格评估中,我们发现它在生物研究和进攻性网络安全方面仍落后于 Mythos 5。有关这些评估的更多信息,请参阅我们的系统卡。
As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.与前代产品 Opus 4.8 一样,我们特意避免在网络任务上训练 Opus 5。然而,由于变得更加通用,该模型在这些任务上的表现仍然有了显著提高,在发现网络安全漏洞方面已接近 Mythos 5。不过,在利用这些漏洞方面,它仍显著落后于 Mythos 5——即在将漏洞转化为实质性网络威胁方面。
This is illustrated by Opus 5’s performance on OSS-Fuzz, an evaluation we’ve developed to assess how well models can find and then exploit vulnerabilities without extensive human guidance. Although Mythos 5 and Opus 5 identify vulnerabilities with similar success, Opus 5’s score on the development of exploits is far behind that of Mythos 5.Opus 5 在 OSS-Fuzz 上的表现说明了这一点,这是我们开发的一项评估,旨在评估模型在没有大量人工指导的情况下发现并利用漏洞的能力。尽管 Mythos 5 和 Opus 5 在识别漏洞方面的成功率相似,但 Opus 5 在开发漏洞利用程序方面的得分远落后于 Mythos 5。

Safeguards for Opus 5Opus 5 的保障措施
Claude Opus 5’s safeguards are designed to allow beneficial uses of the model in both cybersecurity and biology. They are similar to those we applied to Opus 4.8, with the exception of some stronger guardrails on a narrow range of cyber tasks.Claude Opus 5 的保障措施旨在允许在网络安全和生物学领域对模型进行有益的使用。它们类似于我们应用于 Opus 4.8 的措施,只是在极少数网络任务上设置了更强的护栏。
Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation.网络安全。Opus 5 的网络分类器在比例上比 Fable 5 的限制更少。它们允许 Opus 5 在源代码中查找漏洞,但会阻止“基于二进制”的漏洞扫描(一种更容易与恶意行为者关联的方法)、渗透测试和漏洞利用生成。
Based on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5. In Claude.ai, Claude Code, and Claude Cowork, any flagged requests will fall back to Opus 4.8 by default. Fallbacks to Opus 4.8 can also be enabled on the API.根据我们的测试,我们预计分类器的干预频率将比 Fable 5 低约 85%。在 Claude.ai、Claude Code 和 Claude Cowork 中,任何被标记的请求将默认回退到 Opus 4.8。回退到 Opus 4.8 也可在 API 上启用。
Our Cyber Verification Program (CVP) facilitates cybersecurity work that would otherwise be impeded by the model’s safeguards. Enterprises and researchers who are already part of the CVP have immediate access to a version of Opus 5 with fewer security restrictions.我们的网络验证计划 (CVP) 促进了网络安全工作,否则这些工作会受到模型保障措施的阻碍。已经是 CVP 一部分的企业和研究人员可以立即访问具有较少安全限制的 Opus 5 版本。
Biology. Since Opus 5 has a similar suite of safeguards to Opus 4.8, it is now our most capable generally available model for scientific research. Nevertheless, the model still shows important limitations on long-running, autonomous research tasks, which is where we expect AI models to pose the most substantial biology-related risks. (Mythos 5 remains the stronger model for this type of biological work.) As part of this launch, biology-related requests that are blocked on Fable 5 will now route to Opus 5 rather than Opus 4.8.生物学。由于 Opus 5 拥有一套与 Opus 4.8 类似的保障措施,它现在是我们用于科学研究的最强大的通用模型。尽管如此,该模型在长期、自主的研究任务上仍然表现出重要的局限性,而这正是我们预计 AI 模型会带来最重大生物相关风险的地方。(Mythos 5 仍然是此类生物学工作的更强模型。)作为此次发布的一部分,在 Fable 5 上被阻止的生物相关请求现在将路由到 Opus 5,而不是 Opus 4.8。
Getting started入门
Claude Opus 5 is available today on all platforms, priced at $5 per million input tokens and $25 per million output tokens (the same as Opus 4.8). Developers can get started with claude-opus-5 on the Claude API.Claude Opus 5 即日起在所有平台上可用,定价为每百万输入 Token 5 美元,每百万输出 Token 25 美元(与 Opus 4.8 相同)。开发者可以在 Claude API 上使用 claude-opus-5 开始使用。
It’s also offered in Fast mode, where it runs around 2.5 times the default speed. As with Opus 4.8, Fast mode is available at twice Opus 5’s base price on the Claude Platform and through usage credits in Claude Code.它还以“快速模式”(Fast mode)提供,运行速度约为默认速度的 2.5 倍。与 Opus 4.8 一样,快速模式在 Claude 平台上以 Opus 5 基本价格的两倍提供,并通过 Claude Code 中的使用额度提供。
Alongside Opus 5, we’re releasing two updates in beta:除了 Opus 5,我们还发布了两个处于测试阶段的更新:
- Mid-conversation tool changes on the Claude Platform. Within a conversation, developers can now change which tools Claude can use without invalidating the prompt cache.Claude 平台上的对话中途工具更改。在对话中,开发者现在可以更改 Claude 可以使用的工具,而不会使 Prompt 缓存失效。
- Automatic fallbacks on the API. Users can now choose to have requests that are flagged by our safety classifiers on Opus 5 (or Fable 5) automatically route to another model. With automatic fallbacks on, API requests always route to the best available model by default rather than being blocked.API 上的自动回退。用户现在可以选择让被 Opus 5(或 Fable 5)安全分类器标记的请求自动路由到另一个模型。开启自动回退后,API 请求默认总是路由到最佳可用模型,而不是被阻止。
Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access.与之前的 Opus 模型一致,Opus 5 对于通用访问没有数据保留要求。
For more guidance on how to get the best out of Opus 5, see our prompting guide.有关如何充分利用 Opus 5 的更多指导,请参阅我们的提示词指南。
Footnotes脚注
Frontier-Bench v0.1, Effort plot: These results are from an internal run of Frontier-Bench v0.1, on the mini-SWE-agent harness and a GKE backend, mean reward over 5 attempts per task. Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.Frontier-Bench v0.1,努力程度图:这些结果来自 Frontier-Bench v0.1 的内部运行,基于 mini-SWE-agent 工具和 GKE 后端,每个任务 5 次尝试的平均奖励。Opus 4.8 用作 Opus 5 和 Fable 5 在安全分类器拒绝时的回退模型。
Related content相关内容
A research agenda for the Economic Futures Research Fund经济未来研究基金的研究议程
We’re sharing the research agenda for the Anthropic Economic Futures Research Fund.我们正在分享 Anthropic 经济未来研究基金的研究议程。
Read more阅读更多Ask Claude about the Anthropic Economic Index向 Claude 询问 Anthropic 经济指数
We're launching the Anthropic Economic Index connector for Claude, which lets anyone explore real data about AI and work.我们正在为 Claude 推出 Anthropic 经济指数连接器,让任何人都可以探索关于 AI 和工作的真实数据。
Read moreAnthropic is donating another $20 million to Public First ActionAnthropic 向 Public First Action 再捐赠 2000 万美元
Anthropic is contributing an additional $20 million to Public First Action, bringing our total support to $40 million.Anthropic 正在向 Public First Action 再捐赠 2000 万美元,使我们的总支持额达到 4000 万美元。
Read more





