The State of AI Safety人工智能安全现状
Key insights from the latest AI Safety Index | Edition #305来自最新AI安全指数的关键洞察 | 第305期
The Future of Life Institute's July 2026 AI Safety Index is out, offering an updated and detailed picture of the current state of AI safety.生命未来研究所(Future of Life Institute)于2026年7月发布了最新版AI安全指数,提供了对当前人工智能安全状况的更新且详细的描述。
An observation that stood out to me, also highlighted in the UN scientific panel's report I wrote about this week, is the mismatch between the speed of AI capability advancements and the accompanying AI governance mechanisms.一个让我印象深刻的观察(也是本周我写到的联合国科学小组报告中强调的)是:AI能力提升的速度与相应的AI治理机制之间存在脱节。
AI development seems to have outgrown some of the benchmarks built to measure and assess AI capabilities and risks.AI的发展似乎已经超越了部分用于衡量和评估AI能力与风险的基准。
At the same time, there are growing examples of serious AI implications that deserve more attention, such as AI chatbot-related dependence, deskilling, and self-harm, as well as an increasing interweaving of AI companies and the military.与此同时,越来越多的严重AI影响案例值得关注,例如与AI聊天机器人相关的依赖、技能退化、自我伤害,以及AI公司与军工领域日益加深的交织。
This latest AI Safety Index noted that companies have not kept pace and, in various aspects, are actually moving backward.最新的AI安全指数指出,各公司并未跟上步伐,并且在多个方面实际上在倒退。
Below are its key findings:以下为其主要发现:
1. Anthropic, OpenAI, and Google DeepMind stay on top, Meta improves, and xAI deteriorates.1. Anthropic、OpenAI 和 Google DeepMind 保持领先,Meta 有所改善,xAI 状况恶化。
2. European dissonance: Although the European Union is a leader in AI safety regulation, the top European AI company, Mistral, scored dead last on safety.2. 欧洲的失调:尽管欧盟在AI安全监管方面处于领先地位,但欧洲顶尖AI公司Mistral在安全方面排名垫底。
3. Inadequate safety is a global problem. Three companies receive failing grades, one each from the United States (xAI), China (DeepSeek), and Europe (Mistral).3. 安全不足是一个全球性问题。三家公司的安全评级不及格,分别来自美国(xAI)、中国(DeepSeek)和欧洲(Mistral)。
4. Reviewers flagged the industry's pivot to military AI use as an emerging current harm risk.4. 评审者指出,行业转向军事AI应用是一种新兴的当前危害风险。
5. Anthropic, OpenAI, Google DeepMind, and Meta have weakened or voided their pledges to pause unilaterally if redlines are approached, with some citing conditions contingent on competitors.5. Anthropic、OpenAI、Google DeepMind 和 Meta 已削弱或取消其承诺——即当接近红线时单方面暂停,有些公司声称有条件取决于竞争对手。
6. Existential Safety is the weakest domain industry-wide. No company received a grade higher than C-.6. 存在性安全是全行业最薄弱的领域。没有公司获得高于C-的评级。
7. Safety rhetoric outpaces revealed behavior. Across Google DeepMind, OpenAI, and xAI, leadership’s reassuring public messaging diverges from commercial conduct and legislative stance, making stated commitments an unreliable proxy for actual safety practice.7. 安全相关言论超过了实际表现。在Google DeepMind、OpenAI和xAI中,领导层安抚性的公开信息与商业行为和立法立场不一致,使得公开承诺无法可靠反映实际安全实践。
8. Companies are publishing and updating safety frameworks, but these frameworks lack teeth.8. 各公司在发布和更新安全框架,但这些框架缺乏威慑力。












Excellent distillation. Findings 7 and 8 are one auditor's finding restated: a self-published safety framework is a self-assessment, and the assurance industry exists because self-assessments are worthless without an independent grader. The Index quietly supplies that grader, dragging AI safety toward the SOC 2 moment security went through years ago. A pledge voided the moment a competitor moves was never a redline; it is a defection clause with better PR.
Existential safety at a D+! Im sure everything is fine!