The State of AI SafetyAI 安全之现状
Key insights from the latest AI Safety Index | Edition #305探寻最新《AI 安全指数》之核心要义 | 第 305 期
The Future of Life Institute's July 2026 AI Safety Index is out, offering an updated and detailed picture of the current state of AI safety.未来生命研究所(Future of Life Institute)于 2026 年 7 月发布了新版《AI 安全指数》,将当下 AI 安全之态势,剖析得淋漓尽致。
An observation that stood out to me, also highlighted in the UN scientific panel's report I wrote about this week, is the mismatch between the speed of AI capability advancements and the accompanying AI governance mechanisms.我细读之下,深觉一事尤为紧要,亦与我本周所撰联合国科学小组报告之要点不谋而合:AI 演进之速,与治理机制之滞,两者已渐行渐远,脱节甚重。
AI development seems to have outgrown some of the benchmarks built to measure and assess AI capabilities and risks.AI 发展之势如脱缰野马,旧时用以衡量其能力与风险的尺规,已然不够用了。
At the same time, there are growing examples of serious AI implications that deserve more attention, such as AI chatbot-related dependence, deskilling, and self-harm, as well as an increasing interweaving of AI companies and the military.与此同时,AI 带来的隐忧亦如暗流涌动,愈发显眼:世人对聊天机器人的依赖渐深,自身技能反而在退化,甚至诱发自残之祸。更有甚者,AI 巨头与军方势力纠缠愈深,界限愈发模糊。
This latest AI Safety Index noted that companies have not kept pace and, in various aspects, are actually moving backward.最新《AI 安全指数》指出,各家公司非但未能与时俱进,反而在某些关窍上,竟似有倒行逆施之态。
Below are its key findings:其核心发现,列于下文:
1. Anthropic, OpenAI, and Google DeepMind stay on top, Meta improves, and xAI deteriorates.一、Anthropic、OpenAI 与 Google DeepMind 依然稳坐头筹,Meta 渐有起色,而 xAI 则每况愈下。
2. European dissonance: Although the European Union is a leader in AI safety regulation, the top European AI company, Mistral, scored dead last on safety.二、欧洲之迷局:虽说欧盟在 AI 安全监管上一马当先,然其麾下头号 AI 大厂 Mistral,在安全评分上竟是垫底。
3. Inadequate safety is a global problem. Three companies receive failing grades, one each from the United States (xAI), China (DeepSeek), and Europe (Mistral).三、安全之困,已成全球之疾。共有三家公司惨遭不及格,分别来自美国 (xAI)、中国 (DeepSeek) 与欧洲 (Mistral)。
4. Reviewers flagged the industry's pivot to military AI use as an emerging current harm risk.四、评审者警示,业界正大举向军用 AI 倾斜,此乃潜伏之祸患。
5. Anthropic, OpenAI, Google DeepMind, and Meta have weakened or voided their pledges to pause unilaterally if redlines are approached, with some citing conditions contingent on competitors.五、Anthropic、OpenAI、Google DeepMind 与 Meta 皆已弱化甚至背弃了当初的承诺——即在触及红线时单方面暂停研发。更有甚者,竟以竞对之举为由,为己开脱。
6. Existential Safety is the weakest domain industry-wide. No company received a grade higher than C-.六、生存安全(Existential Safety)乃业界最薄弱之环,竟无一家公司能获 C- 以上之评级。
7. Safety rhetoric outpaces revealed behavior. Across Google DeepMind, OpenAI, and xAI, leadership’s reassuring public messaging diverges from commercial conduct and legislative stance, making stated commitments an unreliable proxy for actual safety practice.七、言行不一,莫过于此。Google DeepMind、OpenAI 与 xAI 之高层,对外言辞恳切,实则商业行径与立法立场大相径庭。其口中承诺,已难作为衡量安全实践之准绳。
8. Companies are publishing and updating safety frameworks, but these frameworks lack teeth.八、各家虽皆在修缮安全框架,然终究是纸上谈兵,难有雷霆手段。











Excellent distillation. Findings 7 and 8 are one auditor's finding restated: a self-published safety framework is a self-assessment, and the assurance industry exists because self-assessments are worthless without an independent grader. The Index quietly supplies that grader, dragging AI safety toward the SOC 2 moment security went through years ago. A pledge voided the moment a competitor moves was never a redline; it is a defection clause with better PR.
Existential safety at a D+! Im sure everything is fine!