Skip to main content

I hadn’t heard of Dan Guido until a few months ago, when I came across the video of a talk he gave at [un]prompted, an AI security practitioners’ conference. Dan is the CEO and cofounder of Trail of Bits, a software security research and development firm that works with companies in tech, defense, and finance. But Dan wasn’t talking about security. He was talking about what it takes to make a company AI native, which is close to the center of the bullseye for many of us right now.在几个月前看到 Dan Guido 在 AI 安全从业者会议 [un]prompted 上的演讲视频之前,我从未听说过他。Dan 是软件安全研发公司 Trail of Bits 的首席执行官兼联合创始人,该公司与科技、国防和金融领域的企业均有合作。但 Dan 当时谈论的并非安全,而是如何让一家公司实现“AI 原生化”——这正是我们许多人目前最关注的核心议题。

We’ve been trying to figure out how to do that at O’Reilly, but until I came across Dan’s talk, we didn’t have a structured process. We’ve been building along the lines he laid out ever since. So for this episode of Live with Tim I asked Dan to reprise the talk before we got to the conversation. He was supposed to take twenty minutes, like his original conference talk, but he took thirty-five, and I had to cut him off slightly before the end to make room for questions. That was a tough choice, since everything he had to say was golden.在 O’Reilly,我们一直在探索如何做到这一点,但在看到 Dan 的演讲之前,我们缺乏一个结构化的流程。自那以后,我们一直按照他所规划的路线进行建设。因此,在这一期《Live with Tim》节目中,我邀请 Dan 在我们进行对话之前重温了他的演讲。他原本打算像在会议上那样讲二十分钟,结果讲了三十五分钟,我不得不稍微打断他,以便留出提问时间。这是一个艰难的决定,因为他所说的每一句话都是金玉良言。

Dan opened by reminding us of the current state of play in enterprise AI adoption. In February, Fortune reported on a National Bureau of Economic Research study in which nearly 90% of some 6,000 executives said AI had produced no measurable change in employment or productivity at their firms over three years. People started calling it the new Solow paradox, after Robert Solow’s 1987 line that “you can see the computer age everywhere except in the productivity statistics.”Dan 开场时回顾了目前企业采用 AI 的现状。今年二月,《财富》杂志报道了美国国家经济研究局的一项研究,其中约 6000 名高管中近 90% 表示,AI 在三年内并未给他们公司的就业或生产力带来可衡量的变化。人们开始将其称为“新索洛悖论”,这源于罗伯特·索洛(Robert Solow)1987 年的名言:“你可以在任何地方看到计算机时代,唯独在生产力统计数据中看不到。”

Dan’s belief is that this isn’t evidence that AI doesn’t work. It’s evidence that most companies are deploying AI wrong. They hand out ChatGPT and Claude licenses, and then leadership waits for the magic to happen. It doesn’t.Dan 认为,这并非证明 AI 无效,而是证明大多数公司部署 AI 的方式不对。他们发放了 ChatGPT 和 Claude 的许可证,然后领导层就坐等奇迹发生。但这并不会发生。

Dan started out by describing three levels of AI adoption.Dan 首先描述了 AI 采用的三个阶段。

  1. AI assisted is where everyone starts: “You give people access to ChatGPT, it drafts emails, it summarizes documents. It’s just a productivity tool, and your organization doesn’t change. Your workflows are the exact same as they were before. You just have a little buddy that helps you with a couple of tasks.” “AI 辅助”(AI assisted)是所有人的起点:“你让人们使用 ChatGPT,它能起草邮件、总结文档。它只是一个生产力工具,你的组织结构并没有改变。工作流程和以前完全一样。你只是多了一个能在几项任务上帮你一把的小伙伴。”
  2. AI augmented is where you start redesigning workflows, so that AI does the first pass on a code review and a human does the second. “AI 增强”(AI augmented)阶段,你开始重新设计工作流程,例如让 AI 对代码审查进行第一轮处理,人类进行第二轮处理。
  3. AI native is structural: “That’s where you’ve redesigned the company and its workflows from the ground up, assuming the AI is going to be there and that it’s a core participant. That’s not really a tool. That’s more thinking about AI as teammates.”“AI 原生”(AI native)是结构性的:“在这个阶段,你从零开始重新设计了公司及其工作流程,预设 AI 将会参与其中,并成为核心参与者。这不再仅仅是一个工具,而是将 AI 视为团队成员。”

In his framing, the first of the three is a tool and the last is an operating system. For Trail of Bits, he said that “operating system” has a specific purpose:在他看来,前两者是工具,而最后一种是操作系统。对于 Trail of Bits,他说“操作系统”有特定的目的:

“I want our security expertise to compound as code. Every engagement we do, all the skills, the workflows, everything that we build makes the next engagement faster and better.”“我希望我们的安全专业知识能转化为代码并不断叠加。我们完成的每一次业务、积累的所有技能、工作流程以及我们构建的一切,都能让下一次业务变得更快、更好。”

Employee resistance is the first problem员工抵触是首要问题

Dan confessed how hard it was to get started on the ladder from AI Assisted to AI Native:Dan 坦言,从“AI 辅助”迈向“AI 原生”的阶梯起步非常艰难:

“When I announced last year that we were all in on AI, that we were going to be using it across all of our workflows and redesigning the way the company operates, I’d say only about 5% of the company was with me. 95% was resistant.” About 20% was actively resisting. The other 75% were resisting more passively. “They’ll go along with it in public, but in process they’ll sabotage it. They’ll hope that if they keep their head low, this will pass over them, and that three months from now management’s focus will change and it won’t be a problem anymore, and we can get back to doing what we were doing. That’s where the majority of people land when these initiatives happen.”“去年我宣布我们将全力投入 AI,并将把 AI 应用于我们所有的工作流程并重组公司运营方式时,我想只有约 5% 的员工支持我。95% 的人都在抵触。” 其中约 20% 是积极抵触,另外 75% 则是消极抵触。“他们表面上会配合,但在实际操作中会搞破坏。他们希望只要自己保持低调,这阵风就会过去;三个月后管理层的重心就会转移,不再是个问题,我们就能回到原来的工作方式。当这类倡议出现时,大多数人都会处于这种状态。”

Rather than argue with his employees, Dan studied the literature on why people reject new technology and decided he needed to address four biases against AI: self-enhancing bias, identity threat, opacity, and intolerance for imperfection.Dan 没有与员工争论,而是研究了人们抵制新技术的相关文献,并决定解决针对 AI 的四种偏见:自我增强偏见、身份威胁、不透明感以及对不完美的零容忍。

Self-enhancing bias is the habit of crediting your wins to your own judgment and your losses to circumstance, which is a particular problem for senior people who are strongly attached to the years of experience and intuition that got them to their present position. Opacity is not being able to see how a decision got made. Dan’s observation is that you don’t understand your doctor’s reasoning either, but somehow you trust the doctor but get suspicious of the machine. Dan didn’t mention this work specifically, but intolerance for imperfection seems to refer to Dietvorst, Simmons, and Massey’s work on algorithm aversion, which found that people abandon an algorithm after watching it err once, even when it outperforms the human alternative. Their follow-up paper found that giving people even a slight ability to modify the algorithm’s output is enough to overcome the aversion.自我增强偏见是指将成功归功于自己的判断,将失败归咎于环境的习惯,这对那些极其依赖过往经验和直觉才走到今天的高管来说尤为严重。不透明感是指无法理解决策是如何做出的。Dan 观察到,你同样无法理解医生的推理过程,但你却信任医生,却对机器产生怀疑。虽然 Dan 没有明确提到这项研究,但对不完美的零容忍似乎指的是 Dietvorst、Simmons 和 Massey 关于“算法厌恶”的研究,该研究发现,人们在观察到算法出错一次后就会弃用它,即使它表现得比人类更好。他们的后续研究发现,只要让人们有微小的能力去修改算法的输出,就足以克服这种厌恶。

Dan spent the most time on identity threat. He described a study in which the same kitchen appliance was advertised in two ways: “On one hand, it does the cooking for you. On the other hand, it helps you cook better. It’s the same device. The people who identified as cooks rejected the first version and accepted the second.”Dan 在“身份威胁”上花费了最多时间。他描述了一项研究,同一款厨房电器以两种方式进行广告宣传:“一种是‘它为你做饭’,另一种是‘它帮你做得更好’。设备是一样的,但自认为是厨师的人拒绝了第一种表述,接受了第二种。”

Most knowledge work, Dan argued, and security auditing in particular, is what he called symbolic rather than instrumental. That is, it carries meaning about who you are. “So I have to frame AI as something that makes you a more dangerous auditor,” he said. “Not that it does the audit for you.”Dan 认为,大多数知识工作,特别是安全审计,属于他所说的“象征性”而非“工具性”工作。也就是说,它承载着关于“你是谁”的意义。“所以我必须将 AI 包装成能让你成为更厉害的审计员的东西,”他说,“而不是让它替你完成审计。”

In his work at Trail of Bits, he deliberately built a countermeasure for each bias.在 Trail of Bits,他刻意为每一种偏见构建了对策。

  • Self-enhancing bias is addressed by “an AI maturity matrix” with visible levels, because you can’t claim you’re already good enough when there’s a published ladder that identifies a different set of skills as critical. 自我增强偏见通过“AI 成熟度矩阵”来解决,其中包含可见的级别,因为当有一个公开的阶梯明确指出另一套技能才是关键时,你就不能再声称自己已经足够好了。
  • Identity threat gets skills repositories, where an engineer who writes a hard plugin gets credit for encoding their expertise. Hackathons also change the dynamic from resistance to exploration. I’m putting words in Dan’s mouth here, but I think he’d agree that when experienced developers are called on as mentors in a hackathon, that also reduces their experience of AI as an identity threat. 身份威胁通过技能库来解决,编写复杂插件的工程师因将专业知识编码化而获得认可。黑客马拉松也将动态从抵触转变为探索。虽然这是我的解读,但我认为他会同意:当经验丰富的开发者在黑客马拉松中担任导师时,这也会减少他们将 AI 视为身份威胁的感受。
  • Intolerance for imperfection gets a curated marketplace, sandboxing, and hardened defaults, so everyone’s first experience of AI isn’t a disaster. 对不完美的零容忍通过精心策划的市场、沙箱和强化默认设置来解决,确保每个人第一次使用 AI 的体验都不会是一场灾难。
  • Opacity gets a written AI handbook that clarifies the usage policy and the risk model rather than just saying “trust us.”不透明感通过编写 AI 手册来解决,手册明确了使用政策和风险模型,而不是简单地要求“相信我们”。

Here’s Dan’s slide on “the remedies that actually worked”:这是 Dan 关于“真正有效的补救措施”的幻灯片:

The remedies that actually worked

Returning to one of my hobby horses, this is a kind of mechanism design. In my recent piece on the missing mechanisms of the agentic economy, I argued that we need to start with desired outcomes and ask ourselves what mechanisms will help to produce them. Dan’s approach seems to be really good at this. Most enterprises are treating AI adoption as a procurement problem or a communications problem. Dan treated it as a question of what incentives, defaults, and status ladders produce the behavior you want, given how people actually respond.回到我最喜欢的话题,这是一种机制设计。在我最近关于代理经济缺失机制的文章中,我主张我们应该从期望的结果出发,问自己什么样的机制有助于产生这些结果。Dan 的方法在这方面做得非常好。大多数企业将 AI 采用视为采购问题或沟通问题。Dan 将其视为一个问题:考虑到人们的实际反应,什么样的激励、默认设置和地位阶梯能产生你想要的行为?

The last remedy on Dan’s list is that the CEO has to lead by example. He noted, “I was the first person through the door. My voice as the CEO matters a lot more than people think. The passive 50% of the company that isn’t sure if this initiative is going to be successful, they’re watching to see what leadership actually does, not what it says.”Dan 清单上的最后一个补救措施是 CEO 必须以身作则。他指出:“我是第一个跨入大门的人。作为 CEO,我的声音比人们想象的要重要得多。公司里那 50% 不确定这项倡议是否会成功的消极派,他们正在观察领导层实际上在做什么,而不是在说什么。”

A ladder, not a mandate阶梯,而非指令

Trail of Bits already tracked about 50 engineering skills for performance review, things like Python, git, Rust, and various security auditing capabilities. Dan pulled AI skills out into their own matrix, with four levels, from not engaged through capable and adoptive to transformative. Each of these levels is detailed separately and more specifically for assurance, engineering, sales, and project management.Trail of Bits 已经追踪了约 50 项工程技能用于绩效评估,例如 Python、git、Rust 以及各种安全审计能力。Dan 将 AI 技能提取出来,形成了自己的矩阵,分为四个等级:从“未参与”到“有能力”、“采用”再到“变革”。每个等级都针对保障、工程、销售和项目管理进行了详细且具体的说明。

He noted that “The highest level of the maturity matrix is not somebody who uses AI the most. It’s somebody who invents new ways to work and builds tools with AI. So the identity of the expert shifts from ‘I don’t need AI’ to ‘I’m the one who makes AI useful for the company.’” This was his first important design choice.他指出:“成熟度矩阵的最高级别不是指使用 AI 最多的人,而是指发明新工作方式并用 AI 构建工具的人。因此,专家的身份从‘我不需要 AI’转变为‘我是那个让 AI 对公司有用的人’。” 这是他第一个重要的设计选择。

The second is what level zero means. He said “If you’re at level zero, if you’re not engaged, that means you’re fighting back against the company. If you dismiss AI as hype, if you refuse to use AI for security work, this is a disagreement on principles, not on skills. For people who were stuck in the not engaged category, we had hard conversations, and there were people who left the company.” Levels one through three are a skill issue, and the remedy is time with the tools.第二个是关于零级的定义。他说:“如果你处于零级,即未参与,意味着你在对抗公司。如果你将 AI 斥为炒作,如果你拒绝在安全工作中使用 AI,这是原则上的分歧,而非技能问题。对于那些停留在‘未参与’类别的员工,我们进行了严厉的谈话,有些人因此离开了公司。” 一到三级属于技能问题,补救措施就是投入时间去使用这些工具。

While the slide describing the capability matrix is shown in the preceding video clip, here’s where you can find the full deck so you can study it in more detail.虽然描述能力矩阵的幻灯片在前面的视频片段中有所展示,但你可以在此处找到完整的演示文稿,以便更详细地研究它。

Driving adoption and skills with hackathons通过黑客马拉松推动采用和技能提升

One of the best ways Trail of Bits developed to move people up the ladder was to hold a hackathon every two months. Dan runs them with clear goals rather than as a free-for-all. The focus area and learning objectives are defined in advance and announced a week ahead, with separate instructions for engineers and non-engineers. People work in pairs so everything gets reviewed. There’s a demo session at the end, and then follow-through. (It’s an important part of Dan’s big idea, that you have to build a system by which, in his words, organizational knowledge and capability compounds.) He noted that “In the days afterward we keep one or two people around, and they collect all the reusable artifacts, structure them, and put them into the places they need to be.”Trail of Bits 开发出的提升员工阶梯的最佳方法之一是每两个月举办一次黑客马拉松。Dan 带着明确的目标来组织这些活动,而不是任其自由发挥。重点领域和学习目标提前定义并在一周前公布,针对工程师和非工程师有不同的说明。人们两两配对工作,确保一切都能得到审查。最后有演示环节,随后是后续跟进。(这是 Dan 大构想的重要组成部分,即你必须建立一个系统,用他的话说,让组织知识和能力不断叠加。)他指出:“在随后的几天里,我们会留下一两个人,由他们收集所有可重用的工件,对其进行结构化,并放入它们应该在的地方。”

I asked what people outside of product and engineering actually work on, since the answer for an accountant at a hackathon was not obvious. Dan’s response is that the hackathon isn’t measured in artifacts shipped but in where people sit on the capability ladder the following week. Essentially, he’s running a training program that happens to produce useful output, rather than a production sprint that happens to teach people something.我问他产品和工程以外的人到底在做什么,因为黑客马拉松对会计师来说目标并不明显。Dan 的回答是,黑客马拉松的衡量标准不是交付了多少工件,而是下周人们在能力阶梯上的位置。本质上,他是在运行一个恰好能产生有用产出的培训计划,而不是一个恰好能教会人们某些东西的生产冲刺。

The first hackathon, he told me, was the equivalent of a beach cleanup: “It’s like those companies that send everybody to the beach with a big stick and say, let’s go pick up a bunch of trash and put it away, and then you get the big team photo after with all the contractor bags of garbage. That’s what we did with our public source code repositories.”他告诉我,第一次黑客马拉松相当于海滩清理:“就像那些让每个人拿着大棍子去海滩清理垃圾的公司,最后大家拿着装满垃圾的承包商袋子拍大合照。我们对公共源代码仓库就是这么做的。”

He picked it because open source maintenance is the part of the job that feels like a grind. No new features, just closing issues and stale dependencies on public code where nothing was at risk. “As an open source maintainer, you just get beaten down by the public. This doesn’t work, I can’t use it, this thing sucks. Dozens of issues pointing out flaws you already knew about. It feels burdensome. We wanted people to see that adopting AI would relieve burden.”他选择这个是因为开源维护是工作中感觉最枯燥的部分。没有新功能,只是关闭问题和处理公共代码中过时的依赖项,而这些地方并没有什么风险。“作为开源维护者,你只会受到公众的打击。这个不行,我没法用,这东西太烂了。几十个问题指出的都是你已经知道的缺陷。这感觉很沉重。我们希望人们看到,采用 AI 可以减轻负担。”

The second hackathon was about shipping impactful product updates, but it was also designed to move everyone up the capability ladder by giving up control. Engineers had to run Claude Code in bypass permissions mode, fully autonomous, on public repositories, inside sandboxes the company had prepared in advance. The one they’re running now is about persistent background agents that can be handed a task during an audit and come back with a proof of concept exploit or a draft finding.第二次黑客马拉松是关于交付有影响力的产品更新,但也旨在通过放弃控制权让每个人在能力阶梯上更进一步。工程师必须在公司预先准备好的沙箱内,在公共仓库上以绕过权限模式(bypass permissions mode)完全自主地运行 Claude Code。他们现在正在进行的是关于持久的后台代理,这些代理可以在审计过程中被分配任务,并带回概念验证漏洞利用或初步发现结果。

Here’s a look at Dan’s slack message announcing the hackathon:这是 Dan 在 Slack 上宣布黑客马拉松的消息:

The slack message announcing the second hackathon. (From Dan’s slide deck.)
The slack message announcing the second hackathon. (From Dan’s slide deck.)宣布第二次黑客马拉松的 Slack 消息。(来自 Dan 的演示幻灯片。)

Everything the hackathons produce gets harvested into artifacts.黑客马拉松产生的一切都会被转化为工件。

Trail of Bits runs three skills repositories: an internal one for company workflows, a public one that anyone can use, and a curated one that vets third-party skills before they’re allowed in.Trail of Bits 运行三个技能库:一个用于公司内部工作流程的内部库,一个任何人都可以使用的公共库,以及一个在第三方技能被允许进入前进行审查的精选库。

Publishing skills to the public repository is not just a marketing exercise. “It keeps us honest, and it forces us to write things that other people can use, not just people outside the company but inside too,” Dan said. “It really helps us think about the tribal knowledge that’s baked into the tool.”发布技能到公共仓库不仅仅是一种营销手段。“它让我们保持诚实,并迫使我们编写其他人可以使用的东西,不仅是公司外部的人,也包括公司内部的人,”Dan 说。“这确实有助于我们思考融入工具中的部落知识。”

The curated repository exists because Trail of Bits knows how bad the supply chain is. They’ve published research on how to write malicious skills, and so Dan is not going to tell 130 employees to start downloading code from strangers and running it on their laptops. “If you want adoption, you need a safe supply chain.”存在精选库是因为 Trail of Bits 知道供应链有多糟糕。他们发布了关于如何编写恶意技能的研究,因此 Dan 不会告诉 130 名员工开始从陌生人那里下载代码并在笔记本电脑上运行。“如果你想要采用,你需要一个安全的供应链。”

Turning scar tissue into infrastructure将伤疤转化为基础设施

Perhaps even more important than the skills repository is, as Dan put it, “turning scar tissue into infrastructure.”正如 Dan 所言,比技能库更重要的可能是“将伤疤转化为基础设施”。

“Every single time Claude Code didn’t do something we wanted, we would bake it into a set of global, copy-pasteable defaults. Known good settings, recommended patterns. I call it scar tissue. If I hire somebody new tomorrow, I don’t want them to have to go through the entire discovery process of the last year of Trail of Bits to figure out how to use the tool.”“每当 Claude Code 没有按照我们想要的方式执行时,我们就会将其固化为一套全局的、可复制粘贴的默认设置。已知的良好设置、推荐的模式。我称之为伤疤。如果我明天雇佣一个新人,我不希望他们必须经历 Trail of Bits 过去一年的整个发现过程才能弄清楚如何使用该工具。”

The configuration repository, claude-code-config, is where the accumulated lessons live.配置仓库 claude-code-config 就是积累的经验所在。

Trail of Bits Claude code config

Dan built the first version himself and then opened it to pull requests from the whole company, assigning someone after each hackathon to go collect what people hadn’t contributed on their own. “It’s easier to put out something that’s unpolished than it is to get it perfect on the first try.”Dan 自己构建了第一个版本,然后向全公司开放了拉取请求(pull requests),并在每次黑客马拉松后指派专人收集人们没有主动贡献的内容。“发布一些未经打磨的东西比试图第一次就做到完美要容易得多。”

In short, a big part of the Trail of Bits “enterprise AI operating system” approach is a set of standardized tools and hardened defaults. Standardization isn’t a straitjacket. It’s a foundation.简而言之,Trail of Bits “企业 AI 操作系统”方法的重要组成部分是一套标准化的工具和强化的默认设置。标准化不是紧身衣,而是地基。

On sandboxing, Trail of Bits deliberately didn’t pick a single preferred solution. There’s a devcontainer for developers, dropkit for disposable DigitalOcean droplets, COOP for isolated VMs, and the sandboxing now built into Claude Code for casual users. “The point isn’t that everybody uses the same sandbox,” Dan said. “The point is that everyone has a safe sandbox to use, and that it’s easy for them to do it.”关于沙箱,Trail of Bits 刻意没有选择单一的首选解决方案。开发者有 devcontainer,一次性 DigitalOcean droplet 有 dropkit,隔离虚拟机有 COOP,现在 Claude Code 中还为普通用户内置了沙箱。“重点不是每个人都使用同一个沙箱,”Dan 说,“重点是每个人都有一个安全的沙箱可以使用,而且操作起来非常简单。”

Another of the hardened defaults is procedural. Trail of Bits enforces a seven day cooldown on every package their developers install:另一个强化默认设置是程序性的。Trail of Bits 对开发者安装的每个包强制执行七天的冷却期:

“There are dozens of security companies scanning the internet trying to find a new cool blog post they can write about malicious code hiding on PyPI or npm, and they usually figure out there’s a supply chain issue within hours. So we just delay all the packages that Trail of Bits uses. Generally the malicious stuff gets picked up before we ever get a chance to run it.”“有几十家安全公司在扫描互联网,试图找到一篇关于隐藏在 PyPI 或 npm 上的恶意代码的新博客文章,他们通常在几小时内就能发现供应链问题。所以我们只是延迟了 Trail of Bits 使用的所有包。通常恶意内容在我们有机会运行它之前就会被发现。”

That’s free-riding on a competitive market for security research, and given the speed of today’s market, it’s an elegant solution. There’s a whole class of defenses like this waiting to be found, where the mechanism is not a technical system but a well-chosen delay.这是在利用竞争激烈的安全研究市场进行“搭便车”,考虑到当今市场的速度,这是一个优雅的解决方案。还有一整类像这样的防御措施等待被发现,其机制不是技术系统,而是一个精心选择的延迟。

Data, and DJ Patil’s “Tidy House”数据,以及 DJ Patil 的“整洁的房子”(Tidy House)

The problem we run into most often as we build AI workflows at O’Reilly isn’t the model or the tooling. It’s data. Who has access to which system, which system does that data live in, and who do I ask? In a 500 person company that’s annoying. I wonder what it’s like at a company with 50,000 employees.我们在 O’Reilly 构建 AI 工作流程时遇到的最常见问题不是模型或工具,而是数据。谁有权访问哪个系统,数据存在哪个系统中,我该问谁?在一家 500 人的公司里,这很烦人。我不知道在一家拥有 50,000 名员工的公司里会是什么样。

I told Dan about DJ Patil’s Tidy House framing. He agreed that data access for AI is a big problem. His answer starts with permissions:我向 Dan 介绍了 DJ Patil 的“整洁的房子”框架。他同意 AI 的数据访问是一个大问题。他的回答从权限开始:

“The permissions debt is invisible until an agent hits it. Making data agent legible is a forced permission audit. You have to actually go through and figure out who can access what…. It also raises the stakes for permissions errors. If you overshare information, now an agent inside your company is going to find it instantly. There are a lot of these technical debt sort of things where, with agents, all of it’s becoming due at the same time.”“权限债务在代理触及它之前是不可见的。让数据对代理可读是一次强制性的权限审计。你必须真正去梳理并弄清楚谁可以访问什么……这也提高了权限错误的风险。如果你过度共享信息,现在你公司内部的代理会立即找到它。有很多这种技术债务,随着代理的出现,所有这些债务都在同一时间到期了。”

Every shortcut an organization took with its data over the past twenty years is being called at once, and the companies that can run the audit, make fast decisions about boundaries, and then actually share their data are the ones that will get a force multiplier.组织过去二十年来在数据上走的每一条捷径都在同时被清算,而那些能够进行审计、就边界做出快速决策并真正共享数据的公司,将获得力量倍增器。

Dan is against letting a thousand flowers bloom, because uncoordinated teams create overlap rather than compounding. He’d rather have one centralized foundation, with innovation happening on top of that. He suggested a useful metric for making that work across team boundaries is what fraction of your team’s data did you make reusable for everyone else, and how much of it is being used by teams outside your own.Dan 反对让“百花齐放”,因为不协调的团队会造成重叠,而不是叠加。他宁愿拥有一个集中的基础,并在其之上进行创新。他建议了一个在团队边界上实现这一点的有用指标:你团队的数据中有多少比例是你为其他人制作成可重用的,以及有多少是被你团队之外的团队所使用的。

What post-AI jobs look like后 AI 时代的工作是什么样的

Before the first hackathon, Trail of Bits ran hands-on sessions to teach its operations and go-to-market staff the basics of git and the command line. Not mastery, just enough to be a consumer of the thing. Here we are fifty years into my career and the Unix command line still matters. Dan’s non-technical staff mostly work inside Claude Cowork or Codex Desktop now, but he thinks the command line experience was worth it because they know what’s happening under the hood.在第一次黑客马拉松之前,Trail of Bits 举办了实践课程,教导运营和市场人员 git 和命令行的基础知识。不是精通,只要足以成为工具的使用者即可。我已经工作了五十年,Unix 命令行依然重要。Dan 的非技术人员现在大多在 Claude Cowork 或 Codex Desktop 中工作,但他认为命令行体验是值得的,因为他们知道底层发生了什么。

What happens to a job when the tool can do a lot of what humans used to do? Dan gave the example of his own technical editors. His editors used the hackathons to build the tools that got them out of line editing, including one that turns a public presentation into a blog post in the company’s voice. What the writers do now is consult on how to frame a story so it is effective with a particular audience.当工具可以完成人类过去所做的大部分工作时,工作会发生什么变化?Dan 以他自己的技术编辑为例。他的编辑利用黑客马拉松构建了让他们摆脱行编辑(line editing)的工具,包括一个能将公开演示文稿转换为公司语气的博客文章的工具。现在,作者们的工作是咨询如何构建故事,使其对特定受众有效。

I agree. Human jobs aren’t going away any time soon. This gets heard as optimism when it’s really just observation. AI is going to replace a lot of what we used to do, but it is also going to hand us a large amount of new work, and much of that work hasn’t been understood yet. Quality assurance for agent systems is one of the new jobs. So is skills product management, which is a role that didn’t exist eighteen months ago and now has a headcount at a 130 person security firm.我同意。人类的工作短期内不会消失。当人们听到这句话时,会觉得这是一种乐观,但实际上这只是观察。AI 将取代我们过去所做的很多事情,但它也会交给我们大量新的工作,而其中许多工作尚未被理解。代理系统的质量保证就是其中之一。技能产品管理也是如此,这是一个 18 个月前不存在的角色,现在在一家 130 人的安全公司里已经有了编制。

I asked a question towards the end about how we’re going to know which skills and agents are any good. What Dan has so far is telemetry pulled from developers’ dot files through the company’s device management system, which tells him what gets used and what breaks, plus one AI systems engineer whose job is product management for the skills repository, reviewing incoming pull requests and deprecating overlapping skills.我在最后问了一个问题:我们将如何知道哪些技能和代理是好的。Dan 目前所拥有的是通过公司设备管理系统从开发者的点文件(dot files)中提取的遥测数据,这告诉他什么被使用了,什么坏了;此外还有一名 AI 系统工程师,其工作是技能库的产品管理,负责审查传入的拉取请求并废弃重叠的技能。

What Dan thinks comes next is evaluation. He says: “Once you invest a lot into these agent systems, you need proof that they do the job. The way you do that is you give everybody a performance review. You give them an evaluation data set, a benchmark.”Dan 认为接下来是评估。他说:“一旦你对这些代理系统投入了大量资金,你需要证明它们能胜任工作。实现这一点的办法是给每个人进行绩效评估。你给他们一个评估数据集,一个基准。”

Trail of Bits is now building benchmarks for its core skills. How well can we find bugs in this language? How well can we write a statement of work? Constructing those datasets is real work, with positive and negative cases, and comparisons against the algorithmic tools that already exist.Trail of Bits 现在正在为其核心技能构建基准。我们在这个语言中发现漏洞的能力如何?我们编写工作说明书的能力如何?构建这些数据集是真正的工作,包含正面和负面案例,以及与现有的算法工具进行比较。

Put the reps in多加练习(Put the reps in)

I asked Dan for the top five mistakes he made. He said there was only one. “You need to allocate an appropriate amount of FAFO time. (That’s F Around and Find Out.) A product comes out on Friday. There’s no documentation for it. There’s no training guidance for it. There’s no course on it. You can’t wait until somebody systematizes the knowledge. You just need to do it.”我问 Dan 他犯过的最大五个错误。他说只有一个。“你需要分配适量的 FAFO 时间。(即 F Around and Find Out,摸索并找出答案。)一个产品周五发布。没有文档,没有培训指导,没有课程。你不能等到有人把知识系统化。你只需要直接动手去做。”

Then he gave an analogy to going to the gym.然后他给出了一个去健身房的类比。

The recipe for success成功的秘诀

Dan has a replicable recipe, which he summarized as follows:Dan 有一个可复制的秘诀,总结如下:

  1. Standardize on one agent workflow that you can support.标准化一个你能支持的代理工作流程。
  2. Write an AI handbook so that risk decisions aren’t ad hoc, and that everyone is playing the same game.编写 AI 手册,使风险决策不再是临时的,确保每个人都在同一个规则下行事。
  3. Create a capability ladder that makes clear that improvement is expected.创建一个能力阶梯,明确表示进步是预期的。
  4. Run short adoption sprints that force hands-on usage.进行短期的采用冲刺,强制进行实践使用。
  5. Capture everything as reusable artifacts: skills + configs + a curated supply chain.将一切捕获为可重用的工件:技能 + 配置 + 精选供应链。
  6. Make autonomous agents safe with sandboxing + guardrails + hardened defaults.通过沙箱 + 防护栏 + 强化默认设置,使自主代理变得安全。

The Trail of Bits skills repository is public. So is the curated marketplace, the configuration repository, the devcontainer, dropkit, and COOP (Continuity of Operations planning). He wrote up the whole playbook on The Trail of Bits Blog and gave a version of it to tl;dr sec. He thinks publishing makes the work better because it forces the tribal knowledge out into the open where it can be checked.Trail of Bits 技能库是公开的。精选市场、配置仓库、devcontainer、dropkit 和 COOP(业务连续性规划)也是如此。他在 The Trail of Bits Blog 上写下了整个手册,并将其中一个版本提供给了 tl_dr sec。他认为发布可以让工作变得更好,因为它迫使部落知识暴露在公开场合,以便受到检验。

Which brings me back to the Solow paradox, which seemed to disappear by the late 90s, when US aggregate productivity did finally go up. That didn’t happen because computers got faster. It disappeared because companies figured out how to reorganize themselves around what computers could do, and eventually those organizational recipes spread widely enough to show up in aggregate statistics. The same has to happen today. The current AI discourse is obsessed with model capability and largely uninterested in diffusion. The problem is not that the models are oversold. It’s that almost nobody has done the necessary organizational work, and the few who have are mostly keeping it to themselves.这让我回到了索洛悖论,它似乎在 90 年代末消失了,当时美国总体生产力确实有所提高。这不是因为计算机变快了,而是因为公司弄清楚了如何围绕计算机能做的事情重组自身,最终这些组织食谱传播得足够广泛,从而在总体统计数据中显现出来。今天也必须发生同样的事情。目前的 AI 讨论沉迷于模型能力,而对扩散几乎不感兴趣。问题不在于模型被过度吹捧,而在于几乎没有人做必要的组织工作,而少数做过的人大多将其秘而不宣。

If you want to go beyond the highlight videos shown above, watch Dan’s entire talk here. His slide deck is here. And be sure to check out the Trail of Bits Github repository.如果你想超越上面显示的精彩片段,请在这里观看 Dan 的完整演讲。他的幻灯片在这里。并一定要查看 Trail of Bits 的 Github 仓库。

Post topics: AI & MLAI 与机器学习文章主题:AI 与机器学习