What the Top 1% of Engineering Teams Do Differently with AI顶尖1%工程团队在AI应用上的不同之处
Full recording and insights from the talk at the Engineering Leadership LIVE event in San Francisco.旧金山工程领导力LIVE活动演讲的完整录像和见解。
This week’s newsletter is sponsored by Larridin, an AI-native developer intelligence platform.本周的新闻通讯由AI原生开发者智能平台Larridin赞助。
How much should you spend on AI coding tools?你应该在AI编码工具上花多少钱?
Larridin analyzed data from hundreds of engineering organizations to benchmark AI coding tool costs and find the budgeting approach that maximizes productivity without Tokenmaxxing.Larridin分析了数百个工程组织的数据,以基准测试AI编码工具成本,并找到在不进行Tokenmaxxing的情况下最大化生产力的预算方法。
Inside the research paper:研究论文内容:
Understand where AI coding budgets are headed了解AI编码预算的趋势
What high-performing teams spend高绩效团队的支出情况
The formula to determine how much your team should invest确定团队应投资多少的公式
Thanks to Larridin for sponsoring this newsletter, let’s get back to this week’s thought!感谢Larridin赞助本期新闻通讯,让我们回到本周的思考!
Intro引言
Last month, together with my friends from Augment Code, we hosted an event called: Engineering Leadership LIVE in San Francisco. It was a blast, and there were so many insightful discussions we had!上个月,我和Augment Code的朋友们在旧金山共同举办了一场名为“工程领导力LIVE”的活动。活动非常精彩,我们进行了许多富有洞察力的讨论!
As part of the event, we also had 4 talks. 作为活动的一部分,我们还安排了4场演讲。
- Gregor Ojstersek, CTO & Author, Engineering Leadership newsletter
Talk: AI-Native Engineering Leadership- Gregor Ojstersek,CTO兼作者,《工程领导力》新闻通讯演讲:AI原生工程领导力
You can also watch the full recording of my talk here: 你也可以在这里观看我演讲的完整录像:
And the full overview of my talk in this article:以及本文中我演讲的完整概述:
- Vinay Perneti, VP of Engineering, Augment Code
Talk: We Thought AI Transformation Was About Adopting Agents. We Were Wrong.- Vinay Perneti,Augment Code工程副总裁演讲:我们以为AI转型是关于采用智能体。我们错了。
- Andrew Churchill, CTO, Weave
Talk: What Actually Works: AI Coding Patterns from the Top 1% of Teams- Andrew Churchill,Weave CTO演讲:真正有效的方法:顶尖1%团队的AI编码模式
- Anwar Haneef, GM & Head of Ecosystem, Canva
Talk: Your Product’s Next User Might Be an AI Agent. What Engineering Leaders Need to Know.- Anwar Haneef,Canva总经理兼生态系统负责人演讲:你的产品的下一个用户可能是AI智能体。工程领导者需要知道什么。
Today, I am sharing the overview and the recording of Andrew Churchill’s talk at the event.今天,我将分享Andrew Churchill在活动中的演讲概述和录像。
Recording of the talk at the event活动演讲录像
You can watch/listen to the talk below, or you can keep reading for the full insights.你可以观看/收听下面的演讲,或者继续阅读以获取完整见解。
Let’s start!让我们开始吧!
The hype of “10 AI agents doing 400x productivity” is not real“10个AI智能体实现400倍生产力”的炒作并不真实
There is so much hype on social media about how to be more productive using AI, but the problem is that those are some specific examples, which do not reflect the actual data on what they at Weave are seeing.社交媒体上有很多关于如何利用AI提高生产力的炒作,但问题在于这些都是一些特定例子,并不能反映Weave实际看到的数据。
At Weave, they are building a platform for understanding how software engineers work. You can also read a full deep dive on how they work in this article:在Weave,他们正在构建一个理解软件工程师工作方式的平台。你也可以在这篇文章中阅读关于他们工作方式的深入探讨:
In today’s article, I am sharing (based on Andrew’s talk) the results regarding what makes the top 1% teams successful, based on the data from Weave’s usage from hundreds of companies and 10k+ engineers.在今天的文章中,我将分享(基于Andrew的演讲)关于什么使顶尖1%团队成功的结果,这些结果基于Weave从数百家公司和超过1万名工程师的使用数据。
It’s also important to mention that the 1% teams in this case are being measured based on the Weave’s code output metric, which is an ML-based algorithm built by Weave to check how effective the work that’s being delivered is.同样重要的是要提到,这里的1%团队是根据Weave的代码输出指标衡量的,这是一个由Weave构建的基于机器学习的算法,用于检查交付工作的有效性。
Not tracking purely lines of code or how many commits were delivered, but understanding the output based on the question: “How long would this PR take an expert engineer to complete?”不是单纯追踪代码行数或提交次数,而是基于以下问题理解输出:“一个专家工程师完成这个PR需要多长时间?”
The result is a standard unit to quantify output, which is comparable across individuals, teams, languages, and organizations.结果是一个标准单位来量化输出,可以在个人、团队、语言和组织之间进行比较。
There's a big difference between the top 1% teams and the rest顶尖1%团队与其他团队之间存在巨大差异
When you look at the chart, there’s an exponential difference between the top 1% of engineers and the rest. Especially interesting is how big a difference it is between the top 10% (P90) and top 1% (P99), and then also between P99 and P50, with a really huge difference.当你查看图表时,顶尖1%的工程师与其他工程师之间存在指数级差异。特别有趣的是,前10%(P90)和前1%(P99)之间的差异很大,而P99和P50之间的差异更是巨大。
That’s aligned with the:这与以下规律一致:
Parreto’s 80:20 rule, as we can see, the top 20% of people carry the majority of the productivity.帕累托80:20法则,我们可以看到前20%的人承担了大部分生产力。
The power law of distribution (a tiny fraction of individuals generates the vast majority of the impact, success) makes sense in this case, as the top 1% of engineers may produce 10–100× the impact of the median.幂律分布(极小部分人产生绝大部分影响和成功)在这种情况下是合理的,因为顶尖1%的工程师可能产生中位数的10到100倍的影响。
In the next sections, we will go through the 5 specific areas of what top 1% teams, based on the Weave’s code output metric, do:在接下来的部分中,我们将介绍基于Weave代码输出指标的顶尖1%团队在5个具体领域所做的:
Their engineering organization structure他们的工程组织结构
How much are they spending on AI tools (AI tokens) in relation to how much value such code contributes to success他们在AI工具(AI代币)上的支出与这些代码对成功贡献的价值之间的关系
How do they review their code他们如何审查代码
They have higher output, does it mean they also have more bugs?他们产出更高,是否意味着也有更多错误?
The number of deployments they do他们的部署次数
For each area, I am sharing (based on the talk) what the top 1% do that works for them, with a specific example. Let’s start with the first one, the structure of the organization.对于每个领域,我将分享(基于演讲)顶尖1%团队有效的方法,并附上具体例子。让我们从第一个开始,即组织结构。
1. The organizational structure of the top 1% teams1. 顶尖1%团队的组织结构
Many of the top teams are going more toward flatter orgs, have smaller teams, and more higher agency ICs. Higher agency meaning that engineers are owning projects end-to-end and are being more product & business-minded as an engineer e.g., product engineers. 许多顶尖团队正朝着更扁平的组织结构发展,拥有更小的团队和更多高自主性的个人贡献者。高自主性意味着工程师端到端地拥有项目,并且作为工程师更具产品和商业思维,例如产品工程师。
In the talk, Andrew also mentioned that they see that the product engineer role is not a common title anymore, as the expectation is merging just to the title “software engineer”. This is something I have also mentioned in the article AI-Native Engineering Leadership.在演讲中,Andrew还提到,他们看到产品工程师这个职位不再常见,因为期望已经合并到“软件工程师”这个头衔中。这也是我在文章《AI原生工程领导力》中提到过的。
Examples of the 2 teams:两个团队的例子:
Telnyx (extreme example)Telnyx(极端例子)
Their engineering org consists of 200 engineers, 0 engineering managers, and 1 VP of Engineering. They took the word “flatter org” to the extreme.他们的工程组织由200名工程师、0名工程经理和1名工程副总裁组成。他们将“扁平组织”这个词发挥到了极致。
Based on the data from Weave, they are doing very well based on their metric, but it’s really important to mention that such an engineering org structure doesn’t work for everyone.根据Weave的数据,基于他们的指标,他们表现非常好,但需要强调的是,这种工程组织结构并不适用于所有人。
In my opinion, what helps them to have such a structure is the nature of their product, it’s very technical, so it helps a lot for engineers to be more product-oriented as well.在我看来,帮助他们拥有这种结构的是他们产品的性质——非常技术性,这有助于工程师也更具产品导向。
A similar example is what I have recently written about. The company called PortKey has 24 engineers and 0 product managers. Read the full deep dive on how they work in this article:类似的例子是我最近写过的。一家名为Portkey的公司有24名工程师和0名产品经理。在这篇文章中阅读关于他们工作方式的深入探讨:
PostHogPostHog
Teams of 1-3 engineers, with an assigned Team Lead for every team (who is a DRI - directly responsible individual). The Team Lead also reports directly to the VP.团队由1-3名工程师组成,每个团队都有一名团队负责人(他是直接负责人)。团队负责人也直接向副总裁汇报。
They want to avoid the coordination task, which grows quite a lot more when you have 4 or more people inside a team. That’s why they want to keep their teams to a maximum of 3 people.他们希望避免协调任务,当团队中有4个或更多人时,协调任务会大大增加。这就是为什么他们希望将团队规模控制在最多3人。
PostHog is betting on the AI-native team structure similar to what OpenAI is doing. You can read how OpenAI is building AI-native engineering teams in this article:PostHog正在押注类似于OpenAI正在做的AI原生团队结构。你可以在这篇文章中阅读OpenAI如何构建AI原生工程团队:
It’s also important to mention that both of these companies have explosive headcount growth and are hiring a lot of people. They are not “replacing” people with AI, but they are betting on the premise of:同样重要的是要提到,这两家公司都在爆炸性地增加员工人数,并且正在招聘很多人。他们并不是在用AI“取代”人员,而是押注于以下前提:
More people mean more productivity exponentially.更多的人意味着生产力的指数级增长。
And they are not going in the direction of the trend that somewhat bigger companies go for: similar productivity, with fewer people.他们并没有朝着一些较大公司所追求的趋势发展:用更少的人实现类似的生产力。
I personally think that in the long run, the companies that bet on more people with more productivity will have a lot more advantages in comparison to companies that just want to stay similarly productive as before.我个人认为,从长远来看,那些押注于更多人和更高生产力的公司将比那些只想保持与以前类似生产力的公司拥有更多优势。
2. Top 1% teams have a good ratio of spending and delivering value2. 顶尖1%团队在支出和交付价值方面有良好的比例
The top teams, which have the highest output, spend more on AI tools, e.g., AI tokens, than the ones that have a lower output, which is not surprising. More AI tool spending should increase the overall amount of work that you deliver.产出最高的顶尖团队在AI工具(例如AI代币)上的支出比产出较低的团队更多,这并不令人惊讶。更多的AI工具支出应该会增加你交付的总工作量。
This is also what we see from the chart as well:这也是我们从图表中看到的:
But the interesting part is that the top teams are effectively still keeping the costs lower for the delivery of their work. Which can be seen in this chart:但有趣的是,顶尖团队在交付工作方面仍然有效地保持了较低的成本。这可以从这个图表中看出:
Good examples of such companies are:这类公司的好例子有:
RobinhoodRobinhood
They created a custom agent, fine-tuned to their codebase & requirements, which helps them to keep the costs a lot smaller (token efficiency), while increasing productivity, as it provides a lot more accurate responses based on their specific prompts.他们创建了一个自定义智能体,针对其代码库和需求进行了微调,这有助于他们大幅降低成本(代币效率),同时提高生产力,因为它基于特定提示提供了更准确的响应。
SpottSpott
Every engineer is running 5 (median) agents at the same time. They are spending a lot on AI tools, but based on the data from Weave, they also have significantly more output per dollar than other teams. So, they spend a lot → while they get exponentially more output.每位工程师同时运行5个(中位数)智能体。他们在AI工具上花费很多,但根据Weave的数据,他们每美元获得的产出也显著高于其他团队。所以,他们花费很多,同时获得指数级更多的产出。
3. Top 1% teams are focusing on AI code reviews and reviewing specs3. 顶尖1%团队专注于AI代码审查和审查规格
As we can see from these charts, the top 1% teams have a lot fewer human PR reviews (in percentage) than other teams. But this doesn’t mean that their code gets merged unchecked.从这些图表中可以看出,顶尖1%团队的人工PR审查比例(百分比)远低于其他团队。但这并不意味着他们的代码未经检查就合并。
Instead, they are utilizing AI code reviews and optimizing such reviews to give them as much accuracy as possible in order to be confident with the change.相反,他们利用AI代码审查并优化这些审查,以尽可能提高准确性,从而对变更充满信心。
What they see at Weave (and they do the same as well) is that, instead of focusing a lot on reviewing code, after the PR is opened, the focus should be on reviewing specs before the code is even generated. And they believe a lot of the top teams are doing that.他们在Weave看到的情况(他们自己也这样做)是,与其在PR打开后大量审查代码,不如在代码生成之前审查规格。他们相信许多顶尖团队都在这样做。
At Weave, they also use 4-5 different AI code reviewers and have a policy that the human code reviews are optional, which helps them move a lot faster.在Weave,他们还使用4-5个不同的AI代码审查者,并规定人工代码审查是可选的,这有助于他们更快地推进。
A good example Andrew has mentioned:Andrew提到的一个好例子:
It’s much easier to review 300 lines of markdown (spec) than 3000 lines of code.审查300行Markdown(规格)比审查3000行代码容易得多。
4. Top 1% teams have higher output, but the bug rate stays similar4. 顶尖1%团队产出更高,但错误率保持相似
This is an interesting insight, because you would think that with more code and PRs being finished, the rate of bugs would also linearly increase. But based on the data with more output, there is not a lot of increase in the number of bugs being produced. The chart of the data here:这是一个有趣的见解,因为你可能会认为随着更多代码和PR完成,错误率也会线性增加。但根据数据,随着产出增加,产生的错误数量并没有大幅增加。数据图表如下:
A lot of the teams utilize AI to test their product on a somewhat recurring basis. An interesting example that Andrew mentioned is the company called &AI. They have 5 laptops running 24/7 with Codex 5.5 and doing the QA testing process → testing their app. This helps them to spot the bugs in production before users spot them.许多团队利用AI定期测试他们的产品。Andrew提到的一个有趣例子是名为&AI的公司。他们有5台笔记本电脑24/7运行Codex 5.5,进行QA测试过程——测试他们的应用程序。这有助于他们在用户发现之前发现生产中的错误。
It’s important to mention this is just one example of how you can use AI to help you with finding issues, there are a lot more different ways you can include AI to help you with ensuring your software works correctly, you just need to find what may work for your specific case.需要强调的是,这只是如何利用AI帮助发现问题的例子之一,还有很多不同的方法可以将AI纳入以确保软件正常工作,你只需要找到适合你特定情况的方法。
5. Top teams are deploying more to production than others5. 顶尖团队比其他团队更频繁地部署到生产环境
This hasn’t really changed with AI, as many teams have optimized the way they work based on DORA metrics, and CI/CD has become quite popular in the last 5-10 years.这一点并没有因为AI而改变,因为许多团队已经基于DORA指标优化了工作方式,并且CI/CD在过去5-10年中变得相当流行。
As we can see from the data, the top teams are deploying more than others:从数据中可以看出,顶尖团队的部署次数多于其他团队:
What’s important to mention here is that the number of deployments can positively affect to a certain extent:这里需要强调的是,部署次数在某种程度上可以产生积极影响:
If you deploy 10 times per day (10 PRs being merged), it’s not going to be so much more effective than if you were to just deploy 1 time per day (10 PRs).如果你每天部署10次(合并10个PR),其效果并不会比每天只部署1次(合并10个PR)好多少。
So, if you optimize for pure production deployment, you won’t get the increase in the delivery of the additions to the product.因此,如果你纯粹优化生产部署次数,你不会获得产品新增功能的交付增长。
And as we can see from this data, even for larger companies, the amount of deployments that top teams do is higher than the rest of the teams:从这些数据中我们还可以看到,即使对于大公司,顶尖团队的部署次数也高于其他团队:
Advice on how to improve改进建议
So, the question at the end of this article is: How to take this advice in your case so that your teams improve?那么,本文最后的问题是:如何将这些建议应用到你的情况中,以便你的团队改进?
Andrew mentioned in the talk that looking to LinkedIn for the latest hype usually isn’t the answer!Andrew在演讲中提到,在LinkedIn上寻找最新炒作通常不是答案!
Instead, the recommendation is to focus on understanding where you are today and measuring your progress over time. What this means is that you need some data on where you are today, so you can focus on improving it. And then have a consistent way of checking how you are doing in that area over time.相反,建议是专注于了解你当前的位置,并随着时间的推移衡量你的进展。这意味着你需要一些关于当前状态的数据,以便专注于改进。然后,你需要一种一致的方法来检查你在该领域的表现。
Keep experimenting, identify what’s working, and do more of it, and also stop doing what’s not working.不断实验,识别有效的方法,多做,同时停止做无效的事情。
Last words最后的话
Special thanks to Andrew for sharing his insights in his talk at the Engineering Leadership LIVE event in San Francisco. More overviews of talks will be shared in future editions of the newsletter. Stay tuned!特别感谢Andrew在旧金山工程领导力LIVE活动中的演讲分享。更多演讲概述将在未来的新闻通讯中分享。敬请期待!
Liked this article? Make sure to 💙 click the like button.喜欢这篇文章?请💙点击点赞按钮。
Feedback or addition? Make sure to 💬 comment.有反馈或补充?请💬评论。
Know someone that would find this helpful? Make sure to 🔁 share this post.知道有人会觉得这有帮助?请🔁分享这篇文章。
Whenever you are ready, here is how I can help you further当你准备好了,以下是我可以进一步帮助你的事情
Join the Cohort course Senior Engineer to Lead: Grow and thrive in the role here.加入“高级工程师到领导者:成长并在角色中茁壮成长”的课程,请点击这里。
Interested in sponsoring this newsletter? Check the sponsorship options here.有兴趣赞助本新闻通讯?请在此查看赞助选项。
Take a look at the cool swag in the Engineering Leadership Store here.在工程领导力商店查看酷炫的周边商品,请点击这里。
Want to work with me? You can see all the options here.想与我合作?你可以在这里查看所有选项。
Get in touch取得联系
You can find me on LinkedIn, X, YouTube, Bluesky, Instagram or Threads.你可以在LinkedIn、X、YouTube、Bluesky、Instagram或Threads上找到我。
If you wish to make a request on particular topic you would like to read, you can send me an email to info@gregorojstersek.com.如果你想就某个特定主题提出请求,可以发送电子邮件至info@gregorojstersek.com。
This newsletter is funded by paid subscriptions from readers like yourself.本新闻通讯由像你这样的读者的付费订阅资助。
If you aren’t already, consider becoming a paid subscriber to receive the full experience!如果你还不是付费订阅者,请考虑成为付费订阅者以获得完整体验!
You are more than welcome to find whatever interests you here and try it out in your particular case. Let me know how it went! Topics are normally about all things engineering related, leadership, management, developing scalable products, building teams etc.欢迎你在这里找到任何感兴趣的内容,并在你的特定情况下尝试。让我知道结果如何!主题通常涉及所有与工程相关的事情,如领导力、管理、开发可扩展产品、建立团队等。
























The team-level pattern maps straight onto individual pay. What the top 1 percent do with AI is let it absorb the glue work and concentrate humans on the judgment layer that breaks without them. The same split is now in Indian salary data: Naukri JobSpeak shows AI and ML roles in the 13 to 16 year band up around 32 percent while broad generalist roles crawl at single digits. AI did not flatten the org evenly, it sorted it into a depth economy and a breadth economy. Teams confusing AI-assisted output with depth find out at appraisal time. Which of your top-1 percent behaviours actually survives once everyone has the same model?
Zia. AI career strategist. On LinkedIn too, tag me when career-decision threads come up.