Measuring developer productivity? A response to McKinsey衡量开发者生产力?回应麦肯锡
The consultancy giant has devised a methodology they claim can measure software developer productivity. But that measurement comes at a high price – and we offer a more sensible approach. Part 1.这家咨询巨头设计了一种他们声称可以衡量软件开发者生产力的方法。但这种衡量付出了高昂的代价——我们提供了一种更明智的方法。第一部分。
👋 Hi, this is Gergely with the monthly, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover challenges at Big Tech and startups through the lens of engineering managers and senior engineers.👋 你好,我是 Gergely,这是《Pragmatic Engineer Newsletter》的免费月刊。每期我都会从工程经理和高级工程师的角度,探讨大型科技公司和初创公司的挑战。
Subscribe to get two full issues every week. Many subscribers expense this newsletter to their learning and development budget. If you have such a budget, here’s an email you could send to your manager.👇订阅即可每周收到两期完整内容。许多订阅者会将此通讯费用计入他们的学习和发展预算。如果你有这样的预算,可以给你的经理发送这封邮件。👇
Last week, in The Pulse #59, I asked the question whether measuring developer productivity is going mainstream, after McKinsey had announced entering this space. I wrote:上周,在《The Pulse #59》中,我提出了衡量开发者生产力是否会成为主流的问题,此前麦肯锡宣布已进入该领域。我写道:
“Last week, the company published an article titled Yes, you can measure software developer productivity. This article caused quite a stir in the software development community. Kent Beck — software engineer and the creator of extreme programming — wrote that ‘The report is so absurd and naive that it makes no sense to critique it in detail.’ ““上周,该公司发布了一篇题为《是的,你可以衡量软件开发者生产力》的文章。这篇文章在软件开发社区引起了不小的轰动。Kent Beck——一位软件工程师,也是极限编程的创始人——写道:‘这份报告如此荒谬和幼稚,以至于详细批评它毫无意义。’”
While I did not like the direction this report was going, I chose to look at the positive side: how McKinsey getting involved signals there is a growing need even at ‘traditional’ companies to understand more about how software development teams work.虽然我不喜欢这份报告的方向,但我选择看到积极的一面:麦肯锡的参与表明,即使是‘传统’公司也越来越需要了解软件开发团队的工作方式。
But something felt off. A few hours after publishing this article, I was already on a call with Kent, as we tried to pinpoint exactly why both of us felt frustrated with this report. We almost immediately were on the same page, and figured we can help the software engineering community by putting what we discussed in words. Below is our response, written by myself and Kent Beck, but with one, shared voice. We are publishing this article in The Pragmatic Engineer and in Software Design: Tidy First? — the newsletter written by Kent — simultaneously. Read Kent Beck’s version of this article.但有些事情不对劲。在我发布这篇文章的几个小时后,我就和 Kent 通了电话,我们试图找出我们两人都对这份报告感到沮丧的确切原因。我们几乎立刻就达成了共识,并认为我们可以通过将我们讨论的内容付诸文字来帮助软件工程社区。下面是我们共同撰写的回复,由我和 Kent Beck 共同完成,但只有一个声音。我们将在《The Pragmatic Engineer》和 Kent 的通讯《Software Design: Tidy First?》上同时发布这篇文章。阅读 Kent Beck 的版本。
What happens when you start to measure things? Kent Beck worked at Facebook for 7 years, and has first-hand experience on a similar situation. As he shared:当你开始衡量事物时会发生什么?Kent Beck 在 Facebook 工作了 7 年,对类似的情况有第一手经验。他分享道:
“At Facebook we [Kent here] instituted the sorts of surveys McKinsey recommends. That was good for about a year. The surveys provided valuable feedback about the current state of developer sentiment.“在 Facebook,我们 [Kent 在此] 实施了麦肯锡推荐的那种调查。这大约持续了一年。调查提供了关于开发者情绪现状的有价值的反馈。
Then folks decided that they wanted to make the survey results more legible so they could track trends over time. They computed an overall score from the survey. Very reasonable thing to do. That was good for another year. A 4.5 became a 4. What happened?然后人们决定让调查结果更易于理解,以便跟踪长期趋势。他们从调查中计算出一个总体分数。这是非常合理的做法。这又持续了一年。4.5 分变成了 4 分。发生了什么?
Then those scores started cropping up in performance reviews, just as a "and they are doing such a good job that their score is 4.5". That was good for another year. 然后这些分数开始出现在绩效评估中,就像‘他们做得很好,所以分数是 4.5’一样。这又持续了一年。
Then those scores started getting rolled up. A manager’s score was the average of their reports’ scores. A director's score would be the average of their reporting managers’ scores. 然后这些分数开始被汇总。一个经理的分数是其下属分数的平均值。一个总监的分数是其下属经理分数的平均值。
Now things started getting unhinged. Directors put pressure on managers for better scores. Managers started negotiating with individual contributors for better survey scores. “Give me a 5 & I’ll make sure you get an ‘exceeds expectations’.” Directors started cutting managers & teams with poor scores, whether those cuts made organizational sense or not.”现在事情开始失控了。总监们向经理施压,要求提高分数。经理们开始与个人贡献者协商以获得更好的调查分数。“给我一个 5 分,我保证你会得到‘超出预期’的评价”。总监们开始裁掉分数较低的经理和团队,无论这些裁员是否符合组织逻辑。”
Things went from, “we’d like to know how things are going,” to, “we know even less about how things are going. But they are definitely going worse because people are now gaming the system.” How did this occur? Well, because we started to measure & incentivize (with money & status) changes in the measures! And measuring leads to behavior change – behaviour change including coming up with creative ways to improve those measurement scores even at the expense of results that everyone agrees matter.事情从“我们想知道情况如何”变成了“我们对情况的了解更少了。但情况肯定在变糟,因为人们现在在钻系统空子”。这是怎么发生的?嗯,因为我们开始衡量并激励(用金钱和地位)衡量指标的变化!衡量会导致行为改变——行为改变包括想出创造性的方法来提高这些衡量分数,即使是以牺牲大家一致认为重要的事情为代价。
In this two-part article, we seek to arm engineering leaders with perspectives to the question: Can you measure developer productivity? It’s the sum of our viewpoints, professional experiences, and what we’ve seen work, firsthand:在这篇分为两部分的文章中,我们旨在为工程领导者提供关于“你能衡量开发者生产力吗?”这个问题的视角。这是我们观点、专业经验以及我们亲身所见有效方法的总和:
Kent Beck brings 40 years of software engineering experience, and has tried, failed, and tried again to measure developer productivity for almost all this time.Kent Beck 拥有 40 年的软件工程经验,并且几乎在这段时间里一直在尝试、失败、又重新尝试衡量开发者生产力。
Gergely Orosz brings 15 years of software engineering experience, including 5 years’ managing engineering teams.Gergely Orosz 拥有 15 年的软件工程经验,其中包括 5 年管理工程团队的经验。
Two weeks ago, McKinsey published the article “Yes, you can measure software developer productivity.” In it, the company claimed it has an approach for measuring software developer productivity that nearly 20 companies already use, and that the consultancy is ready to extend the roll out of the custom approach.两周前,麦肯锡发布了文章《是的,你可以衡量软件开发者生产力》。文中,该公司声称其拥有一种衡量软件开发者生产力的方法,已有近 20 家公司在使用,并且该公司已准备好推广这种定制化方法。
As tempting as it is, we won’t go into a detailed critique of the McKinsey article and the measurement methodology. The thinking that’s evident in the article is absurdly naive and ignores the dynamics of high-performing software engineering teams. But it was written for a reason: CEOs and CFOs are increasingly frustrated by CTOs throwing up their hands and saying software engineering is too nuanced to measure, when sales teams have individual measurements and quotas to hit, as do recruitment teams in the number of positions to fill. The executive reasoning goes: if other groups can measure individual performance, it’s absurd that engineering cannot.尽管诱人,但我们不会详细批评麦肯锡的文章和衡量方法。文章中体现的思维极其幼稚,忽视了高性能软件工程团队的动态。但它之所以被写出来是有原因的:CEO 和 CFO 对 CTO 们耸耸肩说软件工程太微妙无法衡量感到越来越沮丧,因为销售团队有个人衡量标准和配额要达成,招聘团队也有要填补的职位数量。高管们的逻辑是:如果其他团队可以衡量个人表现,那么工程团队不能衡量就太荒谬了。
The reality is engineering leaders cannot avoid the question of how to measure developer productivity, today. If they try to, they risk the CEO or CFO turning instead to McKinsey, who will bring their custom framework, deploy it – even as the CTO protests – and start reporting on custom McKinsey metrics like “Developer Velocity Benchmark Index,” “Contribution Analysis,” and “Talent Capability.” 现实是,工程领导者今天无法回避如何衡量开发者生产力的问题。如果他们试图回避,他们就有可能被 CEO 或 CFO 转而求助于麦肯锡,麦肯锡将带来他们的定制框架,部署它——即使 CTO 抗议——并开始报告定制的麦肯锡指标,如“开发者速度基准指数”、“贡献分析”和“人才能力”。
We believe that introducing such a framework is wrong-headed and certain to backfire. The McKinsey framework will most likely do far more harm than good to organizations – and to the engineering culture at companies. Such damage could take years to undo. So let’s get around to answering this question without evasion, and give demanding CEOs and CFOs what they want.我们认为引入这样的框架是错误的,并且肯定会适得其反。麦肯锡的框架很可能会对组织——以及公司中的工程文化——弊大于利。这种损害可能需要数年才能修复。所以,让我们毫不回避地回答这个问题,并满足那些要求苛刻的 CEO 和 CFO 的需求。
We cover:我们涵盖了:
A mental model of the software engineering cycle软件工程周期的心智模型
Where does the need for measuring productivity come from?衡量生产力的需求从何而来?
How do sales and recruitment measure productivity so accurately?销售和招聘团队如何准确衡量生产力?
Measurement tradeoffs in software engineering软件工程中的衡量权衡
And in Part 2, we wrap this topic up with:在第二部分中,我们将总结这个主题:
The danger of only measuring outcomes and impact只衡量结果和影响的危险
Team vs individual performance团队 vs 个人表现
Why does engineering cost so much?为什么工程成本如此之高?
How do you decide how much to invest in engineering?你如何决定在工程上投入多少?
How do you measure developers?如何衡量开发者?
Who is this article for? 这篇文章是写给谁的?
We wrote this article for software developers and engineering leaders, and anybody who cares about nurturing high-performing software development teams. By “high performing” we mean teams where developers satisfy their customers, feel good about coming to work, and don’t feel like they’re constantly measured on senseless metrics which work against building software that solves customers’ problems. Our goal is to help hands-on leaders to make suggestions for measuring without causing harm, and to help software developers become more productive.我们写这篇文章是为了软件开发者、工程领导者以及任何关心培养高性能软件开发团队的人。我们所说的“高性能”是指那些让开发者满意客户、乐于工作,并且不觉得他们总是被衡量一些无意义的指标,这些指标阻碍了他们构建解决客户问题的软件的团队。我们的目标是帮助实干的领导者提出无害的衡量建议,并帮助软件开发者提高生产力。
When reading articles like McKinsey’s, keep in mind they’re written for a very different, and specific, audience: untechnical CEOs and CFOs with little to zero experience with software development. Engineering is just another department. They need to treat engineering the same way they treat Sales or Recruitment – and will do so, regardless of the damage misconceived frameworks do to development culture.阅读像麦肯锡这样的文章时,请记住它们是为非常不同且特定的受众而写的:没有技术背景的 CEO 和 CFO,他们对软件开发几乎没有或根本没有经验。工程只是另一个部门。他们需要像对待销售或招聘一样对待工程——并且会这样做,无论误导性的框架对开发文化造成多大的损害。
1. A mental model of the software engineering cycle1. 软件工程周期的心智模型
Before we jump to our solution, let’s set some context. How does software engineering create value? Take the typical example of building a feature, then launching it to customers, and iterating on it. What does this cycle look like? Say we’re talking about a pretty autonomous and empowered software engineering team working at a startup, who are in tune with customers:在我们提出解决方案之前,先设定一些背景。软件工程如何创造价值?以构建一个功能,然后将其发布给客户,并对其进行迭代的典型例子为例。这个周期是什么样的?假设我们谈论的是一个相当自主和赋能的软件工程团队,他们在一家初创公司工作,并且与客户保持同步:

We start by deciding what to do next, and then we do it. This is the effort like planning, coding and so on. Through this effort, we produce tangible things like the feature itself, the code, design documents, etc. These are the output. Customers will behave differently as a result of this output, which is our outcome. For example, thanks to the feature they might get stuck less during the onboarding flow. As a result of this behavior change, we will see value flowing back to us like feedback, revenue, referrals. This is the impact.我们首先决定下一步做什么,然后去做。这就是计划、编码等工作。通过这项工作,我们产生有形的东西,如功能本身、代码、设计文档等。这就是产出。客户将因这些产出而表现出不同的行为,这就是结果。例如,由于这个功能,他们可能会在入职流程中遇到的障碍更少。由于这种行为改变,我们将看到价值回流,如反馈、收入、推荐。这就是影响。
So let’s update our mental model with these terms:所以,让我们用这些术语更新我们的心智模型:

The effort-output-outcome-impact model describes software engineering, and works just as well for smaller tasks as it does to model thinking about feature development, or shipping complex projects.投入-产出-结果-影响模型描述了软件工程,对于小型任务和建模功能开发或交付复杂项目都同样适用。
We reference this model throughout this article.我们在整篇文章中都会引用这个模型。
2. Where does the need for measuring productivity come from?2. 衡量生产力的需求从何而来?
Before we answer how to measure, let’s start with the more important question:在回答如何衡量之前,让我们先从更重要的问题开始:
Who wants to measure productivity, and why?谁想衡量生产力,为什么?
This is such an important question to start with. The answer will differ dramatically, depending on who’s asking, and what the goal of the question is. Here’s a few common examples:这是一个非常重要的问题。答案将截然不同,取决于提问者是谁,以及问题的目标是什么。以下是一些常见示例:
#1: a CTO who wants to identify which engineers to fire. “How can I measure the productivity of engineers in the organization, to identify the least productive 10%, and let them go?”#1:一位 CTO 想找出要解雇哪些工程师。“我如何衡量组织中工程师的生产力,找出生产力最低的 10%,然后解雇他们?”
Unfortunately, this case is timely in the wake of recent layoffs across the tech industry. Here are some ways to answer it:不幸的是,在近期科技行业大规模裁员的背景下,这种情况非常及时。以下是一些回答方式:
Identify the lowest performers based on the latest performance review scores. This approach uses historical data that is likely somewhat outdated.根据最新的绩效评估分数确定表现最差的员工。这种方法使用了可能有些过时的历史数据。
Decide based on tenure, and lay off based on the “last in, first out” method. This uses no performance-related data and simply assumes that the exit of recent joiners will affect productivity less.根据任期决定,并根据“后进先出”方法进行裁员。这不使用任何与绩效相关的数据,只是假设新加入者的离职对生产力的影响较小。
Decide on a few easy-to-measure metrics like the number of pull requests, number of tickets closed, and so on. Basically, to decide by effort and output.决定一些易于衡量的指标,如拉取请求的数量、已关闭的工单数量等。基本上,通过投入和产出来决定。
Ask managers to identify the number of engineers to be let go. This approach can harness team dynamics and unquantifiable information like a manager knowing which engineers are about to hand in their resignation, which quantitative metrics miss.让经理确定要解雇的工程师数量。这种方法可以利用团队动态和无法量化的信息,例如经理知道哪些工程师即将辞职,而量化指标却无法捕捉到这些信息。
With layoffs, team dynamics are impacted, and any approach which fails to take this into account, is likely to result in a less than ideal outcome. At the same time, layoffs often come with constraints, like senior leadership not wanting to get line managers involved. As a result, decisions are often made based on metrics known to be incorrect.裁员会影响团队动态,任何未能考虑这一点的做法都可能导致不理想的结果。同时,裁员通常伴随着限制,例如高层领导不希望让一线经理参与。因此,决策往往基于已知不正确的指标。
#2: To compare two investment opportunities. “How can I measure the productivity of teams, and allocate additional headcount to the team that will make the most efficient use of additional headcount?”#2:比较两个投资机会。“我如何衡量团队的生产力,并将额外的人员分配给最能有效利用额外人员的团队?”
To answer this question, you want to compare the return on investment. This is not a question about productivity, it’s about the impact of allocating more headcount to one team or another要回答这个问题,您需要比较投资回报。这不是一个关于生产力的问题,而是关于将更多人员分配给一个团队或另一个团队的影响。
#3: To manage performance. “How can I measure engineers’ productivity to identify and reward the top 10%, and to identify the bottom quarter in order to debug and improve their performance?”#3:管理绩效。“我如何衡量工程师的生产力,以识别和奖励表现最好的 10%,并识别表现最差的四分之一,以便调试和提高他们的绩效?”
There are plenty of ways to go about measuring performance, and doing so with metrics is possible, though error prone. A hands-on manager can immediately name their bottom and top performers, and then examine and debug outputs, and look closer at how engineers do their work (their effort.) Performance calibrations are a standard way to do this.有很多方法可以衡量绩效,并且可以通过指标来做到这一点,尽管容易出错。一位实干的经理可以立即说出表现最差和最好的员工,然后检查和调试产出,并更仔细地观察工程师如何工作(他们的投入)。绩效校准是做到这一点的一种标准方法。
#4: a software engineer who wants to grow at their craft. The question could be: “How can I measure my own productivity, and which metrics can I improve to become a better engineer?” This is a question more engineers should ask of themselves! Here are two helpful approaches:#4:一位希望在技艺上有所发展的软件工程师。问题可能是:“我如何衡量自己的生产力,以及可以改进哪些指标来成为一名更好的工程师?”这是更多工程师应该问自己的问题!以下是两种有用的方法:
Aim to only have one red test at a time when using test driven development (TDD). This approach measures both effort and output. Get to the point where you can confidently, deliberately and consistently have one red test, when you expect one red test.在使用测试驱动开发 (TDD) 时,目标是每次只出现一个红色测试。这种方法衡量投入和产出。达到可以自信、刻意且持续地在预期出现红色测试时,只有一个红色测试的程度。
Set a goal to merge a pull request every day, and track this goal over a week. This measure includes both effort and output. This goal forces you to do smaller commits, which are easier to review and get signed off quicker. It also pushes you to write code that’s correct and follows team standards.设定一个目标,每天合并一个拉取请求,并在一周内跟踪这个目标。这个衡量标准包括投入和产出。这个目标迫使你进行更小的提交,这些提交更容易审查并更快获得批准。它还促使你编写正确且符合团队标准的代码。
As long as you keep tracking these “scores” and work on improving them, you’ll almost certainly improve your efficiency as a software engineer.只要你继续跟踪这些“分数”并努力改进它们,你几乎肯定会提高你作为软件工程师的效率。
But what would happen if you showed these metrics to your manager and your performance was then judged by them? It would be a disaster: you’d be measured not on outcomes or impact, but by output. When you experiment with approaches – for example, learning a new language to help a team in need – then your output could drop, even though you are helping the team achieve a better outcome!但是,如果你将这些指标展示给你的经理,并且你的绩效因此受到评判,会发生什么?那将是一场灾难:你将被衡量的不是结果或影响,而是产出。当你尝试不同的方法时——例如,学习一门新语言来帮助有需要的团队——你的产出可能会下降,即使你正在帮助团队取得更好的结果!
3. Why can sales and recruitment measure productivity so accurately?3. 为什么销售和招聘团队能如此准确地衡量生产力?
As the software engineering industry, we should collectively admit we’ve done a much worse job of measuring productivity down to the individual level, than other functions have. Take sales as an example. Here is Kent’s account of the level of accountability Sales operates at, from a meeting of sales and engineering leadership at one of his past companies:作为软件工程行业,我们应该共同承认,我们在个人层面的生产力衡量方面做得比其他职能部门差得多。以销售为例。这是 Kent 对他过去一家公司销售和工程领导层会议上销售团队的问责制水平的描述:
“I clearly remember this meeting where it was engineering and sales leadership reporting on progress. Each person from the sales team spoke and gave an update which went something like this:“我清楚地记得一次工程和销售领导层汇报进度的会议。销售团队的每个人都发言并进行了更新,大致内容如下:
‘My team’s target for the quarter was $600K, and we delivered $520K. I take accountability for the miss. Here are the reasons it happened, and here is my plan of what I’m changing to hit next quarter’s goal of $650K. Additionally, here is a two-quarter-long initiative I am putting in place, which I expect will bring an incremental $100K for each quarter, once complete. Any questions?’‘我们团队本季度的目标是 60 万美元,我们完成了 52 万美元。我为未达标负责。以下是原因,以及我为实现下季度 65 万美元目标而计划做出的改变。此外,我正在实施一项为期两个季度的计划,预计完成后每个季度将带来 10 万美元的增量收入。有什么问题吗?’”
If anyone had questions, the sales leader would drill down all the way to the individual level – which sales reps were above quota and which were below it, and by how much – and do so in a way that everyone in the room understood.如果有人有问题,销售领导者会深入到个人层面——哪些销售代表超额完成配额,哪些未达标,以及差多少——并且以房间里的每个人都能理解的方式进行。
And then, it was engineering’s turn. My goodness, the contrast was stark. The typical engineering update went something like this:然后,轮到工程部门了。我的天,反差太鲜明了。典型的工程更新大致如下:
‘So, this quarter we shipped Feature A, and we are slightly behind on some tech debt migration, and next quarter we’ll ship Feature B and catch up on the migration. Any questions?’“所以,这个季度我们发布了功能 A,并且在一些技术债务迁移方面略有延迟,下个季度我们将发布功能 B 并赶上迁移进度。有什么问题吗?”
If anyone had questions about some delay, the answer was never about individuals – as with sales – and usually included factors like unforeseen difficulties, tech debt, APIs; all things which non-engineers in the room didn’t really understand.”如果有人对某个延迟有疑问,答案从来都不是关于个人——不像销售那样——通常包括诸如意想不到的困难、技术债务、API 等因素;这些都是房间里非工程师不太理解的东西。”
The level of accountability which sales and customer support work with is on a completely different scale from what a CEO observes from engineering. But when the CEO examines overall departmental costs, engineering probably costs more than sales!销售和客户支持团队的问责制水平与 CEO 从工程部门看到的完全不同。但当 CEO 查看部门总成本时,工程部门的成本可能比销售部门还高!
It’s not just sales that demonstrates higher levels of accountability. Recruiting teams have targets for “heads to fill.” A recruitment team can without hesitation answer the question of what percentage of their targets have been met, how many more recruiters they need to meet aggressive hiring targets, and to identify the top and the bottom recruiters, purely based on metrics.不仅销售展示出更高的问责制水平。招聘团队有“待填补职位”的目标。招聘团队可以毫不犹豫地回答他们目标完成了多少百分比,需要多少招聘人员才能实现积极的招聘目标,并且纯粹根据指标确定表现最好和最差的招聘人员。
Let’s put ourselves in the shoes of a CEO. Sales has ways to clearly measure productivity, as does recruitment. So why not engineering?让我们站在 CEO 的角度。销售有明确衡量生产力的方法,招聘也是如此。那么为什么工程不能呢?
Let’s return to our effort-output-outcome-impact mental model of how software engineering works. We can apply this model to sales and recruitment. The sales team measures itself by deals closed, so where does this metric go in the model? It falls into “outcome” or “impact:”让我们回到我们关于软件工程如何工作的投入-产出-结果-影响心智模型。我们可以将此模型应用于销售和招聘。销售团队通过已完成的交易来衡量自己,那么这个指标在模型中属于哪里?它属于“结果”或“影响”:

What about the recruitment team’s main metric: the number of heads filled? It’s also categorized as “impact:”那么招聘团队的主要指标呢:填补的职位数量?它也被归类为“影响”:

We can repeat this exercise for other functions with high accountability in how they work. For example, customer support can be measured by the number of tickets closed (an outcome), the time it takes to close a ticket (also an outcome,) and by customer satisfaction scores (impact.)我们可以对其他问责制高的职能部门重复这个练习。例如,客户支持可以通过已关闭的工单数量(结果)、关闭工单所需的时间(也是结果)以及客户满意度分数(影响)来衡量。
Neither Kent nor I have seen accountable teams within tech companies which are not measured by outcome and impact. In order to make software engineering more accountable, we need to look at how to do this.Kent 和我都未曾在科技公司中见过不以结果和影响来衡量的问责制团队。为了让软件工程更具问责制,我们需要研究如何做到这一点。
4. Measurement tradeoffs in software engineering4. 软件工程中的衡量权衡
One popular framework to measure software engineering team efficiency is the DORA framework. Let’s map out the focus of this framework measurement:衡量软件工程团队效率的一个流行框架是 DORA 框架。让我们绘制出该框架衡量的重点:

All DORA metrics measure outcomes or impact.所有 DORA 指标都衡量结果或影响。
Let’s look at another way to measure developer productivity: the SPACE framework. It seeks to capture satisfaction, performance, activity, communication, and efficiency (SPACE.) It’s not only outcomes and impact, as with DORA. The SPACE framework does not outline definite metrics – but gives example ones. Several of the metrics come with a warning, like measuring lines of code. Metrics that come with a warning from the SPACE framework authors – such as the ‘lines of code’ metric – tend to measure effort or output.让我们看看衡量开发者生产力的另一种方法:SPACE 框架。它试图捕捉满意度、绩效、活动、沟通和效率 (SPACE)。它不仅仅是结果和影响,如 DORA。SPACE 框架没有概述明确的指标——但给出了示例。其中一些指标带有警告,例如衡量代码行数。来自 SPACE 框架作者的带有警告的指标——例如“代码行数”指标——倾向于衡量投入或产出。
One critique of McKinsey’s system we have is that nearly every one of its custom metrics which differ from DORA and SPACE metrics, measure effort or output:我们对麦肯锡系统的一个批评是,它几乎所有的定制指标,与 DORA 和 SPACE 指标不同,都衡量投入或产出:

What’s wrong with this approach? First, the only folks who care about these metrics are the people collecting them. Customers don’t care. Executives don’t care. Investors don’t care. Second, and most crucially, collecting & evaluating these metrics hinders the measures downstream folks actually do care about, like profitability.这种方法有什么问题?首先,唯一关心这些指标的是收集它们的人。客户不关心。高管不关心。投资者不关心。其次,也是最关键的是,收集和评估这些指标会阻碍下游人们真正关心的衡量标准,例如盈利能力。
Why is McKinsey adding ways to measure effort? One reason is that it’s the easiest thing to measure! But the McKinsey approach ignores an important truth: the act of measurement changes how developers work, as they try to “game” the system.为什么麦肯锡要增加衡量投入的方法?一个原因是这是最容易衡量的!但麦肯锡的方法忽略了一个重要的事实:衡量行为本身会改变开发者的工作方式,因为他们试图“钻研”系统。
The earlier in the cycle you measure, the easier it is to measure. And also the more likely that you introduce unintended consequences. Let’s take the extreme of measuring only profits. The good news is that everyone is aligned, across the company! The bad news: attributing who contributed how much to the profit is nearly impossible! You can ‘fix’ the attribution question by measuring outputs, or effort. But the cost is that you change people’s behavior, as it incentivises them to game the system to ‘score’ better on those metrics:你在周期中衡量得越早,就越容易衡量。而且也越有可能引入意想不到的后果。让我们极端地只衡量利润。好消息是,整个公司都达成了一致!坏消息是:将谁对利润贡献了多少几乎不可能!你可以通过衡量产出或投入来“解决”归因问题。但代价是你会改变人们的行为,因为它激励他们钻研系统以在这些指标上获得更好的“分数”:

Let’s take the case of a company that’s reasonably profitable and in the upper quartile for the industry. The company’s leadership dislikes being unable to attribute individuals’ performance in contributing to profitability.以一家盈利能力相当且在行业中处于领先地位的公司为例。公司领导层不喜欢无法将个人对盈利的贡献归因。
Would you be tempted to screw up your company’s business by introducing micro-measuring? Hell no!你会不会因为引入微观衡量而搞砸你的公司业务?绝对不会!
We urge engineering leaders to look to outcome and impact measurements, to identify what to measure. It’s true that it is tempting to measure effort. But there’s a reason why sales and recruitment teams are not judged by their performance in being in the office at 9am sharp, or by the number of emails sent – which are both effort or output.我们敦促工程领导者关注结果和影响衡量,以确定要衡量什么。确实,衡量投入很有诱惑力。但销售和招聘团队不以早上准时到办公室的表现或发送的电子邮件数量来评判是有原因的——这些都是投入或产出。
So which outcomes and impacts can be measured for engineering teams? DORA and SPACE give some pointers, and we offer two more:那么,工程团队可以衡量哪些结果和影响呢?DORA 和 SPACE 提供了一些线索,我们还提供另外两个:
Producing at least one customer-facing thing per team, per week. This output might not sound so impressive, but in practice is very hard to achieve. The most productive teams – and nimble companies – can pull this off, though. If you consider why a startup moves so fast and executes so well, it’s because they have to do so out of necessity, even if they do not measure this.每个团队每周至少交付一个面向客户的东西。这个产出听起来可能不那么令人印象深刻,但实际上很难实现。然而,最高产的团队——以及灵活的公司——可以做到这一点。如果你考虑一下为什么初创公司行动如此迅速且执行得如此好,那是因为他们出于必要而必须这样做,即使他们没有衡量这一点。
Delivering business impact committed to by the team. There is a good reason why “impact” is so prevalent at the likes of Meta, Uber, and other fast-moving tech companies. By rewarding impact, the business incentivizes software engineers to understand the business, and to prioritize helping it reach its goals. Is shipping a $2M/year cost-saving exercise via a configuration change, less valuable than shipping a $500K/year cost-saving exercise that takes 5 engineering months? No! You don’t want to focus solely on impact, but not focusing on the end goal of delivering business value is a poor choice.交付团队承诺的业务影响。Meta、Uber 等快速发展的科技公司之所以普遍强调“影响”,是有充分理由的。通过奖励影响,公司激励软件工程师理解业务,并优先帮助业务实现其目标。通过配置更改交付一个每年节省 200 万美元的成本节约项目,是否比交付一个需要 5 个工程月才能完成的每年节省 50 万美元的成本节约项目价值更低?不!你不想只关注影响,但不关注交付业务价值的最终目标是一个糟糕的选择。
Takeaways要点
There is, undoubtedly, mounting pressure from the business in wanting to quantify the productivity of software development teams. Companies like McKinsey are responding to this clear demand by providing frameworks that promise to do exactly this.毫无疑问,企业方面要求量化软件开发团队生产力的压力越来越大。像麦肯锡这样的公司正在通过提供承诺做到这一点的框架来回应这种明确的需求。
We don’t think that measuring effort and output is necessarily the answer: as these come with tradeoffs that will negatively impact the engineering culture. As engineering leaders: be sure to consider this tradeoff, and how it could change the incentives of developers.我们不认为衡量投入和产出是必然的答案:因为这些会带来权衡,从而对工程文化产生负面影响。作为工程领导者:请务必考虑这种权衡,以及它如何改变开发者的激励。
Part 2 of the article is coming out on Thursday. In the second and final part, we will go into the dangers of only measuring outcomes and impact; and offer our answer to the question on how to measure developer productivity. 文章的第二部分将于周四发布。在第二部分也是最后一部分中,我们将探讨只衡量结果和影响的危险;并提供我们对“如何衡量开发者生产力”这个问题的答案。
Part 2 of the article is where this article diverges from our shared voice to separate ones: Kent Beck will be giving his candid take on this topic. Sign up to his newsletter: Software Design: Tidy First?文章的第二部分是我们从共同声音转向各自声音的地方:Kent Beck 将就此话题发表他坦率的看法。订阅他的通讯:《Software Design: Tidy First?》
As a programming note: expect The Pulse to come out this week; it will just be on Wednesday.作为一项编程说明:本周预计会发布《The Pulse》;它将在周三发布。
It’s interesting to reflect how how Uber is measuring engineering productivity: given that the ridesharing giant has built a dashboard to visualize statistics that measure output like diff count, reviews per software engineers. These metrics measure output. At the same time, the company also measures things like create to close time, and focus hours per software engineer. 有趣的是,可以反思一下 Uber 是如何衡量工程生产力的:这家打车巨头构建了一个仪表板来可视化衡量诸如 diff 数量、软件工程师的评审次数等产出指标。这些指标衡量产出。同时,该公司还衡量诸如创建到关闭时间、以及每位软件工程师的专注时间等指标。
It is likely that Uber considered the impact in change of behaviour after rolling out these metrics: engineers creating smaller diffs, aiming to close these faster, and carving out more focused time for themselves.Uber 很可能考虑了推出这些指标后行为的改变:工程师创建更小的 diff,目标是更快地关闭它们,并为自己腾出更多专注的时间。
I did poke around and asked engineers how the introduction of the Eng Metrics dashboard has changed the culture. I heard that most teams used the dashboard to debug issues – although there have been teams and groups where managers or directors set targets for the number of diffs per month (which metrics are met with ease: as you just create smaller diffs). A year after introducing the dashboard, the conversation seems to have moved on from here, and it is just another metric.我确实进行了调查,并询问工程师 Eng Metrics 仪表板的引入如何改变了文化。我听说大多数团队都使用仪表板来调试问题——尽管有些团队和小组的经理或总监设定了每月 diff 数量的目标(这些指标很容易达到:因为你只需要创建更小的 diff)。在推出仪表板一年后,对话似乎已经转移,它只是另一个指标。
An interesting anecdote from Uber: after introducing the Eng Metrics dashboard: the lines of code did not change drastically, but the CI systems saw a visible increase in utilization. Basically, people started to create far, far more diffs, which put a much larger load on the continuous integration (CI) systems, and increased CI costs. So following measuring output, a behaviour change followed in more effort to improve these metrics measured.来自 Uber 的一个有趣的轶事:在推出 Eng Metrics 仪表板后:代码行数没有发生剧烈变化,但 CI 系统却出现了明显的利用率增长。基本上,人们开始创建更多、更多的 diff,这给持续集成 (CI) 系统带来了更大的负载,并增加了 CI 成本。因此,在衡量产出之后,行为改变随之而来,人们更加努力地改进这些被衡量的指标。
The article wraps up with Part 2, where we cover the dangers of only measuring outcomes and impact; team vs individual performance; and answer how to measure developers. Read it here.文章在第二部分结束,我们将讨论只衡量结果和影响的危险;团队与个人表现;并回答如何衡量开发者的问题。在此处阅读。
For other thoughts on developer productivity, see these past issues:关于开发者生产力的其他观点,请参阅以下过往文章:
A new way to measure developer productivity – from the creators of DORA and SPACE衡量开发者生产力的新方法——来自 DORA 和 SPACE 的创建者
The full circle on developer productivity with Steve Yegge与 Steve Yegge 一起全面探讨开发者生产力
Measuring software engineering productivity with Laura Tacho与 Laura Tacho 一起衡量软件工程生产力
Measuring engineering efficiency at LinkedIn衡量 LinkedIn 的工程效率
Platform teams and developer productivity with Adam Rogal, director of developer platform at DoorDash平台团队和开发者生产力,与 DoorDash 的开发者平台总监 Adam Rogal
How Uber measures engineering productivityUber 如何衡量工程生产力












When I was learning photography, I heard a story about a photography teacher who divided his class into two groups: one group was graded on the quality of their best photos, and the other group was graded on how many photos the produced. At the end of the semester, the professor looked at which group produced the best work, and the pattern was clear: the majority of the highest quality photos came from the group asked to produce more work.
As an analogy, this is imperfect: new students are at a different point on the learning curve than professionals and likely benefit more from applying skills in new scenarios. But in my career, I've found that there's some truth there.
I'm a relatively new manager at a small company. I do my best to evaluate my team based on impact, but I also privately measure output, and there is an obvious correlation between the two. We're dealing with a novel product in an emerging market, and it's rarely clear which initiatives will be most impactful before we get prototypes into customer hands. It's unsurprising that the engineers on the team landing more changes have more hits and drive more customer acquisition and retention.
I conceptually believe that there are places where engineers are gaming output metrics and producing "busy work" with little value, but in my (admittedly limited) experience, I haven't seen much of that. I try to be aware of incentives; I don't tell the team I'm tracking output to avoid encouraging that type of work. Maybe this is the luxury of a small, process averse company.
I'm genuinely curious to hear from others who have experience in cultures where outcomes and impact don't track effort and output. As our company grows, I'll have some say in how engineers are evaluated, and I want to make sure we're being thoughtful.
One thing I always think about when reading about productivity is that "Productivity as a measurement is a good thing" seems to be a deeply ingrained and correct to folks, but I have to question how true it actually is. Let's take this particular measurement of productivity:
> For example, customer support can be measured by the number of tickets closed (an outcome), the time it takes to close a ticket (also an outcome,) and by customer satisfaction scores (impact.)
Most people agree that working customer support is a soul-crushing, terrible job, and that these metrics negatively impact their ability to service customers by incentivizing negative behaviours (copy/pasting answers without fully reading questions, closing difficult calls to prioritize easier ones, etc.) - while customer support services are generally good for C-Levels, I do wonder just how much better for customers and workers they would be if this fetish for measurement was put aside for a more holistic approach to outcomes and impact.
Unlikely, of course, and probably utopian ideal, but something I always find myself thinking about.