CTOs Agree: Cognitive Debt Is the New Technical DebtCTO们一致认为:认知债务是新技术债务

Our Shift CTO Craft Dinner format is built on candor, rather than slides or sponsor pitches. It’s just a group of senior engineering leaders talking about what’s actually happening in their organizations. At our Toronto dinner held on the outskirts of the CTO Craft conference, the theme was AI adoption in engineering我们的Shift CTO晚宴形式基于坦诚,而非幻灯片或赞助商推介。它只是一些高级工程领袖讨论他们组织中实际发生的事情。在多伦多的晚宴上,主题是工程中的AI采纳。
Within the first 5 minutes it was clear we’d spend the evening confirming one uncomfortable truth, similar to the one we heard earlier this spring in London: nobody has truly yet figured it out.在前5分钟内就很明显,我们整个晚上要确认一个令人不安的事实——就像我们今年春天在伦敦听到的那样:没有人真正弄清楚。
Note: We’ll try to figure it out at our Shift developer conference in September on the beautiful Croatian coast – tickets are on sale!注意:我们将在九月于美丽的克罗地亚海岸的Shift开发者大会上尝试弄清楚——门票已开售!
The free-for-all is over免费放任的时代已经结束
Two years ago, the mandate was simple: spend on AI, no questions asked. That era is over.两年前,口号很简单:投入AI,毫无疑问。那个时代已经结束。
What’s replaced it is a harder conversation, one several participants had clearly been having with their CFOs. The question has shifted from are you using AI to what are you getting for it?取而代之的是更艰难的对话,几位参与者显然已经在与他们的CFO进行这种对话。问题已经从“你在使用AI吗”转变为“你从中得到什么”。
The need for an ROI hasn’t changed. If anything, that window of just spending freely is dwindling. For larger organizations, the expectation is return within 12 months, or sooner.对ROI的需求并未改变。事实上,那段可以随意花钱的窗口正在缩小。对大型组织而言,期望在12个月内甚至更快看到回报。
The challenge, as one participant put it, is that the I in ROI is completely unmanaged. Engineering capacity used to mean headcount, something finance could model. Now it means tokens, and nobody controls how many tokens any individual engineer burns on a given day.正如一位参与者所说,挑战在于ROI中的I(投资)完全没有被管理。过去,工程容量意味着人头,财务可以建模。现在,它意味着令牌(tokens),而且没有人控制单个工程师每天消耗多少令牌。
The CFO has no control once they’ve signed the contract over what the actual investment is going to be. And if you can’t tell me the investment, what does projected return even mean?CFO在签署合同后对实际投资额没有控制权。如果你不能告诉我投资额,那么预测回报到底意味着什么?
Several people around the table had run into the same wall: organizations that bought the tools, signed the contracts, and then realized there’s no financial model inside the company to manage what comes next. The comparison kept coming up: early cloud adoption, FinOps before FinOps existed. We’re in that same window. Costs are still a fantasy, usage is still undefined, and the metrics to measure it haven’t been invented yet.桌上的几个人都碰到了同样的壁垒:组织买了工具,签了合同,却发现公司内部没有财务模型来管理后续工作。大家不断提到的对比是:早期的云采纳、FinOps出现之前的FinOps。我们正处在同样的窗口期。成本仍是幻想,使用量仍未定义,衡量它的指标尚未发明。
One participant’s take: pushing for ROI too hard right now might mean measuring the wrong things entirely. The smarter move is to establish a baseline first.一位参与者的观点:现在过于强硬地追求ROI可能会导致完全测量错误的东西。更聪明的做法是先建立基线。
I think eventually it will get to a more predictable state where we can say, approximately this many tokens leads to this much feature value. But I don’t think we can predict that yet.我认为最终会达到一个更可预测的状态,我们可以说,大约这么多令牌会产生这么多功能价值。但我认为我们现在还无法预测。
The more pragmatic response, shared by more than one team: stop the free-for-all, start standardizing. This doesn’t mean we’re telling people to use AI less, but nudging from “use everything” to “use the same things, smarter.”更务实的回应是,多支团队都说:停止免费放任,开始标准化。这并不意味着我们在叫人少用AI,而是从“用所有东西” nudging 到“用同样的东西,更聪明”。

What you’re actually hiring for now你现在实际上在招聘的是什么
The hiring discussion exposed a split in how people in the room think about what an engineer actually is.招聘讨论暴露出在场的人对工程师到底是什么的分歧。
One participant drew a clean line between two different roles that often get conflated: the AI engineer who ships product, and the engineer who owns system design. Language doesn’t matter anymore: Python, Go, Rust, Node. But system design hasn’t changed. Someone still has to think about availability, budgets, and the architectural decisions that AI can’t make for you.一位参与者在两个经常被混淆的角色之间划清了界限:交付产品的AI工程师,以及负责系统设计的工程师。语言已经不重要了:Python、Go、Rust、Node。但系统设计没有改变。仍然有人必须考虑可用性、预算以及AI无法为你做出的架构决策。
The new software engineer is a product leader. Someone thinking about what the product is, not just how it works. But we still need technical people who think about the design. Those are two different things.新一代软件工程师是产品领袖。思考产品是什么,而不仅仅是它如何工作。但我们仍然需要懂技术、思考设计的人。这是两件不同的事。
On the interview side, the consensus leaned toward keeping technical fundamentals, but with caveats. One team hadn’t changed their process yet, still testing for systematic thinking, for the ability to break down ambiguous problems. Others were actively rethinking it. The most interesting take came from someone who’d shifted their interviews toward code review rather than coding, precisely because that’s what engineers actually do now.在面试方面,大家的共识倾向于保留技术基础,但有前提条件。一个团队的流程尚未改变,仍在测试系统性思维、拆解模糊问题的能力。其他团队则在积极重新思考。最有趣的观点来自一位把面试重点从编码转向代码审查的人,因为这正是工程师现在实际做的事。
I changed our interview process to focus on code review, because that’s what we’re actually doing. And implementation is now AI-assisted, however you choose to use your agents.我把我们的面试流程改为聚焦代码审查,因为这才是我们实际在做的事。实现现在是AI辅助的,无论你如何使用你的代理。
If your team can generate code faster than it can review it, you have a bottleneck. The constraint is human judgment, not output.如果你的团队生成代码的速度快于审查的速度,就出现了瓶颈。约束是人类判断,而不是产出。
Team reactions and ‘recalibration moments’团队反应与“重新校准时刻”
Across the table, nobody described a team that was uniformly enthusiastic or uniformly resistant. The reality was messier.在桌子另一侧,没有人描述出一个完全热情或完全抵触的团队。现实更为混乱。
One leader talked about late adopters at his company who finally jumped in after being gently pushed, then hit what he called “recalibration moments“: realizing that whole categories of work that used to take days now take hours, and having to rethink how their schedule works.一位领袖谈到他公司里迟到的采用者,最终在被轻推后跳进来,然后经历了他所谓的“重新校准时刻”:意识到以前需要几天的工作现在只要几个小时,并且必须重新思考他们的时间表。
Another described something more complicated: an engineer who is highly productive with AI, genuinely good at using it, but also deeply skeptical of AI-generated output.另一位描述了更复杂的情况:一名工程师在AI上高度生产力,真的很会使用AI,但对AI生成的输出深感怀疑。
There’s one person who’s really good at using AI, very productive, but also highly sensitive to anything AI has produced. “This is written by AI.” “Okay, but is it good?” “Yes, it’s good. But it’s written by AI.”有个人非常擅长使用AI,产出很高,但对任何AI生成的东西都非常敏感。“这是AI写的。”“好,但它好么?”“好,它是AI写的。”
The concern about AI making people stop thinking came up more than once. The counterargument wasn’t a dismissal. It was a reframe: this is a management problem, not a technology problem. The tools make it easy to be lazy. The job of a leader is to make laziness not worth it.关于AI让人停止思考的担忧出现了不止一次。反驳并不是否定,而是重新框定:这是管理问题,而不是技术问题。工具让人变懒变得容易。领袖的工作是让懒惰变得不值得。
You can choose to make something fast and long and not that good. Or you can use it to iterate and really drill it down to something short. We can all write really long letters now. Whether we should is a different question.你可以选择做得快、长但不太好。或者用它来迭代,真正把它压缩成短小。我们现在都能写很长的信。是否该这么做是另一个问题。
One company ran an anonymous survey and found 90% of engineers at that organization actually want to use AI, higher than expected. What surprised him wasn’t the enthusiasm but what came next: questions about performance management, about promotion criteria, about how individual contribution gets recognized when anyone can now generate code. The adoption had outrun the enablement.一家公司做了匿名调查,发现该组织90%的工程师实际上想使用AI,超出预期。让他惊讶的不是热情,而是接下来出现的问题:关于绩效管理、晋升标准、当任何人都能生成代码时个人贡献如何被认可的疑问。采纳速度已经超过了赋能。
We’re applying AI adoption rapidly, but we haven’t updated the career pathway or rethought the competency matrix. They have every right to ask those questions.我们在快速采纳AI,但还没有更新职业路径或重新思考能力矩阵。他们完全有理由提出这些问题。

The feature debt problem nobody wants to talk about没人愿意谈论的功能债务问题
AI makes it cheap to write code. That is not the same as it being cheap to ship it, or to maintain it. One participant put it cleanly: cognitive debt is the new technical debt.AI让写代码变得廉价。这并不等同于交付或维护也变得廉价。一位参与者简洁地说:认知债务是新技术债务。
The room had seen the same pattern: teams adding features at a pace that would have been impossible two years ago, now dealing with the maintenance overhead that comes with it. Legacy code that was already hard to understand is now harder, because the people who wrote it aren’t being careful. They’re being fast. And internal tools that were never meant to be permanent are now permanent because someone shipped them with three prompts.现场看到相同的模式:团队以两年前不可能的速度添加功能,却现在要处理随之而来的维护负担。原本已经难以理解的遗留代码现在更难,因为写代码的人不再小心,他们在追求速度。而那些本不该永久存在的内部工具,因为有人用了三次提示就交付了,现在却成了永久品。
Those processes, build or buy, is this worth maintaining long term, were there for a reason. But if you can spin something up in an afternoon, it’s easy to skip them. The problem comes later.这些流程——自行构建还是购买、是否值得长期维护——曾经有其原因。但如果你能在一个下午搭建好东西,就很容易跳过它们。问题会在后面出现。
One participant’s response: write whatever you want, but writing it doesn’t mean you’re shipping it. Code is cheap, but launching it isn’t. Keeping that distinction alive in a team is harder than it sounds when management is celebrating every PR.一位参与者的回应:想写什么就写什么,但写了不代表交付。代码便宜,发布却不便宜。当管理层庆祝每一次PR时,在团队中保持这种区分感比听起来要难得多。
The concept of a dedicated “technical health team” came up as a structural answer: engineers who ship features and engineers who delete or refactor them, treated as equally valuable work. Getting buy-in for that from a business incentivized by velocity is the actual challenge.专门的“技术健康团队”概念被提出作为结构性答案:负责交付功能的工程师和负责删除或重构功能的工程师,被视为同等有价值的工作。让业务在追求速度的激励下接受这一点才是真正的挑战。
Who should be shipping PRs?谁应该交付PR?
The most spirited part of the dinner: if AI makes it easy for anyone to write code, should product managers be shipping PRs?晚宴最激烈的部分:如果AI让任何人都能写代码,产品经理应该交付PR吗?
The room was divided. Not on the possibility (everyone agreed the tools make it technically feasible) but on whether it’s the right thing to optimize for.现场意见分歧。不是对可能性的分歧(大家都同意工具在技术上是可行的),而是对是否是正确的优化方向的分歧。
One participant pushed back hard on the narrative of non-engineers shipping to production as a win:一位参与者强烈反对把非工程师交付到生产视为胜利的叙事:
If you’re effectively making engineers into code review monitors, protecting the company from PRs they didn’t write, while the PM gets the credit for shipping, check in on your engineers’ mental health.如果你实际上把工程师变成代码审查监视器,保护公司免受他们未写的PR影响,而PM获得交付的功劳,请关注你工程师的心理健康。
Others were more pragmatic. Enabling non-engineers to contribute on smaller, lower-risk tasks reduces the feedback loop for everyone. Designers who can write a component to spec without waiting for an engineer. Product managers who can fix a copy bug without opening a Jira ticket. That has value, as long as the guardrails are right.其他人更务实。让非工程师在小型、低风险任务上贡献,能缩短所有人的反馈循环。设计师可以在不等待工程师的情况下编写符合规范的组件。产品经理可以在不打开Jira工单的情况下修复文案错误。只要护栏设置得当,这就有价值。
The self-driving car analogy landed well: an average AI-assisted non-engineer is probably better than an average unassisted one. But nobody’s comparing them to the engineers who spent years developing the expertise. The comparison only makes sense within a specific complexity range.自动驾驶汽车的类比很受欢迎:平均水平的AI辅助非工程师可能比平均水平的未辅助者更好。但没有人把他们和花了多年时间培养专业技能的工程师进行比较。比较只在特定复杂度范围内有意义。
If a feature is 80% engineering and 20% product thinking, an engineer can do it. If it’s the reverse, maybe a PM can handle it. What I think is actually happening is that we’re becoming value creators who pick up whatever slice of the work makes sense, regardless of title.如果一个功能80%是工程,20%是产品思考,工程师可以完成。如果相反,也许PM可以处理。我认为实际发生的情况是,我们正成为价值创造者,挑选任何有意义的工作片段,而不在乎头衔。
The harder structural question got no clean answer: if anyone can ship, whose job is it to hold everything together? Several people in the room said the only realistic response is radical team autonomy: small groups of three to five people who own their own decisions, with management’s job shifting to alignment rather than gatekeeping.更硬核的结构性问题没有得到明确答案:如果任何人都能交付,谁来把所有东西粘在一起?现场几位人士认为唯一现实的回应是激进的团队自治:三到五人的小团队拥有自己的决策,管理层的工作转向对齐而不是把关。
Code review is the bottleneck. AI might fix it.代码审查是瓶颈。AI可能会解决它。
One participant had been thinking about this from a CI/CD angle: the problem isn’t that teams can’t generate code, it’s that they can’t review it fast enough. Human review is now the bottleneck, and the solution isn’t more reviewers. It’s smarter triage.一位参与者从CI/CD的角度思考:问题不在于团队不能生成代码,而在于他们审查不够快。人工审查现在是瓶颈,解决方案不是更多审查者,而是更聪明的分流。
His team had been experimenting with confidence scoring on PRs: using AI to assess the risk of a change and surface only the parts that actually need human eyes. A 15,000-line PR with three lines that need human review isn’t a 15,000-line review problem. It’s a three-line problem, if you can trust the rest.他的团队正在尝试对PR进行置信度评分:使用AI评估变更风险,只展示真正需要人工关注的部分。一个有15000行代码的PR,如果只有三行需要人工审查,就不是15000行的审查问题,而是三行的问题,只要你能信任其余部分。
I honestly think you can have a fifteen thousand line PR and say, I need a human to review these three lines. Everything else here is fine. I don’t know how to do that yet, but I know it has to happen.老实说,我认为你可以有一个一万五千行的PR,然后说,我只需要人审查这三行。其他的都没问题。我还不知道怎么做到,但我知道必须实现。
The deeper point was about abstraction layers. We don’t read assembly, we trust compilers. The question is whether you can build enough validation infrastructure, feature flags, observability, acceptance tests, mutation testing, to make a similar trust relationship work with AI-generated code.更深层的点在于抽象层。我们不阅读汇编代码,我们信任编译器。问题是你能否构建足够的验证基础设施、功能标记、可观测性、验收测试、变异测试,使得对AI生成代码也能建立类似的信任关系。
Nobody in the room claimed they’d solved it. Several said they were pushing as hard as they could in that direction.现场没有人声称已经解决。几位表示他们正尽全力向那个方向推进。
One practical suggestion that came out of this: invest in evals now, not later. The cost of building AI features isn’t the hard part. The cost of verification is. If you build a solid eval suite today, you can swap providers, survive model deprecations, or move to open source without starting from scratch.一个实际建议是:现在就投资评估(eval),而不是以后。构建AI功能的成本并不是难点,验证的成本才是。如果今天构建一个坚实的评估套件,你可以换供应商、应对模型淘汰,或迁移到开源而无需从头开始。
The cost of building isn’t as high. The cost of verification is high. When the vendor you’re partnering with changes their model, which they do often, you just run that suite of evals and you’re okay.构建的成本并不高。验证的成本很高。当你合作的供应商更换模型(他们经常这么做),只要运行那套评估,你就没事。

The vendor you’re betting on你所押注的供应商
What happens when Anthropic or OpenAI raises prices, goes down, or gets disrupted?当Anthropic或OpenAI涨价、宕机或被颠覆会怎样?
One participant said it plainly: her entire company has a single point of failure on Anthropic. Engineering understands single points of failure. The people in marketing and finance building on top of Claude do not.一位参与者直言不讳:她的整家公司在Anthropic上只有一个单点故障。工程师懂单点故障,营销和财务的同事却不懂。
So many teams outside engineering are building things, and I don’t think they truly understand what’s underneath it. If all of a sudden we can’t get inference, what happens to marketing? What happens to finance? They’re gonna call engineering.很多非工程团队在构建东西,我觉得他们并不真正了解底层是什么。如果我们突然拿不到推理服务,营销会怎样?财务会怎样?他们会去找工程。
The counterpoint was that competition will keep pricing in check. Open source models are no longer years behind; they’re a couple of versions back at most. You can’t double your prices when customers have alternatives. But the lock-in concern isn’t really about the model itself. It’s about everything built around it. Skills, tooling, internal workflows — those are much harder to migrate than a model endpoint.反方观点是竞争会抑制价格。开源模型不再落后多年,最多也只落后几个版本。当客户有替代品时,你不可能把价格翻倍。但锁定风险并不在模型本身,而在围绕它构建的一切。技能、工具、内部工作流——这些比模型端点更难迁移。
The practical advice: build internal UI wrappers over generic model APIs now, before your teams are locked into specific product interfaces. It’s cheap to do, and it means you can swap the model underneath without rebuilding the interface your teams depend on.实用建议:现在就在通用模型API之上构建内部UI包装层,在团队锁定到特定产品接口之前。这样做成本低,而且可以在不重建团队依赖的界面的情况下更换底层模型。
What’s next? Let’s talk about it at Shift接下来会怎样?我们在Shift上聊聊吧
These reflections, reactions and conversations were part of just one of our Shift CTO dinners that we’re organizing as a lead up to our Shift engineering and AI conference in September. All of our dinner participants have been invited – and so are you!这些思考、反应和对话只是我们组织的众多Shift CTO晚宴之一,作为九月Shift工程与AI大会的前奏。所有晚宴参与者都已受邀——你也一样!
We’re also planning some Engineering leadership programming, but will be talking about topics just like the ones in this article – building AI native engineering teams, growing your developer career, scaling systems – across the conference agenda. I’d love for you to join us – get your tickets and see you in Zadar!我们还在策划一些工程领导力项目,但会在大会议程中讨论类似本文的话题——构建AI原生工程团队、发展开发者职业、系统扩展等。我很期待你加入——获取门票,九月在扎达尔见!
The dinner was held under Chatham House rules. Quotes are used without attribution.晚宴遵循查塔姆宫规则。引用未标明出处。


