Transcript文字记录
Renato Losio: We will chat about the infrastructure challenge behind production AI. A couple of words about today's topic. AI, as we all know, has moved from experiments that we were doing probably a couple of years ago to now probably always on a system that runs whole business operations. As adoption grows, one of the biggest challenges we all face is no longer building just models, but as well how to run them in a reliable way at scale. We have seen as well in the last few months, large organizations like even GitHub, a platform that many of us rely on every day, discuss the challenges they face in scaling their capacities and the pressure they experience under the increased workload they had with AI. If a company at GitHub scale is rethinking its infrastructure, what should the rest of us be learning?Renato Losio:我们将讨论生产环境 AI 背后的基础设施挑战。先简单介绍一下今天的话题。众所周知,AI 已经从几年前的实验阶段,转变为如今几乎全天候运行、支撑整个业务运营的系统。随着采用率的增长,我们面临的最大挑战不再仅仅是构建模型,而是如何在大规模环境下可靠地运行它们。在过去几个月中,我们也看到像 GitHub 这样许多人每天依赖的平台,都在讨论其在扩展容量时面临的挑战,以及在 AI 增加工作负载下所承受的压力。如果像 GitHub 这样规模的公司都在重新思考其基础设施,我们其他人应该从中学习什么?
Today, probably AI is not just increasing the load, but it's really changing the shape of the load we experience in production.如今,AI 可能不仅是在增加负载,它实际上正在改变我们在生产环境中遇到的负载形态。
My name is Renato Losio. I'm a staff editor here at InfoQ. I'm joined today by four experts coming from very different companies, countries, sectors, backgrounds that are going to give their perspective. They're all experts in the field, and answer your questions and give their feedback on all the open topics. Before starting our discussion, I'd like just to give a chance to each one of them to give a short intro.我是 Renato Losio,InfoQ 的资深编辑。今天,来自不同公司、国家、行业和背景的四位专家将加入我的行列,分享他们的观点。他们都是该领域的专家,将回答你们的问题并对所有开放话题提供反馈。在开始讨论之前,我想请每位嘉宾做一个简短的自我介绍。
Luca Bianchi: I'm Luca Bianchi. I'm the CTIO of MESA. I'm focused on developing software for highly regulated sectors. We have products in that domain, and that change of workload is something that we are experiencing, because yesterday, our workload was pretty much databases. Now, they are pretty much tokens. Everything changed in the shape and the amount of workload.Luca Bianchi:我是 Luca Bianchi,MESA 的 CTIO。我专注于为高度监管的行业开发软件。我们在该领域拥有产品,而这种工作负载的变化正是我们正在经历的,因为过去我们的工作负载主要是数据库,而现在主要是 Token。工作负载的形态和数量都发生了变化。
Alex Infanzon: I'm Alex Infanzon. I'm a Solutions Architect with Cockroach Labs, working with the engineering teams to help them solve data infrastructure challenges. I talk every day with AI-native companies, fintechs, and platform teams to help them try to figure out the agentic applications, why they work in a prototype and are failing in production. Lately, the most interesting conversations are not about models, they are about the infrastructure underneath them.Alex Infanzon:我是 Alex Infanzon,Cockroach Labs 的解决方案架构师,与工程团队合作帮助他们解决数据基础设施挑战。我每天都与 AI 原生公司、金融科技公司和平台团队交流,帮助他们找出代理应用在原型阶段有效但在生产环境中失败的原因。最近,最有趣的对话不再是关于模型,而是关于模型底层的基础设施。
Meryem Arik: I'm Meryem. I'm one of the co-founders at Doubleword. We're an inference provider. We've been working in inference and model serving since before ChatGPT came out. Back then, people didn't think it was as important as it is now. Most recently, we've been talking and thinking a lot about tokenomics and the problem of token cost and scalability in inference as it goes to production.Meryem Arik:我是 Meryem,Doubleword 的联合创始人之一。我们是一家推理提供商。在 ChatGPT 问世之前,我们就已经在从事推理和模型服务工作了。当时,人们并不认为这像现在这样重要。最近,我们一直在深入思考 Token 经济学以及推理在进入生产环境时的 Token 成本和可扩展性问题。
Simerus Mahesh: My name is Simerus. I'm on the founding team of a company called Forge. It's a startup based in San Francisco, backed by a few folks like Greylock, Palantir, BoxGroup. Previously, I've worked at a few other companies like Meta, Google, PlayStation, where I've done a lot of infrastructure-related work, pertinent to data centers, or cloud infrastructure, and even just agentic systems. Now, I'm pretty focused on the AI security side of things, and just overall working on the infrastructure for deploying production AI agents.Simerus Mahesh:我是 Simerus,Forge 创始团队成员。这是一家位于旧金山的初创公司,得到了 Greylock、Palantir、BoxGroup 等机构的支持。此前,我曾在 Meta、Google、PlayStation 等公司工作,从事过大量与数据中心、云基础设施甚至代理系统相关的基础设施工作。现在,我非常专注于 AI 安全领域,并致力于为部署生产级 AI 代理构建基础设施。
Unexpected Infra Bottlenecks in Production AI生产环境 AI 中意想不到的基础设施瓶颈
Renato Losio: I'd like to start with a very personal curiosity. It's like, for you, all really experts in the field, what's been the most surprising infrastructure bottleneck you have seen in production AI? Because I have my own experience, but I have a limited one. Do you have any suggestion, anything that you didn't expect to see?Renato Losio:我想从一个非常个人的好奇心开始。对于你们这些真正的领域专家来说,在生产环境 AI 中见过的最令人惊讶的基础设施瓶颈是什么?因为我有自己的经验,但比较有限。你们有什么建议,或者有什么是你们没预料到会看到的吗?
Meryem Arik: Has anything happened that I didn't expect to see? I think everything that's happened has been along our prediction. We had two major predictions when we started the company four years ago. One is that open-source models would have to win for a number of reasons, like cost, performance, privacy being some of them. We had another bet that token cost would be a very big issue, and that people would scale very quickly in the amount they're spending on AI and tokens. Both of those have been proven to be true. I think the thing that has surprised us is how true that's gotten, how quickly.Meryem Arik:有什么是我没预料到会发生的吗?我认为发生的一切都在我们的预测之内。四年前我们创办公司时,有两个主要预测。一是开源模型必将胜出,原因包括成本、性能和隐私等。另一个赌注是 Token 成本将成为一个大问题,人们在 AI 和 Token 上的支出将迅速增加。这两点都被证明是正确的。我认为让我们感到惊讶的是,这一切发生得如此之快。
Jevons Paradox is one of those things that we as an industry talk about and know well, but it's one thing talking about and knowing well, and there's another thing living through it where we maybe anticipated that the token cost would go up, let's say, aggregate 10x every year. It seems like it's more like 100x for most organizations. The scale of change has been much faster than we could have possibly expected. I think we were directionally correct. I think the thing that surprised us is the pace, which makes it much more difficult to think about how you build for production. Because if I was building, let's say, an application for production a year ago, I'm probably building it for an order of magnitude less scale than it needs to be for next year. That adds a lot of challenges and complexity.杰文斯悖论(Jevons Paradox)是我们行业内经常讨论且非常了解的事情,但讨论是一回事,亲身经历则是另一回事。我们可能预料到 Token 成本每年会增长 10 倍,但对大多数组织来说,似乎更像是 100 倍。变化的规模比我们预期的要快得多。我认为我们的方向是正确的,但速度让我们感到惊讶,这使得考虑如何为生产环境构建系统变得更加困难。因为如果我一年前在为生产环境构建应用,我可能预期的规模比明年实际需要的规模要小一个数量级。这增加了许多挑战和复杂性。
Renato Losio: Have you seen anything similar? Do you have any other feedback?Renato Losio:你们见过类似的情况吗?还有其他反馈吗?
Alex Infanzon: Yes. No, definitely. The token cost is increasing, and it's a big challenge. It's breaking the budget of companies, no matter how large or small they are. They are definitely hard to manage. I'm seeing also some other things that are breaking the bank. People used to think that GPU constraints, and that was going to be a cause. What we are seeing as well is not only GPU cost, it's also the energy utilized in data centers. Now, if you want to actually have a data center, you have to plan years in advance because you have to think about locations where there's enough energy to maintain and allow these AI applications to run. That's at the infrastructure level. At the next level up, what we are seeing is that the legacy data layer is becoming very costly to maintain.Alex Infanzon:是的,绝对有。Token 成本正在增加,这是一个巨大的挑战。无论公司规模大小,它都在打破预算。它们确实很难管理。我还看到其他一些正在耗尽资金的事情。人们过去认为 GPU 限制会是一个原因。但我们看到的不仅是 GPU 成本,还有数据中心消耗的能源。现在,如果你想拥有一个数据中心,你必须提前几年规划,因为你必须考虑是否有足够的能源来维持和运行这些 AI 应用。这是在基础设施层面。在更高一层,我们看到遗留数据层变得维护成本极高。
Traditional databases are not well suited for AI high velocity and high constraint demands. With distributed SQL uses, it is one of the answers to address that. We're seeing more and more that databases are used to store all this information and to make AI applications operational. You can store the data of the usage of your agents and then make reports about how your tokens are behaving, what's the memory they are consuming, do traceability, and so forth. The cost of infrastructure is just growing up exponentially these days.传统数据库并不适合 AI 的高速度和高约束需求。分布式 SQL 是解决该问题的方法之一。我们越来越多地看到数据库被用于存储所有这些信息,并使 AI 应用具备可操作性。你可以存储代理的使用数据,然后生成关于 Token 行为、内存消耗、可追溯性等的报告。如今,基础设施成本正在呈指数级增长。
Resource Constraints资源限制
Renato Losio: Alex mentioned as well electrical power. I was thinking in general, what's the first resource you're likely to run out? Is it more really GPU, database capacity, network, engineering time?Renato Losio:Alex 也提到了电力。我通常在想,最容易耗尽的第一资源是什么?是 GPU、数据库容量、网络,还是工程时间?
Luca Bianchi: Actually, the first to run out, I think it is usually the availability of the external system. If you are using internal GPUs, so if you are provisioning GPUs either in cloud or on-prem, it doesn't matter. If you're provisioning GPUs, probably electricity and the shortage of the hardware could be something that you need to take into account. If you are going, like probably most of us, directly towards external providers, that could be a public cloud or other alternatives. One of the most challenging issues that I've been facing when shifting to AI has been the availability of the external systems that could change dramatically just in a number of hours. Let me give you an example. I have been experiencing some regular or recurring issues with databases in the past.Luca Bianchi:实际上,我认为最先耗尽的通常是外部系统的可用性。如果你使用的是内部 GPU,无论是在云端还是本地配置,都没关系。如果你在配置 GPU,电力和硬件短缺可能是你需要考虑的因素。如果你像我们大多数人一样直接转向外部提供商(可能是公有云或其他替代方案),那么在转向 AI 时我面临的最具挑战性的问题之一就是外部系统的可用性,它可能在几个小时内发生剧烈变化。举个例子,我过去在数据库方面遇到过一些常规或反复出现的问题。
You provision your database and you know that your database can scale up or down with a given latency or maybe you can be throttled, but they are quite clear figures that you can handle and you can also mitigate. On the other side, when you are sticking to an AI endpoint, say you are sending a message and you are expecting back in, say, 30 seconds to have an answer, that answer can fail, that answer can arrive in three times or four times the expected time, or that answer can arrive and be truncated. It could depend. I have been experiencing some really nasty issues when we were just moving from, say, noon to 3 p.m. The only difference was that a new region was awake. The U.S.你配置数据库时,知道它可以在给定的延迟下扩展或缩减,或者可能会被限流,但这些数字非常明确,你可以处理并缓解。另一方面,当你依赖 AI 端点时,比如你发送一条消息,期望在 30 秒内得到回复,那个回复可能会失败,可能会以预期的三到四倍时间到达,或者回复到达时被截断。这取决于具体情况。我曾经历过一些非常棘手的问题,当时我们只是从中午移动到下午 3 点。唯一的区别是一个新的区域醒了——美国。
were awake and they were starting to use the same endpoints and the same constrained resources and then we were throttled, just basically shifting a couple of hours later during the day. This is something that is quite new as a challenge when you are dealing with external AI resources.他们醒了,开始使用相同的端点和受限资源,然后我们就被限流了,仅仅是因为一天中晚了几个小时。当你处理外部 AI 资源时,这是一个相当新的挑战。
Renato Losio: What's your experience with infrastructure bottlenecks on a personal level? What do you see as the most challenging part, or what have you faced so far?Renato Losio:你个人在基础设施瓶颈方面的经验如何?你认为最具挑战性的部分是什么,或者你目前面临过什么?
Simerus Mahesh: I think tagging on to what Alex said about energy consumption and power, I think this is a very critical factor when it comes to working with infrastructure and managing data centers. I was on a data center optimization team where I directly optimized for power consumption, for example. I never found this to be the bottleneck, per se, especially for DevEx-related reasons or even production-related reasons. The main bottleneck, I think, was compute. Most companies that I've worked at, based on anecdotal experience, that have their own data centers, their shortage comes from compute, just because of the fact that it takes a lot to run workloads. It's a lot of things that are needed when doing something like this, especially at such a large scale when it comes to vertical scaling, horizontal scaling.Simerus Mahesh:接着 Alex 关于能源消耗和电力的观点,我认为在处理基础设施和管理数据中心时,这是一个非常关键的因素。我曾在数据中心优化团队工作,直接优化过功耗。我从未发现这本身是瓶颈,特别是在开发体验或生产原因方面。我认为主要的瓶颈是计算资源。根据经验,我工作过的大多数拥有自己数据中心的公司,其短缺都来自计算资源,仅仅是因为运行工作负载需要消耗大量资源。在进行此类工作时,特别是在垂直扩展和水平扩展的大规模环境下,需要考虑很多因素。
With the advent of AI and agents, you have things that sometimes keep running even after the response has completed or whatever, because it doesn't follow a regular thread-type model, but more like a process-type model. Let's say, for example, if an agent spins up a sub-agent, it can keep running even if the parent agent just dies for some reason. There are some parallels with operating systems there. Back to the point, compute is the main bottleneck that I've seen. These companies are spending a lot of effort optimizing, configuring their own Kubernetes environments to be able to manage this, handle this. It's an ever-growing system bend. It improves with time, but with AI workloads being more and more unpredictable, I have seen it increase and cause more uncertainty and breakage.随着 AI 和代理的出现,有些东西即使在响应完成后也会继续运行,因为它不遵循常规的线程模型,更像是进程模型。例如,如果一个代理启动了一个子代理,即使父代理因某种原因死亡,它也可能继续运行。这与操作系统有一些相似之处。回到重点,计算资源是我见过的主要瓶颈。这些公司花费大量精力优化和配置自己的 Kubernetes 环境来管理和处理这些问题。这是一个不断增长的系统负担。它会随着时间推移而改善,但随着 AI 工作负载变得越来越不可预测,我看到它在增加,并导致了更多的不确定性和故障。
Planning Capacity for Unpredictable Workloads为不可预测的工作负载规划容量
Renato Losio: It's hard to prepare capacity, but I wonder, how do you plan capacity for a workload, as Luca mentioned, that things can change quickly and you can hardly forecast what the API is going to do. You might have a change that, compared to the past, is not a spike anymore, it's like 10 per or whatever overnight. How can you really plan for that? How can you as well minimize the side effect?Renato Losio:准备容量很难,但我很好奇,正如 Luca 所提到的,当事情变化很快且几乎无法预测 API 会做什么时,你如何为工作负载规划容量?你可能会遇到一个变化,与过去相比不再是峰值,而是一夜之间增长 10 倍或其他什么。你如何真正为此做规划?你又如何最大限度地减少副作用?
Meryem Arik: I think it's so difficult that you should probably not try to do it unless it's your job. I've been working in open-source model inference for a while, and one of the key use cases used to be that you wanted open-source model inference so you could host it yourself and self-host it. Actually, the job of inference has just become too big for most companies to even attempt to self-host at any decent scale, because you end up having to do a full operation of capacity planning and of building an entire inference stack, which is just far more than any business wants to do. I would say that this is best done by the inference companies. It's correct, and what someone mentioned earlier, is that when you end up on multi-tenant endpoints, you do end up with this noisy neighbor effect.Meryem Arik:我认为这太难了,除非这是你的本职工作,否则你不应该尝试这样做。我从事开源模型推理已经有一段时间了,关键用例之一是人们想要开源模型推理以便自己托管。实际上,推理工作对大多数公司来说已经变得太大了,以至于无法在任何体面的规模上尝试自托管,因为你最终必须进行全面的容量规划操作并构建整个推理堆栈,这远超任何企业的意愿。我会说这最好由推理公司来完成。这是正确的,正如之前有人提到的,当你最终使用多租户端点时,确实会产生这种“吵闹邻居”效应。
You'll end up that your endpoints are a little bit too slow when the U.S. wakes up. There's a lot of infrastructure providers, inference providers specifically, who have done a very good job at capacity planning, and they will probably do a better job than you will. As was just mentioned earlier, there just isn't enough compute in the world to satisfy the demand. Even the best capacity planning in the world and the best inference provider in the world still probably doesn't have enough compute. There are still times where there are going to be noisy neighbor effects. I think we've done a relatively good job of this. We mainly serve large volume tasks and large long-running agent tasks. We don't get hit by the same noisy neighbor thing that other people do, but it's very hard.当美国醒来时,你的端点会变得有点慢。有很多基础设施提供商,特别是推理提供商,在容量规划方面做得非常好,他们可能会比你做得更好。正如刚才提到的,世界上没有足够的计算资源来满足需求。即使是世界上最好的容量规划和最好的推理提供商,可能仍然没有足够的计算资源。仍然会有出现“吵闹邻居”效应的时候。我认为我们在这方面做得相对较好。我们主要服务于大容量任务和长时间运行的代理任务。我们不会像其他人那样受到同样的“吵闹邻居”影响,但这非常困难。
I think people that try to do that capacity planning themselves and GPU planning themselves will find it very difficult, unless it's their full-time job. That would be my word of caution.我认为那些尝试自己进行容量规划和 GPU 规划的人会发现这非常困难,除非这是他们的全职工作。这是我的忠告。
Measuring the AI Cost Per Dev衡量每位开发者的 AI 成本
Renato Losio: How are companies deciding the AI cost per developer? What is the average cost that companies should budget for that? Do you have any insight about how to measure the results based on the cost involved?Renato Losio:公司是如何决定每位开发者的 AI 成本的?公司应该为此预算的平均成本是多少?关于如何根据所涉及的成本来衡量结果,你们有什么见解吗?
Alex Infanzon: It depends. It depends on what infrastructure you have procured, what are the models that you are using, what your AI application is using. Cost is very difficult for this. It's so difficult that Uber just recently published this in May that they ran out of budget for the yearly budget. They consumed it in the first four months of the year. I think the cost was something about between $200 and $500. Something about that rate per user. What triggered the consumption was the incentives that they were giving to people to use AI. They were encouraging people to use AI, and this went X number of times higher. Actually, they had games where you were in the leaderboard. If you were using more AI, you were at the top of the leaderboard and so forth.Alex Infanzon:这取决于你采购了什么基础设施,你正在使用什么模型,以及你的 AI 应用正在使用什么。成本对这一点来说非常难以衡量。它太难了,以至于 Uber 最近在五月份发布消息称,他们用完了年度预算。他们在前四个月就消耗完了。我认为成本大约是每用户 200 到 500 美元。触发消耗的是他们给予人们使用 AI 的激励措施。他们鼓励人们使用 AI,这导致使用量增加了 X 倍。实际上,他们甚至有排行榜游戏。如果你使用更多的 AI,你就会排在排行榜前列,等等。
Renato Losio: You're saying basically wrong incentives that were pushing the cost.Renato Losio:你是说基本上是错误的激励措施推高了成本。
Alex Infanzon: You have to make your engineers aware of the cost, and keep in mind that when you are using AI application agents, they are very eager to help the user to answer the question. They figure out ways. If something is not working, they do a loop and then they try again, try again, try a different route. They are creative. With this creativity, there is more token usage, more resources consumed. It can become a nightmare to try to contain the cost. It's not what is the cost right now, but how do you help your organizations, your IT organizations to be responsible and contain that cost? Because as the participant was asking you, it's a problem that is out there today. You are going to run out. Probably Meryem's team and product can help a lot with these kinds of situations.Alex Infanzon:你必须让工程师意识到成本,并记住当你使用 AI 应用代理时,它们非常渴望帮助用户回答问题。它们会想出各种办法。如果某件事行不通,它们会循环尝试,再试一次,尝试不同的路径。它们很有创造力。伴随着这种创造力,会有更多的 Token 使用和更多的资源消耗。试图控制成本可能会变成一场噩梦。现在的成本是多少并不重要,重要的是你如何帮助你的组织、你的 IT 组织负责任地控制成本?因为正如参与者问你的那样,这是一个当今存在的问题。你会耗尽预算。也许 Meryem 的团队和产品可以在这种情况下提供很大帮助。
Meryem Arik: I think the amount that teams are spending on token cost is varying wildly. I don't think there's any issue with spending a huge amount of money on tokens as long as they're productive. For example, I have very good friends at companies like NVIDIA and similar companies whose token spend for people in their team is like $15,000 a month for one team member, and they don't mind it because they're like, actually, we needed to hire for this team. It's improving our velocity. We are doing so much more than we could. These are very productive tokens, and so it's so fine spending $15,000 a month. I know other companies who are spending $200 a month per employee and are complaining. I think Uber's limit or cap was like $1,500 or whatever it is.Meryem Arik:我认为团队在 Token 成本上的支出差异很大。只要他们富有成效,我认为在 Token 上花费大量资金没有任何问题。例如,我在 NVIDIA 等公司有非常好的朋友,他们团队成员的 Token 支出约为每月 15,000 美元,他们并不介意,因为他们认为这实际上是我们团队需要招聘的成本。它提高了我们的速度。我们所做的比我们能做的多得多。这些是非常有成效的 Token,所以每月花费 15,000 美元完全没问题。我知道其他公司每位员工每月花费 200 美元还在抱怨。我认为 Uber 的限额或上限大约是 1,500 美元或其他什么。
The amount that you're spending on tokens, I think you should spend as much as you think you're getting value out of, and if you are getting value out of $50,000 a month per employee, then you should spend that amount. You need to be getting the value out of it. It's like the value capture, and that is proving difficult for some people. There is for some people also this element of token maxing as Alex was saying, like people putting things on loops and they end up doing incredibly stupid things to try and solve basic problems. I have no issue with spending huge amounts on tokens if it's useful. For us, we've spent a lot on tokens, but I think we deploy them on the whole pretty productively, and so I don't mind.我认为你应该在你认为能获得价值的范围内尽可能多地花费在 Token 上,如果你能从中获得每月每位员工 50,000 美元的价值,那么你就应该花那么多钱。你需要从中获得价值。这就像价值捕获,这对某些人来说很难。正如 Alex 所说,有些人也有这种“Token 最大化”的倾向,比如人们把事情放在循环中,最终为了解决基本问题而做一些极其愚蠢的事情。如果 Token 有用,我完全不介意在上面花费巨资。对我们来说,我们在 Token 上花费了很多,但我认为我们总体上部署得相当高效,所以我并不介意。
Alex Infanzon: That is a problem also as well of getting a measure of the ROI. To your point, if you don't get the ROI, then you're just burning money with nothing.Alex Infanzon:这也是衡量投资回报率(ROI)的问题。正如你所说,如果你没有获得 ROI,那么你只是在烧钱,一无所获。
Simerus Mahesh: We're building a governance platform to be able to essentially govern what your agents are doing, and like how much they're spending and stuff. Basically, like to answer part of the participant's question, you would need some observability type wrapper around all the agents that you're deploying, or even running within your organization. There are specific ways you can collect this. For example, for all Claude Code sessions, for all Codex sessions, this telemetry and the logs are literally getting stored on your file system. You can directly access them if you want to do a native solution. Going back to what Alex was saying, governance is a very big issue, I think nowadays. Because first of all, we don't want to let engineers do whatever, because what if they're pasting a bunch of production code, or what if they're pasting in API keys?Simerus Mahesh:我们正在构建一个治理平台,本质上是为了管理你的代理在做什么,以及它们花费了多少钱等。基本上,为了回答参与者的部分问题,你需要围绕你部署或在组织内运行的所有代理建立某种可观测性包装器。你可以通过特定方式收集这些信息。例如,对于所有 Claude Code 会话、所有 Codex 会话,这些遥测数据和日志实际上都存储在你的文件系统上。如果你想使用原生解决方案,可以直接访问它们。回到 Alex 所说的,我认为治理如今是一个非常大的问题。因为首先,我们不想让工程师随心所欲,因为如果他们粘贴了一堆生产代码,或者粘贴了 API 密钥怎么办?
We need some sort of way to govern exactly what's getting fed into these models and stuff. That is a very big layer that me and my team are actually tackling right now. It is a very big problem. An out of the box solution obviously takes a lot of time to make, which is why we're focusing on this. The more custom solutions are going to be a bit hacky, which I think you can probably implement in-house maybe, and get a naïve solution working, but something more robust, probably going to be like an out of box solution for it.我们需要某种方式来治理输入到这些模型中的内容。这是我和我的团队目前正在解决的一个非常大的层面。这是一个非常大的问题。开箱即用的解决方案显然需要很长时间才能制作出来,这就是为什么我们专注于此。更定制化的解决方案会有点“黑客”风格,我认为你可能可以在内部实现并获得一个简单的解决方案,但更稳健的方案可能还是需要开箱即用的产品。
Tracking Company Data in Production在生产环境中跟踪公司数据
Renato Losio: Actually, that brings me to the next question that is related to how to track company data in production and what are some of the challenges for not experimenting with large data in production? Do you have experience as well with some sensitive data in your sector?Renato Losio:这实际上引出了下一个问题,即如何在生产环境中跟踪公司数据,以及在生产环境中不进行大数据实验面临哪些挑战?在你们的行业中,你们是否有处理敏感数据的经验?
Luca Bianchi: It is a quite difficult question, because it is something that we know quite well, or we're supposed to know quite well, the challenges. It is quite difficult to find a solution that could fit all the different use cases. Let me explain a bit more. We know that we need to keep some data reserved and some other data secure. The degree of reservation or security that we need to provide can change based on the kind of customer, on the kind of sector that you are targeting, and also about the kind of data that you are managing. This means that sometimes you don't have one solution that could fit all. Let me give an example. I have had many customers that decided to go directly for self-hosted models, because then they can have all the data, all the data management, all the data processing within their perimeters.Luca Bianchi:这是一个相当困难的问题,因为我们非常了解,或者说应该非常了解这些挑战。很难找到一个适用于所有不同用例的解决方案。让我解释一下。我们知道我们需要保持一些数据保留,另一些数据安全。我们需要提供的保留或安全程度可能会根据客户类型、目标行业以及你管理的数据类型而变化。这意味着有时你没有一个万能的解决方案。举个例子,我有很多客户决定直接使用自托管模型,因为这样他们就可以在自己的边界内拥有所有数据、所有数据管理和所有数据处理。
The problem was that these costs and even the size and the scalability of the system was very difficult to plan beforehand, due to two factors. The first one is that the price of the underlying hardware resources are constantly changing, but the other is that also the accuracy and the models are changing as well. If you need to plan beforehand for the next six months and then say, ok, which kind of model I have to bring in production or something using a Qwen, whatever version, and I need to plan that now. I basically don't know what is going to happen within six months in this time frame. It is very difficult. To have 100% data locality, it is difficult.问题在于,由于两个因素,这些成本甚至系统的规模和可扩展性都很难提前规划。第一个因素是底层硬件资源的价格在不断变化,另一个是准确性和模型也在不断变化。如果你需要提前六个月规划,然后说,好的,我必须在生产环境中引入哪种模型,或者使用 Qwen 的什么版本,我需要现在就规划。我基本上不知道在这段时间内六个月内会发生什么。这非常困难。要实现 100% 的数据本地化是很困难的。
On the other side, an approach that I've seen being used by a lot of companies, and we are leveraging that as well, is to have different kinds of models. Some local models that could handle very sensitive data and not anonymized data, and then an anonymization layer that could strip all the most sensitive data and then leverage on frontier models in order to be able to balance the shift between data security, data locality, and on the other side, cost and forward-looking, because being able to make predictions within six months, to me, is very difficult right now.另一方面,我看到很多公司使用的一种方法(我们也在利用这种方法)是拥有不同类型的模型。一些可以处理高度敏感数据和非匿名数据的本地模型,然后是一个可以剥离所有最敏感数据的匿名化层,然后利用前沿模型,以便能够在数据安全、数据本地化与成本和前瞻性之间取得平衡,因为对我来说,现在在六个月内做出预测非常困难。
Alex Infanzon: I've been noticing for the last couple of years that the database layer is becoming the control plane these days. In the old days, the database was just the repository in the back to get the data. These days, especially in agentic AI, the database has become the control plane. Why is that? Because a lot of the transactions and a lot of the metrics that the agents are generating, the telemetry, is stored in the database, and then you can then produce reports of usage, token utilization, and so forth. Furthermore, what you are also storing is the identity of the agents. You have metadata of the agents stored in the database. Why is this important? Now you have to manage the identity.Alex Infanzon:过去几年我注意到数据库层正在成为如今的控制平面。在过去,数据库只是后台获取数据的存储库。如今,特别是在代理 AI 中,数据库已经成为了控制平面。为什么?因为代理生成的许多事务和许多指标(遥测数据)都存储在数据库中,然后你可以生成使用报告、Token 利用率等。此外,你存储的还有代理的身份。代理的元数据存储在数据库中。为什么这很重要?现在你必须管理身份。
For security reasons, you have to manage the identity and treat the agents with an identity to access what they have access to and revoke access to the agent. Also, the database, the reason it's in the control plane is now you have to track what the agent's doing to have a log so you can audit what's happening in your environment. That can be stored also in the database. In terms of locality, what Luca was mentioning, you need a database that is distributed globally and can provide to you latency times for local applications. That's one thing that with a distributed SQL you can achieve by having a database that is distributed across multiple regions in the world, but it's seen as a single logical database on top of it. That is the idea.出于安全原因,你必须管理身份,并赋予代理一个身份以访问其有权访问的内容,并撤销代理的访问权限。此外,数据库之所以在控制平面中,是因为现在你必须跟踪代理在做什么,以便有一个日志来审计环境中发生的事情。这也可以存储在数据库中。在本地化方面,正如 Luca 所提到的,你需要一个全球分布的数据库,并能为本地应用提供延迟时间。这是分布式 SQL 可以实现的一点,通过拥有一个分布在世界多个区域的数据库,但在逻辑上被视为一个单一的数据库。这就是思路。
Agents accessing data in Italy will have local latencies because they access nodes that are co-located in the region. The most important thing for agents, all this information has to be synchronized and always consistent. That's the key point. You cannot allow for eventual consistency in these types of applications. They have to be consistent. Otherwise, your agents are going to act on wrong information. Rolling back things that agents do with wrong information is going to be tremendously difficult because that triggers things, and it's out of control.在意大利访问数据的代理将具有本地延迟,因为它们访问的是位于该区域的节点。对于代理来说,最重要的事情是所有这些信息必须同步且始终保持一致。这是关键点。你不能允许这些类型的应用出现最终一致性。它们必须是一致的。否则,你的代理将基于错误的信息采取行动。回滚代理基于错误信息所做的事情将极其困难,因为它会触发一系列连锁反应,导致失控。
Rollbacks for AI ApplicationsAI 应用的回滚
Renato Losio: How easy is it actually to roll back an AI capacity cleanly? I'm thinking mostly about traditional applications, but what does it actually mean for an AI application to roll back in this sense?Renato Losio:干净地回滚 AI 容量实际上有多容易?我主要是在考虑传统应用,但在这种意义上,AI 应用回滚实际上意味着什么?
Meryem Arik: What do you mean with regards to capacity planning?Meryem Arik:关于容量规划,你指的是什么?
Renato Losio: If I think about a very simple old, traditional application, I'm thinking, I have my cluster of applications with a number of nodes growing and scaling down and the database that was scaling down capacity, whatever. I actually wonder when I think about a growing number of tokens, growing number of everything, but I think if I want to have an elastic capacity that I pay for what I use, how do you actually scale down? How do I roll back things on AI? Probably I have a completely different mental model I'm coming from. Probably I'm old enough that I'm not used to these new approaches.Renato Losio:如果我想到一个非常简单、古老的传统应用,我会想,我有我的应用集群,节点数量在增长和缩减,数据库容量也在缩减,等等。我实际上很好奇,当我想到越来越多的 Token,越来越多的东西时,如果我想拥有一个我为所用付费的弹性容量,你实际上是如何缩减的?我如何回滚 AI 上的东西?可能我来自一个完全不同的思维模型。可能我年纪大了,不习惯这些新方法。
Meryem Arik: The majority of people who are using AI in production, with a few exceptions of very regulated businesses, interact with them through third-party API providers.Meryem Arik:大多数在生产环境中使用 AI 的人(除了极少数受监管的企业外),都是通过第三方 API 提供商与它们交互的。
Renato Losio: They give away the problem.Renato Losio:他们把问题转嫁出去了。
Meryem Arik: They give away the problem. This is the problem that I have to deal with. I have to deal with the problem of like, how do I scale up and scale down instances very quickly? How do I swap models in and out for each other very quickly? For example, on my stack, let's say I've got 20 different models that I offer and a fixed GPU capacity, how can I offer all of those models in line with the usage that people want? There's a very active research community figuring out how we can do better cold starts, how we can do faster scale-ups. For most people, they've given this problem away to an inference provider, but it's a very real problem that we solve and that we think about.Meryem Arik:他们把问题转嫁出去了。这就是我必须处理的问题。我必须处理如何非常快速地扩展和缩减实例的问题。如何非常快速地为彼此交换模型的问题。例如,在我的堆栈上,假设我提供 20 种不同的模型和固定的 GPU 容量,我如何根据人们的使用需求提供所有这些模型?有一个非常活跃的研究社区正在研究我们如何做得更好,如何实现更快的冷启动,如何实现更快的扩展。对大多数人来说,他们已经把这个问题转嫁给了推理提供商,但这确实是我们解决并思考的一个非常现实的问题。
Simerus Mahesh: I think I can take a little different lens to this, actually. It obviously depends like what you mean by rollback. There are things, obviously, just like rolling back a system prompt, model version, or just even a feature flag, which can be pretty clean, just like standard rollback procedures. Rolling back an AI capability that has already taken action is a lot harder as we move to autonomous agents, coding agents that are working on production and stuff. With agents, the output isn't just text. The agent may have created files, changed configuration, opened pull requests, called cloud APIs, which are not really reversible to some degree, updated state in a database, or a queue, or whatever, or even triggered some downstream workflow on Jenkins or something like CI/CD pipeline, whatever.Simerus Mahesh:我认为我可以从一个稍微不同的角度来看待这个问题。这显然取决于你所说的“回滚”是什么意思。显然,有些事情就像回滚系统提示词、模型版本,甚至只是功能标志,这可以非常干净,就像标准的回滚程序一样。随着我们转向自主代理、在生产环境中工作的编码代理等,回滚已经采取行动的 AI 能力要困难得多。对于代理,输出不仅仅是文本。代理可能已经创建了文件、更改了配置、打开了拉取请求、调用了云 API(在某种程度上是不可逆的)、更新了数据库或队列中的状态,或者甚至触发了 Jenkins 或 CI/CD 流水线等下游工作流。
At this point, I feel like rollback becomes less about reverting a deployment and more like compensating for side effects. I think the right approach is design for rollback before launch itself, like feature flags, dry run models and approval gates, especially with these autonomous agents, coming back to the point of governance, which is a very important player for agents these days. Idempotent operations, similar to ACID compliant databases and stuff. I think the cleanest rollback is often just like preventing irreversible actions from happening automatically in the first place. I don't think there's a clear answer to this because there's so many variable workloads and different possibilities. I think prevention is probably the key thing to note here, especially with autonomous agents.在这一点上,我觉得回滚不再是关于恢复部署,而是关于补偿副作用。我认为正确的方法是在发布之前就设计好回滚,比如功能标志、试运行模型和审批门禁,特别是对于这些自主代理。回到治理这一点,这对如今的代理来说是一个非常重要的角色。幂等操作,类似于 ACID 兼容数据库等。我认为最干净的回滚通常只是防止不可逆的操作自动发生。我认为对此没有明确的答案,因为有太多的可变工作负载和不同的可能性。我认为预防可能是这里需要注意的关键点,特别是对于自主代理。
Luca Bianchi: I totally agree because it is quite difficult to roll back the work that an agent has done. That's quite strange because the history of the conversation is quite easy to be recovered and to be saved. The problem is that the reasoning part and the reason why an agent decided to use a given tool and not another, it is something that is not easily rollbackable. I can roll back the effect of an action. You have done this query to the database, I can roll back that query and do whatever I want. It is very difficult to roll back the reasoning process that led to that query or led to the usage of that tool.Luca Bianchi:我完全同意,因为回滚代理所做的工作非常困难。这很奇怪,因为对话历史记录很容易恢复和保存。问题在于推理部分,以及代理决定使用给定工具而不是另一个工具的原因,这是不容易回滚的。我可以回滚一个动作的效果。你对数据库执行了这个查询,我可以回滚该查询并做任何我想做的事。回滚导致该查询或导致使用该工具的推理过程是非常困难的。
It's a bit of uncertainty that we have in agentic systems and we need to deal with them, and to deal with them with guardrails maybe or just focusing on the effects of the actions.这是我们在代理系统中存在的不确定性,我们需要处理它们,也许通过护栏,或者仅仅关注动作的效果。
Alex Infanzon: I totally agree with you. We are working with a very large credit card provider at the moment, and the reason I brought this in is because once you have agents, multi-agents working, one agent can trigger one thing on stale data or wrong data and pass that along to another agent, and that agent is going to do its own thing and maybe trigger two more agents doing actions, downstream actions, and they are going to persist the results that they created into somewhere. All that information and all these records are going to be wrong because they were derived from a wrong assumption from the initial agent. We are working for this credit card company. What we are working on is we are partnering with two other companies, one called DBOS, and that's a product that basically helps track the workflow of the agents.Alex Infanzon:我完全同意你的观点。我们目前正在与一家非常大的信用卡提供商合作,我之所以提到这一点,是因为一旦你有了代理、多代理在工作,一个代理可能会在陈旧数据或错误数据上触发一件事,并将其传递给另一个代理,而那个代理会做它自己的事情,并可能触发另外两个代理执行下游动作,它们会将它们创建的结果持久化到某个地方。所有这些信息和所有这些记录都将是错误的,因为它们源于初始代理的错误假设。我们正在为这家信用卡公司工作。我们正在做的是与另外两家公司合作,一家叫 DBOS,这是一个本质上帮助跟踪代理工作流的产品。
The workflow is actually stored in the database and it can be rolled back and then take action. You know the exact steps that the agents took so you can trace back and fix things. The other company that we are working with is called Memori, and they are the ones responsible for storing and persisting the memory state of the agents into the database. Also, to be able to figure out how to write and provision better prompts using what the agents already learned that is stored in the database, and then things that they've done in the past, inject that into new prompts and then continue executing. That is also very important. The key thing is by leveraging vector search within the database you can actually look for similarities, things that had happened in the memory that are similar and that they are persisted in the database.工作流实际上存储在数据库中,可以回滚并采取行动。你知道代理采取的确切步骤,因此你可以追溯并修复问题。我们合作的另一家公司叫 Memori,他们负责将代理的内存状态存储和持久化到数据库中。此外,为了能够弄清楚如何使用代理已经学到的存储在数据库中的内容来编写和配置更好的提示词,以及它们过去所做的事情,将这些注入到新的提示词中,然后继续执行。这也非常重要。关键在于,通过利用数据库内的向量搜索,你可以实际寻找相似之处,即内存中发生过且持久化在数据库中的相似事物。
Renato Losio: Search is not really technically a rollback.Renato Losio:搜索在技术上并不是真正的回滚。
Real-Time Saga Pattern and Rollbacks实时 Saga 模式和回滚
Participant 1: Would or could real-time Saga pattern that builds up over the conversation help with being able to roll back?参与者 1:随着对话建立的实时 Saga 模式是否会或能够有助于实现回滚?
Alex Infanzon: The Saga pattern is good if you have a way to trace the steps of every single agent, every single action that they did, and record those updates. Then when the model fails or the agent fails, then you can revert and start tracing back and undoing things. That's the only way that you can do. In the past, all that was in the log of the database. It was easy just to go to the log and roll back a transaction. Today, it's more complicated than that because you are not rolling back a transaction. You are rolling back a series of actions that occur on different agents that have different states. That is the problem.Alex Infanzon:如果你有一种方法来跟踪每个代理的步骤、它们执行的每一个动作,并记录这些更新,那么 Saga 模式是很好的。然后当模型失败或代理失败时,你就可以恢复并开始追溯和撤销事情。这是你唯一能做的方法。过去,所有这些都在数据库的日志中。去日志回滚事务很容易。今天,情况比这复杂得多,因为你不是在回滚一个事务。你是在回滚一系列发生在具有不同状态的不同代理上的动作。这就是问题所在。
Renato Losio: As Luca mentioned, they could as well be not reversible. You might have to compensate, you might have to act as soon as you had said.Renato Losio:正如 Luca 所提到的,它们也可能是不可逆的。你可能必须进行补偿,你可能必须在你说完话后立即采取行动。
Participant 2: Actually, I was thinking about what the implication of that is. We mentioned before the privacy and we didn't talk too much about security. I was actually wondering, who is getting the worst of this? Is it more the data folks, the platform one, the on-call one, security people? Who is getting the hardest part right now with capacity? What I feel is that we say capacity is limited. We have limited resources. We might scale 10 per in a few hours or a few days, whatever. Even large companies like GitHub had challenges. What's next? Who is the one to blame? Who is having the hardest part right now?参与者 2:实际上,我一直在思考这意味着什么。我们之前提到了隐私,但没有过多谈论安全。我实际上想知道,谁在承受最糟糕的情况?是数据人员、平台人员、值班人员,还是安全人员?谁现在在容量方面承受着最艰巨的部分?我的感觉是,我们说容量是有限的。我们有有限的资源。我们可能在几个小时或几天内扩展 10 倍。即使像 GitHub 这样的大公司也面临挑战。接下来是什么?该责怪谁?谁现在承受着最艰巨的部分?
Simerus Mahesh: I've done a lot of SRE work, software engineering work, production engineering work. I think I can confidently say that it's still the SRE, like the on-call SRE people, or even the platform engineers, but mainly the on-call SRE people. The reason is that mainly it's because AI workloads are changing shape very quickly. In a normal product, traffic growth is usually tied to just user growth or just a known launch. With AI systems, the same number of users can suddenly produce so much more traffic or load onto your production systems. Just because of the fact that a change prompt or a new tool call that was added for your agent, or just even increased context can just make such a more aggressive agent a loop. It can increase your demand by tenfold for your infrastructure.Simerus Mahesh:我做过很多 SRE 工作、软件工程工作、生产工程工作。我可以自信地说,仍然是 SRE,即值班的 SRE 人员,或者是平台工程师,但主要是值班的 SRE 人员。原因主要是 AI 工作负载的形态变化非常快。在普通产品中,流量增长通常与用户增长或已知发布相关。对于 AI 系统,相同数量的用户可能会突然对你的生产系统产生多得多的流量或负载。仅仅因为提示词的改变、为代理添加了新的工具调用,或者仅仅是增加了上下文,就可以使代理变得更激进、更循环。它可以使你的基础设施需求增加十倍。
This basically just goes to show that production can break even when user traffic is normal. Typically, obviously, the first people that come to mind when trying to fix these sorts of things are the SREs. Even the request count might not be alarming, but each request is now doing a lot more work behind the scenes. Just to illustrate the point, the agent can run longer, call more tools, create more state. The on-call person isn't just dealing with more traffic, they're dealing with a workload that the cost and the behavior can just change overnight, like drastically. Thus, the lack of sleep.这基本上表明,即使在用户流量正常时,生产环境也可能崩溃。通常,在试图修复这些问题时,首先想到的人显然是 SRE。即使请求计数可能并不令人担忧,但每个请求现在在后台做了更多的工作。为了说明这一点,代理可以运行更长时间、调用更多工具、创建更多状态。值班人员不仅是在处理更多的流量,他们还在处理一种成本和行为可能在一夜之间发生剧烈变化的工作负载。因此,睡眠不足。
Alex Infanzon: I would say that the database is always guilty until proven otherwise. It's always the database which is the problem. You can see this when an engineer decides to change the embeddings of the database and the data platform team has to migrate the schema with 40 million rows or more than that. That's a huge problem. Or if the security team flags a governance gap in agent access, it's the data platform team that needs to build the audit trail. As Simerus was mentioning, the SRE team gets paged at 2 a.m. because the agent loop is hammering the database. Guess who is going to be in the call? It's the DBA. I've been a DBA for many years, so I know that you are always the person responsible until you can prove that they are not.Alex Infanzon:我会说数据库总是“有罪,直到证明无罪”。问题总是出在数据库上。当工程师决定更改数据库的嵌入,而数据平台团队必须迁移拥有 4000 万行或更多数据的模式时,你就能看到这一点。这是一个巨大的问题。或者,如果安全团队标记了代理访问中的治理差距,那么需要构建审计跟踪的就是数据平台团队。正如 Simerus 所提到的,SRE 团队在凌晨 2 点接到寻呼,因为代理循环正在猛烈攻击数据库。猜猜谁会接到电话?是 DBA。我做了很多年的 DBA,所以我知道,在你证明不是你之前,你总是那个负责任的人。
Renato Losio: Are we saying the old rule that the DBA is always the first culprit?Renato Losio:我们是在说那个老规矩吗:DBA 永远是第一个替罪羊?
Alex Infanzon: Yes, you are always a culprit. Your database is slow. Your database is not good. The data is wrong. Yes, I know. You are always guilty until proven innocent. Yes, that is the key thing. It is, as we are saying, very complex to do this. The data platform team has my genuine sympathy for the work they've done. On the other hand, I would say databases are becoming the control plane again of agentic AI, because now we need to keep this state consistent across all of that, all the agents in your applications.Alex Infanzon:是的,你总是罪魁祸首。你的数据库很慢。你的数据库不好。数据是错的。是的,我知道。在证明无罪之前,你总是被视为有罪。是的,这就是关键所在。正如我们所说,要做到这一点非常复杂。数据平台团队为他们所做的工作赢得了我真诚的同情。另一方面,我想说数据库正在重新成为代理式 AI(agentic AI)的控制平面,因为现在我们需要在所有这些应用中的所有代理之间保持状态的一致性。
Renato Losio: Do you agree? Do you see always database as the main culprit?Renato Losio:你同意吗?你是否总是认为数据库是主要的罪魁祸首?
Luca Bianchi: Yes. I think that it could be, if not the main, just one of the most important. Because in this world where we have a lot of moving parts that can be replaced quite easily with maybe some generated code, everything hits the database. Everything at the end of the day, you are going to be slapped in the face by your database, or your poorly designed database, or the performances that are not sufficient, or the data is not structured well enough to be retrieved. There is something quite different from the past. This is the fact that, in the past, we used to plan a database for humans. I can remember a very famous book from Hernandez, "Database Design for Mere Mortals". The idea was to allow programmer developers to properly design a database for the data to be retrieved and shown to humans.Luca Bianchi:是的。我认为即使不是主要的,它也是最重要的因素之一。因为在这个拥有许多易于被生成的代码替换的活动部件的世界里,一切最终都会触及数据库。归根结底,你会被你的数据库、设计糟糕的数据库、性能不足的数据库,或者结构不够好而无法检索的数据所打击。这与过去有很大的不同。事实是,在过去,我们通常为人类设计数据库。我记得 Hernandez 写过一本非常著名的书,《Database Design for Mere Mortals》(普通人的数据库设计)。其理念是让程序员开发人员能够正确地设计数据库,以便将数据检索并展示给人类。
Now the logics are changing. In the near future, we probably will have agents retrieving data from databases, maybe from a number of databases. Some constraints may be related to the fact that we cannot keep in our cognitive load too many databases or too many different databases for the same team or whatsoever, maybe are becoming less important. I definitely agree databases are where everything fails.现在逻辑正在改变。在不久的将来,我们可能会有代理从数据库中检索数据,甚至是从多个数据库中检索。一些限制可能与我们无法在认知负荷中容纳太多数据库,或者同一个团队无法处理太多不同数据库等事实有关,这些可能变得不那么重要了。我绝对同意数据库是所有问题爆发的地方。
Alex Infanzon: I have to clarify, the database is not failing. You are blaming the database. It's not failing. They're always saying, it's the database. To your point, there are so many agents, so many things accessing the database that a DBA is not capable of keeping up with the demand. That's why, at least in Cockroach Labs, we are working on making the database more intelligent. We are adding AI capabilities to our database. In real time, the AI agent can be monitoring, why is performance stalling? Why are my writes slowing down? Why is my latency going up? All those things can be done automatically by the AI inside the database. At the beginning of the 2000s, Oracle was talking about autonomous databases.Alex Infanzon:我必须澄清一下,数据库并没有崩溃。是你把责任推给了数据库。它没有崩溃。他们总是说,是数据库的问题。正如你所说,有太多的代理,太多的东西在访问数据库,以至于数据库管理员(DBA)无法跟上需求。这就是为什么至少在 Cockroach Labs,我们致力于让数据库变得更智能。我们正在为数据库添加 AI 功能。实时地,AI 代理可以监控:为什么性能停滞了?为什么我的写入变慢了?为什么我的延迟在增加?所有这些都可以由数据库内部的 AI 自动完成。在 2000 年代初,Oracle 就在谈论自治数据库。
Renato Losio: It's quite some time ago.Renato Losio:那已经是相当久以前的事了。
Alex Infanzon: We are at the realization now of that capability by putting agents inside the database. That's from the infrastructure point of view.Alex Infanzon:我们现在通过在数据库内部放置代理来实现这一功能。这是从基础设施的角度来看的。
Infrastructure - Wisdom from the Past, and the Present基础设施——来自过去与现在的智慧
Renato Losio: Actually, I have a different question in that sense. It's like, if we look back at the past, what's one piece of conventional infrastructure wisdom that used to be true and is not true anymore with AI? Do you have any feeling of what has really changed, because now we are saying, ok, the database is still blamed as the main culprit. Something else has changed in infrastructure, apart from provisioning that we know that there's not enough capacity out there. Anything else that we can say has really changed significantly?Renato Losio:实际上,我对此有一个不同的问题。如果我们回顾过去,有什么传统的架构智慧在 AI 时代不再适用了?你是否感觉到有什么真正发生了变化?因为现在我们说,好吧,数据库仍然被指责为主要罪魁祸首。除了我们知道容量不足的配置问题外,基础设施中还有什么其他东西发生了重大变化吗?
Meryem Arik: I think something that's changed pretty significantly is the unit economics of running software has changed really radically. It used to be the case that people loved investing in software businesses, because they had great margins, and they were super scalable and all of these things, because the underlying infrastructure cost was not the driver. Now that's not true. Now your infrastructure costs are so expensive, and will be a huge part of your cost base as a business going forward. The days of 70%, 80% margins, it's just not possible anymore. Infrastructure is actually a key cost driver, or maybe the key cost driver in a business. That's a new phenomenon that we're not used to seeing in software businesses.Meryem Arik:我认为发生重大变化的一点是,运行软件的单位经济效益发生了根本性的改变。过去,人们喜欢投资软件业务,因为它们有很高的利润率,而且非常易于扩展,所有这些都是因为底层基础设施成本不是驱动因素。现在情况不同了。现在你的基础设施成本非常昂贵,并且在未来将成为你业务成本基础的很大一部分。70%、80% 利润率的日子已经一去不复返了。基础设施实际上是一个关键的成本驱动因素,甚至可能是业务中最关键的成本驱动因素。这是我们在软件业务中不常见到的新现象。
Simerus Mahesh: AI agents specifically, and just like AI workloads in general, I think have made the idea of if the demand spikes, then just autoscale your way through it, like obsolete. I think that definitely worked better when the workload was mostly human driven, bounded by just normal user behavior, click patterns, database access patterns. With AI agents, now one user action can trigger a loop of model calls, tool calls, code execution even, because it has sandbox environments, retries, database reads, database writes, cloud API calls. If that loop is inefficient, or misconfigured, autoscaling doesn't actually solve the problem, it can actually amplify it by turning a product bug, or just like a small inefficiency into a huge infrastructure bill, or like even outage, if effects cascade downstream.Simerus Mahesh:特别是 AI 代理,以及一般的 AI 工作负载,我认为它们使得“如果需求激增,就通过自动扩容来解决”的想法变得过时了。我认为当工作负载主要是由人类驱动,受限于正常的行为、点击模式、数据库访问模式时,这种方法确实更有效。有了 AI 代理,现在一个用户操作可能会触发一系列模型调用、工具调用,甚至是代码执行,因为它有沙盒环境、重试、数据库读写、云 API 调用。如果这个循环效率低下或配置错误,自动扩容并不能解决问题,反而可能通过将产品错误或小规模的低效放大,导致巨额的基础设施账单,甚至在影响级联到下游时导致停机。
The new wisdom is, before you scale AI workloads, you first need to bound them, put limits around runtime execution, tool calls, retries, your multi-tenancy environment setup, your sandbox environment for security and isolation, and blast radius overall. Elasticity and cloud compute is still super useful. The cloud is genuinely a great place. I think bounded, like autonomy has to come first with these AI agent workloads.新的智慧是,在扩展 AI 工作负载之前,首先需要对其进行限制,为运行时执行、工具调用、重试、多租户环境设置、用于安全和隔离的沙盒环境以及整体爆炸半径设置边界。云计算的弹性仍然非常有用。云确实是一个好地方。我认为对于这些 AI 代理工作负载,必须首先实现有界限的自主性。
Alex Infanzon: I agree with you, just one thing that I've been thinking a lot about in recent years is, I used to think that I should do my capacity planning for my expected growth, but with agentic AI, that's really impossible. I think that the better idea is to plan for elasticity, make your architecture elastic. Otherwise, you cannot predict what is the actual workload that these agents are going to generate. That is a huge problem. I cannot predict that workload. I cannot say with this amount of database infrastructure, I can do that. I need to be able to allow my database to be elastic. Think about that elasticity. To your point, Simerus, definitely what you need to have is guardrails in the agents to constrain the uses of requirements. To me, elasticity is going to be key to allow my agentic application to grow or shrink as needed.Alex Infanzon:我同意你的观点。近年来我一直在思考的一件事是,我过去认为我应该为预期的增长进行容量规划,但对于代理式 AI 来说,这真的不可能。我认为更好的主意是为弹性进行规划,让你的架构具有弹性。否则,你无法预测这些代理将产生什么样的实际工作负载。这是一个巨大的问题。我无法预测那种工作负载。我不能说有了多少数据库基础设施,我就能做到这一点。我需要能够让我的数据库具有弹性。考虑一下那种弹性。正如你所说,Simerus,你绝对需要在代理中设置护栏来约束需求的使用。对我来说,弹性将是让我的代理式应用根据需要增长或缩减的关键。
Simerus Mahesh: I agree. Also, I do think that elasticity should come second versus like bounded, you should bound your agents first and ensure that it's secure, and it doesn't have a blast radius that's too big. Because with elasticity, like when you use AWS, you essentially have infinite compute, it's just a matter of cost in the user's perspective. What this can typically do is trigger more and more compute, depending on what you use for autoscaling. If you use Karpenter to generate more nodes, or just create more pods within your Kubernetes environments, you create more and more compute and containers running your programs for things that could be like misconfigurations, or something for behavior that you don't want. I think that's very scary, because that can rack up your AWS bills, for example, or just even lead to just slowness.Simerus Mahesh:我同意。此外,我确实认为弹性应该排在有界限之后,你应该先限制你的代理,并确保它是安全的,且不会有太大的爆炸半径。因为有了弹性,当你使用 AWS 时,从用户的角度来看,你本质上拥有无限的计算能力,这只是成本问题。这通常会导致触发越来越多的计算,具体取决于你用于自动扩容的工具。如果你使用 Karpenter 来生成更多节点,或者在 Kubernetes 环境中创建更多 Pod,你就会创建越来越多的计算和容器来运行你的程序,而这些程序可能存在配置错误或你不希望的行为。我认为这非常可怕,因为这可能会增加你的 AWS 账单,或者甚至导致系统变慢。
Renato Losio: There are basically two problems that I see people complaining about the lead one. One side is that your AI workload can trigger that and you suddenly have a huge bill. On the other side, cloud provider due to the capacity as well that Meryem mentioned before, are putting soft and hard limits in place that most of the time are not enough to address what you want to have. If you're using Bedrock, or if you're using anything accessing your model, but suddenly you don't have enough capacity there. You have the two problems at the same time, people that complain that we don't have enough capacity, on the other side people that don't put enough constraint on the infrastructure that they actually provision.Renato Losio:我看到人们抱怨的主要有两个问题。一方面是你的 AI 工作负载可能会触发这种情况,导致你突然收到巨额账单。另一方面,正如 Meryem 之前提到的,云提供商由于容量限制,正在实施软限制和硬限制,这些限制大多数时候不足以解决你想要实现的目标。如果你正在使用 Bedrock,或者使用任何访问模型的服务,但突然发现容量不足。你同时面临这两个问题:人们抱怨容量不足,而另一方面,人们没有在他们实际配置的基础设施上设置足够的约束。
Alex Infanzon: Yes, database vendors like CockroachDB, what we're trying to do to minimize is separate compute from storage, that gives you the elasticity of compute. If you need more compute resources, you add more compute, or if you need more storage, you add more storage, try to segregate those two, so you can better manage your infrastructure.Alex Infanzon:是的,像 CockroachDB 这样的数据库供应商,我们正在努力做的是将计算与存储分离,这为你提供了计算的弹性。如果你需要更多的计算资源,你就增加计算;如果你需要更多的存储,你就增加存储,尽量将两者分开,这样你就能更好地管理你的基础设施。
Actionable Insights可操作的见解
Renato Losio: I'm an attendee in this roundtable, I enjoyed and you convinced me of the importance of the topic. I'd like to say, what can I do as a practitioner, what should be an action item that I can take care of. It can be reading a book, can be reading an article, can be provision an instance, can be set some guardrail? What's the advice you can give to someone that they can implement tomorrow?Renato Losio:我是这次圆桌会议的参与者,我很享受这次讨论,你们说服了我这个话题的重要性。我想问,作为一名从业者,我能做些什么?我应该采取什么行动?可以是读一本书、读一篇文章、配置一个实例,或者设置一些护栏?你们能给明天就能实施的人什么建议?
Luca Bianchi: Starting from tomorrow, reconsider your architecture in light of the adoption of AI and how it impacts the infrastructure. I think that architecture is the first and foremost point that you should consider starting from tomorrow. I think some evergreen books such as "Evolutionary Architectures" from Neal Ford, are something that are really worth reading, because they can tell you some tenets, some principles that you could apply even in this changing world. Expect things to change. Expect new databases with shared compute to be released pretty soon. Expect new agents. Expect new models. The question is, how do you stay in business with that? The only answer that I could find is building an architecture that is able to evolve, that is able to adapt and use this new advancement as soon as they become available.Luca Bianchi:从明天开始,根据 AI 的采用及其对基础设施的影响,重新审视你的架构。我认为架构是你从明天开始应该考虑的首要问题。我认为一些常青书籍,例如 Neal Ford 的《Evolutionary Architectures》(演进式架构),非常值得一读,因为它们可以告诉你一些即使在这个不断变化的世界中也可以应用的原则。预料到事情会发生变化。预料到很快就会发布带有共享计算的新数据库。预料到会有新的代理。预料到会有新的模型。问题是,你如何在这种情况下保持业务?我能找到的唯一答案是构建一个能够演进、能够适应并利用这些新进展的架构。
Meryem Arik: The biggest piece of advice that I would give is for anyone who hasn't tried open-source models in, let's say, 6 months or 12 months, try the latest open-source models. They have made huge capability improvements over the last months, like GLM 5.2. They are incredibly good, much cheaper. The latency is really good depending on the inference provider. Open-source models, you should definitely try them and have them.Meryem Arik:我能给出的最大建议是,对于任何在过去 6 到 12 个月内没有尝试过开源模型的人,去尝试一下最新的开源模型。它们在过去几个月里取得了巨大的能力提升,比如 GLM 5.2。它们非常出色,而且便宜得多。根据推理提供商的不同,延迟也非常理想。你应该一定要尝试并使用开源模型。
Simerus Mahesh: Going off of what Luca was saying, but not entirely. I don't think you should start by redesigning your entire architecture. That's something that is constantly evolving and something you can't really do in one sitting, because migration is a very big thing. I think I would pick one production AI workflow, preferably a critical one to either developers internally, or even just production itself. Then trace what actually happens when a user or a developer or whatever triggers it. Map the full path, like model calls, tool calls, database queries. Then just basically ask yourself questions regarding this path related to like, what's the maximum runtime? What is the maximum number of tool calls? What happens if a dependency slows down? What happens if a model retry goes wrong or whatever? I think the goal is to find the places where the system is unbounded.Simerus Mahesh:接着 Luca 的话,但不完全相同。我不认为你应该从重新设计整个架构开始。那是一个不断演进的东西,你无法一次性完成,因为迁移是一件大事。我认为我会选择一个生产环境中的 AI 工作流程,最好是对内部开发人员或生产本身至关重要的一个。然后追踪当用户、开发人员或任何触发它的人操作时实际发生了什么。映射完整的路径,比如模型调用、工具调用、数据库查询。然后基本上问自己关于这条路径的问题,比如:最大运行时是多少?最大工具调用次数是多少?如果依赖项变慢了会发生什么?如果模型重试出错会发生什么?我认为目标是找到系统不受限制的地方。
You obviously may not fix everything, but I think you can identify usually one or two concrete things that you can then evolve upon and base the start of your agentic system or your AI system change from. I think that's a good first step.你显然可能无法修复所有问题,但我认为你通常可以确定一两件具体的事情,然后在此基础上进行演进,并以此作为你的代理系统或 AI 系统变革的起点。我认为这是一个很好的第一步。
Alex Infanzon: For me, it's basically be more rigorous in evaluating your data infrastructure. Ask your vendor to show you how the data infrastructure works under failure, not under a predefined load. The benchmarks, as of today, it's TPC-C, they are very constrained. They work perfectly on a system that is up and running 100% of the time, but your vendor should show you now what happens if the network partitions, or what happens if a node goes down, or what happens if somebody asks me to change my schema. Can I change a schema online? What happens if a whole region fails? What happens if I'm updating my database software?Alex Infanzon:对我来说,基本上是在评估数据基础设施时要更加严格。要求你的供应商向你展示数据基础设施在故障下的工作方式,而不是在预定义的负载下。目前的基准测试(如 TPC-C)非常受限。它们在一个 100% 时间运行的系统上工作得非常完美,但你的供应商现在应该向你展示:如果网络分区了会发生什么?如果一个节点宕机了会发生什么?或者如果有人要求我更改模式会发生什么?我可以在线更改模式吗?如果整个区域发生故障会发生什么?如果我正在更新数据库软件会发生什么?
All this, you need to ask your vendor to show how the benchmark should be showing you what happens when you are under pressure, because AI agents are going to stress your infrastructure, and things are going to happen, we know that.所有这些,你需要要求你的供应商展示基准测试应该如何向你展示当你处于压力下时会发生什么,因为 AI 代理会给你的基础设施带来压力,我们知道事情总会发生。
See more presentations with transcripts查看更多带有文字记录的演示
/presentations/ai-infrastructure-scaling-architecture/en/slides/slideinfoqlivejune-1782274088149.jpg)


/sponsorship/rsc/dbfb37c5-a393-4740-85e7-8497be3ac7c0/cover/Ad_InfoQ_Q2_The-Rise-of-Client-Side-Risk_750x960-1775219700934.jpg)
/sponsorship/rsc/52318350-f698-4619-9bd5-4d0e77e65a52/cover/HarnessRSC8Deploment-1776926578638.jpg)
/sponsorship/rsc/4af8f505-0249-48af-85d0-9fe26f15ad12/cover/AkkaGartner-1752831577847.jpg)
/sponsorship/rsc/87743dd9-e2a2-4ef6-b20e-d99f70eadc42/cover/PayaraValueAdding-1764918816856.jpg)
/sponsorship/rsc/b9061931-c59f-4442-b6fd-53f609e10fa5/cover/AblyStateful-1774946229111.jpg)
/sponsorship/rsc/78c3e50a-f549-47ef-a4b0-a85dfc7bab76/cover/HarnessRSCAINative-1776927376871.jpg)
/sponsorship/rsc/6f649b0b-d769-4f7e-8531-b6b817864fee/cover/HarnessAISoftwareEng-1779173260344.jpg)
/filters:no_upscale()/sponsorship/topic/d60483d5-683c-43bc-a560-7a70de67bd33/HarnessWebinarJuly16-RSB-1780677692800.png)