The smartest AI model ever released will complete one agentic knowledge-work task for about $31 in LLM inference cost, and it gets the whole task right 3% of the time, which works out to roughly $1,000 for a single fully-correct result. It turns out “smartest” and “most economical” are not the same model. But for the first time we can know which model is which.有史以来最智能的 AI 模型,在 LLM 推理成本上完成一项代理知识工作任务大约需要 31 美元,并且它在 3% 的时间里能正确完成整个任务,这意味着完成一项完全正确的任务大约需要 1,000 美元。事实证明,“最智能”和“最经济”并非同一模型。但这是我们第一次能够区分它们。
The numbers come from AA-Briefcase, an agentic knowledge-work benchmark Artificial Analysis published on June 18. It is the first clean third-party measurement of what one unit of agentic work costs, on long-horizon tasks built by people who do this work at Google, McKinsey, and BCG. It arrives in the middle of an argument that has been going in circles for weeks. One side says the labs are losing billions and the whole industry runs on subsidy. The other says inference is high-margin and the labs are printing money. Both are right, because the labs have two different business models: flat-rate subscriptions, and token-based API pricing.这些数据来自 AA-Briefcase,这是 Artificial Analysis 于 6 月 18 日发布的代理知识工作基准测试。这是首次对代理工作单位成本进行第三方独立测量,测试的是由谷歌、麦肯锡和 BCG 的从业人员构建的长周期任务。它出现在一个已经持续数周的争论中。一方认为,实验室正在亏损数十亿美元,整个行业都依赖补贴。另一方则认为,推理利润丰厚,实验室正在印钞票。双方都说对了,因为实验室有两种不同的商业模式:固定费率订阅和按 token 計費的 API 定價。
Both OpenAI and Anthropic have flat-rate plans, and the brutal economics of these plans say they have to stop, because agentic workloads have broken the model. If you build agents and run them on a $200 monthly plan, you’re the problem for the labs. The question is: when the subsidies end (and it’s when, not if) where do you go?OpenAI 和 Anthropic 都提供固定费率套餐,而这些套餐残酷的经济学表明它们必须停止,因为代理工作负载已经颠覆了这种模式。如果你构建代理并在每月 200 美元的套餐上运行它们,那么你就是这些实验室的问题所在。问题是:当补贴结束时(是时候,而不是是否),你将何去何从?
Flat-rate AI subscriptions are subsidizing agentic workloads固定费率 AI 订阅正在补贴代理工作负载
Nobody is hiding the data. SemiAnalysis bought every consumer tier and ran each to its weekly limit on long-horizon coding and agentic work, and the result is now widely quoted: a fully utilized $200 ChatGPT Pro plan, valued at published API rates, would cost up to $14,000 a month, a 70x gap. Measured against true serving cost, which SemiAnalysis puts near a quarter of list price, the same maxed-out user burns roughly $3,500 in compute against $200 in revenue, a 17.5x gap. This cost to the labs is concentrated on the heaviest users, making a flat plan a cross-subsidy: the person who asks three questions a day funds the engineer running agents overnight.没有人隐藏数据。SemiAnalysis 购买了所有消费者套餐,并将其运行到每周的上限,用于长周期编码和代理工作,结果现在被广泛引用:一个充分利用的每月 200 美元的 ChatGPT Pro 套餐,按公布的 API 费率计算,每月成本高达 14,000 美元,差距为 70 倍。与 SemiAnalysis 估计的接近标价四分之一的实际服务成本相比,同一个最大用户消耗约 3,500 美元的计算成本,收入为 200 美元,差距为 17.5 倍。实验室的这些成本集中在重度用户身上,使得固定套餐成为一种交叉补贴:每天提三个问题的人资助了夜间运行代理的工程师。
What ended that cross-subsidy as a viable business model is a change in the workload. A plan priced for chat is now running build-server work. Microsoft Research found agentic coding tasks consume roughly 1,000 times the tokens of a standard query, and the frontier models on AA-Briefcase spend over 100,000 output tokens on a single task.使这种交叉补贴作为可行商业模式终结的是工作负载的变化。一个为聊天设计的套餐现在正在运行构建服务器工作。微软研究院发现,代理编码任务消耗的 token 量大约是标准查询的 1,000 倍,而 AA-Briefcase 上的前沿模型在单个任务上花费超过 100,000 个输出 token。
One way out for the labs might be if costs per token dropped fast enough to escape their burn. And per-token cost is dropping fast, on the order of 9x to 900x per year depending on the capability. But total spend climbs anyway because agents consume orders of magnitude more tokens than the chat usage the plans were priced for. Agents are not straining the subscription model, they have dismantled it.实验室的一个出路可能是,如果每个 token 的成本下降得足够快,以至于它们能够摆脱亏损。每个 token 的成本正在迅速下降,每年下降 9 倍到 900 倍,具体取决于能力。但总支出仍然在增加,因为代理消耗的 token 量比套餐定价所依据的聊天使用量要大几个数量级。代理并没有给订阅模式带来压力,它们已经摧毁了它。
Why API-priced LLM inference has different economics为什么按 API 定价的 LLM 推理具有不同的经济性
At the metered API layer the economics invert. We have some data on this too. According to leaked fiscal 2025 statements, reported by OpenAI critic Ed Zitron and independently verified by the Financial Times, OpenAI’s implied gross margins improved from 28% in 2024 to 43% in 2025 (the 43% figure includes subscription revenue, implying even better API-only margins). In March 2025, during its Open Source Week, DeepSeek disclosed a theoretical 545% cost-profit margin on its serving stack, with the caveat that real revenue runs lower because most usage is free or discounted. The 2026 echo came from the analyst scaling01, who argued that if GLM-5.2 sells at a profit at $4.40 per million output tokens while closed APIs charge multiples of that, the closed labs could run margins north of 90%.在计量 API 层,经济性发生了逆转。我们也有一些关于这方面的数据。根据 OpenAI 批评者 Ed Zitron 报道并经《金融时报》独立验证的泄露的 2025 财年报表,OpenAI 的隐含毛利率从 2024 年的 28% 提高到 2025 年的 43%(43% 的数据包括订阅收入,这意味着 API 独有的利润率更高)。2025 年 3 月,在其开源周期间,DeepSeek 公布了其服务堆栈理论上 545% 的成本利润率,但需要注意的是,实际收入较低,因为大部分使用是免费或打折的。2026 年的说法来自分析师 scaling01,他认为如果 GLM-5.2 以每百万输出 token 4.40 美元的价格销售并盈利,而闭源 API 的收费是其数倍,那么闭源实验室的利润率可能超过 90%。
So the two camps were never in conflict. Subscription economics and API economics are different businesses, and only one is subsidized. The flat plan loses money because of who uses it and how hard, not because a token is expensive to serve. The API pricing isn’t in any trouble.因此,这两个阵营从未发生过冲突。订阅经济学和 API 经济学是不同的业务,只有一种是补贴的。固定套餐亏损是因为使用它的人以及使用强度,而不是因为 token 的服务成本高昂。API 定价没有任何问题。
Why AI agent pricing is moving from subscriptions to usage-based billing为什么 AI 代理定价正从订阅转向按使用量计费
It is inevitable because the capital behind it has a visible horizon. Epoch projects that aggregate cash capex across the major hyperscalers, growing about 70% a year, overtakes their operating cash flow, growing about 23% a year, around the third quarter of 2026, the point where combined free cash flow reaches zero. Past that line the buildout runs on outside capital, which will create pressure to cut costs or raise price.这是不可避免的,因为其背后的资本有一个可见的期限。Epoch 预测,到 2026 年第三季度左右,主要超大规模厂商的总资本支出(年增长约 70%)将超过其运营现金流(年增长约 23%),届时自由现金流将为零。在此之后,建设将依赖外部资本,这将带来削减成本或提高价格的压力。
The coming repricing will not hit every provider at once. SemiAnalysis’s utilization work shows OpenAI’s top tier reaching zero gross margin near 5.7% utilization against roughly 10% for Anthropic, so OpenAI is the most exposed and the likeliest to move first. Anthropic is not waiting. Its new metered credit caps for agent usage took effect on June 14. The all-you-can-eat plan is being repriced.即将到来的重新定价不会同时影响所有提供商。SemiAnalysis 的利用率研究表明,OpenAI 的最高套餐在 5.7% 的利用率下接近零毛利率,而 Anthropic 的利用率约为 10%,因此 OpenAI 受到的影响最大,最有可能率先行动。Anthropic 并没有等待。其新的代理使用量计量信用额度上限已于 6 月 14 日生效。无限畅用套餐正在重新定价。
So we’re rapidly approaching a future when you have to pay the real cost of inference to run your agents. But as I’ve written before, you can’t do that math based on how much tokens cost. You have to do it based on what outcomes cost. And the last time I wrote about that, we didn’t have those numbers. Now we do.因此,我们正迅速接近一个未来,届时你将不得不支付运行代理的实际推理成本。但正如我之前写过的,你无法根据 token 的成本来计算。你必须根据结果的成本来计算。而上次我写这个的时候,我们还没有这些数字。现在我们有了。
How to calculate cost per successful task for AI agents如何计算 AI 代理的每次成功任务成本
Cost per successful task is the cost of running a model on a task divided by the rate at which the model completes that task correctly. It is the number that matters for agentic workloads, because a cheaper model only saves money if it still produces the outcome you need.每次成功任务成本是指运行模型完成任务的成本除以模型正确完成该任务的速率。这是代理工作负载的关键数字,因为更便宜的模型只有在仍然能产生你所需结果的情况下才能节省成本。
AA-Briefcase ran frontier models against 91 private tasks across four multi-week projects and reported the cost of each directly.AA-Briefcase 在四个为期数周的项目中对 91 个私有任务运行了前沿模型,并直接报告了每个任务的成本。
| Model | Cost per task | AA-Briefcase Elo | Cost vs GLM-5.2 |
|---|---|---|---|
| Claude Fable 5 | ~$31 | 1587 | 12.9x |
| Claude Opus 4.8 (max) | $10.40 | 1356 | 4.3x |
| GLM-5.2 (max), open weights | $2.40 | 1266 | 1x |
Cost per task varies by more than 800x across every model tested, from over $31 for the leader down to about $0.04 for a quantized DeepSeek variant that never reaches frontier quality. But the number that matters even more is the success rate. Fable 5 leads the benchmark and satisfies every rubric criterion on 3% of tasks, and on 31 of the 91 tasks no model scores above 50%. So the real cost per outcome is cost per successful task, cost divided by the rate the work is done right, and for Fable that is about $1,000 for one fully-correct result in this benchmark.在测试的所有模型中,每次任务成本的差异超过 800 倍,从领先模型的 31 美元以上到大约 0.04 美元(一个量化后的 DeepSeek 变体,但从未达到前沿质量)。但更重要的数字是成功率。Fable 5 在基准测试中领先,在 91 个任务中有 3% 的任务满足所有评分标准,并且在 31 个任务中,没有模型的得分高于 50%。因此,实际的每次结果成本是每次成功任务成本,即成本除以工作正确完成的速率,对于 Fable 来说,在这个基准测试中,一次完全正确的结果大约需要 1,000 美元。
But in the AA-Briefcase benchmark, the only model with meaningful successful outcomes was Fable. To figure out your own cost per successful task, you need to run evals on your own workloads. Arize AX helps teams trace agent runs, evaluate success, and compare model changes by outcome cost.但在 AA-Briefcase 基准测试中,唯一具有有意义的成功结果的模型是 Fable。要计算你自己的每次成功任务成本,你需要对你自己的工作负载运行评估。Arize AX 帮助团队跟踪代理运行,评估成功率,并通过结果成本比较模型变化。
Use evals to switch to cheaper AI models without losing quality使用评估在不损失质量的情况下切换到更便宜的 AI 模型
The table above shows you the decision you haven’t made yet. Artificial Analysis names open-weight GLM-5.2 and DeepSeek V4 Pro as the strongest price/performance options on the board, with GLM-5.2 landing about 90 Elo below Opus 4.8 for less than a quarter of the cost and outranking GPT-5.5 xhigh while costing less to run. On a metered bill, paying 4.3x for Opus or 12.9x for Fable 5 over GLM-5.2 is a choice you would have to justify per task. On a flat plan you never see it, so you never make it.上表显示了你尚未做出的决定。Artificial Analysis 将开源的 GLM-5.2 和 DeepSeek V4 Pro 列为当前最强的价格/性能选项,GLM-5.2 的 Elo 评分比 Opus 4.8 低约 90 分,但成本不到其四分之一,并且优于 GPT-5.5 xhigh,同时运行成本更低。在按量计费的情况下,为 Opus 支付 GLM-5.2 的 4.3 倍或为 Fable 5 支付 12.9 倍,你需要为每个任务进行辩护。在固定套餐下,你永远看不到它,所以你永远不会做出选择。
That is what flat rate plans have been buying you: the freedom to ignore price-performance tables and run whatever model scores highest. The day the meter turns on, the 4.3x and the 12.9x stop being invisible and start being actual bills, and cost per successful task will dictate your new stack.这就是固定费率套餐一直为你提供的:可以忽略价格-性能表并运行得分最高的模型的自由。当计量表启动的那一天,4.3 倍和 12.9 倍将不再是隐形的,而是实际的账单,每次成功任务成本将决定你的新技术栈。
The rational move is the one the benchmark already points to: down the capability curve to the efficient open-weight models, and onto backends you can fine-tune and self-host. Open weights keep collapsing in price, with DeepSeek V4-Pro at $0.435 in and $0.87 out per million tokens, and teams are already cutting real workloads on them, as Decagon did to take voice-agent cost down roughly 6x.理性的选择是基准测试已经指明的方向:沿着能力曲线向下,转向高效的开源模型,并转向你可以进行微调和自托管的后端。开源模型价格持续下跌,DeepSeek V4-Pro 的每百万 token 输入成本为 0.435 美元,输出成本为 0.87 美元,团队已经在上面削减实际工作负载,就像 Decagon 所做的那样,将语音代理成本降低了约 6 倍。
The hard part is not finding a cheaper model. It is switching to one without losing the success rate that justified the expensive model in the first place. That is an evaluation problem before it is a cost problem, and it is the one worth solving now, while you can still A/B a downgrade against a subsidized baseline. We wrote the playbook for exactly this move: how to ditch your frontier model for an SLM.困难之处不在于找到更便宜的模型。而在于在不损失最初证明昂贵模型合理性的成功率的情况下切换到更便宜的模型。这首先是一个评估问题,然后才是成本问题,而现在值得解决的是这个问题,因为你仍然可以对降级与补贴基线进行 A/B 测试。我们已经为这一举措制定了操作手册:如何用 SLM 替换你的前沿模型。
What to do before AI model subsidies end在 AI 模型补贴结束之前该做什么
Subsidies are ending, you are the party being subsidized, and the meter is coming. So run the expensive models while choosing the best result is still free, and build as though the meter is already on. Know your cost per successful task, keep your workloads portable enough to move down the curve, and validate the switch with evals before the bill forces it. The cheapest your agent will ever be to run is today. Spend that head start on getting ready.补贴即将结束,你就是被补贴的一方,计量表即将到来。所以,在选择最佳结果仍然免费的时候运行昂贵的模型,并像计量表已经启动一样进行构建。了解你的每次成功任务成本,保持你的工作负载足够便携,以便沿着曲线向下移动,并在账单迫使你这样做之前通过评估验证切换。你的代理运行成本永远不会比今天更低。利用这个先发优势来做好准备。