The Rise of Intelligence Ownership

What is the common denominator among companies succeeding with AI-first strategies?

Justinas Zaliaduonis · Joris Zilinskis · Fabian Hildesheim · Joel Hainzl · Gediminas Pazera
July 27, 2026

Cost vs. quality on catalog integrity目录完整性中的成本与质量权衡

Scatter plot of cost per 1,000 listings versus quality: frontier models cost $19 to $172 per thousand at 70 to 76 percent quality; the fine-tuned 9B reaches 87 percent at about 50 cents per thousand.
The article in one picture: on the same catalog-review workflow, with the same tools, images, and scorer, our GRPO fine-tune of a 9B open-source model (pink) beats every frontier configuration we tested, at $0.50 per 1,000 listings: 40× cheaper than the least expensive frontier setup and ~340× cheaper than the most expensive. 一图读懂本文:在相同的目录审查工作流中,使用相同的工具、图像和评分器,我们通过 GRPO 微调的 9B 开源模型(粉色)击败了我们测试过的所有前沿模型配置,且成本仅为每 1,000 条目 0.50 美元:比最便宜的前沿方案便宜 40 倍,比最昂贵的方案便宜约 340 倍。

Part I第一部分

Everyone asked the same question每个人都在问同一个问题

Since ChatGPT launched in 2022, business leaders have been asking the same question: what can AI do for us? The answer began with low-risk tasks: summarizing documents, drafting emails, producing first drafts that a human would edit. 自 2022 年 ChatGPT 发布以来,企业领导者们一直在问同一个问题:AI 能为我们做什么?最初的答案集中在低风险任务上:总结文档、起草电子邮件、生成由人工润色的初稿。

It quickly moved into higher-value cognitive work, such as software development and content generation, and grew into more ambitious projects, like attempts to build an AI company brain, a system connected to internal knowledge, data, and tools that could coordinate work and eventually operate parts of the business autonomously. AI 很快进入了更高价值的认知工作领域,例如软件开发和内容生成,并发展出更宏大的项目,比如尝试构建“企业 AI 大脑”——一个连接内部知识、数据和工具的系统,旨在协调工作并最终自主运营部分业务。

While a lot of time, energy and tokens have been invested in AI adoption, measurable outcomes have barely been achieved at scale. However, some companies embraced being AI-first and saw enormous gains in productivity, revenue, and cost, while others lagged behind or failed to change their organizations enough to reach high ROI. 尽管在 AI 采用上投入了大量的时间、精力和 Token,但几乎没有取得可规模化的显著成效。然而,一些公司积极拥抱“AI 优先”战略,在生产力、营收和成本方面获得了巨大收益;而另一些公司则落后了,或者未能进行足够的组织变革以实现高投资回报率(ROI)。

Recent data from corporate expense management platform Ramp reveals a stark contrast in performance: the top quartile of companies investing in AI saw their revenue more than double between November 2022 and December 2025, while businesses with zero AI expenditure experienced a mere 15% increase. 企业费用管理平台 Ramp 的最新数据显示了绩效上的鲜明对比:在 2022 年 11 月至 2025 年 12 月期间,AI 支出排名前四分之一的公司营收翻了一番以上,而零 AI 支出的企业仅增长了 15%。

Top AI adopters outperform across industries顶尖 AI 采用者在各行业中表现卓越

top quartile of AI spendersAI 支出排名前 25% 的公司 businesses with zero AI spend零 AI 支出的企业
100125150175200Nov 202220232024Dec 2025AI-heavy: revenue ×2+zero AI spend: +15%
Same three years, same economy: in Ramp's data across its customer base, the heaviest AI adopters more than doubled revenue while businesses spending nothing on AI grew about 15% (indexed, Nov 2022 = 100; curve drawn from the reported endpoints). The rest of this article is about what the AI-heavy group actually did. 同样的三年,同样的经济环境:根据 Ramp 客户群的数据,AI 采用率最高的公司营收翻了一番以上,而零 AI 支出的企业增长约 15%(以 2022 年 11 月为 100 点基准;曲线根据报告的终点绘制)。本文其余部分将探讨这些高 AI 投入群体具体做了什么。

There are many reasons why AI has done wonders for some companies while others have struggled to see the return on their investment, but research primarily points in five directions. AI 为一些公司创造了奇迹,而另一些公司却难以看到投资回报,原因有很多,但研究主要指向五个方向。

01

Redesign the process, not just the task重塑流程,而不只是任务

Becoming AI-first means rethinking how the work is structured, not dropping a model into a workflow built around people: what gets approved, who reviews what, and which handoffs still need a human. Where the process stays untouched, legacy bottlenecks absorb the productivity gains before they reach the P&L. In McKinsey's 2025 survey of organizations using gen AI, workflow redesign was the attribute most correlated with EBIT impact, and only 21% of them had redesigned any workflow at all.成为“AI 优先”意味着重新思考工作结构,而不是将模型直接嵌入现有的以人为中心的工作流中:需要明确哪些环节需要审批、谁来审核什么、以及哪些交接点仍然需要人工介入。如果流程保持不变,遗留的瓶颈会在生产力提升反映到损益表(P&L)之前就将其消耗殆尽。在麦肯锡 2025 年针对生成式 AI 使用组织的调查中,流程重塑是与 EBIT(息税前利润)影响关联度最高的因素,但仅有 21% 的组织进行了流程重塑。

02

Incentivize experimentation激励试验

Models, tooling and best practices change weekly, so last quarter's setup is rarely still the right one. That only gets picked up if people are rewarded for trying things and reporting what failed, not just for shipping. Technical teams are the natural place to start, since they see the same problems recur across functions and can tell which of them a model can actually take over.模型、工具和最佳实践每周都在变化,因此上个季度的配置很少依然适用。只有当员工因尝试新事物并汇报失败经验(而不仅仅是因交付成果)而受到奖励时,这种创新才能持续。技术团队是开展试验的最佳起点,因为他们能发现跨职能部门重复出现的问题,并能判断哪些问题是模型真正可以接管的。

03

Provide tailored business context提供量身定制的业务背景

Prompt engineering and retrieval can inject business context at call time, but doing it well is its own engineering program: getting to the data, enforcing access controls on what each request may see, building retrieval that surfaces the right evidence, and managing a context window that models use unevenly as it grows.提示词工程(Prompt Engineering)和检索技术可以在调用时注入业务背景,但要做好这一点本身就是一项工程挑战:包括获取数据、对每个请求实施访问控制、构建能呈现正确证据的检索系统,以及管理模型在上下文窗口增长时使用效率不均衡的问题。

04

Measure usage and impact衡量使用情况与影响

Every AI line item eventually meets the CFO question: what did this change, and was it worth it? In most deployments, nobody can answer it: there is no infrastructure to track the model's performance, decision costs, or impact on efficiency, and self-reported time savings are often inaccurate. Without a scored evaluation on your own data, a "vibe evaluation" is the ceiling of what you can claim, and a hard budget to defend.每一项 AI 开支最终都会面临 CFO 的质询:它带来了什么改变?是否值得?在大多数部署中,没人能回答这个问题:缺乏基础设施来追踪模型性能、决策成本或对效率的影响,而自报的时间节省往往不准确。如果没有基于自身数据的评分评估,所谓的“感觉评估”就是你所能声称的上限,且很难在预算中辩护。

05

Set clear business goals within the AI budget在 AI 预算内设定明确的业务目标

AI brought a pricing model most companies were not used to. Paying per token instead of per seat makes costs scale with usage, which makes it hard to lay out a cost plan or estimate the capital efficiency gains for internal workloads. Uber went through its annual engineering budget in four months, and Microsoft cancelled most of its Claude licenses to bring costs back under control. Today's prices also understate the problem, since most AI labs are subsidizing token costs to capture market share, and frontier model prices are expected to rise.AI 带来了一种大多数公司不习惯的定价模式。按 Token 而非按席位付费使得成本随使用量扩展,这导致很难规划成本或评估内部工作负载的资本效率提升。Uber 在四个月内就花光了年度工程预算,微软取消了大部分 Claude 许可证以控制成本。目前的价格还掩盖了问题,因为大多数 AI 实验室都在通过补贴 Token 成本来抢占市场份额,而前沿模型的价格预计将会上涨。

In this article we give a detailed overview of the deployment technique the winning group keeps converging on: fine-tuning open-source models with reinforcement learning. We cover how it addresses the last three challenges above, and how it turns knowledge only your organization has (namely data, tools, and processes) into a model no vendor API can match at a fraction of the cost. 在本文中,我们将详细介绍获胜群体趋向采用的部署技术:通过强化学习对开源模型进行微调。我们将涵盖它如何解决上述后三个挑战,以及如何将仅贵司拥有的知识(即数据、工具和流程)转化为任何供应商 API 都无法比拟的模型,同时只需极低的成本。

TL;DR摘要

2.2×2.2 倍

Revenue growth of the top quartile of AI spenders between November 2022 and December 2025 in Ramp's data. Companies with zero AI spend grew about 15% over the same three years, in the same economy: the heavy adopters grew eight times as much.根据 Ramp 的数据,2022 年 11 月至 2025 年 12 月期间,AI 支出排名前四分之一的公司营收增长。在同样的三年和经济环境下,零 AI 支出的公司增长约 15%:高采用率公司的增长是其八倍。

1

Playbook the winners converge on: an open-source model, proprietary task data, and reinforcement learning against a scored copy of the workflow. Bridgewater's trained model makes ~30% fewer mistakes than the best frontier model, Harvey's legal agent beats GPT-5.5 and Claude Opus 4.8 on its own rubrics, and Intercom's Fin Apex resolves more support issues at lower cost.获胜者趋向采用的策略:开源模型、专有任务数据,以及针对工作流评分副本的强化学习。Bridgewater 的训练模型比最好的前沿模型减少了约 30% 的错误,Harvey 的法律代理在自有指标上击败了 GPT-5.5 和 Claude Opus 4.8,而 Intercom 的 Fin Apex 以更低的成本解决了更多的支持问题。

87.3%87.3%

Share of the maximum achievable score our GRPO-trained 9B open-source model reached on catalog review, vs 76.9% for the best frontier configuration: a 13.5% relative improvement over the frontier, and 36% over its own untrained base (64.2%). The five frontier models, even with optimized prompts, plateaued within a tenth of a point of each other; the trained specialist cleared that ceiling.我们经 GRPO 训练的 9B 开源模型在目录审查中达到的最高可实现分数占比,而最好的前沿配置为 76.9%:相对于前沿模型提升了 13.5%,相对于其自身未经训练的基础模型(64.2%)提升了 36%。五款前沿模型即使经过提示词优化,表现也都在 0.1 分的范围内停滞不前;而经过训练的专业模型则突破了这一上限。

68×68 倍

Cost advantage per reviewed listing: $0.50 per 1,000 with the specialist vs $34 with the strongest frontier model, and still 40× cheaper than the least expensive frontier option. At roughly 40 million decisions a day, that is about $7M a year instead of $500M, a 98% cost reduction.单次审查条目的成本优势:专业模型为每 1,000 条目 0.50 美元,而最强的前沿模型为 34 美元,且仍比最便宜的前沿选项便宜 40 倍。按每天约 4,000 万次决策计算,这相当于每年约 700 万美元,而非 5 亿美元,实现了 98% 的成本削减。

Part II第二部分

What the winners do differently获胜者有何不同之处

Most of the companies pulling ahead in the AI race made the same discovery: owning your intelligence wins on both performance and cost. A model trained to complete your specific workflows in your specific environment is very likely to outperform a general-purpose model that has never seen inside your company. Additionally, since you do not need to pack as many general-purpose capabilities into a model that is meant to operate in a specific environment, you can often get away with a smaller model that is orders of magnitude cheaper to run. 大多数在 AI 竞赛中领先的公司都有一个相同的发现:拥有自己的智能在性能和成本上都能获胜。一个经过训练以完成特定环境下的特定工作流的模型,极有可能胜过从未深入了解过贵公司的通用模型。此外,由于不需要在旨在特定环境下运行的模型中塞入过多的通用能力,你通常可以使用一个规模更小、运行成本低几个数量级的模型。

Owning your intelligence does not mean cancelling the ChatGPT or Claude subscription. Most workflow automation still starts with frontier models, and that is the right first move: it establishes a baseline for what is technically possible, and every call generates the data (inputs, decisions, corrections) that a specialist model later trains on. Once the automation leaves the prototyping stage, the priority flips to cost and performance at volume, and that is where fine-tuning open-source models with reinforcement learning comes in. In addition to that, your Fable 5 or ChatGPT model can call the specialist model to handle the parts of the workflow that require your internal knowledge, and the specialist model can call the frontier model for tasks that require high general ability. 拥有自己的智能并不意味着取消 ChatGPT 或 Claude 的订阅。大多数工作流自动化仍然始于前沿模型,这是正确的初步举措:它为技术可行性建立了基准,并且每次调用都会生成数据(输入、决策、修正),供后续的专业模型训练使用。一旦自动化脱离原型阶段,优先级就会转向规模化后的成本和性能,这时就是通过强化学习微调开源模型发挥作用的时候了。此外,你的 Fable 5 或 ChatGPT 模型可以调用专业模型来处理需要内部知识的流程环节,而专业模型也可以在需要高通用能力的任务时调用前沿模型。

How a model learns to operate in your environment模型如何学习在你的环境中运作

Diagram: a task goes to an AI model running on your infrastructure; the model interacts with tools and data, a rubric scores the outcome, and the reward signal updates the model.
The model interacts with your tools and your data, a rubric scores the outcome, and the reward signal updates the model. Repeated across tens of thousands of tasks, the workflow's judgment settles into the weights, while changing facts (prices, inventory, policy text) stay in tools where they belong. 模型与你的工具和数据交互,评分系统对结果进行评估,奖励信号随后更新模型。通过数万次任务的重复,工作流的判断逻辑会固化到模型权重中,而不断变化的事实(价格、库存、政策文本)则保留在它们所属的工具中。

Over the past two years this has hardened into a playbook: an open-source model, proprietary task data, and a reinforcement-learning stage against a scored version of the workflow. Below, we discuss three scenarios where this approach has been applied to real-world tasks. 在过去两年中,这已演变成一套成熟策略:开源模型、专有任务数据,以及针对工作流评分版本的强化学习阶段。下面,我们讨论三个将此方法应用于实际任务的场景。

Bridgewater Associates is one of the largest hedge funds in the world. Its analysts sift a constant stream of articles, filings, and emails, judging which documents are relevant to the firm's investment thesis and where boilerplate content begins. The catch is that relevant means relevant by Bridgewater's internal judgment, and no amount of prompting got frontier models to absorb that judgment reliably. So, the company decided to train an open-source model on labels from its own expert investors. The trained model makes roughly 30% fewer mistakes than the best frontier model, at a fraction of the inference cost. Bridgewater Associates 是全球最大的对冲基金之一。其分析师需要筛选源源不断的文章、文件和电子邮件,判断哪些文档与公司的投资论点相关,以及模板化内容从何处开始。难点在于“相关”是指符合 Bridgewater 内部的判断标准,而无论如何编写提示词,前沿模型都无法可靠地吸收这种判断。因此,该公司决定利用其专家投资者的标签来训练一个开源模型。该训练模型比最好的前沿模型减少了约 30% 的错误,且推理成本仅为其一小部分。

Harvey builds AI agents for law firms. Its hardest workloads are long-horizon: transaction due diligence and legal memo drafting, where the agent navigates large document sets, errors compound across steps, and even the best frontier models at maximum reasoning effort kept falling short of the quality bar. Harvey ran reinforcement learning on an open-weight model and got a legal agent that outperforms both GPT-5.5 and Claude Opus 4.8 on its rubrics. Harvey 为律师事务所构建 AI 代理。其最艰巨的工作负载是长期任务:交易尽职调查和法律备忘录起草。在这种任务中,代理需要处理庞大的文档集,错误会在步骤间累积,即使是最好的前沿模型在最大推理努力下也无法达到质量标准。Harvey 对一个开源权重模型进行了强化学习,最终获得了一个在自有指标上优于 GPT-5.5 和 Claude Opus 4.8 的法律代理。

Intercom is a customer-service platform whose AI agent, Fin, resolves almost two million customer issues a week. At that volume the problem is unit economics: frontier per-call pricing adds up fast, and every point of resolution rate matters. So Intercom's AI group post-trained its own vertical support model, Fin Apex, on billions of customer-service interactions. Intercom reports that it resolves more issues than the best frontier models while being cheaper to run. Intercom 是一家客户服务平台,其 AI 代理 Fin 每周解决近 200 万个客户问题。在如此大的规模下,问题在于单位经济效益:前沿模型的单次调用定价累积极快,且每一个解决率的百分点都至关重要。因此,Intercom 的 AI 团队在数十亿次客户服务交互的基础上,对其垂直支持模型 Fin Apex 进行了后训练。Intercom 报告称,该模型解决的问题比最好的前沿模型更多,且运行成本更低。

The same shape repeats well beyond these three. The appendix collects eight more deployments, with what each model was trained to do and what changed once it shipped. 同样的模式在上述三家公司之外也反复出现。附录收集了另外八个部署案例,展示了每个模型被训练执行的任务以及上线后的变化。

One deployment pattern repeats across these cases: reinforcement learning pushes an open-source model past the frontier on a specific set of workflows, at a fixed and dramatically lower cost per call. The deployment typically starts by using prompt and context engineered frontier models to establish the strongest default baseline. Those frontier traces and learnings are then reused to pack the focused capability into a compact model the company owns. 这些案例中重复出现一种部署模式:强化学习使开源模型在特定工作流集上超越了前沿模型,且单次调用成本固定且大幅降低。部署通常始于使用提示词和上下文工程化的前沿模型来建立最强的默认基准。随后,这些前沿模型的轨迹和经验被复用,将核心能力封装进公司自有的紧凑模型中。

Part III第三部分

Case Study: Catalogue Integrity Agent案例研究:目录完整性代理

One of the central challenges of running an e-commerce platform is keeping the product catalog trustworthy. Every listing must land in the correct category of the platform's taxonomy, and the attributes that power search, filters, recommendations, and downstream operations must be accurately extracted from its images and description. 运营电子商务平台的核心挑战之一是保持产品目录的可信度。每个商品条目必须落在平台分类体系的正确类别中,且驱动搜索、筛选、推荐和下游运营的属性必须从其图像和描述中准确提取。

+15%+15%our fine-tuned 9B open-source model reviews e-commerce product listings more accurately than the best frontier model we tested我们微调后的 9B 开源模型在审查电子商务产品条目时的准确性优于我们测试过的所有前沿模型
68×cheaper per listing when our fine-tuned model does the reviewing instead of that frontier model当使用我们微调的模型而非前沿模型进行审查时,每个条目的成本更低

Platforms staff this work with teams of catalog-integrity analysts, which grow together with the catalog: more listings mean more categories to know, more attributes to check, and more ambiguous edge cases to judge. Inconsistent decisions propagate: products become harder to find, recommendations deteriorate, and policy violations slip through. Miss a counterfeit and you expose customers and brands to fraud; over-flag and you build an expensive review queue that frustrates legitimate sellers. 平台通常配备目录完整性分析师团队来处理这项工作,团队规模随目录增长:更多的条目意味着需要了解更多的类别、检查更多的属性,以及判断更多模糊的边缘案例。不一致的决策会导致恶性循环:产品变得难以搜索,推荐质量下降,政策违规行为漏网。漏掉一个假冒商品会使客户和品牌面临欺诈风险;过度标记则会建立昂贵的审查队列,令合规卖家感到沮丧。

Let's start with the scale. eBay alone carries about 2.5 billion live listings, and Shopify's catalog absorbs more than 10 million product updates a day. Walmart has said that doing its AI-assisted catalog work with people alone would have taken roughly 100× the headcount. 让我们从规模谈起。仅 eBay 就拥有约 25 亿个在售条目,Shopify 的目录每天吸收超过 1,000 万次产品更新。沃尔玛曾表示,如果仅靠人工进行 AI 辅助的目录工作,大约需要 100 倍的人力。

Getting it wrong costs revenue: 71% of shoppers say they have returned a product because it didn't match its listing. But the reviewing itself is also expensive. A mid-size marketplace generates on the order of 10 million listing creates and edits a day; reviewed with a frontier model, that workload costs roughly half a billion dollars a year, while a fine-tuned specialist does the same job for around $10M: 出错会损失营收:71% 的购物者表示曾因产品与描述不符而退货。但审查本身也很昂贵。中型市场每天产生约 1,000 万次条目创建和编辑;如果使用前沿模型审查,该工作负载每年成本约为 5 亿美元,而微调后的专业模型完成同样的工作仅需约 1,000 万美元:

One workload, two price tags同一工作负载,两种价格标签

frontier API per call前沿 API 单次调用 fine-tuned specialist, self-hosted微调后的专业模型,自托管
$0M$100M$200M$300M$400M$500MFrontier API, ~30M calls/day≈ $500M / yrFine-tuned specialist≈ $10M / yr
The ~50× gap between the two bars is the entire value proposition made concrete: at ingestion scale, the frontier-per-call architecture does not merely cost more, it fails the budget outright. This is why the companies running this exact workflow do not run it on rented models. 两个柱状图之间约 50 倍的差距是整个价值主张的具体体现:在摄入规模下,前沿模型按次调用的架构不仅成本更高,而且根本无法通过预算审核。这就是为什么运行此工作流的公司不会使用租用的模型。

The companies actually running this workflow confirm the math. Shopify classifies products with fine-tuned open models at roughly 40 million inferences a day, a volume it says commercial APIs can't economically serve. Inspired by the industry leaders, we ran the playbook ourselves: an agent that examines a listing, searches the product taxonomy, checks the brand, retrieves the attribute schema, and commits a structured decision, escalating to human review when the evidence is thin or the risk too high. 实际运行此工作流的公司证实了这一计算。Shopify 使用微调后的开源模型进行分类,每天进行约 4,000 万次推理,其表示商业 API 无法经济地提供此规模的服务。受行业领袖启发,我们亲自执行了该策略:一个检查条目、搜索产品分类、核对品牌、检索属性模式并提交结构化决策的代理,当证据不足或风险过高时,会自动升级至人工审查。

One episode, end to end一个端到端的处理过程

Product photo: black PU-coated work gloves

ListingWork gloves, PU-coated, black, size 10条目:工作手套,PU 涂层,黑色,10 号

Claimed brandAmazonBasics声称品牌:AmazonBasics

RegionUS地区:美国

1search_taxonomy → Safety Work Gloves1. 搜索分类 → 安全工作手套

2lookup_brand → registered, not protected2. 查询品牌 → 已注册,未受保护

3get_attribute_schema → brand, color, material, size3. 获取属性模式 → 品牌、颜色、材质、尺寸

Commits category, attributes, and verdict allowed✓ 提交类别、属性和允许的判定结果

The agent works through the listing the way an analyst would: it finds the right category, verifies the claimed brand, pulls the attribute schema for that category, and commits its decision. A clean listing ends in an approval; a suspicious one must end in a flag. The model learns to weigh the two mistakes asymmetrically, because missing a real violation costs it 7× more than raising a false alarm. 该代理像分析师一样处理条目:查找正确的类别、验证所称品牌、拉取该类别的属性模式并提交决策。合规的条目以批准结束;可疑的条目则必须标记。模型学会了不对称地权衡两种错误,因为漏掉真正的违规行为造成的代价是误报的 7 倍。

Part IV第四部分

A digital twin the model can practice in模型可以练习的数字孪生

A model can only practice a workflow it can actually perform, fail at, and retry. So we rebuilt the catalog workflow as a digital twin of an e-commerce platform: a simulated environment with the same listings, the same tools, and the same stakes as the real thing. The raw material is the Amazon Berkeley Objects dataset of real product images and listing data, which we turned into 177,767 review episodes: one listing each, with image, title, description, claimed brand, and region. Into that stream we planted controlled policy cases, mismatched images, conflicting brand claims, and deliberately legitimate claims as hard negatives, so every episode has a known correct answer to score against. 模型只能练习它能够实际执行、失败并重试的工作流。因此,我们重建了目录工作流作为电子商务平台的数字孪生:一个模拟环境,拥有与真实平台相同的条目、工具和风险。原始素材是 Amazon Berkeley Objects 数据集的真实产品图像和条目数据,我们将其转化为 177,767 个审查片段:每个片段包含一个条目,带有图像、标题、描述、所称品牌和地区。我们在该流中植入了受控的政策案例、不匹配的图像、冲突的品牌声明以及故意伪装成合法声明的“硬负样本”,以便每个片段都有已知的正确答案用于评分。

Inside the twin, the model works the way an analyst would. It can search a taxonomy of roughly 13,000 categories, check whether a brand is registered and protected, and retrieve the required attributes for a chosen category; then it must commit a category, the attributes, and a policy decision. A scorer grades every episode, rewarding correct outputs and penalizing missed violations, unsupported attributes, invalid categories, and wasted tool calls. The penalties encode business priorities directly: a missed violation costs 7× more than a false alarm. 在孪生系统中,模型的工作方式与分析师一样。它可以搜索约 13,000 个类别的分类体系,检查品牌是否已注册并受保护,并检索所选类别所需的属性;然后它必须提交类别、属性和政策决策。评分器对每个片段进行评分,奖励正确的输出,并惩罚漏掉的违规行为、不支持的属性、无效类别以及浪费的工具调用。惩罚直接编码了业务优先级:漏掉违规的代价是误报的 7 倍。

The digital twin, end to end端到端的数字孪生

TOOLS search_taxonomy() lookup_brand() get_attribute_schema() listing agent decision score reward 0.3·category + 0.3·attributes + 0.4·policy − tool_overage
Every part of the real workflow has a counterpart in the twin: listings flow in, the agent consults the same tools an analyst would, and a scorer grades each decision. The reward flowing back is what turns practice into learning. 真实工作流的每个部分在孪生系统中都有对应:条目流入,代理咨询分析师使用的相同工具,评分器对每个决策进行评分。反馈回来的奖励就是将练习转化为学习的过程。

Benchmarking the frontier models前沿模型基准测试

Before any training run, we measured how far the frontier could get on its own. We benchmarked five frontier models (GPT-5.5, GPT-5.6-sol, Gemini 3.1 Pro, Claude Opus 4.8, and Claude Fable 5) on 200 stratified validation episodes with identical tools, images, scorer, and turn budget, both with a plain prompt and with optimized prompt instructions: 2,800 characters of extraction conventions, lookup procedure, and worked examples, tuned the way a prompt engineer would tune a production system. The chart below shows both configurations, with our final GRPO-trained model added for comparison. 在任何训练运行之前,我们都测量了前沿模型在自身能力下能达到什么水平。我们在 200 个分层验证片段上对五款前沿模型(GPT-5.5、GPT-5.6-sol、Gemini 3.1 Pro、Claude Opus 4.8 和 Claude Fable 5)进行了基准测试,使用相同的工具、图像、评分器和调用预算,分别在普通提示词和优化后的提示词指令下进行测试:包括 2,800 个字符的提取惯例、查询程序和工作示例,其调优方式与提示词工程师调优生产系统的方式一致。下图显示了两种配置,并添加了我们最终 GRPO 训练的模型进行对比。

The best frontier configuration reached 76.9% of the achievable score; the trained 9B reached 87.3%. The gap is not intelligence. A frontier model starts every episode from zero: it has never seen this store's taxonomy, doesn't know its inventory conventions, which attribute values count as supported, or how the platform wants corner cases resolved, and it has to reconstruct all of that on the fly, from whatever fits in the prompt, on every single call. Optimized instructions can compress some of that knowledge, but the corner cases that decide the score are exactly the ones no instruction prompt can enumerate. 最好的前沿配置达到了可实现分数的 76.9%;训练后的 9B 模型达到了 87.3%。差距不在于智能。前沿模型在每个片段中都是从零开始:它从未见过该商店的分类体系,不知道其库存惯例、哪些属性值被视为支持,或者平台希望如何解决边缘案例,它必须在每次调用时从提示词中容纳的内容里实时重建所有这些信息。优化后的指令可以压缩部分知识,但决定分数的边缘案例恰恰是任何指令提示词都无法穷举的。

Model benchmarks模型基准测试

Bar chart: the fine-tuned Qwen3.5-9B scores 87.3% of the achievable ceiling, above every frontier model; frontier bars show zero-shot scores in solid teal with a lighter extension for the gain from optimized prompt instructions; untrained open-source bases score lower.
Each frontier bar shows the model's zero-shot score in solid teal; the lighter extension is what optimized prompt instructions added. Optimization helped, but only so much: the optimized configurations converged within a tenth of a point of each other, and the strongest zero-shot extractor (Gemini) got worse with the optimized instructions, so its bar shows no gain. The extra instructions weren't free either: they inflated input-token bills by 28–55% depending on the model, on every call, forever. Prompted task knowledge is rented per call; trained task knowledge is bought once and lives in the weights. 每个前沿模型柱状图显示了模型在零样本(zero-shot)下的分数(深青色);较浅的延伸部分是优化后的提示词指令增加的分数。优化有所帮助,但非常有限:优化后的配置在 0.1 分的范围内收敛,而最强的零样本提取器(Gemini)在使用优化指令后表现反而变差,因此其柱状图没有显示增长。额外的指令也并非免费:它们在每次调用中都使输入 Token 账单增加了 28–55%,且是永久性的。提示词任务知识是按次租用的;训练后的任务知识是一次性购买并存在权重中的。

Part V第五部分

Training details训练细节

The training setup was modest. We rented two RTX PRO 6000 GPUs, one generating rollouts and one applying gradient updates, with the open-source prime-rl framework running the infrastructure. The full run was 1,000 optimizer steps, took about three and a half days, and cost roughly $500 in GPU time. 训练设置非常简易。我们租用了两块 RTX PRO 6000 GPU,一块用于生成 rollout,一块用于应用梯度更新,并使用开源的 prime-rl 框架运行基础设施。整个运行过程为 1,000 个优化器步长,耗时约三天半,GPU 时间成本约为 500 美元。

Most of that was not needed to catch the frontier. The model crossed the frontier band after roughly 250 steps, about a day of training; the remaining steps were spent squeezing out maximum performance. The final model scores 0.626 on the benchmark, 87.3% of the achievable ceiling and about ten points above the best frontier configuration. 大部分时间并不需要用来追赶前沿模型。模型在约 250 步(约一天的训练)后就跨越了前沿模型区间;剩余的步数用于榨取最大性能。最终模型在基准测试中得分为 0.626,达到了可实现上限的 87.3%,比最好的前沿配置高出约 10 个百分点。

The gain comes from teaching the model the semantics of its environment: over thousands of scored episodes it learns how the store's taxonomy, tools, and policies relate to the reward, and converges on the optimal distribution of actions over the environment's action space. A general model spreads its capacity across everything; the specialist spends all of it on the specific task at the expense of general ability. 收益来自于教授模型其环境的语义:通过数千个评分片段,它学习了商店的分类体系、工具和政策如何与奖励相关联,并在环境的动作空间中收敛于最优的动作分布。通用模型将其能力分散在所有事物上;而专业模型则将其全部投入到特定任务中,以牺牲通用能力为代价。

The result also isn't a dead end. When a more capable open-source base model ships, the recipe transfers: as a rule of thumb, a stronger base yields a stronger post-trained specialist. And the deployed model generates its own training data as it works: logged decisions can be distilled into supervised fine-tuning sets, making each re-training cheaper and better-informed than the last. 结果也不是死胡同。当性能更强的开源基础模型发布时,配方是可以迁移的:经验法则是,更强的基础模型会产生更强的后训练专业模型。而且部署的模型在工作时会生成自己的训练数据:记录的决策可以提炼成监督微调集,使得每次重新训练都比上一次更便宜、更明智。

Training run训练运行

0.450.500.550.600.650.70step 02004006008001000frontier plateau 0.50–0.545max achievable 0.7179B: 0.671 @ 1,000
Every point is a measured evaluation from the W&B training monitor (two sampled rollouts per episode). The 9B run climbs from roughly 0.50 to 0.671 at step 1,000, clearing the frontier band before step 250. The monitor's sampled decoding runs slightly above the strict benchmark harness, where the final 9B adapter scores 0.626. 每个点都是来自 W&B 训练监视器的测量评估(每个片段采样两个 rollout)。9B 模型运行从约 0.50 攀升至第 1,000 步的 0.671,在第 250 步之前就清理了前沿模型区间。监视器的采样解码运行略高于严格的基准测试工具,最终 9B 适配器得分为 0.626。

So far, we have discussed how reinforcement learning helps a language model navigate your company environment more accurately. But one of the most important aspects of this paradigm is that the accuracy also comes cheaper than running frontier models. The cost-vs-quality chart that opens this article puts both on one picture, task score against cost per thousand listings, and the takeaway for a budget owner is simple: with everything else on the chart you are choosing between quality and cost. The trained specialist ends that trade-off, delivering the best score we measured at close to the lowest price. 到目前为止,我们讨论了强化学习如何帮助语言模型更准确地导航你的公司环境。但这种范式最重要的方面之一是,其准确性比运行前沿模型更便宜。本文开篇的成本与质量对比图将两者放在了一张图上,即任务分数与每千个条目的成本,对于预算负责人来说,结论很简单:图表上的其他所有方案你都在质量与成本之间做选择。训练后的专业模型终结了这种权衡,以接近最低的价格提供了我们测量到的最好分数。

Fine-tuning adds ~23 points over the base 9B at the same ~$0.50 per 1,000 listings: 40× cheaper than the least expensive frontier configuration (Gemini, $19/1k) and ~340× cheaper than the most expensive (GPT-5.5-pro, $172/1k). The 2,800 characters of prompt instructions also raised GPT-5.5's measured cost by a third, a prompt tax paid on every future call. The specialist's instructions live in its weights. At Shopify-scale volume of roughly 40 million decisions a day, the gap between $34 and $0.50 per thousand is about $500M a year versus $7M. 微调在每 1,000 条目约 0.50 美元的成本下,比基础 9B 模型提升了约 23 分:比最便宜的前沿配置(Gemini,19 美元/1k)便宜 40 倍,比最昂贵的(GPT-5.5-pro,172 美元/1k)便宜约 340 倍。2,800 个字符的提示词指令也使 GPT-5.5 的测量成本增加了三分之一,这是在每次未来调用中支付的“提示词税”。专业模型的指令存在于其权重中。在 Shopify 规模每天约 4,000 万次决策的体量下,每千次 34 美元与 0.50 美元之间的差距约为每年 5 亿美元与 700 万美元。

Wondering where these curves sit for your workflow? Book a 30-minute audit →想知道这些曲线在你的工作流中处于什么位置?预约 30 分钟审计 →

Part VI第六部分

What can owning AI do for you?拥有 AI 能为你做什么?

Here is the whole experiment in three sentences. We took one high-volume e-commerce workflow, reviewing product listings, and rebuilt it as a digital twin. We let a small open-source model practice inside it for a few days and roughly $500 of GPU time. The trained specialist ended above every frontier configuration we measured, at a fraction of their price per decision. 这里用三句话总结整个实验。我们选取了一个高容量的电子商务工作流——审查产品条目,并将其重建为数字孪生。我们让一个小型的开源模型在其中练习了几天,耗费约 500 美元的 GPU 时间。训练后的专业模型最终超过了我们测量过的所有前沿配置,且单次决策价格仅为其一小部分。

The recipe transfers to any work your business repeats at volume. Look for the places where people or API calls turn information into decisions all day: routing tickets, extracting fields from documents, checking submissions against policy, classifying products, approving or flagging transactions. What unites them is that every decision can be checked: there is a rule, a schema, or an expert who can say whether it was right. That check is the whole trick. If a decision can be scored, a model can practice it; if it can only be debated, it cannot. 该配方可迁移到贵司任何重复进行的高容量工作中。寻找那些人员或 API 调用整天将信息转化为决策的地方:路由工单、从文档中提取字段、根据政策检查提交内容、分类产品、批准或标记交易。它们的共同点是每个决策都可以被检查:存在规则、模式或专家可以判断其是否正确。这种检查就是全部诀窍。如果决策可以评分,模型就可以练习它;如果决策只能争论,那就无法练习。

It is just as important to know where this is the wrong tool. Two questions settle most cases: how frequently the task runs, and whether its outcome is verifiable. 了解在何处使用错误的工具同样重要。两个问题可以解决大多数情况:任务运行的频率以及结果是否可验证。

Right tool for the right-shaped problem针对正确形状问题的正确工具

prompt-optimized frontier model fine-tuned specialist owned, practiced, cheap at scale frontier model frontier model + human review verifiable non-verifiable rare task frequency frequent
The upper right is where fine-tuning pays: frequent work with verifiable outcomes. Everything else is better rented: a prompt-optimized frontier model covers verifiable but rare work, and a human stays in the loop where outcomes can't be verified. When the problem is changing facts rather than judgment, retrieval is the fix at any volume. The checklist below condenses the upper-right conditions into something you can hold your own workflows against: 右上角是微调产生回报的地方:具有可验证结果的频繁工作。其他所有工作最好采用租用方式:提示词优化的前沿模型涵盖可验证但罕见的工作,而在结果无法验证的地方则保留人工介入。当问题是事实变更而非判断时,无论容量如何,检索都是解决方案。下面的清单将右上角的条件浓缩成了你可以用来衡量自己工作流的准则:

Your workflow is a fine-tuning candidate if at least one point applies…如果至少符合以下一点,你的工作流就是微调候选对象……

  • It happens at high volume, often enough that per-decision cost and errors compound into real money它以高容量发生,频率高到单次决策成本和错误会累积成真金白银
  • Every outcome can be checked by a rule, test, or rubric, with no person in the loop每个结果都可以通过规则、测试或评分标准进行检查,无需人工介入
  • Your experts agree on what a correct answer looks like你的专家对于什么是正确答案达成一致
  • A capable model already succeeds sometimes, just not reliably enough有能力的模型有时已经能成功,只是不够可靠
  • The right answer can't be a lucky guess正确答案不能是幸运的猜测
  • It spans several steps: reasoning, tool calls, then a committed decision它跨越多个步骤:推理、工具调用,然后是提交决策
  • It runs on your own tools, schemas, and policies它运行在你自己的工具、模式和政策上
  • Different mistakes carry different costs不同的错误带有不同的成本
  • Sensitive data cannot leave infrastructure you control敏感数据不能离开你控制的基础设施

That last one is also the safety argument: a model you train runs inside your own boundary, so prompts and records never reach a vendor. One point is enough to be worth a conversation, and working out which ones apply is what the call is for. 最后一点也是安全性的论点:你训练的模型运行在你自己的边界内,因此提示词和记录永远不会到达供应商。符合一点就值得一谈,而通话的目的就是为了弄清楚哪些点适用。

If that describes one of your workflows, the question is no longer whether it's the right fit for you, but which of your decisions you should stop renting. Owning your intelligence is the kind of advantage that compounds: the model, the evaluation, and the data stay yours and improve together, on your terms. If you would like to see what that looks like for your business, we would be glad to work through it with you: 如果这描述了你的一个工作流,问题就不再是它是否适合你,而是你应该停止租用哪些决策。拥有自己的智能是一种会复合的优势:模型、评估和数据保持为你所有,并根据你的条款共同改进。如果你想看看这对你的业务意味着什么,我们很乐意与你一起探讨:

See these numbers on your workflow在你的工作流中查看这些数据

Sounds like something that would help your business? Book a call and we'll help you evaluate the potential business impact and determine the business ROI, free of charge. 听起来这会对你的业务有所帮助?预约通话,我们将免费帮你评估潜在的业务影响并确定业务 ROI。

Book a free expert call预约免费专家通话

Want to try the models yourself? They're on Hugging Face.想亲自尝试这些模型吗?它们在 Hugging Face 上。

Appendix附录

The track record: task-trained specialists vs. prompted frontier models业绩记录:任务训练专业模型 vs. 提示词前沿模型

Beyond the flagships, the same shape repeats across industries. Every result below is the company's own reported figure against the frontier model it was replacing or competing with; each company name links to the write-up it came from. 除了旗舰案例外,同样的模式在各行业中反复出现。以下每个结果都是公司报告的对比数据,针对其正在取代或竞争的前沿模型;每个公司名称都链接到其来源文档。

CompanyWhat the model doesResultBusiness payoff
Cognition Writes and fixes software in production Beats GPT-5.5 on a standard coding benchmark Frontier-level output at a fraction of the serving cost, streaming 1,000 tokens per second
AT&T Summarizes 900k support calls a day and flags personal data 17% better than GPT-4o at catching personal data Reviews fraud cases 12× faster at GPT-4o accuracy; saves millions of dollars annually
LinkedIn Matches candidates to job openings 4% more accurate than the GPT model it replaced 75× cheaper than GPT-4, 6× cheaper than GPT-4o
Ambience Healthcare Assigns medical billing codes More accurate than 18 board-certified physicians on a gold-panel test 12 points ahead of prompted o4-mini, the OpenAI model it was built from
Phonely Answers customer phone calls 99.2% accuracy vs GPT-4o's 94.7% Responses 73% faster than its previous GPT-4o setup; one customer replaced 350 human agents in a month
OpenPipe Handles email and support tickets 93% on support QA where OpenAI's o3 scored 50% 64× cheaper and 5× faster to run than o3
Perplexity Answers search queries with sources Matches GPT-4o in blind user tests 10× the speed of GPT-4o at a fraction of the price
Checkr Classifies criminal-record entries in background checks Beats GPT-4 on the hardest cases 5× cheaper and 30× faster than its previous GPT-4 setup

Sources & notes来源与注释

  1. Ramp AI-spend and revenue data, Nov 2022 – Dec 2025: Ramp CEO Eric Glyman on X, July 2026.Ramp AI 支出和营收数据,2022 年 11 月 – 2025 年 12 月:Ramp CEO Eric Glyman 在 X 上的发言,2026 年 7 月。
  2. Uber: Fortune, May 2026, the entire 2026 AI budget spent in four months. Microsoft: BuildMVPFast, 2026, on the cancelled Claude Code licenses.Uber:Fortune,2026 年 5 月,2026 年全年 AI 预算在四个月内花光。微软:BuildMVPFast,2026 年,关于取消的 Claude Code 许可证。
  3. Thinking Machines Lab × Bridgewater AIA Labs, Learning to Replicate Expert Judgment in Financial Tasks, June 2026.Thinking Machines Lab × Bridgewater AIA Labs,学习在金融任务中复制专家判断,2026 年 6 月。
  4. Ramp × Prime Intellect, How Ramp Used RL to Beat Frontier Models at Spreadsheet Search, July 2026.Ramp × Prime Intellect,Ramp 如何利用 RL 在电子表格搜索中击败前沿模型,2026 年 7 月。
  5. Shopify Engineering, Leveraging multimodal LLMs at scale.Shopify Engineering,大规模利用多模态 LLM。
  6. Checkr, LLMOps Micro-Summit presentation on fine-tuned charge classification.Checkr,关于微调费用分类的 LLMOps 微型峰会演示。
  7. Catalog sizes: eBay investor relations (~2.5B live listings, Dec 2025); Instacart Engineering (1.4B items, 1,500+ retailers); Marketplace Pulse (420M+ Walmart.com products; Walmart does not disclose).目录规模:eBay 投资者关系(约 25 亿在售条目,2025 年 12 月);Instacart Engineering(14 亿项,1,500+ 零售商);Marketplace Pulse(4.2 亿+ Walmart.com 产品;沃尔玛未披露)。
  8. Walmart Q2 FY2025 earnings call, via CIO Dive: 850M+ catalog data points created or improved with generative AI, estimated at ~100× current headcount if done manually.沃尔玛 2025 财年第二季度财报电话会议,来自 CIO Dive:通过生成式 AI 创建或改进了 8.5 亿+ 目录数据点,估计如果手动完成则需要约 100 倍的当前人力。
  9. Salsify 2025 Consumer Research (n=1,910, US/UK) and Akeneo 2023 Global B2C Survey. Both are vendor-commissioned surveys by product-information-management companies; treat directionally.Salsify 2025 消费者研究(n=1,910,美/英)和 Akeneo 2023 全球 B2C 调查。两者均为产品信息管理公司委托进行的供应商调查;仅供参考方向。
  10. The Wall Street Journal, AI giants are handing out tons of free computing power to grab startup share: the token-cost subsidies referenced in Problem 01.华尔街日报,AI 巨头正在发放大量免费算力以抢占初创公司份额:问题 01 中提到的 Token 成本补贴。
  11. Anthropic Engineering, Effective context engineering for AI agents.Anthropic Engineering,针对 AI 代理的有效上下文工程。
  12. Liu et al., Lost in the Middle: How Language Models Use Long Contexts, arXiv 2307.03172.Liu 等人,迷失在中间:语言模型如何使用长上下文,arXiv 2307.03172。
  13. METR, AI usage survey, May 2026: self-reported AI gains overestimate measured time savings by roughly 40 percentage points.METR,AI 使用调查,2026 年 5 月:自报的 AI 收益比测量的时间节省高出约 40 个百分点。
  14. Sentra, What is a company brain? The 2026 guide.Sentra,什么是企业大脑?2026 年指南。
  15. Case write-ups linked in the text: Harvey × Applied Compute, Intercom Fin Apex, Cognition SWE-1.7, Adaptive ML × AT&T, LinkedIn EON, Ambience Healthcare, Phonely (VentureBeat), OpenPipe ART·E, Perplexity Sonar, Instacart, DoorDash × Applied Compute, Shopify Flow agent.文中链接的案例报告:Harvey × Applied Compute、Intercom Fin Apex、Cognition SWE-1.7、Adaptive ML × AT&T、LinkedIn EON、Ambience Healthcare、Phonely (VentureBeat)、OpenPipe ART·E、Perplexity Sonar、Instacart、DoorDash × Applied Compute、Shopify Flow 代理。