Introduction简介

At Grab, analytics sits close to almost every decision that matters. Our north star is the democratisation of intelligence, ensuring that anyone making a business call has immediate access to trustworthy answers.在 Grab,分析工作与几乎每一项重要决策都息息相关。我们的核心目标是实现“智能民主化”,确保每一位做出业务决策的人都能即时获得可靠的答案。

Over the last two years, model capability has crossed a threshold enabling this shift. Agents now do in minutes what used to take a week: preparing the data, writing queries, running deep analysis and developing insights for business opportunities, designing experiments and interpreting the results, drafting the commentary that follows, and more. Our throughput is no longer rate-limited by how fast an individual can write code, build a deck, or run a deep-dive. It is rate-limited by how fast we can frame the right problem, judge the right answer, and influence the right decision.过去两年间,模型能力的突破推动了这一转变。以前需要一周才能完成的工作——包括准备数据、编写查询、进行深度分析、挖掘业务机会、设计实验并解读结果,以及撰写后续评论等——如今代理(Agents)只需几分钟即可完成。我们的产出效率不再受限于个人编写代码、制作幻灯片或进行深度调研的速度,而是取决于我们界定问题、判断答案以及推动决策的速度。

As autonomy climbs, an analyst’s impact moves from producing the artefact to owning the question and the call behind it, and the role evolves to become part builder, part advisor, part strategist, owning the loop rather than running it. That unlocks two things at once: work we already do, faster and at lower marginal cost, and work we could never staff before, sitting beside every product manager, business owner, and operator at the moment they decide.随着自主性的提升,分析师的影响力已从单纯的“产出成果”转向“把控问题与决策”。这一角色正在演变为集构建者、顾问和战略家于一身,通过掌控整个流程循环而非仅仅执行流程来创造价值。这同时实现了两方面的突破:一是让现有工作以更快的速度、更低的边际成本完成;二是让我们能够处理以往无法覆盖的工作,在每一位产品经理、业务负责人和运营人员做决策时提供实时支持。

The ladder能力阶梯

We were heavily inspired by Dan Shapiro’s framing of five levels for AI coding. We use a similar ladder that defines how much of the loop an agent should own and where human judgement stays for every analytics loop.我们深受 Dan Shapiro 关于人工智能编程五个等级的启发,并采用了一套类似的阶梯模型,用于定义每个分析循环中代理应承担的责任范围以及人工判断的介入点。

One distinction runs across every level: who owns the loop, and where human judgement is required.每个等级都遵循一个核心原则:谁拥有该循环的控制权,以及在何处需要人工判断。

Level What Human role Agent role
L2 AI-Assisted Owns and executes every step; uses AI to draft, suggest, summarise Drafts SQL, suggests a visualisation
L3 Human plans, agent owns steps, human reviews Frames the question, picks the metric, the segment, and the comparison frame, reviews evidence, owns the recommendation Discovers data, writes and runs the query, sanity checks, drafts the write-up, flags caveats
L4 Agent plans, agent owns workflows, human reviews Sets intent and guardrails; reviews at gates (anomaly, novel scope, sensitive cut); owns the stakeholder relationship and sign-off Orchestrates discovery through query, analysis, validation, narrative and publish; runs validation, escalates exceptions
L5 End-to-end autonomous Sets objectives, quality bars, risk thresholds, escalation rules; reviews exceptions only Detects anomalies and opportunities, runs the loop, surfaces insight, evolves the metric layer, context and skills

Human judgement remains at every level, and autonomy never removes accountability. Humans own problem framing, canonical metric definitions, the causal story behind a move, business-case assumptions, the go/no-go, and the stakeholder relationship. A higher level means more of the mechanical loop sits with the agent and more human attention concentrates on the ambiguous, high-stakes work.在任何等级下,人工判断都不可或缺,自主性永远不能替代责任。人类负责界定问题、定义标准指标、分析决策背后的因果逻辑、设定业务假设、决定是否执行以及维护利益相关者关系。等级越高,意味着代理承担的机械性循环工作越多,而人类的精力则越集中于那些模糊且高风险的工作。

Making the climb攀登阶梯

Five core capabilities move a workflow up the ladder. They also gate the climb in order: L3 needs execution and certified context, L4 needs gates and agentic review good enough that reviewing only at gates is honest, L5 needs a learning loop that closes.五项核心能力推动工作流程在阶梯上不断提升。这些能力按顺序构成了晋升的关卡:L3 需要执行力和经过认证的上下文;L4 需要设置关卡和足够可靠的代理审查机制,以确保仅在关键节点进行人工复核是合理的;L5 则需要形成闭环的学习机制。

  • Execution: A stack that runs the loop end to end rather than a notebook/workflow a human drives.执行力:一套能够端到端运行整个循环的系统,而非依赖人工驱动的笔记本或工作流。
  • Knowledge: Metrics certified at the right grain, discoverable in our catalogue, grounded in context an agent can read. Ambiguous definitions cause most analytics slop.知识:指标需在合适的粒度下经过认证,在我们的目录中可被发现,并基于代理可读取的上下文。模糊的定义是导致分析工作混乱的主要原因。
  • Control: Repeatable expectations become mechanical checks, while human review handles what a rule cannot.控制:将可重复的预期转化为机械化的检查,而人工复核则处理规则无法覆盖的情况。
  • Review and governance: Agents check their own output against the gates and escalate on defined triggers. We govern definitions, targets, risk and exceptions.审查与治理:代理根据预设关卡检查自身输出,并在触发特定条件时进行升级处理。我们负责治理定义、目标、风险和异常情况。
  • Learning: When an agent fails the same way twice, we encode the fix into context documents, golden datasets, evals and gates.学习:当代理两次犯下同样的错误时,我们将修复方案编码进上下文文档、黄金数据集、评估体系和关卡中。

What this looks like in practice实践中的应用

What follows is a set of explorations from the last two years. Some run in production today, while others are still teaching us where the limits are.以下是我们过去两年的探索成果。其中一些已投入生产环境,另一些则仍在帮助我们探索能力的边界。

Loops that run end to end端到端运行的循环

Spartan is our end-to-end agentic analytics workflow, embedded across surface areas (like Slack), and most of its usage comes from people who are not analysts. On any given day, the Slack channel enables a range of analytics actions: from ads salespeople pulling spend breakdowns for a named merchant, to campaign managers sizing audiences for a target segment, and country teams asking why a number moved week on week. All of it in plain business language.Spartan 是我们的端到端代理分析工作流,已嵌入 Slack 等多种界面,其大部分用户并非分析师。在任何一天,该 Slack 频道都能支持一系列分析操作:从广告销售人员提取特定商家的支出明细,到营销经理为目标细分市场评估受众规模,再到国家团队查询某项指标周环比变动的原因。所有这些操作均以通俗的业务语言完成。

Figure 1. Index architecture across our knowledge base.图 1. 跨知识库的索引架构。

Two requests from July best demonstrate how it works. A commercial manager asked why revenue fell in the Philippines mid-market segment in the last two weeks of June. Separately, a product manager asked for a summary of a frequency-cap experiment on the ads surface. Both arrived as natural language questions in Slack and took entirely different routes through the system.七月份的两个请求很好地展示了其工作原理。一位商务经理询问为何六月最后两周菲律宾中端市场的收入下降;与此同时,一位产品经理要求总结一项针对广告页面的频次上限实验。这两个请求均以自然语言形式出现在 Slack 中,但在系统中走出了完全不同的路径。

The router reads the first as a root-cause question and sends it down the diagnostic path. It identifies the best analysis framework for ads revenue, which is codified knowledge of how the metrics in that domain relate to each other, which dimensions are worth decomposing, and what counts as a meaningful move. Then it works through segment, market and campaign type against certified metrics to isolate what changed. The second question never touches the data lake. The router reads it as an experiment question, selects the experiment skill, pulls the pre-computed scorecard and the test’s own metadata from our experiment platform, and summarises the read rather than recomputing it. This is powered through 50+ skills and 120+ analysis frameworks that sit behind that routing decision. Underlying that is an index that tells the agent what to search, context that tells it how to query, and a framework that tells it how to think. Because the frameworks are shared rather than living in an analyst’s head, the interpretation compounds instead of being re-derived every time someone asks.路由系统将第一个请求识别为根本原因分析,并将其导向诊断路径。它会确定最适合广告收入的分析框架,即关于该领域指标如何关联、哪些维度值得拆解以及何为显著变动的编码知识。随后,它通过已认证的指标对细分市场、地区和活动类型进行分析,从而锁定变动原因。第二个请求则完全无需触及数据湖。路由系统将其识别为实验类问题,选择实验技能,从我们的实验平台提取预计算的记分卡和测试元数据,并直接总结结果而非重新计算。这一过程由 50 多种技能和 120 多种分析框架支撑,并由路由决策调度。其底层是一个告诉代理搜索什么的索引、提供查询方法的上下文,以及指导思考方式的框架。由于这些框架是共享的,而非仅存在于分析师脑中,因此解读经验能够不断积累,无需每次有人提问时都重新推导。

The second example of such a loop is Scarlet, which powers near-self-healing pipelines (L4). When a pipeline fails, an agent runs the root-cause analysis, triages, and then either fixes it or hands it to the team that owns the upstream problem. It escalates when the failure sits outside its documented runbooks or the pre-defined gates fire.此类循环的第二个例子是 Scarlet,它支持近乎自愈的流水线(L4)。当流水线失败时,代理会进行根本原因分析、分类,然后自行修复或将其移交给负责上游问题的团队。如果失败情况超出了其文档记录的运行手册或预定义关卡的范围,它会进行升级处理。

Figure 2. Scarlet in action on Slack.图 2. Slack 中的 Scarlet 运行实况。

Context that maintains itself自我维护的上下文

Context sets an agent’s ceiling. An agent that does not know a metric’s grain, its exclusions, and its caveats will guess and confidently produce wrong outputs at speed and at scale.上下文决定了代理的能力上限。如果一个代理不了解指标的粒度、排除项和注意事项,它就会盲目猜测,并以极高的速度和规模产出错误的结论。

Realising the criticality of this, we have dedicated platform investment, as well as dedicated functional bandwidth to generate context docs.意识到这一点至关重要,因此我们投入了专门的平台资源,并分配了专门的职能带宽来生成上下文文档。

Figure 3. ContextIQ.图 3. ContextIQ。

Context goes out of date faster than anyone maintains it by hand, so we build the maintenance into our workflows. We built ContextIQ, and its Context Lifecycle Manager, to treat context as something with a lifecycle rather than a document somebody wrote once. A newer skill of ours reads an instrumentation spec alongside the existing context, proposes the SQL changes that follow from it, and updates the context document in the same pass. We work the problem from the other direction too. When we categorise an agent failure in production, we patch the context document behind it.上下文的过时速度远超人工维护的速度,因此我们将维护工作内置于工作流中。我们构建了 ContextIQ 及其“上下文生命周期管理器”,将上下文视为一个有生命周期的对象,而非一次性文档。我们的一项新技能可以同时读取仪表盘规范和现有上下文,提出相应的 SQL 修改建议,并同步更新上下文文档。我们还从另一个方向解决问题:当我们在生产环境中对代理的失败进行分类时,我们会直接修补其背后的上下文文档。

Two analysts recently used our internal agents to understand how the packaging fee is stored as a configuration. Having found the answer, the agent opened a merge request that committed both a certified-context table reference and a golden-dataset test case, so the next agent to ask the same question would find the answer already documented and the check already in place. One of the analysts spotted a false positive in it. The agent corrected itself and reopened the merge request. That is the learning capability working as designed, and it happened without anyone setting out to demonstrate it.最近,两位分析师利用我们的内部代理了解包装费如何作为配置存储。找到答案后,代理自动发起了一个合并请求(Merge Request),提交了经过认证的上下文表引用和黄金数据集测试用例。这样,下一位提出相同问题的代理就能直接找到已记录的答案和已就绪的检查项。其中一位分析师发现了其中的误报,代理随即进行了自我修正并重新开启了合并请求。这就是学习能力按预期发挥作用的体现,而且这一切是在没有刻意演示的情况下发生的。

Loops that run unattended无人值守的循环

The step from L3 to L4 is mostly the step from interactive to scheduled, and it is where we go down the path of autonomous execution, because no human is watching at the moment the work runs.从 L3 到 L4 的跨越主要是从“交互式”转向“定时执行”,这也是我们迈向自主执行之路的关键,因为工作运行时没有人类在场监控。

We have built root-cause analysis (RCA) as a platform capability, and it powers our automated metric and OKR commentaries, which are published through automated agents (configurable cadence). It judges whether a move is meaningful against standard deviation over six months and year on year, walks the metric tree to find which country, segment or funnel stage carried it, and correlates the operational metrics that moved alongside. Importantly, it also scans internal context for what teams changed on the ground, such as delivery fee and incentive moves, merchant visibility shifts, and experiments shipped in the same period. It also compares the movement against the same period in the previous year, which separates a seasonal effect from a real one and enables it to report a Songkran (Thai New Year) dip as amplified rather than merely expected. All of it is grounded in our own context documents, which keeps the narrative about the business rather than generic model output. The analytics owner is tagged on every report, and edits sync back so corrections land in the system.我们已将根本原因分析(RCA)构建为平台能力,并支持我们的自动化指标和 OKR 评论,这些评论通过自动化代理按配置的频率发布。它会根据六个月及同比的标准差来判断变动是否具有显著性,遍历指标树以找出是哪个国家、细分市场或漏斗阶段导致了变动,并关联同时变动的运营指标。重要的是,它还会扫描内部上下文,了解团队在地面上所做的变动,例如配送费和激励措施的调整、商家可见性的变化以及同期上线的实验。它还会将变动与去年同期进行对比,从而区分季节性影响与真实变动,使系统能将宋干节(泰国新年)期间的下滑报告为“影响加剧”而非仅仅是“预期内”。所有这些都基于我们自己的上下文文档,确保叙述内容围绕业务本身,而非通用的模型输出。分析负责人会被标记在每一份报告上,编辑内容会同步回系统,确保修正信息得到记录。

Figure 4. OKR commentary shared through RCA agent.图 4. 通过 RCA 代理分享的 OKR 评论。

Analysts as builders分析师即构建者

The clearest evidence that our centre of gravity has moved is BriX, an internal portal we built and run ourselves.我们的重心已发生转移,最明显的证据就是 BriX——一个由我们自行构建和运营的内部门户。

Figure 5. Home page of BriX.图 5. BriX 主页。

The premise is to configure once, host everywhere. We configure a system prompt, a set of context files, a model, the MCP connections and an interface once, and what comes out is a purpose-built analytics surface for a particular team or job. Each one inherits certified data, permissions and reusable agent skills rather than being wired up from scratch, and it runs wherever the work already happens: in Slack, invoked from inside an IDE, or on a schedule with nobody watching. We have grown usage more than 10x since September 2025, with strong retention, and every function at Grab now has users on it. Our aim is to put L3 workflows in the hands of people who are not advanced users.其核心理念是“一次配置,随处托管”。我们只需配置一次系统提示词、一组上下文文件、模型、MCP 连接和界面,即可为特定团队或工作任务生成专用的分析界面。每一个界面都继承了已认证的数据、权限和可复用的代理技能,无需从零开始搭建,并且可以在工作发生的任何地方运行:在 Slack 中、在 IDE 内部调用,或按计划自动执行。自 2025 年 9 月以来,我们的使用量增长了 10 倍以上,留存率很高,Grab 的每一个职能部门现在都有人在使用它。我们的目标是将 L3 工作流交到非高级用户手中。

We run it without a product manager, a technical programme manager or a designer. Our data engineers own the product, the platform, the support queue and the eval loop, with Claude Design doing the interface work and the builders triaging their own bugs. In the first half of this year, they shipped 31 production deployments, 283 merge requests and 60 features.我们运营该平台时无需产品经理、技术项目经理或设计师。我们的数据工程师拥有该产品、平台、支持队列和评估循环的完全所有权,由 Claude Design 负责界面工作,构建者们则自行处理错误。今年上半年,他们完成了 31 次生产部署、283 次合并请求和 60 项功能开发。

Three of our apps show the range:我们的三个应用展示了其广泛的应用范围:

  • Insights Lab is the general-purpose surface: a stakeholder asks for a metric, a breakdown or a root-cause in natural language, and the agent loads a specialist skill and answers off certified metrics rather than from memory.Insights Lab 是通用型界面:利益相关者用自然语言询问指标、拆解分析或根本原因,代理会加载专业技能,并基于已认证的指标而非记忆进行回答。
  • We built Funnelytics to enable easy understanding of our consumer funnels. A funnel question used to mean an analyst writing the query and then assembling the view in Tableau or Power BI, and doing it again the next time someone wanted a slightly different path through the app. Now a stakeholder picks the events they care about and Funnelytics queries the raw event stream, builds the Sankey and funnel views, and writes the summary. If they cannot find the right instrumentation, which happens often on products still being redesigned, a live debugger lets them tap through the app on their own phone and watch the events fire.我们构建了 Funnelytics 以便轻松理解我们的消费者漏斗。过去,漏斗问题意味着分析师需要编写查询,然后在 Tableau 或 Power BI 中组装视图,下次有人想要稍微不同的路径时又得重来一遍。现在,利益相关者只需选择他们关心的事件,Funnelytics 就会查询原始事件流,构建桑基图和漏斗视图,并撰写总结。如果他们找不到合适的埋点(这在处于重构阶段的产品中很常见),实时调试器允许他们在自己的手机上操作应用,观察事件触发情况。
  • Monte (like Monte Carlo) runs simulations to put a probability on a business outcome. You give each uncertain input a range rather than a single value, and it runs ten thousand scenarios to return the likelihood of hitting a target.Monte(类似于蒙特卡洛模拟)运行模拟来为业务结果赋予概率。用户为每个不确定输入提供一个范围而非单一数值,它会运行一万次场景模拟,得出达成目标的可能性。
Figure 6. Interface of Insights Lab and Funnelytics. 图 6. Insights Lab 和 Funnelytics 的界面。

Outside the portal, the same instinct shows up in smaller ways. Our analysts have been building more bespoke tools that enable better workflows for themselves and stakeholders.在门户之外,同样的直觉也体现在更小的改进中。我们的分析师一直在构建更多定制工具,为自己和利益相关者实现更好的工作流。

The path forward未来之路

In February, 44% of the tickets our analysts closed were mechanical (data preparation, alerting, reporting); by June, that share had fallen to 30%. That capacity was redirected to other higher-leverage work, such as building new workflows to enable stakeholder self-serve, as well as more time spent on generating deeper insights for business opportunities.二月份,我们分析师关闭的工单中有 44% 是机械性的(数据准备、报警、报告);到了六月,这一比例降至 30%。这些产能被重新分配到其他高杠杆工作中,例如构建新的工作流以实现利益相关者的自助服务,以及投入更多时间挖掘深度的业务机会。

Figure 7. Comparison of percentage of tickets closed in Q1 vs Q2.图 7. 第一季度与第二季度关闭工单百分比对比。

Importantly, our cycle times reduced by ~33%.重要的是,我们的周期时间缩短了约 33%。

Figure 8. Comparison of time taken to resolve a ticket in Q1 vs Q2.图 8. 第一季度与第二季度解决工单耗时对比。

The sharpest version of this sits in a Slack channel where self-serve agents are enabled. In March, an analyst had to step into half of them; by May, it was under a quarter. The share answered with no human involvement rose from 53% to 67% for metric questions, 63% to 90% for data pulls, and 50% to 81% for SQL requests. Just under three in four of the threads were started by someone outside the analytics team, and 85% of them got a first response inside a minute. Nearly every thread is logged as a ticket on the team’s board, and roughly two-thirds of the data exploration tickets on that board now arrive through the channel rather than through an analyst, and are solved by our data agents. For the ~230 tickets that arrived via the channel, if we apply a conservative assumption of 1–2 days per ticket, that is 230 to 470 business days of stakeholder asks that would have been in the backlog.最显著的变化发生在启用了自助服务代理的 Slack 频道中。三月份,分析师需要介入其中一半的请求;到了五月,这一比例降至四分之一以下。指标类问题的无人值守回答率从 53% 升至 67%,数据提取从 63% 升至 90%,SQL 请求从 50% 升至 81%。近四分之三的对话线程由分析团队以外的人发起,其中 85% 在一分钟内得到了首次响应。几乎每个线程都被记录为团队看板上的工单,看板上约三分之二的数据探索工单现在通过该频道而非直接找分析师解决,并由我们的数据代理完成。对于通过频道提交的约 230 个工单,如果我们保守估计每个工单需要 1-2 天,那么这相当于 230 到 470 个工作日的利益相关者需求被从积压任务中释放出来。

None of these arrived on a roadmap. They came from analysts who saw a loop worth automating and built it, which is why the climb is uneven. These have been strong proof points for us to believe our investments are working, and many of these workflows are starting to operate at scale. We will keep experimenting and iterating, and we expect to get a fair amount of it wrong. An analyst who owns a loop, sets its quality bar and reviews its exceptions is doing a different job from one who answers questions. Most of our team is somewhere in that transition today, and we truly believe it is changing what analytics is at Grab.这些成果并非来自路线图规划,而是源于那些发现循环值得自动化并亲手构建的分析师,这就是为什么攀登过程是不均匀的。这些已成为我们相信投资方向正确的有力证明,其中许多工作流已开始规模化运作。我们将继续实验和迭代,并预料到会犯不少错误。一个拥有循环、设定质量标准并审核异常情况的分析师,与一个只负责回答问题的分析师,其工作性质截然不同。我们团队的大多数人目前正处于这一转型过程中,我们坚信这正在改变 Grab 的分析工作定义。

Join us加入我们

Grab is a leading superapp in Southeast Asia, operating across the deliveries, mobility, and digital financial services sectors. Serving over 900 cities in eight Southeast Asian countries: Cambodia, Indonesia, Malaysia, Myanmar, the Philippines, Singapore, Thailand, and Vietnam. Grab enables millions of people every day to order food or groceries, send packages, hail a ride or taxi, pay for online purchases or access services such as lending and insurance, all through a single app. We operate supermarkets in Malaysia under Jaya Grocer and Everrise, which enables us to bring the convenience of on-demand grocery delivery to more consumers in the country. As part of our financial services offerings, we also provide digital banking services through GXS Bank in Singapore and GXBank in Malaysia. Grab was founded in 2012 with the mission to drive Southeast Asia forward by creating economic empowerment for everyone. Grab strives to serve a triple bottom line. We aim to simultaneously deliver financial performance for our shareholders and have a positive social impact, which includes economic empowerment for millions of people in the region, while mitigating our environmental footprint.Grab 是东南亚领先的超级应用,业务涵盖配送、出行和数字金融服务。服务覆盖东南亚八个国家的 900 多个城市:柬埔寨、印度尼西亚、马来西亚、缅甸、菲律宾、新加坡、泰国和越南。Grab 每天让数百万人能够订餐或购买杂货、寄送包裹、打车或预订出租车、支付在线消费或获取贷款和保险等服务,这一切都在一个应用中完成。我们在马来西亚通过 Jaya Grocer 和 Everrise 运营超市,这使我们能够为该国更多消费者带来按需杂货配送的便利。作为金融服务的一部分,我们还通过新加坡的 GXS Bank 和马来西亚的 GXBank 提供数字银行服务。Grab 成立于 2012 年,使命是通过为每个人创造经济赋能来推动东南亚向前发展。Grab 致力于实现三重底线:我们旨在同时为股东提供财务业绩、产生积极的社会影响(包括为该地区数百万人提供经济赋能),并减轻我们的环境足迹。

Powered by technology and driven by heart, our mission is to drive Southeast Asia forward by creating economic empowerment for everyone. If this mission speaks to you, join our team today!以科技为动力,以真心为驱动,我们的使命是通过为每个人创造经济赋能来推动东南亚向前发展。如果这一使命引起了您的共鸣,请立即加入我们的团队!