A couple of months ago, I sat across from my nine-year-old daughter’s teachers at a parent-teacher conference. They were kind but concerned. She takes her time on assignments, they said, she’s often deep in thought. How would she do on timed tests next year? I told them I wasn’t worried. What they described as a problem is, to me, one of the most important things she can learn: the ability to take a hard problem and reason through it from beginning to end. In a world optimized for efficiency, qualities like patience, perseverance, and attention to detail are not deficiencies. They are the foundation of sound judgment, and this is the most valuable skill set.数月之前,我与九岁女儿之师于家长会上相对而坐。师长和蔼,却难掩忧虑:言她做功课慢,常凝神沉思。若逢明年限时测验,不知如何应对?我闻言却无半分担忧。他们口中的弊病,于我而言恰是她此生最需修习的功课——能遇难题,自始至终于逻辑中抽丝剥茧,直至水落石出。世人皆求高效,然耐心、毅力、注重细节这等品性,绝非缺漏,反而是成事之根基,更是这世上最珍贵的本事。
The more time I spend working with AI, the more convinced I become that what matters most for her future isn’t how quickly she can answer. It’s whether she has the judgment to know when an answer can be trusted.我越是与AI打交道,便越坚信:她未来最紧要的,绝非答题速度。而是能否明辨:何时所得之答,可信可托。
I’ve spent decades at Microsoft watching this tension play out: first building tools for other developers, then working across AI as models moved from research curiosities to systems deployed at scale. Now we’re building Microsoft IQ, where we’re exploring how an organization’s collective intelligence can become its greatest advantage. Through every one of those chapters, one thing has remained true: it’s never enough for a system to be powerful; it must also be trustworthy.我在微软数十载,亲见这层张力周而复始:起初为其他开发者铸就工具,后随AI模型从研究阶段的珍奇玩好,渐次落地为大规模部署的系统,遂又投身AI全域。如今我们正打造Microsoft IQ,欲探一组织之集体智慧,何以成为其最大倚仗。历数这些岁月,唯有一事亘古不变:系统光有威能远远不够,更须值得信赖。
Trust is what turns assistance into delegation. When we can trust an agent to do what we intend, within the limits we set, we can hand off the work we never wanted to spend our lives on: the repetitive tasks that drain attention, the mundane work that fills a day without moving anything meaningful forward, the dangerous work humans should not have to do, the work too vast for any individual or team. Agents should take on that toil, extend our reach, and give us back our time for the work that calls for something only humans bring.信任,方能使辅助转为委派。若我们能信任一智能体,依我等设定的边界,成我等所愿之事,那便可把那些本不该耗去一生光阴的活计,尽数交托:耗神耗力的重复劳作,填满整日却无半分进展的琐碎杂务,人类本不必以身犯险的危险差事,还有任何个人或团队都力有不逮的浩繁工作。智能体当承此劳苦,拓我等所能及之边界,把时间还予我等,去做那些 unmistakably 属于我们的工作——思考、决断、创造、互相关照。
My daughter doesn’t know any of this yet. But by the time she’s grown, most of the work that rewards speed and repetition will be work we delegate. What will matter then is exactly what gave her teachers pause: the patience to stay with a hard problem, reason through it, and decide when she’s reached a conclusion she can trust. The very thing they feared might hold her back could be exactly what the next era prizes most.我女儿尚不知这些道理。但待她长大成人,那些奖赏速度与重复的活计,多半已尽数委派给智能体。届时最紧要的,恰恰是她师长们曾忧心忡忡之事:遇难题时沉心静气的耐心,抽丝剥茧推演到底的毅力,以及何时能断定所得结论可信的决断力。他们唯恐会拖累她后腿的品性,说不定正是下一个时代最珍视的宝贝。
So no, I’m not worried about the timed test. I hope she grows up in a world where software carries the toil and people are freed for the work that is unmistakably ours—to think, to judge, to create, to care for one another. That is the future I want agents to make real. But my hope is not evidence it will happen. The future I just described depends on a single question: can we trust agents to do the work? Trust is earned one task at a time. So, I went looking for evidence of where it’s been earned, and where it hasn’t.所以,我半分也不为限时测验担忧。我只盼她成长于这样一个世道:繁重劳苦尽数由软件担下,人类得以腾出双手,去做那些 unmistakably 属于我们的工作——思考、决断、创造、互相关照。那便是我期望智能体兑现的未来。但我的希冀,却不是它必将成真的佐证。方才所述之未来,全系于一个根本问题:我们能否信任一智能体,托付它完成工作?信任是一桩一桩攒出来的。所以我遍寻证据,看哪些地方信任已立,哪些地方尚且空白。
We partnered with MIT Technology Review Insights on new research that draws directly from the technical leaders building this frontier: not the people talking about it, but the people doing it. We surveyed 300 technical experts across AI, data, and cloud domains, spanning 12 industries and 4 regions of the world, asking them to rank their confidence across 101 of the top tasks. What we got back is the 2026 Agent Confidence Index, an honest map of where agents are delivering real value, so our community can see what’s working and move forward together with conviction.我们与 MIT Technology Review Insights 联手开展新调研,直接采信身处前沿、亲手打造这些技术的领军者之言——不是那些空谈之人,而是实干的践行者。我们共调研了三百位技术专家,横跨AI、数据、云三大领域,覆盖全球十二个行业、四大区域,请他们就一百零一项核心任务,逐一给出信心评分。最终所得,便是2026智能体信心指数。这是一份如实标注智能体已创造真实价值的舆图,愿我等社区同仁能借此看清哪些路径行得通,携手笃定前行。
Learn from where confidence is highest取经信心最盛之处
Across the 101 tasks measured, average confidence already lands at 64 out of 100, and thirty tasks clear 70. The highest scores cluster on work that is both predictable and draining: the late nights, the interruptions, the low-value repetition. Automated report generation leads at 83.5. Boilerplate code generation for new features sits at 82.5—the hours a developer no longer spends rewriting the same patterns, freed for the work that challenges them. Certificate expiration monitoring and renewal, at 81.5, ends the scramble that pulls engineers off high-stakes problems for something entirely routine. Real-time data stream monitoring follows at 80.5, and release note generation from commit history at 79.5—the manual end-of-sprint commit review, gone. This is where frontier teams are already delegating to agents, regularly.所衡量的百零一项任务中,平均信心得分已达六十四分(百分制),其中三十项超过七十分。高分任务集中于可预期、耗神耗力的工作:熬夜值守、突发打扰、低价值重复劳动。其中,自动报告生成以八十三点五分居首;新功能模板代码生成得分八十二点五分——开发者无需再耗费工时重写相同范式,可将精力腾出,投入更具挑战性的工作;证书过期监控与续订得分八十一分半,彻底免去了工程师为常规事务抛下高优难题的窘迫;实时数据流监控得分八十分半,提交历史自动生成发布说明得分七十九点五分——那些手动逐条审阅提交记录、熬到迭代结束的苦差,自此消失无踪。这便是前沿团队早已常态化委派给智能体的工作。
The pattern holds across every discipline. In developer and AI workflows it extends to API client maintenance and code identification; in cloud operations, to ticket routing and cost optimization; in data, to anomaly detection. Wherever it sits in the stack, this is work technical teams now trust agents to own.此规律遍及所有领域。于开发者与AI工作流中,延伸至API客户端维护、代码识别;于云运维中,延伸至工单路由、成本优化;于数据领域,延伸至异常检测。无论身处技术栈哪一层,这类工作技术团队如今皆敢委派给智能体全权处置。
What matters most here isn’t what the data says about the tasks; it’s what it says about the people delegating them. When technical experts believe in something deeply enough to hand it real work, that belief ripples outward. It becomes the recommendation they make to their leadership, the solution they build for their customers, and the culture they create for their teams.此处最紧要的,绝非数据对任务的评价;而是数据对委派者的映照。当技术专家对某事物深信不疑到愿托付实活,这份信念便会层层扩散:化为向决策层提出的建议,化为为客户打造的解决方案,化为为团队铸就的文化。

Even the toughest agent tasks are gaining traction纵是最难啃的智能体任务,如今也渐获认可。
Here’s what strikes me most: the tasks ranked lower on the index are still high in absolute terms. Service mesh configuration and troubleshooting sits at 37.5, database schema migration scripting at 46.5, memory leak detection at 48.5. These sit at the very frontier, the interconnected, high-stakes work where investment and innovation are concentrated right now.最令我惊异的是:指数排名靠后的任务,绝对分值依然不低。服务网格配置与排障得分三十七点五,数据库schema迁移脚本编写得分四十六点五,内存泄漏检测得分四十八点五。这些任务正处于最前沿,是当下投资与创新集中的互联互通、高风险工作。
Consider what they demand. Service mesh configuration touches many systems at once. Database migration carries real stakes, requiring precision across data, application, and infrastructure layers at the same time. Memory leak detection means diving deep into a system’s behavior under load, accounting for conditions that shift from one deployment to the next. These are the challenges that have separated great engineers from exceptional ones—and even here, experts see agents helping. Not carrying the work alone, but contributing where it used to be unthinkable. That confidence is still climbing, and that’s telling.且看这些任务的要求:服务网格配置需同时联动诸多系统;数据库迁移风险极高,需数据、应用、基础设施三层同时精准操作;内存泄漏检测则要潜入系统负载下的深层行为,考量不同部署间不断变化的条件。这些挑战向来是区分优秀工程师与顶尖工程师的分水岭——而即便在此处,专家们也认为智能体可助一臂之力。并非要智能体独担大任,而是在那些原本不敢想象的地方出力。而这股信心仍在攀升,其中意味,不言自明。
We’re shipping new capabilities constantly to support this momentum. Database migration tooling in GitHub Copilot now covers not just scripts but the full application and infrastructure migration story. The Azure Site Reliability Engineering (SRE) Agent brings decades of experience operating Azure at scale and deep profiling capabilities directly into memory analysis and performance diagnosis.我们正不断推出新能力,为这股势头添柴加火。GitHub Copilot 中的数据库迁移工具,如今已覆盖脚本之外的全应用、全基础设施迁移流程。Azure站点可靠性工程(SRE)智能体将数十年大规模运营Azure的经验与深度性能分析能力,直接注入内存分析与性能诊断之中。
Why human judgment remains paramount为何人类判断力仍居首位?
When we asked technical experts how they’re navigating agent adoption, 59% named “keeping humans in the loop” as their top priority—ahead of better observability, ahead of governance documentation, and ahead of everything else. That’s a mark of maturity. Teams moving forward with clarity treat agent oversight as non-negotiable, regardless of how capabilities evolve.我们曾问技术专家,他们在推进智能体落地时首要考量为何,百分之五十九的受访者将“keeping humans in the loop”列为首要优先级,排在可观测性提升、治理文档完善等所有选项之前。这便是成熟的标志。思路清晰的团队,无论能力如何迭代,都将智能体监督视为不可商议的底线。
The boundary itself is straightforward. Agents excel at well-specified, high-volume, reversible work: they synthesize data, automate known workflows, and surface anomalies at a speed and scale no human team could match. The moment a decision becomes high-stakes, context-dependent, or hard to undo, a human signs off. That isn’t a limitation of the technology; it’s the architecture of a trustworthy system.权责边界本身并不复杂。智能体擅长处理边界清晰、规模庞大、可逆的工作:它们能整合数据、自动化已知工作流、以人类团队望尘莫及的速度与规模排查异常。但凡决策涉及高风险、强语境依赖,或是难以撤回,便需人类签字拍板。这并非技术的短板,而是可信赖系统的固有架构。
What’s changing, and what remains underappreciated, is the skill it takes to draw that boundary well: the discipline of full-lifecycle evaluations and guardrails. Success means measuring agent output against intent and keeping behavior inside your business strategy. It’s new territory for most engineering teams, and it’s becoming table stakes for modern software faster than most organizations realize. The good news: the same tools generating the work can help you build the harness. Ask GitHub Copilot to write the evals and it will. Frontier teams are already doing this, and it’s why they’re pulling ahead.变与不变之间,最未被重视的,是划定好这条边界所需的本事:全生命周期评估与防护栏的纪律。所谓成功,便是将智能体的产出与初衷对照衡量,令其行为始终不偏离商业战略的轨道。这对多数工程团队而言是新领域,且它成为现代软件准入门槛的速度,比大多数组织预想的更快。好消息是:生成这些工作的工具,本身就能帮你打造约束框架。你只需让GitHub Copilot编写评估用例,它便能做到。前沿团队早已付诸实践,这便是他们领跑的原因。

Agents are opening career doors for engineering智能体正为工程领域打开职业新通路。
Across system reliability and site operations, evaluations and quality assurance, and data pipeline management, 80% or more of respondents see meaningful career opportunity ahead. We believe this is one of the most significant moments in the history of building software, not because agents replace what technical people do, but because what’s left when they take on the toil is the work that defines a career: the judgment calls, the architectural vision, the reasoning to navigate complexity under pressure. That fluency will define the next generation of technical leadership.于系统可靠性、站点运维、评估与质量保证、数据管道管理等领域,百分之八十以上的受访者都看到了极具意义的职业机遇。我们相信,这是软件发展史上最具分量的时刻之一,并非因为智能体取代了技术人的工作,而是因为它们担下繁重劳苦后,剩下的正是定义职业价值的工作:决断、架构视野、压力下驾驭复杂性的推理能力。这份融会贯通的能力,将定义下一代技术领导力的格局。
We’re living this shift at Microsoft, right alongside our customers. Junior developers are using agents to explore codebases on their own and arriving at mentoring conversations with sharper, more sophisticated questions. Senior engineers are covering more ground because the repetitive work that used to fill their days is now delegated, and the work that’s left is harder, interesting, and consequential. Both are growing into more capable versions of themselves. For me, that’s the outcome I’ve always believed technology could deliver.我们与客户并肩,在微软亲身经历这场变革。初级开发者用智能体自主探索代码库,参与指导交流时提出的问题愈发尖锐、老练;资深工程师的覆盖范围更广,因为往日填满日程的重复工作已尽数委派,余下的工作更难、更有趣、也更具分量。二者皆在成长,成为更强大的自己。于我而言,这正是我始终相信技术能带来的成果。

An integrated approach to intelligence and trust智慧与信任的融合之道
Designing more sophisticated agent systems has made one thing clear: agents thrive in well-integrated environments, working best when your whole stack draws on a single source of truth. The high-confidence tasks are the ones we’ve already figured out; the meaningful frontier is the harder, interconnected work, and that’s exactly where observability, governance, security, and unified intelligence have to operate as one.设计更精密的智能体系统,愈发印证一个道理:智能体在高度整合的环境中才能如鱼得水,唯有整个技术栈均基于单一可信数据源时,方能发挥最大效用。高信心任务早已被我们攻克;真正有意义的前沿,是更难、更互联的工作,而这正是可观测性、治理、安全、统一智能必须合而为一的地方。
Microsoft IQ brings your enterprise context into a single, continuous intelligence layer. Within it, Work IQ builds semantic understanding of how your business operates across email, calendar, meetings, chats, files, people, and collaboration patterns. Such depth of knowledge is the reason technical teams choose us, and it’s what drives my focus and passion in learning how people actually work so their agents get them. My colleague Kim Manis, CVP of Product for Microsoft Fabric, has written specifically about what this means for data professionals, and the integral role of Fabric IQ.Microsoft IQ 将企业上下文整合为单一连续智能层。其中,Work IQ 可构建语义层面的理解,涵盖邮件、日程、会议、聊天、文件、人员、协作模式等企业运营全场景。正是这般深厚的知识积累,让技术团队选择我们,也驱动着我专注且热情地研究人类的真实工作模式,好让智能体更懂用户。我的同事、微软Fabric产品副总裁Kim Manis曾专门撰文,谈及这对数据专业人士的意义,以及Fabric IQ不可或却的作用。
It’s all part of the Microsoft Agent Platform, which is becoming the operating system for enterprise AI at scale. From building in GitHub and contextualizing with Microsoft IQ, to running in Microsoft Foundry and governing in Microsoft Agent 365, Microsoft is uniquely positioned to help customers bring together data, models, agents, and human judgment into a continuously improving and secure system.这一切,皆是微软智能体平台(Microsoft Agent Platform)的一部分。该平台正逐步成为大规模企业AI的操作系统。从在GitHub上构建、用Microsoft IQ添加上下文,到在Microsoft Foundry上运行、用Microsoft Agent 365治理,微软具备独一无二的优势,可助力客户将数据、模型、智能体、人类判断力融为一体,打造持续迭代、安全可靠的系统。
Frontier transformation is being led by builders like you.前沿变革,正由诸位这样的构建者引领。
Next steps:下一步:
- Download The 2026 Agent Confidence Index from our partners at MIT Technology Review Insights. It is a free, ungated deep dive into all 101 tasks, broken out by role and workflow, with the patterns and reasoning behind where confidence is strongest and the frontier is expanding.可至我们的合作伙伴 MIT Technology Review Insights 处下载《2026智能体信心指数》。这份报告完全免费、无需门槛,深度拆解全部一百零一项任务,按角色与工作流分类,不仅标注了信心最强的领域,也阐释了前沿边界扩张背后的规律与逻辑。
- Join us at the AI Engineering World’s Fair (June 29-July 2) where our very own Pablo Castro will keynote, and our teams will offer 16 breakout sessions and 4 labs. Swing by the Microsoft booth as well to explore an interactive 3D visualization of the Index data. We want to hear what’s working for you.诚邀诸位参与AI工程世界博览会(六月廿九至七月二日),我司Pablo Castro将担任主旨演讲嘉宾,团队亦将举办十六场分论坛、四场实操实验室。亦可莅临微软展台,体验指数数据的交互式三维可视化。我们期待聆听诸位在实际应用中的经验之谈。
- Learn more about Microsoft IQ and how it connects across Work IQ, Fabric IQ, Foundry IQ, and the newly announced Web IQ. You can catch up on all the developer innovation from Microsoft Build through Satya Nadella’s keynote, Kyle Daigle’s blog post, and the Microsoft Build CLI.欲深入了解Microsoft IQ,以及它如何与Work IQ、Fabric IQ、Foundry IQ、全新发布的Web IQ互联互通,可查阅微软Build大会的相关内容:Satya Nadella的主旨演讲、Kyle Daigle的博客文章,以及微软Build命令行工具。
What’s Working in Agentic AI智能体AI的落地实践
The 2026 Agent Confidence Index report reveals where agents are trusted, the challenges they face, and what leaders should do next《2026智能体信心指数》报告揭示了智能体的信任分布、面临的挑战,以及领导者接下来的行动方向