Economic Research经济研究

Agentic coding and persistent returns to expertise 代理式编码与专业知识的持续回报

2026年6月16日2026年6月16日
Read in PDF以 PDF 阅读
Agentic coding and persistent returns to expertise

Key findings关键发现

  • Building on prior work, we introduce a framework for studying interactive agentic coding based on a privacy-preserving analysis of ~400,000 Claude Code sessions from between October 2025 and April 2026. We evaluate the composition of tasks, human-AI collaboration, and success rates.在先前工作的基础上,我们引入了一个框架,用于研究交互式代理式编码,基于对2025年10月至2026年4月约40万次Claude Code会话的隐私保护分析。我们评估了任务组成、人机协作以及成功率。
  • In a typical session, people make most of the planning decisions (what to do) and Claude makes most of the execution decisions (how to do it). The greater domain expertise a person brings to a session, the more work Claude does per instruction. On coding tasks, every major occupation succeeds––accomplishes what the person set out to do, with verifiable evidence like passing tests or committed work––at nearly the same rate as software engineers, on average.在典型会话中,人们做大部分规划决策(做什么),Claude负责大部分执行决策(怎么做)。个人带来的领域专业度越高,Claude在每条指令上完成的工作就越多。在编码任务中,每个主要职业的成功率——即实现用户设定目标并有可验证的证据(如通过测试或提交代码)——平均与软件工程师几乎相同。
  • The more domain expertise a person has, the more often the session ends in success—though the gap between intermediate and expert users is modest. Over the seven months we observe, the share of sessions spent debugging fell by nearly half, and usage shifted toward more end-to-end agentic use: deploying and running code, analyzing data, and writing non-code documents.个人的领域专业度越高,会话成功结束的频率越高——尽管中级用户与专家用户之间的差距有限。在我们观察的七个月里,调试会话的比例下降了近一半,使用方式也转向更多端到端的代理式使用:部署和运行代码、分析数据以及撰写非代码文档。
  • Over those seven months, the value of the typical task, which we estimate through a comparison to freelance job postings, rose in almost every kind of work—about 25% on average.在这七个月里,我们通过与自由职业岗位的比较估算的典型任务价值几乎在所有工作类型中都有所上升——平均约提升25%。

Introduction引言

Agentic coding has taken off. The share of GitHub projects with coding agent activity has more than doubled since late 2025,1 and Claude Code users now spend an average of 20 hours per week using the tool.2 Can people without formal coding experience successfully direct an agent through complex technical work? And what will rapid adoption and improvement of these tools mean for knowledge work broadly? While we don’t have full answers to these questions yet, we look to Claude Code usage data for early signals.代理式编码正在快速发展。自2025年底以来,GitHub 项目中出现编码代理活动的比例翻了一番以上,Claude Code 用户现在平均每周使用该工具 20 小时。没有正式编码经验的人能否成功指挥代理完成复杂技术工作?这些工具的快速采纳和改进将对知识工作产生何种广泛影响?虽然我们尚未对这些问题给出完整答案,但我们从 Claude Code 的使用数据中寻找早期信号。

This report provides evidence on how Claude Code is used in practice, based on a privacy-preserving analysis of ~400,000 interactive sessions from ~235,000 people between October 2025 and April 2026. It builds on prior work focused on measures of autonomy in Claude Code sessions, and how Claude Code is changing work at Anthropic.3 Here, we introduce a framework for describing interactive AI coding-assistant usage: what kind of work is being done, who is doing it, and whether it succeeds. We focus on Claude Code usage through a command-line interface (CLI), Claude.ai, or the Claude Code desktop app.4 By tracking how agentic coding usage changes as models get more capable, we can better understand how these tools affect the labor market for coding professionals and knowledge workers.本报告基于对2025年10月至2026年4月约40万次交互式会话、约23.5万人的隐私保护分析,提供了 Claude Code 实际使用情况的证据。它在先前关注 Claude Code 会话自主性度量以及 Claude Code 如何改变 Anthropic 工作的研究基础上进行扩展。我们在此引入一个框架来描述交互式 AI 编码助理的使用:正在进行何种工作、由谁完成以及是否成功。我们重点关注通过命令行界面(CLI)、Claude.ai 或 Claude Code 桌面应用的使用。通过追踪代理式编码使用随模型能力提升的变化,我们可以更好地理解这些工具对编码专业人士和知识工作者劳动力市场的影响。

What happens on Claude Code may be a preview of where knowledge work is headed, as agents become embedded in non-coding work. We find that Claude is handling more complex and more valuable tasks. At the same time, there remains a clear division of labor in agentic coding: People decide what to build, and the agent decides how to build it.Claude Code 上发生的事情可能预示着知识工作未来的走向,因为代理正嵌入非编码工作中。我们发现 Claude 正在处理更复杂且更有价值的任务。同时,代理式编码仍然存在明确的劳动分工:人决定做什么,代理决定怎么做。

We also see evidence that domain expertise, and not coding proficiency, amplifies effective use of the tool. In particular, domain experts succeed more often, and more easily recover from errors and misunderstandings. However, the gap between experts and intermediates is modest—suggesting that proficiency in a domain is enough to use the tool almost as effectively as those with deep mastery.我们还看到证据表明,领域专业度而非编码熟练度提升了工具的有效使用。具体而言,领域专家更常成功,并且更容易从错误和误解中恢复。然而,专家与中级用户之间的差距有限——这表明只要具备领域熟练度,就能几乎像深度掌握者一样有效使用该工具。

These findings give us an early read on possible transitions in the labor market. In our data, success is determined by how well a person understands the problem they are trying to solve, not whether they’re trained in coding. If these patterns hold across the economy, it suggests that while agentic coding tools may be absorbing some implementation-heavy work, they are also rewarding those with firm understanding of the problems they solve on the job. Coding agents are not substituting for domain expertise—the more understanding a worker brings to an agent, the more quality work the agent is able to do.这些发现为劳动力市场可能的转变提供了早期洞察。在我们的数据中,成功取决于个人对所要解决问题的理解程度,而不是是否受过编码训练。如果这些模式在整个经济中成立,这意味着虽然代理式编码工具可能吸收了一些实现层面的工作,但它们也在奖励那些对工作中解决的问题有深入理解的人。编码代理并未取代领域专业度——工作者带给代理的理解越多,代理能够完成的高质量工作就越多。


The division of labor
劳动分工

What people use Claude Code for人们使用 Claude Code 的用途

To understand what people are using Claude Code for, we classify each session into one of nine work modes—the single activity that best describes what the session is trying to accomplish.5 Four modes involve writing or maintaining code directly: building something new, fixing something broken, testing code, and orchestrating other agents or automated pipelines. Another category is operating software—deploying, configuring, running pipelines, monitoring systems. Two categories are more about working out what to do: understanding how an existing system works, and planning a change before making it. And two take actions unrelated to code, or where code is incidental to the final product: analyzing data, and communicating via presentations and other prose-based documents.为了了解人们使用 Claude Code 的具体目的,我们将每个会话分类为九种工作模式之一——即最能描述该会话目标的单一活动。四种模式直接涉及编写或维护代码:构建新东西、修复损坏的代码、测试代码以及编排其他代理或自动化流水线。另一类是操作软件——部署、配置、运行流水线、监控系统。两类更侧重于弄清要做什么:理解现有系统的工作原理,以及在实施前进行规划。还有两类涉及与代码无关的行动,或代码仅是最终产品的附属:数据分析以及通过演示文稿和其他文字文档进行沟通。

About 56% of sessions consist of writing (25%), fixing (26%), or testing and orchestrating code (5%). Operating software comprises 17%, while 14% of sessions are planning or exploring, and 13% produce analysis or prose (Figure 1).约 56% 的会话属于写代码(25%)、修复(26%)或测试与编排代码(5%)。操作软件占 17%,而 14% 的会话用于规划或探索,13% 产生分析或文字(图 1)。

Figure 1: The nine modes of work
Each interactive session is classified into the single mode that best describes what it is trying to accomplish.
图 1:九种工作模式 每个交互式会话被归类为最能描述其目标的单一模式。


We classify each session by having a model read its transcript, then using our privacy-preserving analysis tool, we check them against telemetry that's recorded automatically for every session, including whether any lines of code were added or deleted. The two sources have high agreement—for instance, more than 90% of sessions our classifier labeled as creating or modifying code showed code changes in the telemetry. See the Appendix for details.
我们让模型读取会话记录进行分类,然后使用我们的隐私保护分析工具,将其与自动记录的遥测数据(包括是否有代码行被添加或删除)进行比对。两者高度一致——例如,超过 90% 被分类为创建或修改代码的会话在遥测中显示了代码变更。详情见附录。

Who decides what谁决定什么

How autonomous is Claude Code? Capability evaluations suggest the ceiling is high and rising: on benchmarks such as METR's time-horizon evaluations, frontier models can now complete software tasks that would take a person hours, autonomously working through obstacles along the way. But what does usage actually look like in practice? Here, we look at how much steering is done by the person and by Claude in real sessions.Claude Code 的自主程度如何?能力评估表明上限很高且在上升:在 METR 的时间视野评估等基准上,前沿模型现在可以自主完成需要人类数小时才能完成的软件任务,并在过程中自行克服障碍。但实际使用情况到底是怎样的?在这里,我们观察真实会话中人和 Claude 各自的引导程度。

We investigate this question from two angles. First, we focus on the extent to which people are entrusting decisions to Claude, and second we look at how many actions they give to Claude. To understand the division of decision-making in a session, we build a privacy-preserving decision attribution classifier based on the content of a session. We ask a classifier to list all the meaningful decisions in a session. We separate these decisions into planning (what to do, which approach to take, what counts as done) and execution (which files to change, what code to write, what language to write in, which commands to run). The classifier then attributes each decision to Claude or to the user, giving every session two numbers: the user's share of planning decisions and the user's share of execution decisions.我们从两个角度探讨此问题。首先,关注人们将决策委托给 Claude 的程度;其次,关注他们交给 Claude 的行动数量。为了解会话中的决策分配,我们基于会话内容构建了一个隐私保护的决策归因分类器。分类器列出会话中的所有有意义决策,并将其分为规划(做什么、采用哪种方法、何为完成)和执行(修改哪些文件、写什么代码、使用何种语言、运行哪些命令)。随后,分类器将每个决策归因于 Claude 或用户,给每个会话两个数值:用户在规划决策中的占比和在执行决策中的占比。

On average, people make about 70% of the planning decisions but only 20% of the execution decisions (Figure 2). In practice, there is a clear division of labor in agentic coding––people decide what to build, and the agent decides how to build it.平均而言,人们做出约 70% 的规划决策,但仅 20% 的执行决策(图 2)。实际上,代理式编码中存在明确的劳动分工——人决定要构建什么,代理决定如何构建。

To understand the delegation of actions in a session, we look at the session’s structure instead of its content. A Claude Code session involves Claude and the user going back and forth trading prompts (from the user) and actions (taken by Claude)––the user writes a prompt and Claude goes off and does some work, and then the user writes another prompt, and so forth. In a typical session, there are about four such turns. In our historical data from October to April, each prompt the user sends sets off a chain of around 10 actions taken by Claude on average––and sometimes over 100.6 In each turn, Claude reads files, edits code, runs commands, and writes on average 2,400 words of output.为了了解会话中行动的委派情况,我们关注会话的结构而非内容。Claude Code 会话涉及 Claude 与用户来回交换提示(用户提供)和动作(Claude 执行)——用户写提示,Claude 执行工作,然后用户再写提示,如此循环。典型会话约有四轮往返。在我们从十月到四月的历史数据中,每个用户提示平均触发约 10 条 Claude 动作——有时超过 100 条。在每轮中,Claude 会读取文件、编辑代码、运行命令,平均输出约 2,400 字。

How much Claude does between check-ins largely tracks who is making the decisions. When the user keeps control of execution (i.e. makes over 80% of execution decisions), Claude takes fewer actions per turn (about eight actions). And when Claude takes control of planning (i.e. makes over 80% of planning decisions), it takes on the highest number of actions (about 16).Claude 在检查点之间完成的工作量大致取决于谁在做决策。当用户保持对执行的控制(即执行决策占比超过 80%)时,Claude 每轮的动作较少(约八个)。而当 Claude 主导规划(即规划决策占比超过 80%)时,它每轮的动作最多(约 16 个)。

Figure 2: Claude's share of planning and execution decisions
Distribution across sessions of the share of planning decisions (what to do) and execution decisions (how to do it) attributed to Claude rather than the user. In the typical session, the user makes about 70% of planning decisions while Claude makes about 80% of execution decisions.

图 2:Claude 在规划和执行决策中的占比 会话中规划决策(做什么)和执行决策(怎么做)归属 Claude 而非用户的分布。在典型会话中,用户做出约 70% 的规划决策,而 Claude 完成约 80% 的执行决策。

Level of expertise专业水平

From each transcript, Claude rates the user's apparent expertise at the task on a five-point scale from novice to expert. The expertise classifier looks for three signals: how precisely the user frames their directions, what they ask Claude to verify, and whether the user tends to correct Claude or Claude tends to correct the user. Note that expertise is capturing something quite different from job title or general ability, and, crucially, it is task-specific. A senior engineer asking their first Rust question is a beginner at Rust. An accountant who has never used Python, but tells Claude exactly which reconciliation rules a Python script must enforce and catches the edge case it mishandles at month-end close, is an expert at that task.从每个会话记录中,Claude 会对用户在任务中的显性专业度进行五级评分(从新手到专家)。专业度分类器寻找三类信号:用户指令的精确度、他们要求 Claude 验证的内容,以及用户是更常纠正 Claude 还是 Claude 更常纠正用户。请注意,专业度捕捉的内容与职称或一般能力截然不同,且关键在于任务特定性。比如,一位资深工程师第一次提问 Rust,仍算是 Rust 初学者;一位从未使用 Python 的会计师若能准确告诉 Claude 需要执行的对账规则并在月末关闭时捕获边缘案例,则在该任务上属于专家。

The table below shows how we defined each expertise level in the classifier along with an example request from a public dataset of coding agent sessions, SWE-chat. The conversation categorized as Novice gives generic instructions with no implied domain-specific knowledge. The Expert conversation conveys deep knowledge of the codebase and technical environment.下表展示了分类器中每个专业度等级的定义,并附有来自公开编码代理会话数据集 SWE‑chat 的示例请求。标记为新手的对话给出通用指令且不暗示领域知识;标记为专家的对话则展示对代码库和技术环境的深度了解。

Table 1: Expertise classifier
The examples paraphrase, anonymize and condense real sessions labeled by our classifiers. Many of the sessions used in the table come from a public dataset of agentic coding sessions, SWE-chat.
表 1:专业度分类器 示例对真实会话进行改写、匿名化和浓缩,来源于公开的代理式编码会话数据集 SWE‑chat。


We quantify how expertise relates to Claude’s output and activity per prompt. In typical novice sessions, each prompt sets off about five Claude actions and roughly 600 words of output, while expert sessions set off action chains more than twice as long (12 actions) carrying five times the output (3,200 words) (Figure 3). This gap between novice and expert sessions appears within every kind of work and every band of task value.
我们量化了专业度与 Claude 输出及每个提示的动作数量之间的关系。典型的新手会话每个提示触发约五个 Claude 动作,输出约 600 字;而专家会话的动作链长度是其两倍以上(12 个动作),输出是其五倍(3,200 字)(图 3)。这种新手与专家之间的差距在所有工作类型和任务价值区间均有体现。

These measures complement the autonomy measures in our prior report on Claude Code, which tracked how long the agent runs and how often people approve its actions automatically. Our decision attribution measure, by contrast, captures who makes the substantive decisions in a session as a whole, while our measures of output and actions per prompt measure how much autonomous activity from Claude each human prompt sets off.这些度量补充了我们先前报告中对 Claude Code 自主性的衡量——后者追踪代理运行时长以及人们自动批准其动作的频率。相比之下,决策归因度量捕捉的是整个会话中谁在做实质性决策,而每提示的输出和动作度量则衡量每个人类提示触发的 Claude 自主活动量。

Figure 3: Claude does more per prompt for more expert users
Claude produces more actions (left bar) and text output per prompt (right bar) for more expert users. Boxes span the interquartile range (split at the median). Whiskers represent the 5th to 95th percentile. White dots are geometric means. Both upward trends are statistically significant (p < 0.001), as is each adjacent-level step, and they remain significant (at +9% actions and +13% output per expertise level) in a regression controlling for work mode, task value, month, occupation, and model family, with standard errors clustered by user.
图 3:专家用户每提示获得更多 Claude 动作 Claude 为更专家的用户产生更多动作(左柱)和文本输出(右柱)。箱体表示四分位范围(中位数分割),须须表示第 5 至第 95 百分位,白点为几何均值。两条上升趋势均具统计显著性(p < 0.001),以及每相邻层级的差异,并在控制工作模式、任务价值、月份、职业和模型系列的回归中仍保持显著(每提升一级专业度,动作增加 9%,输出增加 13%),标准误按用户聚类。

Who uses Claude Code, and for what谁在使用 Claude Code,使用目的是什么

The users用户

To understand who is doing this work, we infer each user's occupation from the session transcript, mapping it to one of 23 major groups in the Bureau of Labor Statistics’ Standard Occupational Classification (SOC) taxonomy. The classifier is instructed to rely only on signals such as the project context the agent loads at the start of a session, the names and structure of their files, any artifacts they reference (e.g., legal filings, clinical data, financial reports, a curriculum, etc.) and vocabulary they use.7 It is explicitly instructed not to treat the act of coding as evidence of a coding profession. A session is classified into the coding SOC code (Computer and Mathematical Occupations) only when there is clear signal that software or data work is the user’s job. A session in which a lawyer builds a script to automatically flag missing clauses across a folder of contracts is mapped into Legal Occupations, even if the session’s work is primarily software. The session is left unclassified when there is no signal about the user’s occupation.为了解谁在从事这些工作,我们从会话记录中推断每位用户的职业,并将其映射到美国劳工统计局标准职业分类(SOC)中的 23 大类。分类器仅依据会话开始时加载的项目上下文、文件名和结构、引用的文档(如法律文件、临床数据、财务报告、课程材料等)以及使用的词汇等信号进行判断,明确不将编码行为本身视为编码职业的证据。只有当软件或数据工作明显是用户的职业时,才会将会话归入计算机与数学职业(SOC 代码)。例如,律师编写脚本自动标记合同缺失条款的会话仍归入法律职业,即使工作主要涉及软件。若无法获取职业信号,则会话保持未分类。

We were able to infer occupation in about 70% of sessions. Within this set, Computer and Mathematical Occupations, a category which encompasses most software-related jobs, is unsurprisingly the largest group. The next largest are Business and Financial Operations; Arts, Design, and Media; Management; and Life, Physical, and Social Sciences. The fastest-growing non-software occupation groups in our sample are management, sales, and legal occupations.我们约在 70% 的会话中成功推断出职业。在这部分中,计算机与数学职业——涵盖大多数软件相关工作——自然是最大组。其次是商业与金融运营、艺术、设计与媒体、管理以及生命、物理与社会科学。增长最快的非软件职业包括管理、销售和法律职业。

The work工作内容

The composition of the work done with Claude Code changed substantially between October 2025 and April 2026. The clearest change is that the share of sessions spent fixing broken code fell from 33% to 19% (Figure 4). In its place, we saw a greater share of the work that surrounds code. Operating software grew from 14% to 21% of sessions. Writing and data analysis roughly doubled, from about 10% to 20% of sessions.2025年10月至2026年4月期间,Claude Code 所完成工作的构成发生了显著变化。最明显的变化是,修复破损代码的会话比例从 33% 降至 19%(图 4)。取而代之的是围绕代码的工作比例上升。软件运营的会话比例从 14% 增至 21%。写作和数据分析的比例大致翻倍,从约 10% 上升至 20%。

The tasks themselves also grew more valuable. We approximate each session's economic value by asking what the work would cost on a freelance marketplace, calibrated against a public dataset of real postings. By this measure, the estimated value of the average session rose by 27% between October and April. The rise holds across many kinds of work. Building, operating, and fixing-type tasks all grew more valuable by roughly a third or more (about 43%, 34%, and 32% respectively). These price estimates are coarse, so we use them primarily to compare tasks to one another over time, not as dollar values to be read literally.8 For details about the construction of the task estimator, see the Appendix.任务本身的价值也在提升。我们通过将工作在自由职业市场的报价与公开的真实岗位数据进行校准,估算每个会话的经济价值。按此衡量,平均会话价值在十月至四月间上升了 27%。这种增长在多种工作类型中均有体现。构建、运营和修复类任务的价值分别增长约 43%、34% 和 32%。这些价格估算较为粗略,主要用于比较不同任务随时间的相对变化,而非作为精确的美元数值。详情见附录。

Figure 4: The composition and value of Claude Code work, October 2025 to April 2026
Share of sessions in each work mode over the seven-month window. The share of sessions fixing broken code fell from 33% to 19%, while operating software, analyzing data, and writing documents grew.

图 4:Claude Code 工作的构成与价值(2025年10月‑2026年4月) 七个月窗口内各工作模式的会话占比。修复破损代码的会话比例从 33% 降至 19%,而软件运营、数据分析和文档写作的比例则上升。

Success depends on what the user brings成功取决于用户带来的东西

The estimated value of a task is one way to get a sense of how Claude Code is helping people do their work. Another angle is to look at how many sessions are successful, and what characteristics of a session are linked to success. Across all our measures of success, we see a clear pattern: the more expertise a person exhibits in a session, the higher the likelihood of success. Most of the gain is concentrated at the lower end of the expertise scale––the gap between novice sessions and intermediate sessions is bigger than the gap between intermediate and expert.任务价值是评估 Claude Code 帮助人们完成工作的一个视角。另一角度是查看有多少会话成功,以及哪些会话特征与成功相关。所有成功度量均显示出明确的模式:用户在会话中表现出的专业度越高,成功的可能性越大。收益主要集中在专业度的低端——新手会话与中级会话之间的差距大于中级与专家之间的差距。

Before turning to the characteristics of successful sessions, we should be precise about how we measure success. We do not observe users’ real-world outcomes, and we cannot ask them directly whether they got what they wanted out of Claude. Instead, we rely on two complementary transcript-based measures. The first, judged success, comes from a classifier that reads the full transcript and decides whether the person succeeded in doing what they set out to do (with options: succeeded, partially succeeded, failed, no clear goal). Two companion classifiers then rate the strength of the evidence for that judgment to determine verified success. A success signal classifier looks for verifiable evidence of success. In particular, it looks for git activity like commits and pull requests matching the work, as well as test suites passing, and explicit affirmation from the user. It scores the session from "no signal" to “weak signal” (1) to "multiple hard signals” (5). A parallel failure signal scores the evidence that things went wrong—errors, failed tests, retries, the user pushing back on the output. Verified success requires both that the session is judged successful and there is at least one hard verifiable signal of success. For the following analysis, which is focused on the degree of success or failure in a session, we exclude sessions classified as having “no clear goal,” which comprise about 7.7% of our full sample.在讨论成功会话的特征之前,我们需要明确成功的衡量方式。我们无法观察用户的真实世界结果,也不能直接询问他们是否满意 Claude 的输出。因此,我们采用两种互补的基于会话记录的度量。第一种“判断成功”来自一个分类器,它阅读完整记录并判断用户是否实现了预定目标(选项:成功、部分成功、失败、目标不明确)。随后,两 个伴随分类器评估该判断的证据强度,以确定“已验证成功”。成功信号分类器寻找可验证的成功证据,尤其是 git 提交、pull request、通过的测试以及用户的明确肯定。它将会话评分从“无信号”到“弱信号”(1)再到“多个强信号”(5)。相对应的失败信号分类器记录错误、测试失败、重试、用户对输出的负面反馈等证据。已验证成功要求会话被判断为成功且至少有一个强验证信号。后续分析聚焦于会话的成功或失败程度,排除被标记为“目标不明确”的会话,这类会话约占完整样本的 7.7%。

The returns to expertise专业度的回报

So what kinds of sessions are most successful? It turns out that the expertise rating of a session, described above, matters a great deal for the success of a session.那么,哪些会话最成功?事实证明,上文描述的会话专业度评分对成功率影响巨大。

One might worry that expertise isn't the real driver—perhaps experts simply pick different tasks, or differ in other ways. Throughout this section, we partially address this worry by comparing sessions doing the same kind of work, at the same estimated value, in the same month, on the same subject, from people in the same broad occupation group, and ask how outcomes differ by the person’s rated expertise.有人可能担心专业度并非真正驱动因素——也许专家只是选择了不同的任务,或在其他方面有所不同。在本节中,我们通过比较在相同工作类型、相同估计价值、相同月份、相同主题、相同大类职业的会话,来部分检验这一担忧,并观察不同专业度评分的结果差异。

Table 2: Definitions of success and failure derived from classifiers
The examples paraphrase and summarize real sessions from a public dataset of agentic coding interactions, SWE-chat, labeled by our classifiers.
表 2:基于分类器的成功与失败定义 示例对公开的代理式编码交互数据集 SWE‑chat 中的真实会话进行改写和概括,由我们的分类器标记。

Across all of our success measures, the more expertise a person exhibits in a session, the more likely it is that the session succeeds. A novice-rated session reaches our strictest measure, verified success, 15% of the time and at least partial success 77% of the time. A session rated intermediate or up reaches verified success 28-33% of the time and partial success 91-92% of the time (Figure 5).在所有成功度量中,用户在会话中表现出的专业度越高,成功的可能性越大。新手评分的会话在最严格的已验证成功度量下成功率为 15%,至少部分成功率为 77%。中级及以上评分的会话已验证成功率为 28‑33%,部分成功率为 91‑92%(图 5)。

In each measure, most of the gain comes from moving from novice to intermediate; between intermediate and expert, the slope decreases. In the Appendix, we give details about the regressions behind Figure 5.在每项度量中,收益主要来源于从新手到中级的跃迁;中级到专家之间的斜率下降。附录中提供了图 5 背后的回归细节。

Figure 5: Expertise and how sessions end
Session outcomes by the user's rated expertise at the task, on a five-point scale from novice to expert. The left panel includes all sessions. The middle and right panels restrict to sessions that hit trouble (failure signals > 3) and show the share that still end in various definitions of success and failure. Each point is an adjusted rate––we estimate the differences between expertise levels by comparing only sessions that share the same work mode, the same task-value band, the same month, the same task subject, and the same kind of user (software-related occupation or not). Details about the regressions behind these points are in the Appendix. Whiskers are confidence intervals on sample means (most are too small to be visible in this plot). These plots exclude sessions judged by the success outcome classifier to have no clear goal.

图 5:专业度与会话结局 用户在任务上的专业度评分(五级,从新手到专家)与会话结果的关系。左图包含所有会话;中、右图限制在出现问题(失败信号 > 3)的会话,并展示在不同成功/失败定义下的比例。每个点为调整后率——我们仅比较在相同工作模式、相同任务价值区间、相同月份、相同任务主题以及相同用户类型(软件相关职业或非软件)下的会话。回归细节见附录。误差线为样本均值的置信区间(大多数点的误差线太小而不可见)。这些图排除了被成功结果分类器判定为目标不明确的会话。

A similar gradient appears in sessions that run into challenges along the way. We say a session hits trouble when the failure signal records verified evidence of failure. This could be an error, a failed test, multiple attempts to do the same thing, or the user expressing frustration or dissatisfaction. Among sessions that hit trouble, the share that are verified successes rises from 4% for novice-rated sessions to 15% for expert-rated ones, accounting for all the controls described above (Figure 5). Looking at the looser measures, we find that the share of at least partial success is 60% for novice and 80-81% for intermediate through expert sessions.在出现挑战的会话中也呈现类似的梯度。我们将出现问题的会话定义为失败信号记录到已验证的失败证据,这可能是错误、测试失败、多次尝试同一操作或用户表达的沮丧/不满。在出现问题的会话中,已验证成功的比例从新手的 4% 上升到专家的 15%,已控制上述所有变量(图 5)。在宽松的度量下,至少部分成功的比例为新手 60%,中级至专家 80‑81%。

We also track the inverse relationship––expertise versus various measures of failure. Note that in this analysis, the sessions judged as failures are those that do not even partially succeed. We say a troubled session is abandoned if it is judged as failed and zero lines of code are written: 19% of sessions where the user appears to be a novice end abandoned, against 5-7% for everyone else. In other words, the least experienced users are more likely to give up when they are struggling to get the outcome they are after. Part of the value of expertise appears to be the ability to steer the agent in the right direction.9我们还追踪了相反的关系——专业度与各种失败度量的关系。需要注意的是,此分析中被判定为失败的会话是指根本未达到部分成功的会话。我们将出现问题且被判定为失败且未写入任何代码行的会话称为“放弃”。新手会话中有 19% 被放弃,而其他用户仅为 5‑7%。换言之,经验最少的用户在遇到困难时更可能放弃。专业度的价值部分体现在能够将代理引导至正确方向的能力上。


Occupation may matter less than expertise
职业可能不如专业度重要

People in software-related occupations reach verified success in about 30% of their sessions overall, while users from other professions reach verified success about 26% of the time. Among sessions that produce code (i.e., sessions that add or modify at least one line of code), those numbers are 34% and 29% respectively (Figure 6). The gap between software-related occupations and other occupations narrows under our looser definition of success––with both groups reaching at least partial success in code-producing sessions 89% and 88% of the time, respectively. That five-point gap is small, and it has neither widened nor narrowed over seven months, even as the success rates in both groups increased. In code-producing sessions, every one of the ten largest occupations in our dataset lands within seven points of software engineers in terms of their success. Management occupations are highest on verified success, slightly above the software engineering occupations. Their higher verified success rates may reflect management skills that transfer to directing an agent. But they may also partly reflect our measurement: verification rests partially on explicit confirmation in the transcript, and managers may be more likely to communicate when they get what they ask for.10从事软件相关职业的用户在所有会话中约有 30% 达到已验证成功,而其他职业的用户约为 26%。在产生代码的会话(即至少添加或修改一行代码)中,这两个数字分别为 34% 和 29%(图 6)。在我们更宽松的成功定义下,两组在产生代码的会话中至少部分成功的比例分别为 89% 和 88%,差距仅为五个百分点,且在七个月内既未扩大也未缩小,尽管两组的成功率均有所提升。在产生代码的会话中,数据集中排名前十的职业与软件工程师的成功率相差不超过七个百分点。管理类职业在已验证成功率上最高,略高于软件工程师。其更高的已验证成功率可能反映了管理技能在指挥代理时的转移效应,也可能部分源于我们的测量方式:验证部分依赖于会话中的明确确认,而管理者更倾向于在获得所需结果时进行沟通。


Figure 6: Verified and judged success rates in coding sessions by inferred occupation
Share of sessions meeting strict definitions of success––judged success and verified success––among sessions that add or change at least one line of code, by the user's inferred occupational group, for the ten largest groups. Every group is within seven percentage points of software/math users (SOC Code Computer and Mathematical Occupations). Error bars are 95% confidence intervals computed on distinct accounts.

图 6:按推断职业划分的编码会话已验证和判断成功率 在添加或修改至少一行代码的会话中,按用户推断的职业组划分的已验证成功率和判断成功率。十个最大职业组均在软件/数学职业(SOC 代码计算机与数学职业)之内七个百分点范围内。误差条为基于不同账户的 95% 置信区间。

Looking ahead展望未来

The results in this report offer an emerging picture of how agentic coding amplifies some forms of knowledge and skills, while substituting for others. In sessions that produce code, every major occupation succeeds at rates within a few points of those in software-related occupations. It appears that coding agents are making a coding background less relevant to successful programming.本报告的结果提供了一个初步图景,展示了代理式编码如何放大某些知识和技能,同时替代其他技能。在产生代码的会话中,各主要职业的成功率与软件相关职业相差仅几个百分点。看起来,编码代理正在降低编码背景对成功编程的相关性。

At the same time, successful sessions are more likely to exhibit domain expertise. Sessions rated expert reach verified success more than twice as often as those rated novice, and when a session hits trouble, novices abandon the session at several times the rate of everyone else. The shape of the collaboration gives this picture more color—domain experts are able to direct Claude to do more work with each instruction they give. So, the ability to steer Claude toward success comes more from command of a domain than from the ability to write code. A person with such command, in any field, may now be able to do technical work they previously could not. A person without any such expertise will get far less from the same tool. And the gains come mostly from competence, not mastery––a working grasp of the domain captures most of the benefit, while deep specialization adds only a bit more beyond that.与此同时,成功的会话更可能体现领域专业度。专家评分的会话的已验证成功率是新手的两倍以上;当会话出现问题时,新手的放弃率是其他人的数倍。合作模式进一步丰富了这一图景——领域专家能够用每条指令让 Claude 完成更多工作。因此,引导 Claude 成功的能力更多来源于对领域的掌握,而非编码能力。任何领域的拥有此类掌握的人现在都可能完成以前无法完成的技术工作;缺乏此类专业度的人则从同一工具中获益甚少。收益主要来自于能力而非精通——对领域的工作性把握已捕获大部分收益,深度专精仅带来少量额外提升。

These findings are preliminary. As in most of our research, we cannot measure real-world outcomes, like whether code written in a session is actually used or discarded thereafter, or whether it produces an economically valuable artifact. In addition, the non-interactive usage this report excludes is a substantial share of activity. Developing a framework to measure it is a priority for future work. And all of our classifications of sessions depend on a model's reading of the transcript. In the Appendix, we show that our classifiers track independent telemetry in expected directions, and agree with a strong reference model on the majority of sessions. But classifiers remain challenging to validate at scale, and Claude Code sessions add further difficulty, as they may be too long and complex for human labels to serve as ground truth.这些发现仍属初步。与我们的大多数研究一样,我们无法衡量真实世界的结果,例如会话中编写的代码是否实际被使用或随后被丢弃,亦或是否产生了经济价值的成果。此外,本报告排除了非交互式使用,这在整体活动中占有相当份额。开发衡量该部分的框架是未来工作的重点。我们所有的会话分类均依赖模型对记录的阅读。附录中展示了我们的分类器在预期方向上与独立遥测数据保持一致,并在多数会话上与强基准模型保持一致。但在大规模上验证分类器仍具挑战,Claude Code 会话的长度和复杂度也使得人工标签难以作为真实标签。

The picture in this report will be updated as the models, the users, and the division of labor between them change. We hope that these measures will allow us to track consequential shifts as they happen. For instance, if the returns to expertise begin to decrease over time, that would suggest that models are starting to supply the essential judgment that users currently bring, and that the gains from these tools are broadening beyond domain experts. If the share of coding sessions completed successfully by users outside software occupations continues to grow, it could indicate that software production is becoming a part of ordinary work in every field, rather than the product of a single occupation. These shifts would change who benefits from agentic coding, and by how much, and would have implications for what is most valued in the labor market.本报告的图景将随模型、用户以及他们之间的劳动分工变化而更新。我们希望这些度量能够让我们实时追踪重要的转变。例如,如果专业度的回报随时间下降,这将表明模型开始提供用户目前带来的关键判断,工具的收益正向更广泛的用户群扩展。如果非软件职业用户成功完成编码会话的比例继续增长,这可能意味着软件生产正成为各行业普通工作的组成部分,而不再是单一职业的专属。此类转变将改变谁能从代理式编码中受益以及受益程度,并对劳动力市场的价值取向产生影响。

Appendix附录

Available here.
可在此获取。

Citation 引用

@online{hitzig2026agentic,
 author = {Zoe Hitzig and Maxim Massenkoff and Eva Lyubich and Shaoyi Zhang and Ryan Heller and Peter McCrory},
 title = {Agentic coding and persistent returns to expertise},
 date = {2026-06-16},
 year = {2026},
 url = {https://www.anthropic.com/research/claude-code-expertise},
}

Acknowledgements致谢

With acknowledgements to: Jake Eaton, Sarah Pollack, Hanah Ho, Szymon Sacher, Anton Korinek, Santi Ruiz, Kerry Persen, Ankur Rathi, Alex Tamkin, Heather Whitney, Cat Wu, Kacie Jenkins, Jennifer Martinez, Amie Rotherham, Boris Cherny, Eleanor Dorfman, Miles McCain, and Jack Clark.致谢以下人员:Jake Eaton、Sarah Pollack、Hanah Ho、Szymon Sacher、Anton Korinek、Santi Ruiz、Kerry Persen、Ankur Rathi、Alex Tamkin、Heather Whitney、Cat Wu、Kacie Jenkins、Jennifer Martinez、Amie Rotherham、Boris Cherny、Eleanor Dorfman、Miles McCain 和 Jack Clark。


Footnotes脚注

  1. A first study, covering 128,000 public repositories, detected coding-agent activity in an estimated 16-23% of projects as of the end of October 2025. A follow-up study using the same methodology found adoption rates more than twice as high among projects created after that period. Detection of agentic coding activity relies on agent co-authorship tags and configuration files, which likely undercount actual usage.第一项研究覆盖了 12.8 万个公共仓库,估计截至2025年10月底约有 16‑23% 的项目出现编码代理活动。后续使用相同方法的研究发现,创建于此期间之后的项目的采纳率翻了一番以上。检测代理式编码活动依赖于代理共同作者标签和配置文件,可能低估了实际使用情况。
  2. Note that this measures hours in which Claude Code was actively running, not the user’s hands-on time typing to Claude.请注意,此处衡量的是 Claude Code 实际运行的小时数,而非用户手动输入 Claude 的时间。
  3. In addition, Sarkar (2026) and Baumann et al. (2026) have offered lenses through which to understand agentic coding, by studying Cursor IDE sessions and publicly available sessions, respectively.此外,Sarkar(2026)和 Baumann 等(2026)分别通过研究 Cursor IDE 会话和公开可得的会话,提供了理解代理式编码的视角。
  4. Note that we exclude Claude Code usage that runs through third party integrated developer environments, and software development kits. We also therefore exclude sessions in “headless” mode where a user runs a single prompt in the CLI via claude -p “<prompt>” . We exclude this usage since it differs in two key ways––much of it is programmatic, with Claude Code embedded in automated tools and pipelines rather than conversing with a user, and even when a user is present, we do not see a user’s session end-to-end the way we do on the surfaces we include.请注意,我们排除了通过第三方集成开发环境和软件开发工具包运行的 Claude Code 使用。因此,也排除了在“无头”模式下用户通过 CLI 使用 claude -p “<prompt>” 运行单个提示的会话。我们排除此类使用是因为它在两个关键方面不同——大部分是程序化的,Claude Code 嵌入在自动化工具和流水线中,而非与用户对话;即使有用户在场,我们也无法像在本报告所覆盖的界面上那样看到用户的完整会话。
  5. All classifiers in this report use Claude Sonnet 4.6 unless otherwise noted. Details about the classifiers, including their exact full text and validation results, can be found in the Appendix. 本报告中所有分类器均使用 Claude Sonnet 4.6,除非另有说明。分类器的完整文本和验证结果详见附录。
  6. The tail of actions per prompt is long. About 2% of sessions average more than 100 actions per prompt, about 1 in 270 average more than 200, and about 1 in 2,300 average more than 500.每提示的动作数量分布呈长尾。约 2% 的会话平均每提示超过 100 条动作,约 1/270 的会话平均超过 200 条,约 1/2,300 的会话平均超过 500 条。
  7. Like all measures in this report, these inferences are produced using our privacy-preserving analysis tool. No researcher reads individual transcripts, occupation labels are never linked to identifiable users, and we only observe aggregates over a minimum number of distinct users.如本报告中的所有度量,这些推断均使用我们的隐私保护分析工具生成。没有研究人员阅读单个记录,职业标签从不与可识别用户关联,我们仅在满足最小不同用户数量的聚合数据上进行观察。
  8. The estimation approach we take here is intended to get at relative differences in the value of sessions, not absolute value. The dollar amount is based on comparisons to the freelancer market—not salaried work—and comes from an ultimately fuzzy match between the Claude Code session and the job posting. Since the relative estimates will remove any consistent bias from these issues, we place more emphasis there.我们采用的估算方法旨在捕捉会话价值的相对差异,而非绝对价值。美元金额基于对自由职业市场的比较——而非薪资工作——并且最终是通过 Claude Code 会话与岗位发布之间的模糊匹配得出。由于相对估算会消除这些问题的一致性偏差,我们更侧重于相对变化。
  9. Conditioning on trouble selects different sessions for different users. Experts hit trouble less often overall, so the troubled sessions they do have are likely to be on harder problems—using the price estimate of the session as a proxy for the complexity of the session, we see that the average estimated value of a troubled session roughly doubles from the bottom of the expertise scale to the top. Part of the gap in recovery rates may therefore reflect that novices get stuck on routine problems while experts get stuck on challenging hard problems.对出现问题的会话进行条件筛选会导致不同用户的会话不同。专家总体上较少出现问题,因此他们遇到的问题往往更具挑战性——使用会话的价格估算作为复杂度的代理,我们发现出现问题的会话的平均估值从专业度底部到顶部大约翻倍。因此,恢复率的差距部分可能反映了新手卡在常规问题上,而专家卡在更具挑战性的难题上。
  10. Even if the model misclassifies managers, the signals relied upon to determine that the user is a likely manager—perhaps in how tasks are delegated and specified—tend to be associated with greater success. In other words, perhaps acting like a manager confers greater success.即使模型错误地将某些人归类为管理者,决定用户可能是管理者的信号——也许是任务的委派和指定方式——也往往与更高的成功率相关。换言之,表现得像管理者可能带来更高的成功率。



Related content相关内容

Project Fetch: Phase two项目 Fetch:第二阶段

We report results from our latest test of whether Claude can help Anthropic employees perform sophisticated robotics tasks. We found that Claude Opus 4.7, operating without human assistance, was about 20 times faster than the fastest human team at all tasks completed by participants less than a year ago.我们报告了最新测试的结果,检验 Claude 是否能帮助 Anthropic 员工执行复杂的机器人任务。我们发现 Claude Opus 4.7 在无人协助的情况下,比一年前参与者中最快的人类团队快约 20 倍。

Read more阅读更多

Paving the way for agents in biology为生物学中的代理铺路

Read more

Measuring LLMs’ impact on N-day exploits衡量 LLM 对 N 天漏洞的影响

In cybersecurity, a large fraction of real-world harm comes from N-days: vulnerabilities that have already been publicly disclosed, but only patched on some devices. In this post, we evaluate how much large language models can accelerate and automate the process of developing N-day exploits.在网络安全领域,现实世界危害的大部分来源于 N 天漏洞:这些漏洞已公开披露,但仅在部分设备上得到修补。在本文中,我们评估大型语言模型在加速和自动化开发 N 天漏洞方面的能力。

Read more