Subscribe to the Frontier Red Team newsletter订阅前沿红队通讯
Get updates on our latest red-teaming research and findings.获取我们最新的红队研究和发现更新。
Michael Ilie, C. Daniel Freeman, and Kevin K. Troy
In August 2025, we ran an experiment to see how much Claude could help Anthropic employees—who were not robotics experts—perform sophisticated (and amusing) tasks with an off-the-shelf robotic quadruped (henceforth, a robodog). We called this Project Fetch. We found that access to our state-of-the-art model at the time (Claude Opus 4.1) helped one team substantially outperform the other, who had to rely only on the internet and their own ingenuity. The Claude-enabled team got more done, faster.2025年8月,我们进行了一项实验,看看 Claude 能在多大程度上帮助非机器人专家的 Anthropic 员工使用现成的四足机器人(以下简称 robodog)完成复杂且有趣的任务。我们将其称为项目 Fetch。结果显示,使用当时最先进的模型(Claude Opus 4.1)的一支团队显著超越了只能依赖互联网和自身创造力的另一支团队。借助 Claude 的团队完成的工作更多,速度更快。
Before we dragged our colleagues to a warehouse for the experiment, we double checked whether Opus 4.1 could do the tasks entirely on its own. Unquestionably, it could not. Much like our team without Claude, it got hung up on the preliminary task of figuring out how to connect to the robot.在把同事们拉到仓库进行实验之前,我们先确认 Opus 4.1 能否完全自行完成这些任务。答案显而易见:不能。它和没有 Claude 的团队一样,卡在了如何连接机器人这一步。
But AI models are moving fast—even faster than the runaway robodog that almost rammed into one of our human teams back in August.但 AI 模型的进展非常快——甚至比去年8月那只差点撞上我们人类团队的失控 robodog 还快。
We figured it was time to revisit Project Fetch to see if our newer models could outperform the previous generation. Not only did they do that, but Claude Opus 4.7—operating without human assistance—was about 20 times faster than the fastest human team at all tasks completed by our participants less than a year ago.我们认为是时候重新审视项目 Fetch,看看更新的模型是否能超越前一代。结果不仅如此,Claude Opus 4.7 在无人类帮助的情况下,完成任务的速度约为一年内最快人类团队的 20 倍。
This doesn’t mean that LLMs have now solved robotics. Far from it. The latest Claude models still struggled with using the robot to precisely move the beach ball—the “fetching” part of Project Fetch. And none of the tasks in these experiments implicate the more challenging, low-level elements of robotic control, such as developing a specific actuation policy. However, once again, we are seeing a pattern whereby first, models are helpful to humans. Then, humans are helpful to models. Finally, models are largely able to do things themselves. We have seen this in cybersecurity and now the same dynamics are starting to take shape at the intersection of AI and the physical world.这并不意味着大语言模型已经解决了机器人技术。事实恰恰相反。最新的 Claude 模型仍然在精确移动沙滩球——即项目 Fetch 中的“取回”环节——上表现不佳。而且这些实验并未涉及更具挑战性的低层机器人控制元素,例如制定特定的驱动策略。不过,我们再次看到一个模式:首先,模型对人类有帮助;随后,人类对模型有帮助;最终,模型基本能够自行完成任务。我们在网络安全领域已经看到这种趋势,现在同样的动态正开始在 AI 与物理世界的交叉点上显现。
The original Project Fetch had teams of Anthropic employees (randomly assigned to work with or without Claude) do the following steps: operate the robodog using the manufacturer-provided controller, connect to the robodog’s video and lidar sensors, write and operate a program to manually control the robodog, develop a way to monitor the robodog’s path through space, write a program to detect the beach ball, and finally put it all together to autonomously retrieve the ball.最初的项目 Fetch 让 Anthropic 员工(随机分配是否使用 Claude)完成以下步骤:使用制造商提供的控制器操作 robodog,连接 robodog 的视频和激光雷达传感器,编写并运行手动控制 robodog 的程序,开发监测 robodog 在空间中路径的方法,编写检测沙滩球的程序,最后将所有环节整合,实现自主取回球体。
For this autonomous update, we couldn’t ask Claude to use a physical controller, nor did we evaluate the time it took a researcher to use the Claude-programmed controller to retrieve the ball (though we did confirm that it worked as intended). On the remaining subset of tasks, we ran three trials of Opus 4.7 using adaptive thinking with effort set to maximum in Claude Code. We measured the elapsed time for each objective and qualitatively assessed the models’ success.在这次自主更新中,我们无法让 Claude 使用实体控制器,也没有评估研究人员使用 Claude 编程的控制器取回球体所需的时间(尽管我们确认其工作正常)。在剩余任务子集上,我们使用 Claude Code 将努力程度设为最大,运行 Opus 4.7 的自适应思考模式,进行三次试验。我们记录了每个目标的耗时,并对模型的成功情况进行定性评估。
The role of our researcher was limited to plugging a laptop running Claude Code into the robodog, entering the initial prompt, approving commands, and approving the model to go to the next task.研究人员的角色仅限于将运行 Claude Code 的笔记本电脑插入 robodog,输入初始提示,批准指令,以及批准模型进入下一个任务。
Very simply: on every task that was completed by at least one human team in August, Opus 4.7 completed the same task at least ten times faster.1 If you consider the four tasks that were completed by both human teams, Opus 4.7 was, on average, more than 37 times faster than Team Claude-less and more than 18 times faster than Team Claude.非常简单:在 8 月至少有一支人类团队完成的每项任务上,Opus 4.7 的完成速度至少快十倍。如果只看两支人类团队都完成的四项任务,Opus 4.7 平均比无 Claude 团队快 37 倍以上,比有 Claude 团队快 18 倍以上。

The table compares the speed of the original teams (Team Claude and Team Claude-less) to Opus 4.7 on all of the tasks we tested as part of Phase Two.下表比较了原始团队(Claude 团队和无 Claude 团队)与 Opus 4.7 在第二阶段测试的所有任务上的速度。

Whereas the humans struggled to choose between multiple different approaches to interface with the dog’s sensors, Opus 4.7 was able to quickly identify the best path. Much of the code it wrote was effective on the first try (which was not the case for Team Claude or Team Claude-less in the original experiment). Indeed, we can see evidence of Opus 4.7’s efficiency when we look at the volume of code it generated: it was as or more successful than both human teams while producing almost ten times less code than Team Claude.人类在选择与机器人传感器交互的多种方法时往往犹豫不决,而 Opus 4.7 能迅速找出最佳路径。它编写的大部分代码一次就能成功(这在原实验中的两个人类团队并非如此)。事实上,从代码量上可以看出 Opus 4.7 的高效:它的成功率与两支人类团队持平或更高,却产生的代码量几乎只有 Claude 团队的十分之一。

Opus 4.7 was not perfect. For example, it defaulted to using an outdated object detection algorithm. But even then, it was able to work around this and arrive at an effective solution.Opus 4.7 并非完美。例如,它默认使用了过时的目标检测算法。但即便如此,它仍能绕过该问题,找到有效的解决方案。
We observed little within-task variance (in absolute terms) on completion times for steps the model finished. (Though the aforementioned suboptimal algorithm selection is likely why one of the beach ball detection trials took substantially longer than the others.) Overall, for the tasks in this experiment within its capability envelope, Claude is now quite reliable. (See the next section for an analysis of what Claude is still unable to do.)模型完成的各步骤的耗时在任务内部的方差很小(绝对值上)。(前文提到的次优算法选择可能是导致一次沙滩球检测耗时显著更长的原因。)总体而言,在本实验能力范围内,Claude 现在相当可靠。(下一节将分析 Claude 仍无法完成的任务。)

It is worth underscoring (as we did in our previous post) that this progress is not the result of a concerted effort to improve the robotics capabilities of our models. These improvements, like so many others in the history of LLM development, have emerged from much more general scaling.值得强调的是(正如我们在之前的帖子中所说),这些进展并非源于针对机器人能力的专项投入。正如 LLM 发展史上许多其他突破一样,这些提升源自更为通用的规模化。
When using their hands, and with some practice, our humans were able to pilot the robodogs to gently nudge a beach ball back to the home base (a patch of fake grass) where the robots started. This required the ability to quickly perceive if the ball had gone off course, how that error related to the previous command, where the ball was now, and then how to adjust future inputs to more precisely move the ball. This is a kind of closed loop at which people excel (at least after making some mistakes and learning from them).当使用双手并经过一定练习后,我们的人类能够让 robodog 轻轻碰撞沙滩球,使其回到起始的假草坪。这需要快速感知球是否偏离轨道、偏差与上一次指令的关系、球当前所在位置,以及如何调整后续输入以更精确地移动球体。这是一种闭环控制,人类在这方面表现出色(至少在经历了一些错误并从中学习后)。
In our Phase Two experiments, Claude struggled to capture this subtlety. Like the humans who reached the phase of needing to write a program for autonomous beach ball retrieval, Claude was able to move the robot behind the ball and position it to knock the ball back to the starting point. But the efforts to do so were poorly controlled and (again, like our human participants) not successful.在第二阶段实验中,Claude 未能捕捉到这种细微差别。像需要编写自主取球程序的人类一样,Claude 能将机器人移动到球后方并尝试把球击回起点,但控制不佳,未能成功(与我们的人类参与者情况相同)。
One of our researchers with more robotics experience than our Phase One volunteers successfully accomplished the task of programming autonomous fetching. With more time and additional scaffolding, we think it is very likely that current generations of Claude could do the same. What we will be watching for next, though, is the ability of the models to accomplish this final task with the same speed and reliability they displayed on the other elements of Project Fetch.一位拥有比第一阶段志愿者更多机器人经验的研究员成功完成了自主取球的编程任务。若给予更多时间和额外的支撑,我们认为当前一代 Claude 完全有可能实现同样的成果。接下来我们将关注的,是模型能否以同样的速度和可靠性完成这一最终任务。
Writing about Phase One, we emphasized how LLMs could provide uplift to non-expert humans needing to use robots. This is even more true now than before. Models now complete what was previously pair-programming work between humans and models much more quickly by themselves, which means that people can more quickly transition to controlling and using the robots. And for some tasks, a human in the loop controlling the robot may still outstrip the AI model with its (virtual) hand on the D-pad.在第一阶段的文章中,我们强调了 LLM 能为非专家使用机器人提供提升。现在这种提升更为明显。模型现在能够自行完成此前需要人机配合的工作,这意味着人们可以更快地转向控制和使用机器人。而在某些任务中,仍然由人类在环中操作机器人可能会比 AI 模型更有优势。
What is interesting and different is that we now seem much closer to a world where models will be able to use off-the-shelf physical tools with relative ease—at least for limited purposes. This is similar to how AI models used existing software editing tools like string-replace when they made the transition to more agentic coding. We are plausibly entering the early era of physical agentic AI.有趣且不同的是,我们现在似乎更接近于一个模型能够相对轻松使用现成物理工具的世界——至少在有限的用途上。这类似于 AI 模型在转向更具代理性的编码时,使用现有的软件编辑工具(如字符串替换)。我们可能正进入物理代理 AI 的早期阶段。
More research is needed to understand models’ ability to make these physical tools more bespoke, whether by writing control policies tailored to particular tasks or by designing robotic systems. And there may be substantial barriers to this more generalized vision of physically capable and adaptable language models. But as we have seen, apparently large distances in model capability can be traversed quickly. Models building their own software tools might have seemed outlandish not long ago, but it is happening. It would be unwise to rule out the same trajectory in hardware.仍需更多研究来了解模型将这些物理工具定制化的能力,无论是编写针对特定任务的控制策略,还是设计机器人系统。而实现这种更通用的、具备物理能力和适应性的语言模型可能面临重大障碍。但正如我们所见,模型能力的巨大跨越可以迅速实现。模型自行构建软件工具在不久前还显得离奇,如今已在发生。对硬件走同样轨迹的可能性也不应被轻易排除。
Updated Jun 18: Corrected the date of the first phase of Project Fetch. 更新于 6 月 18 日:更正了项目 Fetch 第一阶段的日期。
In cybersecurity, a large fraction of real-world harm comes from N-days: vulnerabilities that have already been publicly disclosed, but only patched on some devices. In this post, we evaluate how much large language models can accelerate and automate the process of developing N-day exploits.在网络安全领域,现实世界的大量危害来源于 N 天漏洞:这些漏洞已公开披露,但仅在部分设备上得到修补。在本文中,我们评估大型语言模型在加速和自动化开发 N 天漏洞过程中的作用。
Read moreGet updates on our latest red-teaming research and findings.获取我们最新的红队研究和发现更新。