The Layers of AI experienceAI体验的层次
Designing beneath the surface表面之下的设计
It’s hard to imagine that it has been only 3 ½ years since ChatGPT was released to the public. We're still so early in the process of understanding how this generative material works, how it incorporates into tasks and journeys, and how it changes what we build and for whom.很难想象ChatGPT向公众发布仅仅三年半时间。我们仍处于理解这种生成式材料如何工作、如何融入任务和流程、以及如何改变我们构建的内容和对象的过程早期。
What is clear though, is we are entering a renaissance of new design roles and opportunities for design influence.但有一点是明确的:我们正在进入一个设计新角色和设计影响力机会的复兴期。

The introduction of generative AI into digital products has upended the interaction model that anchored much of the previous era of design. Great AI products are multi-dimensional. Small changes to one can have an outsized impact on the whole.生成式AI进入数字产品,颠覆了之前设计时代所依赖的交互模型。优秀的AI产品是多维度的。对其中一个维度的微小改变可能对整体产生巨大影响。
Spare me the “design is dead” takes. Design is more important than ever. However, the form of our roles and work is evolving, just as it has before, to meet the new challenges and opportunities presented by our changing medium.别跟我说“设计已死”那套。设计比以往任何时候都更重要。然而,我们角色和工作形式正在演变,就像以前一样,以应对不断变化的媒介带来的新挑战和机遇。
Past is prologue过去皆是序章
The concept of design as a multi-layered domain is not new. What is changing is how complex the system is as a whole, and how deep into the system design can influence.设计作为多领域的概念并不新鲜。变化的是整个系统的复杂性,以及设计能够影响的系统深度。
Deterministic design确定性设计
In the early web, work was often segmented by role. Visual designers owned website UIs; information architects owned site navigation and structure; business stakeholders owned requirements; etc. This reflected the waterfall nature of product development, and often led to workflows where each discipline optimized for their own area of focus rather than the outcomes of the product as a whole.在早期互联网中,工作通常按角色划分。视觉设计师拥有网站UI;信息架构师拥有网站导航和结构;业务利益相关者拥有需求等。这反映了产品开发的瀑布式特性,并常常导致每个学科优化自身关注领域,而非产品整体成果的工作流程。
In 2000, Jesse James Garrett published his seminal essay, The Elements of User Experience, describing an alternative model. He explored how websites were actually composed of multiple planes, each dependent on the others. For example, navigation reflected product strategy, while usability issues in the interface might reveal weaknesses in the underlying architecture.2000年,Jesse James Garrett发表了他的开创性文章《用户体验要素》,描述了一种替代模型。他探讨了网站实际上是如何由多个相互依赖的层面组成的。例如,导航反映了产品策略,而界面中的可用性问题可能揭示底层架构的弱点。
Myopic optimization ultimately harmed the user experience. Garrett argued that designers needed to understand the experience produced by the system as a whole, rather than limiting their responsibility or influence to the layer they directly controlled:短视的优化最终损害了用户体验。Garrett认为设计师需要理解系统整体产生的体验,而不是将责任或影响力局限于自己直接控制的层面:
The user experience development process is all about ensuring that no aspect of the user’s experience with your site happens without your conscious, explicit intent. This means taking into account every possibility of every action the user is likely to take and understanding the user’s expectations at every step of the way through that process用户体验开发过程就是要确保用户与网站交互的每一个方面都经过你有意识、明确的意图。这意味着要考虑用户可能采取的每一个行动的每一种可能性,并理解用户在该过程中每一步的期望。
— Jesse James Garrett, The Elements of User Experience——Jesse James Garrett,《用户体验要素》
The result was a highly deterministic model of design, where the team was responsible for understanding the user’s goals, mapping their journeys, and coordinating decisions across all five planes in order to intentionally shape the final product.结果是一个高度确定性的设计模型,团队负责理解用户目标,绘制用户旅程,并在所有五个层面上协调决策,以有意地塑造最终产品。
Anticipatory design预期性设计

Twenty years later, Jamie Mill revisited this framework as The Elements of Product Design. 二十年后,Jamie Mill重新审视了这一框架,提出了《产品设计要素》。
As products became more algorithmic and adaptive to user data and behavior, it became more difficult to design for every possible use case. 随着产品变得更加算法化并适应于用户数据和行为,为每种可能的用例进行设计变得越来越困难。
Mill’s updated model applied a wider lens and considered the many influences that shape the user experience. Beyond the “solution space” of the product itself, Garrett’s original focus, Mill also considered the “problem space”, where discovery practices reveal user needs and behavior, as well as “the real world,” accounting for constraints, incentives, and existing mental models that shape how the product is understood and used.Mill的更新模型采用了更广阔的视角,考虑了影响用户体验的多种因素。除了Garrett最初关注的“解决方案空间”之外,Mill还考虑了“问题空间”(发现实践揭示用户需求和行为的领域),以及“现实世界”(包括约束、激励和塑造产品理解与使用的现有心智模型)。
This new interpretation reflected an evolution in how we understood the role of design, and who participates in it. Mill recognized that many of the facets that influence how people use and value a product are managed by decisions made outside of the design team, and that product design therefore needed to account for a wider domain of ownership.这种新解释反映了我们对设计角色以及谁参与设计的理解演变。Mill认识到,影响人们使用和评价产品的许多方面是由设计团队之外的决策所管理的,因此产品设计需要考虑更广泛的所有权领域。
This presents product design as more explicitly outcome-oriented than strictly deterministic. The work of design is not merely to define the final delivery, but to anticipate the less predictable conditions around it and facilitate a process that leads to better outcomes for users. 这将产品设计呈现为更明确地以结果为导向,而非严格确定性。设计工作不仅仅是定义最终交付物,而是要预见其周围不太可预测的条件,并促进一个为用户带来更好结果的过程。
The contribution of both Garrett and Mill is that they made the dimensionality of good design tangible. Garrett showed that designers needed to extend their focus beyond the layer of the product they controlled. Mill extended that responsibility beyond the product itself, showing that experience design is also shaped by the user’s context, the product’s domain, and the broader system in which it operates.Garrett和Mill的贡献在于他们使优秀设计的维度变得具体。Garrett表明设计师需要将关注点扩展到他们控制的产品层面之外。Mill将这一责任扩展到产品本身之外,表明体验设计还受到用户环境、产品领域以及其运行所在的更广泛系统的影响。
Probabilistic design and AI experience概率性设计与AI体验

With the advent of generative AI, product systems have become more complex. In some ways, this is an extension of algorithmic products, which already introduced dynamic, personalized experiences. But with AI systems it is no longer only the algorithm that introduces variability; the underlying model itself is probabilistic, creating behaviors and emergent patterns that cannot always be reduced to explicit rules, states, or predefined paths. 随着生成式AI的出现,产品系统变得更加复杂。在某些方面,这是算法产品的延伸,后者已经引入了动态的个性化体验。但在AI系统中,不仅算法引入了可变性;底层模型本身就是概率性的,产生无法总是简化为明确规则、状态或预定义路径的行为和涌现模式。
As a result, every interaction within these products may include traces of decisions, biases, references, and dependencies from the model, its training data, and its available tools, plus any outside context introduced into the interaction. 因此,这些产品中的每一次交互都可能包含来自模型、其训练数据和可用工具的决策、偏见、引用和依赖的痕迹,再加上引入到交互中的任何外部上下文。
We cannot control for every outcome directly through the interface, but we can design the conditions that shape a model’s generation. In that regard, the work of design looks less like specifying every expected state, as Garrett’s model encouraged, and instead closer resembles system design, identifying and manipulating the leverage points in a system1 that exist in the layers below the surface.我们无法通过界面直接控制每个结果,但我们可以设计塑造模型生成的条件。从这个角度看,设计工作看起来不像Garrett模型所鼓励的那样指定每一个预期状态,而是更接近系统设计,识别并操纵系统中存在于表面之下的杠杆点[1]。
We need full-stack designers我们需要全栈设计师
I do not mean that term in the traditional, engineering sense. Designers don’t need to be machine learning engineers, policy experts, or model researchers to build effective AI products. It does mean we need to be multilingual, able to fluently discuss how each layer beneath the interface impacts the user experience, and how to intervene when necessary.我并不是指传统工程意义上的术语。设计师不需要成为机器学习工程师、政策专家或模型研究员来构建有效的AI产品。但这确实意味着我们需要多语言能力,能够流利地讨论界面下的每一层如何影响用户体验,以及在必要时如何进行干预。
Garrett asked designers to look beyond the surface layer they controlled. Mill asked designers to look beyond the product and into the conditions that shaped how it was understood and used. AI asks designers to go one layer deeper again: into the model, the harness, the context, the policies, and the emergent behaviors that produce the experience before it ever reaches the interface.Garrett要求设计师将目光投向自己控制的表面层之外。Mill要求设计师将目光投向产品之外,进入塑造产品理解和使用方式的条件中。AI要求设计师再深入一层:进入模型、约束、上下文、策略以及在体验到达界面之前产生体验的涌现行为。

The layers of AI UXAIUX的层次
AI experience is composed of a set of highly interdependent layers that collectively shape how a product behaves. As the user interacts with the system, each layer may change in form and purpose. Early on, interactions depend heavily on direct instruction from the user. Over time, however, the system takes over, managing the user’s needs through its context of the problem, running independent, constrained by its harness, governing model, and user oversight. AI体验由一组高度相互依赖的层次组成,这些层次共同塑造产品的行为方式。随着用户与系统的交互,每一层的形式和目的都可能发生变化。早期,交互很大程度上依赖于用户的直接指令。然而,随着时间的推移,系统接管,通过其对问题的上下文来管理用户需求,独立运行,受其约束、治理模型和用户监督的限制。
By understanding how each component influences the end experience, designers can better locate where interventions will be most effective at delivering value, supporting human needs, and making the system more legible, accountable, and safe.通过理解每个组件如何影响最终体验,设计师可以更好地定位干预措施在哪些方面最有效,以提供价值、支持人类需求,并使系统更易理解、更负责任和更安全。
The User Interface layer用户界面层
AI design discourse is still heavily weighted toward the surface, exploring the dynamics of chat interfaces along with familiar and novel patterns that connect generative interactions with heuristics and paradigms. AI设计讨论仍然严重偏向表面,探索聊天界面的动态以及将生成式交互与启发式和范式联系起来的熟悉和新颖模式。
This isn’t surprising. The interface is where most people first encounter AI, and generative systems often require an initial input before an interaction can begin. 这并不奇怪。界面是大多数人第一次接触AI的地方,而生成式系统通常需要初始输入才能开始交互。
User interfaces are not going to disappear, but their role changes the deeper into a session a user progresses, supporting the system rather than driving it. It’s likely we’ll see their function and form continue to evolve with the rise of agentic systems, wearables, and other non-traditional products.用户界面不会消失,但随着用户进入会话的深入,它们的角色会发生变化,从驱动系统转向支持系统。我们很可能会看到它们的功能和形式随着自主系统、可穿戴设备和其他非传统产品的兴起而继续演变。
Early in the user journey, AI requires direction from people, guiding its goals, constraints, and other instructions. Users may provide this through, workflows, inline actions, connected services, and other inputs.在用户旅程的早期,AI需要来自人的指导,以引导目标、约束和其他指令。用户可以通过工作流、内联操作、连接的服务和其他输入来提供这些。
We’ve taken to calling these prompts, but prompting is really only one surface for instructing the model. Inline actions, ambient nudges, and user-defined workflows offer a palette of alternatives. A product that relies strictly on prompts has a ceiling for engagement, since it’s inefficient (and annoying) to write long, specific, context-rich instructions with every turn.我们习惯称这些为提示,但提示实际上只是指导模型的一种表面形式。内联操作、环境提示和用户定义的工作流提供了多种替代方案。一个严格依赖提示的产品在参与度上会有上限,因为每次对话都写长篇、具体、上下文丰富的指令既低效又烦人。
In any case, we expect AI products to build context about us over time so they can anticipate our needs rather than wait to be told. The faster a model accurately grasps the user's intent, the faster the system becomes an augmenting utility. When this sub-surface system is working well, the model can act with more autonomy, and the purpose of the interface leans toward oversight, allowing the user to manage and orchestrate the model without requiring constant intervention.无论如何,我们期望AI产品能随着时间积累关于我们的上下文,以便它们能够预见我们的需求,而不是等待被告知。模型越快地准确掌握用户意图,系统就越快成为增强型工具。当这种表面下的系统运作良好时,模型可以更自主地行动,界面的目的转向监督,允许用户管理和编排模型,而无需持续干预。
While traditional systems focus onboarding and early interactions on helping the user learn the product, introducing more advanced features through progressive disclosure as the journey progresses, onboarding into AI products looks less like people learning how to use the system, and more like the system learning how to interpret the user. The better the system’s understanding of the person, the less complication needs to appear in the interface. We’re moving towards progressive autonomy.传统系统在早期交互中专注于帮助用户学习产品,通过渐进式披露随着旅程推进引入更高级功能,而AI产品的入门看起来更像是系统学习如何解读用户,而不是人们学习如何使用系统。系统对用户的理解越好,界面中需要出现的复杂性就越少。我们正在走向渐进式自主。
This is why the debate about AI interfaces cannot be reduced to whether chat is a good or bad surface to anchor on. The right interface depends on the context surrounding the interaction, like how familiar the user is with the domain, how much the AI knows about them, how sensitive the situation is, and how much confidence the system has in its response.这就是为什么关于AI界面的辩论不能简化为聊天是否是好的表面形式。正确的界面取决于交互周围的上下文,比如用户对领域的熟悉程度、AI对用户的了解程度、情况的敏感性以及系统对其响应的信心程度。
As that context changes, the interface may need to evolve as well, even for similar touchpoints. The same task for the same user might require direct instruction early on, but eventually could be served through an autonomous backend process guarded by evals once the system had earned the user’s trust. 随着上下文的变化,界面可能也需要演变,即使对于类似的接触点也是如此。同一用户的同一任务可能在早期需要直接指令,但最终当系统赢得用户信任后,可以通过受评估保护的后台自主流程来提供服务。
Chat can still be an effective surface for this, and should not be discounted, but it’s not a stable state. Interfaces may instead begin to resemble instrument panels, allowing direct inputs but not requiring it.聊天仍然可以是一个有效的表面,不应被轻视,但它不是一个稳定状态。界面可能开始类似于仪表盘,允许直接输入但不强求。
Interface design is therefore becoming less about choosing a single pattern for the use case and more about matching the surface to the state of the relationship between the user and the model at any given time. Behind the scenes, designers need to consider the artifacts an agent may use for shared interactions; the evaluation tools that track the model’s accuracy and flag issues; and the surfaces where people can view and adjust memory, skills, and instructions.因此,界面设计不再是选择单一的用例模式,而是将表面与用户和模型之间在任何时刻的关系状态相匹配。在幕后,设计师需要考虑代理可能用于共享交互的工件;跟踪模型准确性并标记问题的评估工具;以及人们可以查看和调整记忆、技能和指令的表面。
The Context layer上下文层
Below any AI interface sits the context that provides the model with clues about the user’s intent, needs, constraints, and ecosystem. 任何AI界面之下都有上下文,为模型提供关于用户意图、需求、约束和生态系统的线索。
A well-constructed context keeps an AI experience from having to start cold every time a person asks for help. It guides the system to reference what matters about the user, details about their task, plus any surrounding conditions without forcing the person to repeat themselves. We design it deliberately through context engineering, which helps the determine what information should be collected or passed through across interactions.精心构建的上下文可以使AI体验在每次有人请求帮助时不必从头开始。它引导系统引用关于用户的重要信息、任务细节以及周围条件,而无需用户重复。我们通过上下文工程有意地设计它,这有助于确定应该跨交互收集或传递哪些信息。
In that sense, this layer operates as the engine for an AI-powered experience. 从这个意义上说,这一层作为AI驱动体验的引擎运行。
In early interactions, the user helps the model establish an understanding of them through explicit inputs, like descriptions of their goals and concerns, or through imported third-party content and data. Almost immediately, the system begins to generate inferred context from the person’s behavioral patterns, integrated systems, historical interactions, and content, forming the foundation of a working context layer that it can use to interpret future requests.在早期交互中,用户通过显式输入(如对目标和关注的描述)或通过导入的第三方内容和数据,帮助模型建立对用户的理解。几乎立即,系统开始从用户的行为模式、集成系统、历史交互和内容中生成推断的上下文,形成工作上下文层的基础,用于解释未来的请求。
Picture this like slowly exploring a map in a video game. At first most of the terrain is hidden, but as you move through it, the map begins to reveal its topography, its buildings, hazards, and boundaries, becoming more useful as you go. 想象一下在视频游戏中慢慢探索地图。起初大部分地形是隐藏的,但随着你的移动,地图开始显示其地形、建筑、危险和边界,变得越来越有用。
Context works similarly, as the system learns how the user works and what they care about, plus how they prefer the system to interact with them. Gradually, as this surface reveals itself, the AI is able to work more proactively with less direct input from the user, reducing the need for constant instruction as the experience becomes more adaptive and personalized. 上下文的工作原理类似,系统学习用户的工作方式、关心什么以及喜欢系统如何与它们交互。逐渐地,随着这个表面的显现,AI能够更主动地工作,减少用户的直接输入,随着体验变得更加适应性和个性化,减少了持续指令的需求。
An agent should learn, for example, that I prefer certain meetings on Thursday afternoons, that John should usually be invited, that I like shorter drafts for executives, or that a support escalation should be handled with more caution than a routine status update.例如,代理应该了解到我更喜欢在周四下午开某些会议,通常应该邀请John,我更喜欢为高管准备简短草稿,或者支持升级应该比常规状态更新更谨慎地处理。
But a system needs help knowing what context to keep and what to discard. Too much context, through long context windows and bloated memory files, burns through token budgets and can degrade results, a failure commonly called context rot. Too little or unmaintained context allows the system to become inconsistent, unpredictable, or dependent on constant user intervention. 但系统需要知道哪些上下文该保留、哪些该丢弃。过多的上下文,通过长上下文窗口和臃肿的记忆文件,会消耗token预算并降低结果质量,这种失败通常称为上下文腐烂。太少或未维护的上下文会导致系统变得不一致、不可预测或依赖用户持续干预。
Neither situation is good, and both become more serious when the system remembers personal details it should not have, forgets things it should know, or carries forward the wrong context from a user or session.两种情况都不好,当系统记住了不该有的个人细节、忘记了应该知道的事情,或从用户或会话中携带了错误的上下文时,问题会更加严重。
This problem becomes more consequential as agents take a more active role in driving workflows and interacting with data and content on a person’s behalf. Designers need to consider the agent’s experience as well: how it receives context, how it manages goals, when it collaborates with the user, and how visible its actions need to be for review.随着代理在代表用户推动工作流、与数据内容交互方面发挥更积极的作用,这个问题变得更加重要。设计师还需要考虑代理的体验:它如何接收上下文,如何管理目标,何时与用户协作,以及其行动需要多大的可见性以供审查。
The UI and Context layers therefore need to be designed in tight harmony, with consideration for both people and AI agents, how they interact with the user and each other, and how their individual workflows intersect across journeys.因此,UI层和上下文层需要紧密协调设计,同时考虑人和AI代理,它们如何与用户以及彼此交互,以及它们各自的工作流如何在旅程中交叉。
The Harness layer约束层
As experiences become more headless, more of the product’s work moves out of the visible interface and into context-aware, autonomous background processes. AI systems therefore require an operational layer around the model for processing information and coordinating their actions within defined constraints. This serves as the model’s harness, enabling it to complete tasks independently while remaining governed by permissions and user preferences that promote security and more predictable outcomes.随着体验变得更加无头化,更多的产品工作从可见界面转移到上下文感知的自主后台流程中。因此,AI系统需要一个围绕模型的操作层,用于处理信息并在定义的约束内协调行动。这作为模型的约束,使其能够独立完成任务,同时受促进安全和更可预测结果的权限和用户偏好的约束。
It may seem like this layer is the domain of developer experience or application architecture, but model harnesses are increasingly part of the user experience as well. They shape what the system can know, what it can do, how consistently it behaves, and how much control users have over autonomous work.这看起来可能是开发者体验或应用架构的领域,但模型约束也越来越成为用户体验的一部分。它们塑造系统能知道什么、能做什么、行为一致性如何以及用户对自主工作的控制程度。
In practice, this is the difference between an AI that you can chat with and one that operate as a true collaborator by finding and managing data, drafting responses, routing information, and coordinating actions in pursuit of a goal.在实践中,这就是你可以与之聊天的AI与能够通过查找和管理数据、起草响应、路由信息以及协调行动以实现目标来作为真正协作者的AI之间的区别。
There’s no singular form that a harness might take. It can be relatively simple, orchestrating a single agent’s workflows. Or it can manage a more complex agentive orchestration, where a central agent within the harness deploys and oversees the work of multiple sub-agents in pursuit of a single outcome. In either case, the system is composed of multiple components, designed to coordinate capabilities, manage dependencies, and structure how work moves through the broader AI system.约束没有单一形式。它可以相对简单,编排单个代理的工作流。或者它可能管理更复杂的代理编排,其中约束内的中央代理部署并监督多个子代理的工作以追求单一结果。无论哪种情况,系统都由多个组件组成,旨在协调能力、管理依赖关系以及安排工作如何通过更广泛的AI系统进行。
Connectors determine access rules for the model. People need visibility into what data the system has permission to view and manipulate, in what context, and under what conditions. They also need ways to observe access patterns over time and modify rules when needed. This can follow familiar permission patterns for microphone, camera, location, or contacts, where the reason for access is clear when the permission is requested. 连接器确定模型的访问规则。人们需要了解系统有权查看和操作哪些数据、在什么上下文中以及在什么条件下。他们还需要观察随时间变化的访问模式并在需要时修改规则的方法。这可以遵循麦克风、相机、位置或联系人的熟悉权限模式,在请求权限时访问原因明确。
However, connectors also introduce new UX concerns because access to third-party systems changes the model’s context, bringing external content and data into interactions in ways people may not expect. Designers need to make these relationships visible so people understand not only what is connected, but how those connections shape outputs and ongoing behavior.然而,连接器也引入了新的用户体验问题,因为对第三方系统的访问改变了模型的上下文,以人们可能没有预料的方式将外部内容和数据带入交互中。设计师需要使这些关系可见,以便人们不仅了解连接了什么,还了解这些连接如何塑造输出和持续行为。
Tools determine what actions AI can take within the data and context it has access to. These might include reading and writing emails, updating records, or triggering a workflow. If tool permissions are too loose, models can take actions that lead to unintended consequences downstream, which the user may not discover until after the fact. Conversely, if permissions are too restricted, it’s difficult for the agent to perform advanced capabilities without constant user intervention. This may be useful early in the user journey, but over time may lead to missed expectations of performance. 工具确定AI可以在其有权访问的数据和上下文中采取哪些行动。这些可能包括读写电子邮件、更新记录或触发工作流。如果工具权限过于宽松,模型可能采取导致下游意外后果的行动,用户可能事后才发现。相反,如果权限过于严格,代理很难在无需用户持续干预的情况下执行高级功能。这在用户旅程早期可能有用,但随着时间的推移可能会导致性能期望未达成。
The flexibility of tool use affects how much independence the user grants or expects from the model, which directly impacts the quality of outcomes the system can deliver through advanced use. Designers need to construct the product system so tool use is appropriate to the context and risk of the situation where it’s called, calibrating autonomy over time, and surfacing more advanced functionality in a way that leads to engagement instead of mistrust.工具使用的灵活性影响用户授予或期望模型的独立性,这直接影响系统通过高级使用所交付的结果质量。设计师需要构建产品系统,使工具使用适合调用的情境和风险,随时间校准自主性,并以促进参与而非不信任的方式呈现更高级功能。
Skills provide models with reusable working knowledge, such as methods and rules for processing information, required formats and criteria, and overall task instructions. Designers may help determine which skills should be available out of the box, balancing functionality with comprehension. By mapping the journeys and services that underpin the AI interaction, designers can also help determine when to introduce new skills, and how to teach users to construct their own so they understand the downstream effects.技能为模型提供可重复使用的工作知识,例如处理信息的方法和规则、所需格式和标准以及整体任务指令。设计师可以帮助确定哪些技能应该开箱即用,平衡功能性与理解度。通过映射支撑AI交互的旅程和服务,设计师还可以帮助确定何时引入新技能,以及如何教用户构建自己的技能,以便他们理解下游影响。
Since skills have an opinionated impact on the model’s behavior and output, designers should ensure users have visibility and control over which skills are active, what assumptions they contain, and how they are likely to affect the model’s results. When implemented gracefully, skills can help users feel empowered and in control. Otherwise, they can become confusing or overwhelming, particularly to earlier users who haven’t learned the model well enough to understand how to manage it.由于技能对模型的行为和输出有倾向性影响,设计师应确保用户能够看到并控制哪些技能活跃、它们包含什么假设以及它们可能如何影响模型的结果。优雅地实现技能可以帮助用户感到有能力并处于控制之中。否则,它们可能会变得混乱或压倒性,特别是对于还未能充分理解模型如何管理的早期用户。
Agents are autonomous systems that combine skills, tools, and data access, pointed at specific goals to produce outcomes with increasing independence and coordination. They work within loops of delegated responsibility, taking on tasks that extend beyond single interactions or isolated capabilities. Agentic UX is emerging as a discipline in itself because these systems often involve multiple coordinated processes operating across layers of autonomy, introducing new challenges around orchestration, oversight, and emergent behavior.代理是结合了技能、工具和数据访问的自主系统,针对特定目标以日益独立和协调的方式产生结果。它们在委托责任的循环中工作,承担超出单一交互或孤立能力的任务。代理性用户体验正在成为一门学科,因为这些系统通常涉及多个跨自主层协调的流程,带来编排、监督和涌现行为方面的新挑战。
This increase in autonomy underpins the changes at the context and UI layers, as the user experience shifts from directing actions for a single model to supervising agentic systems. A good agent experience makes autonomous work feel orchestrated, allowing users to observe and interrupt the model when needed without requiring them to micromanage every step. Designers need to consider not only how users define goals and constraints, but also how agents coordinate actions, manage objectives, and maintain alignment across multiple surfaces.这种自主性的增加支撑了上下文层和UI层的变化,用户体验从指挥单一模型的动作转向监督代理系统。良好的代理体验使自主工作感觉像经过编排,允许用户在需要时观察和中断模型,而无需微观管理每一步。设计师需要考虑的不仅是用户如何定义目标和约束,还有代理如何协调行动、管理目标以及在多个表面维持一致性。
Together, connectors, tools, skills, and agents form the operational surface of AI systems. They define the boundary between human intent and machine execution.连接器、工具、技能和代理共同构成了AI系统的操作表面。它们定义了人类意图与机器执行之间的边界。
The Model layer模型层
When most people hear the word model, they typically think about recognizable flagship systems like GPT, Claude, Gemini, Grok, and others. To a lay user, these systems may appear interchangeable, but the landscape of AI models is far broader, covering small and large models; general-purpose or vertical; and open and proprietary options. 当大多数人听到“模型”这个词时,他们通常会想到知名的旗舰系统,如GPT、Claude、Gemini、Grok等。对于普通用户,这些系统可能看起来可互换,但AI模型的领域要广泛得多,包括小型和大型模型;通用或垂直;以及开放和专有选项。
For designers, the point is that these differences aren’t arbitrary or only technical. Each model carries distinct characteristics into the end product, like changes in tone or personality, tolerances for risk or ambiguity, general reliability, and other traits that could be good or bad depending on the circumstances. These differences persist across labs and providers, and between different models produced by the same entity.对于设计师来说,关键在于这些差异并非随意或仅技术性的。每个模型都给最终产品带来不同的特征,比如语气或个性的变化、对风险或模糊性的容忍度、整体可靠性以及其他根据情况可能好或坏的特征。这些差异在实验室和提供商之间持续存在,并且同一实体产生的不同模型之间也存在差异。
Models are first and foremost a reflection of their training, including the data, tuning, learned weights, and reinforcement methods used to shape their character. This in turn impacts what it knows by default, how it responds in different situations and contexts, what it avoids, and which assumptions it carries into each interaction. A model trained to focus on reasoning is a poor solution for fast-moving, low-risk environments where latency is costly, just as a faster model may produce a more fluid experience, but with less nuance or reliability. Understanding these differences helps designers anticipate how the end experience will shift depending on the model selected.模型首先反映了其训练,包括用于塑造其特性的数据、调优、学习权重和强化方法。这反过来影响它默认知道什么、在不同情境和环境中如何响应、它避免什么以及它带入每次交互的假设。一个专注于推理的模型不适合快速、低风险、对延迟敏感的环境,就像更快的模型可能产生更流畅但不够细致或可靠的体验一样。理解这些差异有助于设计师预测根据所选模型最终体验将如何变化。
The training of each model also impacts its capabilities, which define what a model is specifically designed to do. Depending on how it was built, a model may excel at reasoning or speed; it may perform better on certain tasks like writing, coding, tool use, or multimodal understanding; and it may therefore work better in conjunction with different tools and domains than others. Designers can influence task design, which determines the work that should be delegated to the model, what should stay with the user, and how the interface can narrow the task so the model can perform well.每个模型的训练也影响其能力,这些能力定义了模型专门设计做什么。根据构建方式,模型可能在推理或速度上表现出色;在某些任务(如写作、编码、工具使用或多模态理解)上表现更好;因此可能更适合与不同的工具和领域结合使用。设计师可以影响任务设计,这决定了应该委托给模型的工作、什么应该留给用户,以及界面如何缩小任务范围以便模型表现良好。
A powerful model can still be a poor fit for a user’s specific need if its capabilities don’t match. Determining whether and how to offer users choices around model delegation is a sensitive aspect of the user experience. It’s not reasonable to expect users to have a broad understanding of the model landscape in order to get good results, but pre-set modes and other parameters can disguise this level of control through the interface.一个强大的模型如果其能力不匹配,仍然可能不适合用户的特定需求。确定是否以及如何为用户提供有关模型委派的选择是用户体验的一个敏感方面。期望用户对模型领域有广泛了解才能获得好结果是不合理的,但预设模式和其他参数可以通过界面隐藏这种控制级别。
Alternatively, model behavior can be designed by defining these tradeoffs up front. A reasoning model can be configured to accept different effort levels, swapping depth and accuracy against latency and cost depending on the circumstances. Or, a creative model can be programmed to accept a different number of turns in its generation, where a smaller number might enable draft mode, giving users the ability to iterate while managing token spend. Latency, verbosity, confidence, refusal patterns, creativity, consistency, and reasoning depth are behaviors that can be tuned, contributing to the distinct feel of the product in use.或者,可以通过预先定义这些权衡来设计模型行为。推理模型可以配置为接受不同的努力水平,根据情况交换深度和准确性与延迟和成本。或者,创意模型可以被编程为接受其生成中的不同步骤数,其中较少步骤可能启用草稿模式,使用户能够在管理token支出的同时进行迭代。延迟、冗长、置信度、拒绝模式、创造力、一致性和推理深度是可以调整的行为,有助于形成使用中的独特产品感觉。
Because models are the primary material of AI products, designers require enough fluency around their attributes to reason about their tradeoffs. They do not need to train the models themselves, but the better they understand how models behave, the more effectively they can harness them for different tasks and ensure the product leverages their strengths and constrains their risks.由于模型是AI产品的主要材料,设计师需要对其属性有足够的流利度,以推理其权衡。他们不需要自己训练模型,但越了解模型的行为方式,就能越有效地利用它们完成不同任务,并确保产品发挥其优势并限制其风险。
The Governance layer治理层
The first four layers describe how AI experiences are composed. The lower layers of governance and emergence shift the frame from composition to operation. These are often not owned by the product team, but they directly affect the conditions in which AI products are deployed and used.前四层描述了AI体验的组成。较低层的治理和涌现将框架从组成转向操作。这些通常不是由产品团队拥有,但直接影响AI产品部署和使用的条件。
Policies, regulations, standards, and preferences are all examples of outside forces that directly or indirectly govern the user experience. Each layer is affected in some form, from data retention preferences that impact context storage; to a company’s philosophy reflecting in model behavior; to standards that determine what the team evaluates to be acceptable performance.政策、法规、标准和偏好都是直接影响或间接支配用户体验的外部力量的例子。每一层都会以某种形式受到影响,从影响上下文存储的数据保留偏好,到反映在模型行为中的公司理念,再到决定团队评估为可接受性能的标准。
As a result, governance cannot be treated as separate from the product, even if many of its underlying decisions live in legal, compliance, security, or executive decision-making. Every distinct combination of these decisions can product meaningfully different experiences for two people using the same model.因此,治理不能被视为与产品分离,即使其许多基础决策存在于法律、合规、安全或执行决策中。这些决策的每一个不同组合可能为使用同一模型的两个不同人产生截然不同的体验。
Consider a product that uses a model from Anthropic versus OpenAI. Each company makes different choices about model design and training, as well safety. Those choices show up in the product as interaction patterns: what the system will answer, how cautious it feels, when it refuses, how it explains boundaries, and how much control product teams have over behavior. 考虑一个使用Anthropic与OpenAI模型的产品。每家公司对模型设计和训练以及安全性做出不同选择。这些选择在产品中表现为交互模式:系统会回答什么、感觉有多谨慎、何时拒绝、如何解释边界以及产品团队对行为的控制程度。
Designers cannot treat these constraints as arbitrary. They shape the product and its interactions with every touch points. For example, a model that refuses to take certain action is an interaction, compared with a model that is eager to act. 设计师不能将这些约束视为随意。它们通过每一个接触点塑造产品及其交互。例如,一个拒绝采取某些行动的模型与一个渴望行动的模型相比,本身就是一种交互。
The hardest constraints that need to be accounted for are rules, which include explicit policies, laws, and restrictions that the product must respect. Rules define the boundaries of the system and shape what the product cannot do or must do, like when to disclose information or ask for permission, and where it has to stop.需要考虑到的最困难的约束是规则,包括产品必须遵守的明确政策、法律和限制。规则定义系统的边界,塑造产品不能做什么或必须做什么,比如何时披露信息或请求许可,以及它必须在哪里停止。
Less severely enforced are standards, which define optimal behavior and outcomes. These translate principles like accuracy, fairness, accessibility, safety, and more into criteria that the system can be designed and evaluated against. Customers may enforce standards contractually, but even when they are not hard rules, they provide a useful framework for tuning the model, harness, and product.不太严格执行的是标准,它们定义最佳行为和结果。这些将准确性、公平性、可访问性、安全性等原则转化为可以设计和评估系统的标准。客户可能通过合同执行标准,但即使它们不是硬性规则,它们也为调整模型、约束和产品提供了有用的框架。
Finally, while unenforceable, preferences generate gates and incentives that shift the behavior of models and training systems over time. When Sam Altman publicly announced that GPT-4o had become “too sycophant-y and annoying” and that the company was prioritizing adjustments, he was responding to a mismatch between the model’s tuned personality and what many users wanted from it. At a smaller scale, a user’s preferences for a model’s voice and tone, or its saved memories and autonomy settings will effect how the system behaves within the product experience.最后,尽管不可执行,偏好会产生随时间改变模型和训练系统行为的大门和激励。当Sam Altman公开表示GPT-4o变得“过于谄媚和烦人”并且公司正在优先进行调整时,他是在回应用户期望与模型调整个性之间的不匹配。在较小规模上,用户对模型声音和语调的偏好,或者其保存的记忆和自主性设置,将影响系统在产品体验中的行为方式。
Designers can influence governance directly through service and policy design or direct advocacy, or indirectly by inspiring the preferences of others. Brand and communication design is a particularly effective tool for amplifying how different preferences and regulations may result in different outcomes within AI products.设计师可以通过服务设计和政策设计或直接倡导直接影响治理,或者通过激发他人的偏好间接影响。品牌和传播设计是放大不同偏好和法规如何在AI产品中导致不同结果的特别有效工具。
At a minimum, designers need to understand the governance framework they are operating, which may require different first-time UX, preferences, expectations, and interactions through the use of the product itself.至少,设计师需要了解他们操作的治理框架,这可能需要通过产品本身的使用而不同的首次用户体验、偏好、期望和交互。
Emergence涌现
Finally, AI experiences are affected by emergence, the unexpected behaviors that arise when probabilistic systems operate in real-world contexts.最后,AI体验受到涌现的影响,即当概率系统在现实世界环境中运行时出现的意外行为。
Plainly speaking, there is more we don't know about these models, and more we don't know that we don't know, than what we can confidently explain. Understanding their behavior is necessary to build for and with them. But we often don't know what they are capable of until we put them into play, at which point the unexpected behavior may already have played out in our customer’s use of it.简单地说,我们对这些模型未知的领域,以及我们不知道自己不知道的领域,比我们能自信解释的要多。理解它们的行为对于为它们构建产品以及与之合作是必要的。但我们通常直到把它们投入使用时才知道它们的能力,而此时意外行为可能已经在客户使用中表现出来了。
Models behave differently across sessions and contexts, as well as across users, tools versions, permissions, and more. This variance can be a strength when used intentionally. In creative tools, for example, variation allows the system to generate alternative paths for exploration, or break out of an anchor that is running stale.模型在不同会话和上下文之间表现不同,也在不同用户、工具版本、权限等之间表现出差异。当有意使用时,这种差异可以成为一种优势。例如,在创意工具中,变化允许系统生成替代路径供探索,或者打破僵化的锚点。
At the same time, variance makes AI products harder to debug. A weird generation could be the result of the model’s training, its harness, outside sources or context, or simply the path the interaction took. Designers and product teams may turn to tools like evals, traces, and other observability tools to understand where the behavior drift is originating from and where the harness needs adjustment. But even then, we are not likely to fully dissect these inner workings any time soon.同时,差异使得AI产品更难调试。一个奇怪的生成可能是由于模型的训练、约束、外部源或上下文,或者仅仅是因为交互所走的路径。设计师和产品团队可能会求助于评估、追踪和其他可观察性工具来理解行为漂移的来源以及约束需要调整的地方。但即使如此,我们也不太可能在短期内完全剖析这些内部运作。
This gets stranger when models develop behavior with unclear or indirect origins, like “OpenAI’s “goblins” incident.2 A small personality-training incentive around common tropes related to geeky personalities eventually showed up in a broader pattern of models mentioning goblins in completely irrelevant moments. 当模型发展出起源不明确或间接的行为时,情况变得更加奇怪,比如OpenAI的“地精”事件[2]。一个关于极客性格常见套路的小型个性训练激励最终在模型在完全不相关的时刻提到地精的更广泛模式中出现。
While funny, this also revealed how small changes at the model level can cascade into visible product behavior in ways teams can’t anticipate or proactively respond to. Other examples are cases where models tend to do well on complex tasks but more poorly on simpler ones, or when models seem to glitch out from random tokens.3虽然有趣,但这揭示了模型层面上的小变化如何以团队无法预见或主动应对的方式级联为可见的产品行为。其他例子包括模型在复杂任务上表现良好但在简单任务上表现较差的情况,或者模型似乎因随机令牌出现故障的情况[3]。
Randomness is a necessary part of these experiences, and in fact is a feature and not a bug: uncertainty is part of what gives generative systems their value. The goal of design is not to eliminate variance or unknown behaviors, but rather to design the conditions that either minimize or mitigate these effects, and make them easier to observe, diagnose, and correct where possible.随机性这些体验的必要部分,实际上是一个特性而非缺陷:不确定性是赋予生成系统价值的一部分。设计的目标不是消除差异或未知行为,而是设计最小化或减轻这些影响的条件,并使它们更易于观察、诊断和在可能的情况下纠正。
A few first principles can guide this work. Observability helps teams see what the system is doing and how it’s behaving. Interpretability helps people understand why the system appears to be following a particular path. Provenance helps teams work backward from a generation to identify how it was formed.一些基本原则可以指导这项工作。可观察性帮助团队看到系统在做什么以及如何表现。可解释性帮助人们理解为什么系统似乎遵循特定路径。溯源帮助团队从生成结果反向工作,以识别它是如何形成的。
That makes emergence distinct from the other layers. It is not something designers configure directly. It is something they design around, monitor for, and respond to as the system encounters conditions the team could not fully predict.这使得涌现与其他层不同。它并非由设计师直接配置。它是设计师围绕其进行设计、监控并响应的事物,因为系统遇到了团队无法完全预测的条件。
What this means for design这对设计意味着什么
The expectation going forward should not be that every designer works across every layer. Full-stack AI designers need to have a general fluency across all inputs into the experience, so they can influence, mitigate, or receive the impacts that upstream work has on the end experience.未来的期望不应该是每个设计师都跨越所有层工作。全栈AI设计师需要对体验的所有输入具有一般流利度,以便他们能够影响、减轻或接收上游工作对最终体验的影响。
In his 2017 Design in Tech Report, John Maeda caught onto the trend of designers becoming more technical, but at the same time, he captured the value of traditional design. These roles are complementary, not competitive.在他2017年的《设计科技报告》中,John Maeda捕捉到了设计师变得更加技术化的趋势,但同时也抓住了传统设计的价值。这些角色是互补的,而非竞争的。

Technical designers are more likely to focus on the Model, Harness, and Context levels, and the role is more than just “design engineer”. Other designers may focus further down the stack, as we might see design-specific titles pop up in policy making and emergent research (these roles exist today, under titles like “business designer”, but have not reached critical mass within the industry). And of course, classical design will remain, but its workflows, tools, and outputs will evolve.技术型设计师更可能专注于模型、约束和上下文层面,这个角色不仅仅只是“设计工程师”。其他设计师可能专注于更底层的堆栈,因为我们可能会看到设计特定的职位出现在政策制定和新兴研究中(这些角色今天已经存在,头衔如“商业设计师”,但在行业内尚未达到临界规模)。当然,经典设计将保留,但其工作流、工具和输出将演变。
In part 2 of this series, I’ll explore this trend in design roles in more depth, how it relates to research around the influence of design spanning decades, and how we can prepare ourselves for the future of our work.在本系列的第2部分,我将更深入地探讨设计角色的这一趋势,它如何与跨越数十年的设计影响力研究相关,以及我们如何为未来工作做好准备。
⁂ Emily⁂ Emily
- In this way, AI design shares much in common with Systems Thinking, recognizing that the lack of direct control requires us instead of multiply our force and influence by first studying the system to determine where the leverage of our effort can be applied to the greatest effect. ↩
- OpenAI identified that GPT-5.5 expressed an “odd affinity for goblin metaphors,” seemingly a relic from the training data used to create the default “geeky” voice that users could choose from as the default persona. ↩
- This example shows the AI app Poke sending random, unintelligible messages to a user in the flow of a conversation. Separately, Google AI results once returned 3 pages of the word “there” (and only that word) while searching for showtimes at the New York City planetarium. ↩

