构建有效的 AI 代理

Building Effective AI Agents

Anthropic Anthropic · Anthropic · 2024-12-19 · Anthropic Engineering ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们与数十个跨行业构建 LLM 代理的团队合作过。始终如一的是,最成功的实现使用的是简单、可组合的模式,而不是复杂的框架。在过去的一年里,我们与数十个跨行业构建大型语言模型(LLM)代理的团队合作过。始终如一的是,最成功的实现并没有使用复杂的框架或专门的库。相反,他们使用简单、可组合的模式进行构建。在这篇文章中,我们分享了从与客户合作以及自己构建代理中学到的经验,并为开发者提供了构建有效代理的实用建议。

We've worked with dozens of teams building LLM agents across industries. Consistently, the most successful implementations use simple, composable patterns rather than complex frameworks. Over the past year, we've worked with dozens of teams building large language model (LLM) agents across industries. Consistently, the most successful implementations weren't using complex frameworks or specialized libraries. Instead, they were building with simple, composable patterns. In this post, we share what we’ve learned from working with our customers and building agents ourselves, and give practical advice for developers on building effective agents.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 14)

全文 · Full text(逐段中英对照)

构建高效智能体 Building effective agents

我们与数十个跨行业构建 LLM 智能体的团队合作过。最成功的实现始终使用简单、可组合的模式,而非复杂的框架。

We've worked with dozens of teams building LLM agents across industries. Consistently, the most successful implementations use simple, composable patterns rather than complex frameworks.

在过去一年中,我们与数十个跨行业构建大语言模型(LLM)智能体的团队合作过。最成功的实现始终没有使用复杂的框架或专门的库,而是采用简单、可组合的模式进行构建。

Over the past year, we've worked with dozens of teams building large language model (LLM) agents across industries. Consistently, the most successful implementations weren't using complex frameworks or specialized libraries. Instead, they were building with simple, composable patterns.

在这篇文章中,我们将分享从客户合作及自身构建智能体过程中学到的经验,并为开发者提供构建高效智能体的实用建议。

In this post, we share what we’ve learned from working with our customers and building agents ourselves, and give practical advice for developers on building effective agents.

什么是智能体? What are agents?

"智能体"可以有多种定义。一些客户将智能体定义为完全自主的系统,能够长时间独立运行,使用各种工具完成复杂任务。另一些客户则用这个词来描述更遵循预设工作流程的规范性实现。在 Anthropic,我们将所有这些变体归类为智能体式系统,但在工作流和智能体之间划出了一个重要的架构区别:

"Agent" can be defined in several ways. Some customers define agents as fully autonomous systems that operate independently over extended periods, using various tools to accomplish complex tasks. Others use the term to describe more prescriptive implementations that follow predefined workflows. At Anthropic, we categorize all these variations as agentic systems, but draw an important architectural distinction between workflowsandagents:

* 工作流是通过预定义代码路径来编排 LLM 和工具的系统。

* Workflows are systems where LLMs and tools are orchestrated through predefined code paths.

* 而智能体则是 LLM 动态地指导自身过程和工具使用的系统,保持对如何完成任务的控制。

* Agents, on the other hand, are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks.

下面,我们将详细探讨这两种类型的智能体式系统。在附录 1(“实践中的智能体”)中,我们描述了客户发现使用这类系统特别有价值的两个领域。

Below, we will explore both types of agentic systems in detail. In Appendix 1 (“Agents in Practice”), we describe two domains where customers have found particular value in using these kinds of systems.

何时(以及何时不)使用智能体 When (and when not) to use agents

在使用大语言模型构建应用时,我们建议尽可能找到最简单的解决方案,仅在需要时才增加复杂性。这可能意味着根本不构建智能体系统。智能体系统通常以延迟和成本换取更好的任务性能,您应该考虑这种权衡何时有意义。

When building applications with LLMs, we recommend finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all. Agentic systems often trade latency and cost for better task performance, and you should consider when this tradeoff makes sense.

当需要更多复杂性时,工作流为定义明确的任务提供了可预测性和一致性,而智能体则在需要大规模灵活性和模型驱动决策时是更好的选择。然而,对于许多应用来说,通过检索和上下文示例优化单个大语言模型调用通常就足够了。

When more complexity is warranted, workflows offer predictability and consistency for well-defined tasks, whereas agents are the better option when flexibility and model-driven decision-making are needed at scale. For many applications, however, optimizing single LLM calls with retrieval and in-context examples is usually enough.

何时以及如何使用框架 When and how to use frameworks

有许多框架可以更轻松地实现智能体式系统,包括:

There are many frameworks that make agentic systems easier to implement, including:

* Rivet,一个拖放式图形用户界面(GUI)LLM 工作流构建器;以及

* Rivet, a drag and drop GUI LLM workflow builder; and

* Vellum,另一个用于构建和测试复杂工作流的 GUI 工具。

* Vellum, another GUI tool for building and testing complex workflows.

这些框架通过简化标准的底层任务(如调用 LLM、定义和解析工具以及链式调用)使入门变得容易。然而,它们通常会创建额外的抽象层,这些抽象层可能会掩盖底层的提示和响应,使其更难调试。它们也可能诱使人们增加复杂性,而更简单的设置就足够了。

These frameworks make it easy to get started by simplifying standard low-level tasks like calling LLMs, defining and parsing tools, and chaining calls together. However, they often create extra layers of abstraction that can obscure the underlying prompts ​​and responses, making them harder to debug. They can also make it tempting to add complexity when a simpler setup would suffice.

我们建议开发者直接使用 LLM API:许多模式可以用几行代码实现。如果你确实使用框架,请确保你理解底层代码。对底层机制的错误假设是客户错误的常见来源。

We suggest that developers start by using LLM APIs directly: many patterns can be implemented in a few lines of code. If you do use a framework, ensure you understand the underlying code. Incorrect assumptions about what's under the hood are a common source of customer error.

请参阅我们的食谱以获取一些示例实现。

See our cookbook for some sample implementations.

构建模块、工作流与智能体 Building blocks, workflows, and agents

在本节中,我们将探讨在生产环境中常见的智能体系统模式。我们从基础构建模块——增强型大语言模型——开始,逐步增加复杂度,从简单的组合工作流到自主智能体。

In this section, we’ll explore the common patterns for agentic systems we’ve seen in production. We'll start with our foundational building block—the augmented LLM—and progressively increase complexity, from simple compositional workflows to autonomous agents.

构建模块:增强型 LLM Building block: The augmented LLM

智能体系统的基本构建模块是经过检索、工具和记忆等增强的 LLM。我们当前的模型能够主动使用这些能力——生成自己的搜索查询、选择合适的工具,并决定保留哪些信息。

The basic building block of agentic systems is an LLM enhanced with augmentations such as retrieval, tools, and memory. Our current models can actively use these capabilities—generating their own search queries, selecting appropriate tools, and determining what information to retain.

我们建议关注实现的两个关键方面:根据你的具体用例定制这些能力,并确保它们为你的 LLM 提供一个简单且文档完善的接口。虽然实现这些增强有多种方式,但一种方法是通过我们最近发布的模型上下文协议,该协议允许开发者通过简单的客户端实现与不断增长的第三方工具生态系统集成。

We recommend focusing on two key aspects of the implementation: tailoring these capabilities to your specific use case and ensuring they provide an easy, well-documented interface for your LLM. While there are many ways to implement these augmentations, one approach is through our recently released Model Context Protocol, which allows developers to integrate with a growing ecosystem of third-party tools with a simple client implementation.

在本文的剩余部分,我们将假设每次 LLM 调用都能访问这些增强能力。

For the remainder of this post, we'll assume each LLM call has access to these augmented capabilities.

工作流:提示链 Workflow: Prompt chaining

提示链将任务分解为一系列步骤,其中每次 LLM 调用处理前一次调用的输出。您可以在任何中间步骤添加程序化检查(参见下图中的“门”),以确保流程仍在正轨上。

Prompt chaining decomposes a task into a sequence of steps, where each LLM call processes the output of the previous one. You can add programmatic checks (see "gate” in the diagram below) on any intermediate steps to ensure that the process is still on track.

何时使用此工作流:此工作流非常适合任务可以轻松且清晰地分解为固定子任务的情况。主要目标是通过使每次 LLM 调用成为更简单的任务,以延迟换取更高的准确性。

When to use this workflow: This workflow is ideal for situations where the task can be easily and cleanly decomposed into fixed subtasks. The main goal is to trade off latency for higher accuracy, by making each LLM call an easier task.

提示链有用的示例:

Examples where prompt chaining is useful:

* 生成营销文案,然后将其翻译成另一种语言。

* Generating Marketing copy, then translating it into a different language.

* 编写文档大纲,检查大纲是否符合某些标准,然后根据大纲编写文档。

* Writing an outline of a document, checking that the outline meets certain criteria, then writing the document based on the outline.

工作流:路由 Workflow: Routing

路由对输入进行分类,并将其导向专门的后续任务。这种工作流实现了关注点分离,并能够构建更专门的提示。没有此工作流,针对一种输入进行优化可能会损害其他输入的性能。

Routing classifies an input and directs it to a specialized followup task. This workflow allows for separation of concerns, and building more specialized prompts. Without this workflow, optimizing for one kind of input can hurt performance on other inputs.

何时使用此工作流:路由适用于存在不同类别且这些类别最好分别处理,并且分类可以由 LLM 或更传统的分类模型/算法准确处理的复杂任务。

When to use this workflow: Routing works well for complex tasks where there are distinct categories that are better handled separately, and where classification can be handled accurately, either by an LLM or a more traditional classification model/algorithm.

* 将不同类型的客户服务查询(一般问题、退款请求、技术支持)导向不同的下游流程、提示和工具。

* Directing different types of customer service queries (general questions, refund requests, technical support) into different downstream processes, prompts, and tools.

* 将简单/常见问题路由到更小、更具成本效益的模型(如 Claude Haiku 4.5),将困难/不寻常的问题路由到能力更强的模型(如 Claude Sonnet 4.5),以优化最佳性能。

* Routing easy/common questions to smaller, cost-efficient models like Claude Haiku 4.5 and hard/unusual questions to more capable models like Claude Sonnet 4.5 to optimize for best performance.

工作流:并行化 Workflow: Parallelization

LLM 有时可以同时处理一个任务,并通过编程方式聚合它们的输出。这种工作流称为并行化,有两种关键变体:

LLMs can sometimes work simultaneously on a task and have their outputs aggregated programmatically. This workflow, parallelization, manifests in two key variations:

* 分段:将任务分解为独立的子任务并行运行。

* Sectioning: Breaking a task into independent subtasks run in parallel.

* 投票:多次运行同一任务以获得多样化的输出。

* Voting: Running the same task multiple times to get diverse outputs.

何时使用此工作流:当分解的子任务可以并行化以提高速度,或者需要多种视角或多次尝试以获得更高置信度的结果时,并行化是有效的。对于涉及多个考虑的复杂任务,LLM 通常在每个考虑由单独的 LLM 调用处理时表现更好,从而允许专注于每个特定方面。

When to use this workflow: Parallelization is effective when the divided subtasks can be parallelized for speed, or when multiple perspectives or attempts are needed for higher confidence results. For complex tasks with multiple considerations, LLMs generally perform better when each consideration is handled by a separate LLM call, allowing focused attention on each specific aspect.

并行化有用的示例:

Examples where parallelization is useful:

* 实现护栏,其中一个模型实例处理用户查询,而另一个实例筛选不当内容或请求。这通常比让同一个 LLM 调用同时处理护栏和核心响应效果更好。

* Implementing guardrails where one model instance processes user queries while another screens them for inappropriate content or requests. This tends to perform better than having the same LLM call handle both guardrails and the core response.

* 自动化评估以评估 LLM 性能,其中每个 LLM 调用评估模型在给定提示下性能的不同方面。

* Automating evals for evaluating LLM performance, where each LLM call evaluates a different aspect of the model’s performance on a given prompt.

* 审查代码中的漏洞,其中几个不同的提示审查代码并在发现问题时标记它。

* Reviewing a piece of code for vulnerabilities, where several different prompts review and flag the code if they find a problem.

* 评估给定内容是否不当,使用多个提示评估不同方面或要求不同的投票阈值以平衡误报和漏报。

* Evaluating whether a given piece of content is inappropriate, with multiple prompts evaluating different aspects or requiring different vote thresholds to balance false positives and negatives.

工作流:编排器-工作器 Workflow: Orchestrator-workers

在编排器-工作器工作流中,一个中心 LLM 动态地分解任务,将其委派给工作器 LLM,并综合它们的结果。

In the orchestrator-workers workflow, a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results.

何时使用此工作流:此工作流非常适合复杂任务,在这些任务中你无法预测所需的子任务(例如,在编码中,需要更改的文件数量以及每个文件中更改的性质可能取决于任务)。虽然它在拓扑上相似,但与并行化的关键区别在于其灵活性——子任务不是预定义的,而是由编排器根据具体输入确定。

When to use this workflow: This workflow is well-suited for complex tasks where you can’t predict the subtasks needed (in coding, for example, the number of files that need to be changed and the nature of the change in each file likely depend on the task). Whereas it’s topographically similar, the key difference from parallelization is its flexibility—subtasks aren't pre-defined, but determined by the orchestrator based on the specific input.

编排器-工作器有用的示例:

Example where orchestrator-workers is useful:

* 每次对多个文件进行复杂更改的编码产品。

* Coding products that make complex changes to multiple files each time.

* 涉及从多个来源收集和分析信息以获取可能相关信息的搜索任务。

* Search tasks that involve gathering and analyzing information from multiple sources for possible relevant information.

工作流:评估器-优化器 Workflow: Evaluator-optimizer

在评估器-优化器工作流中,一个 LLM 调用生成响应,而另一个 LLM 在循环中提供评估和反馈。

In the evaluator-optimizer workflow, one LLM call generates a response while another provides evaluation and feedback in a loop.

何时使用此工作流:当我们有明确的评估标准,并且迭代优化能带来可衡量的价值时,此工作流特别有效。两个适用迹象是:首先,当人类给出反馈时,LLM 的响应可以明显改进;其次,LLM 能够提供此类反馈。这类似于人类作者在撰写精炼文档时可能经历的迭代写作过程。

When to use this workflow: This workflow is particularly effective when we have clear evaluation criteria, and when iterative refinement provides measurable value. The two signs of good fit are, first, that LLM responses can be demonstrably improved when a human articulates their feedback; and second, that the LLM can provide such feedback. This is analogous to the iterative writing process a human writer might go through when producing a polished document.

评估器-优化器有用的示例:

Examples where evaluator-optimizer is useful:

* 文学翻译,其中翻译 LLM 可能最初无法捕捉细微差别,但评估 LLM 可以提供有用的批评。

* Literary translation where there are nuances that the translator LLM might not capture initially, but where an evaluator LLM can provide useful critiques.

* 复杂的搜索任务,需要多轮搜索和分析以收集全面信息,评估器决定是否需要进一步搜索。

* Complex search tasks that require multiple rounds of searching and analysis to gather comprehensive information, where the evaluator decides whether further searches are warranted.

智能体 Agents

随着 LLM 在关键能力上的成熟——理解复杂输入、进行推理和规划、可靠地使用工具以及从错误中恢复——智能体正在生产中涌现。智能体从人类用户的指令或互动讨论开始工作。一旦任务明确,智能体就会独立规划和操作,并可能返回人类以获取更多信息或判断。在执行过程中,智能体在每个步骤从环境中获取“地面实况”(例如工具调用结果或代码执行)以评估其进展至关重要。然后,智能体可以在检查点或遇到障碍时暂停以获取人类反馈。任务通常在完成后终止,但通常也包括停止条件(例如最大迭代次数)以保持控制。

Agents are emerging in production as LLMs mature in key capabilities—understanding complex inputs, engaging in reasoning and planning, using tools reliably, and recovering from errors. Agents begin their work with either a command from, or interactive discussion with, the human user. Once the task is clear, agents plan and operate independently, potentially returning to the human for further information or judgement. During execution, it's crucial for the agents to gain “ground truth” from the environment at each step (such as tool call results or code execution) to assess its progress. Agents can then pause for human feedback at checkpoints or when encountering blockers. The task often terminates upon completion, but it’s also common to include stopping conditions (such as a maximum number of iterations) to maintain control.

智能体可以处理复杂的任务,但它们的实现通常很简单。它们通常只是基于环境反馈循环中使用工具的 LLM。因此,清晰而周到地设计工具集及其文档至关重要。我们在附录 2(“提示工程你的工具”)中扩展了工具开发的最佳实践。

Agents can handle sophisticated tasks, but their implementation is often straightforward. They are typically just LLMs using tools based on environmental feedback in a loop. It is therefore crucial to design toolsets and their documentation clearly and thoughtfully. We expand on best practices for tool development in Appendix 2 ("Prompt Engineering your Tools").

何时使用智能体:智能体可用于开放式问题,在这些问题中很难或不可能预测所需的步骤数,并且无法硬编码固定路径。LLM 可能会运行多个轮次,你必须对其决策有一定程度的信任。智能体的自主性使其成为在可信环境中扩展任务的理想选择。

When to use agents: Agents can be used for open-ended problems where it’s difficult or impossible to predict the required number of steps, and where you can’t hardcode a fixed path. The LLM will potentially operate for many turns, and you must have some level of trust in its decision-making. Agents' autonomy makes them ideal for scaling tasks in trusted environments.

智能体的自主性意味着更高的成本和潜在的复合错误。我们建议在沙盒环境中进行广泛测试,并设置适当的防护措施。

The autonomous nature of agents means higher costs, and the potential for compounding errors. We recommend extensive testing in sandboxed environments, along with the appropriate guardrails.

以下示例来自我们自己的实现:

The following examples are from our own implementations:

* 一个用于解决 SWE-bench 任务的编码智能体,涉及基于任务描述对多个文件进行编辑;

* A coding Agent to resolve SWE-bench tasks, which involve edits to many files based on a task description;

* 我们的“计算机使用”参考实现,其中 Claude 使用计算机完成任务。

* Our “computer use” reference implementation, where Claude uses a computer to accomplish tasks.

组合与定制这些模式 Combining and customizing these patterns

这些构建模块并非规定性的。它们是开发者可以根据不同用例进行塑造和组合的常见模式。与任何大语言模型功能一样,成功的关键在于衡量性能并迭代实现。重复一遍:只有当复杂性确实能改善结果时,才应考虑增加复杂性。

These building blocks aren't prescriptive. They're common patterns that developers can shape and combine to fit different use cases. The key to success, as with any LLM features, is measuring performance and iterating on implementations. To repeat: you should consider adding complexity only when it demonstrably improves outcomes.

总结 Summary

在 LLM 领域的成功不在于构建最复杂的系统,而在于为你的需求构建_正确_的系统。从简单的提示开始,通过全面评估进行优化,仅在简单方案不足时添加多步骤智能体系统。

Success in the LLM space isn't about building the most sophisticated system. It's about building the right system for your needs. Start with simple prompts, optimize them with comprehensive evaluation, and add multi-step agentic systems only when simpler solutions fall short.

在实现智能体时,我们尝试遵循三个核心原则:

When implementing agents, we try to follow three core principles:

1. 保持智能体设计的简洁性。 2. 通过明确展示智能体的规划步骤来优先考虑透明度。 3. 通过详尽的工具文档和测试,精心设计智能体-计算机接口(ACI)。

1. Maintain simplicity in your agent's design.

框架可以帮助你快速入门,但在转向生产环境时,不要犹豫减少抽象层,使用基本组件进行构建。遵循这些原则,你可以创建不仅强大,而且可靠、可维护且受用户信任的智能体。

2. Prioritize transparency by explicitly showing the agent’s planning steps.

3. Carefully craft your agent-computer interface (ACI) through thorough tool documentation and testing.

Frameworks can help you get started quickly, but don't hesitate to reduce abstraction layers and build with basic components as you move to production. By following these principles, you can create agents that are not only powerful but also reliable, maintainable, and trusted by their users.

互动版:图/公式 + 针对本篇提问 →