Agents
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→许多人认为智能体是人工智能的终极目标。Stuart Russell 和 Peter Norvig 的经典著作《人工智能:一种现代方法》(Prentice Hall, 1995)将 AI 研究领域定义为“对理性智能体的研究和设计”。基础模型前所未有的能力为之前难以想象的智能体应用打开了大门。这些新能力使得开发自主、智能的智能体成为可能,它们可以作为我们的助手、同事和教练。它们可以帮助我们创建网站、收集数据、规划旅行、进行市场调研、管理客户账户、自动化数据录入、为我们准备面试、面试候选人、谈判交易等。可能性似乎是无限的,这些智能体的潜在经济价值巨大。本节将从智能体的概述开始,然后继续讨论决定智能体能力的两个方面:工具和规划。智能体以其新的操作模式,也有新的失败模式。本节将以如何评估智能体以捕捉这些失败的讨论结束。
Intelligent agents are considered by many to be the ultimate goal of AI. The classic book by Stuart Russell and Peter Norvig, _Artificial Intelligence: A Modern Approach_ (Prentice Hall, 1995), defines the field of AI research as “_the study and design of rational agents._” The unprecedented capabilities of foundation models have opened the door to agentic applications that were previously unimaginable. These new capabilities make it finally possible to develop autonomous, intelligent agents to act as our assistants, coworkers, and coaches. They can help us create a website, gather data, plan a trip, do market research, manage a customer account, automate data entry, prepare us for interviews, interview our candidates, negotiate a deal, etc. The possibilities seem endless, and the potential economic value of these agents is enormous. This section will start with an overview of agents and then continue with two aspects that determine the capabilities of an agent: tools and planning. Agents, with their new modes of operations, have new modes of failure. This section will end with a discussion on how to evaluate agents to catch these failures.
许多人认为智能体是人工智能的终极目标。Stuart Russell 和 Peter Norvig 的经典著作《人工智能:一种现代方法》(Prentice Hall, 1995)将 AI 研究领域定义为“对理性智能体的研究与设计”。
Intelligent agents are considered by many to be the ultimate goal of AI. The classic book by Stuart Russell and Peter Norvig, _Artificial Intelligence: A Modern Approach_ (Prentice Hall, 1995), defines the field of AI research as “_the study and design of rational agents._”
基础模型前所未有的能力为之前难以想象的智能体式应用打开了大门。这些新能力使得开发自主、智能的智能体成为可能,它们可以作为我们的助手、同事和教练。它们可以帮助我们创建网站、收集数据、规划旅行、进行市场调研、管理客户账户、自动化数据录入、为我们准备面试、面试候选人、谈判交易等。可能性似乎是无限的,这些智能体的潜在经济价值巨大。
The unprecedented capabilities of foundation models have opened the door to agentic applications that were previously unimaginable. These new capabilities make it finally possible to develop autonomous, intelligent agents to act as our assistants, coworkers, and coaches. They can help us create a website, gather data, plan a trip, do market research, manage a customer account, automate data entry, prepare us for interviews, interview our candidates, negotiate a deal, etc. The possibilities seem endless, and the potential economic value of these agents is enormous.
本节将从智能体概述开始,然后继续讨论决定智能体能力的两个方面:工具和规划。智能体以其新的操作模式,也带来了新的失败模式。本节将以关于如何评估智能体以捕捉这些失败的讨论结束。
This section will start with an overview of agents and then continue with two aspects that determine the capabilities of an agent: tools and planning. Agents, with their new modes of operations, have new modes of failure. This section will end with a discussion on how to evaluate agents to catch these failures.
本文改编自《AI 工程》(2025)的智能体章节,并进行了少量编辑以使其成为独立文章。
_This post is adapted from the Agents section of AI Engineering (2025) with minor edits to make it a standalone post._
1. AI 驱动的智能体是一个新兴领域,目前尚无成熟的理论框架来定义、开发和评估它们。本节是基于现有文献构建框架的最大努力尝试,但将随着领域的发展而演变。与本书其他部分相比,本节更具实验性。我从早期审稿人那里得到了有益的反馈,也希望从这篇博文的读者那里获得反馈。
1. AI-powered agents are an emerging field with no established theoretical frameworks for defining, developing, and evaluating them. This section is a best-effort attempt to build a framework from the existing literature, but it will evolve as the field does. Compared to the rest of the book, this section is more experimental. I received helpful feedback from early reviewers, and I hope to get feedback from readers of this blog post, too.
2. 就在本书出版之前,Anthropic 发表了一篇关于构建有效智能体的博文(2024 年 12 月)。我很高兴看到 Anthropic 的博文和我的智能体章节在概念上是一致的,尽管术语略有不同。然而,Anthropic 的文章侧重于孤立的模式,而我的文章涵盖了事物为何以及如何运作。我也更侧重于规划、工具选择和失败模式。
2. Just before this book came out, Anthropic published a blog post on Building effective agents (Dec 2024). I’m glad to see that Anthropic’s blog post and my agent section are conceptually aligned, though with slightly different terminologies. However, Anthropic’s post focuses on isolated patterns, whereas my post covers why and how things work. I also focus more on planning, tool selection, and failure modes.
3. 本文包含大量背景信息。如果觉得过于深入,可以随意跳过。
3. The post contains a lot of background information. Feel free to skip ahead if it feels a little too in the weeds!
“智能体”一词已在许多不同的工程语境中使用,包括但不限于软件智能体、智能体、用户智能体、对话智能体和强化学习智能体。那么,究竟什么是智能体?
The term _agent_ has been used in many different engineering contexts, including but not limited to a software agent, intelligent agent, user agent, conversational agent, and reinforcement learning agent. So, what exactly is an agent?
智能体是任何能够感知其环境并对其环境采取行动的事物。《人工智能:一种现代方法》(1995)将智能体定义为任何可以通过传感器感知其环境并通过执行器对其环境采取行动的事物。
An agent is anything that can perceive its environment and act upon that environment. _Artificial Intelligence: A Modern Approach_ (1995) defines an agent as anything that can be viewed as perceiving its environment through sensors and acting upon that environment through actuators.
这意味着智能体由其运行的环境和它可以执行的动作集来表征。
This means that an agent is characterized by the _environment_ it operates in and _the set of actions_ it can perform.
智能体可以运行的环境由其用例定义。如果开发一个智能体来玩游戏(例如《我的世界》、围棋、《刀塔》),那么该游戏就是它的环境。如果你希望智能体从互联网抓取文档,那么环境就是互联网。自动驾驶汽车智能体的环境是道路系统及其邻近区域。
The _environment_ an agent can operate in is defined by its use case. If an agent is developed to play a game (e.g., _Minecraft,_ Go, _Dota_), that game is its environment. If you want an agent to scrape documents from the internet, the environment is the internet. A self-driving car agent’s environment is the road system and its adjacent areas.
AI 智能体可以执行的动作集通过它可以访问的工具得到增强。许多你日常交互的生成式 AI 应用都是拥有工具访问权限的智能体,尽管是简单的工具。ChatGPT 就是一个智能体。它可以搜索网络、执行 Python 代码和生成图像。RAG 系统也是智能体——文本检索器、图像检索器和 SQL 执行器就是它们的工具。
The _set of actions_ an AI agent can perform is augmented by the _tools_ it has access to. Many generative AI-powered applications you interact with daily are agents with access to tools, albeit simple ones. ChatGPT is an agent. It can search the web, execute Python code, and generate images. RAG systems are agents—text retrievers, image retrievers, and SQL executors are their tools.
智能体的环境与其工具集之间存在强依赖关系。环境决定了智能体可能使用哪些工具。例如,如果环境是国际象棋游戏,智能体唯一可能的动作就是合法的象棋走法。然而,智能体的工具库存限制了它可以运行的环境。例如,如果机器人的唯一动作是游泳,那么它将局限于水环境。
There’s a strong dependency between an agent’s environment and its set of tools. The environment determines what tools an agent can potentially use. For example, if the environment is a chess game, the only possible actions for an agent are the valid chess moves. However, an agent’s tool inventory restricts the environment it can operate in. For example, if a robot’s only action is swimming, it’ll be confined to a water environment.
图 6-8 展示了 SWE-agent(Yang 等人,2024)的可视化,这是一个基于 GPT-4 构建的智能体。它的环境是带有终端和文件系统的计算机。它的动作集包括导航仓库、搜索文件、查看文件和编辑行。
Figure 6-8 shows a visualization of SWE-agent (Yang et al., 2024), an agent built on top of GPT-4. Its environment is the computer with the terminal and the file system. Its set of actions include navigate repo, search files, view files, and edit lines.
AI 智能体旨在完成通常由用户提供的任务。在 AI 智能体中,AI 是处理任务、规划一系列动作以实现该任务并确定任务是否已完成的大脑。
An AI agent is meant to accomplish tasks typically provided by the users. In an AI agent, AI is the brain that processes the task, plans a sequence of actions to achieve this task, and determines whether the task has been accomplished.
让我们回到上面 Kitty Vogue 示例中带有表格数据的 RAG 系统。这是一个具有三个动作的简单智能体:
Let’s return to the RAG system with tabular data in the Kitty Vogue example above. This is a simple agent with three actions:
“预测 Fruity Fedora 未来三个月的销售收入”,智能体可能会执行以下动作序列:
"Project the sales revenue for Fruity Fedora over the next three months"
1. 推理如何完成此任务。它可能决定要预测未来销售,首先需要过去五年的销售数据。智能体的推理可以显示为中间响应。
, the agent might perform the following sequence of actions:
2. 调用 SQL 查询生成来生成获取过去五年销售数据的查询。
1. Reason about how to accomplish this task. It might decide that to predict future sales, it first needs the sales numbers from the last five years. An agent’s reasoning can be shown as intermediate responses.
3. 调用 SQL 查询执行来执行此查询。
2. Invoke SQL query generation to generate the query to get sales numbers from the last five years.
4. 推理工具输出(SQL 查询执行的输出)以及它们如何帮助销售预测。它可能决定这些数字不足以做出可靠的预测,可能是因为缺失值。然后它决定还需要过去营销活动的信息。
3. Invoke SQL query execution to execute this query.
5. 调用 SQL 查询生成来生成过去营销活动的查询。
4. Reason about the tool outputs (outputs from the SQL query execution) and how they help with sales prediction. It might decide that these numbers are insufficient to make a reliable projection, perhaps because of missing values. It then decides that it also needs information about past marketing campaigns.
7. 推理这些新信息足以帮助预测未来销售。然后它生成预测。
5. Invoke SQL query generation to generate the queries for past marketing campaigns.
8. 推理任务已成功完成。
7. Reason that this new information is sufficient to help predict future sales. It then generates a projection.
与非智能体用例相比,智能体通常需要更强大的模型,原因有二:
8. Reason that the task has been successfully completed.
* 复合错误:智能体通常需要执行多个步骤才能完成任务,随着步骤数量的增加,整体准确性会下降。如果模型每步的准确性为 95%,那么 10 步后准确性将降至 60%,100 步后准确性仅为 0.6%。
Compared to non-agent use cases, agents typically require more powerful models for two reasons:
* 更高的风险:通过访问工具,智能体能够执行更有影响力的任务,但任何失败都可能带来更严重的后果。
* Compound mistakes: an agent often needs to perform multiple steps to accomplish a task, and the overall accuracy decreases as the number of steps increases. If the model’s accuracy is 95% per step, over 10 steps, the accuracy will drop to 60%, and over 100 steps, the accuracy will be only 0.6%.
需要许多步骤的任务可能需要时间和金钱来运行。一个常见的抱怨是智能体只擅长消耗你的 API 额度。然而,如果智能体可以自主运行,它们可以节省人类时间,使其成本物有所值。
* Higher stakes: with access to tools, agents are capable of performing more impactful tasks, but any failure could have more severe consequences.
给定一个环境,智能体在环境中的成功取决于它可以访问的工具及其 AI 规划器的强度。让我们首先研究模型可以使用的不同类型的工具。接下来我们将分析 AI 的规划能力。
A task that requires many steps can take time and money to run. A common complaint is that agents are only good for burning through your API credits. However, if agents can be autonomous, they can save human time, making their costs worthwhile.
Given an environment, the success of an agent in an environment depends on the tool it has access to and the strength of its AI planner. Let’s start by looking into different kinds of tools a model can use. We’ll analyze AI’s capability for planning next.
一个系统不需要访问外部工具就能成为智能体。然而,没有外部工具,智能体的能力将受到限制。模型本身通常只能执行一个动作——LLM 可以生成文本,图像生成器可以生成图像。外部工具使智能体能力大大增强。
A system doesn’t need access to external tools to be an agent. However, without external tools, the agent’s capabilities would be limited. By itself, a model can typically perform one action—an LLM can generate text and an image generator can generate images. External tools make an agent vastly more capable.
工具帮助智能体感知环境并对其采取行动。允许智能体感知环境的动作是_只读动作_,而允许智能体对环境采取行动的动作是_写动作_。
Tools help an agent to both perceive the environment and act upon it. Actions that allow an agent to perceive the environment are _read-only actions_, whereas actions that allow an agent to act upon the environment are _write actions_.
智能体可以访问的工具集就是它的工具清单。由于智能体的工具清单决定了它能做什么,因此仔细考虑给智能体提供哪些工具以及多少工具非常重要。更多的工具赋予智能体更强的能力。然而,工具越多,理解和有效利用它们的难度就越大。需要通过实验来找到合适的工具集,这将在后面的“工具选择”部分讨论。
The set of tools an agent has access to is its tool inventory. Since an agent’s tool inventory determines what an agent can do, it’s important to think through what and how many tools to give an agent. More tools give an agent more capabilities. However, the more tools there are, the more challenging it is to understand and utilize them well. Experimentation is necessary to find the right set of tools, as discussed later in the “Tool selection” section.
根据智能体的环境,有许多可能的工具。以下是您可能想要考虑的三类工具:知识增强(即上下文构建)、能力扩展以及让智能体对其环境采取行动的工具。
Depending on the agent’s environment, there are many possible tools. Here are three categories of tools that you might want to consider: knowledge augmentation (i.e., context construction), capability extension, and tools that let your agent act upon its environment.
我希望到目前为止,本书已经让你相信,为模型提供相关上下文对其响应质量至关重要。一个重要的工具类别包括那些有助于增强智能体知识的工具。其中一些已经讨论过:文本检索器、图像检索器和 SQL 执行器。其他潜在工具包括内部人员搜索、返回不同产品状态的库存 API、Slack 检索、电子邮件阅读器等。
I hope that this book, so far, has convinced you of the importance of having the relevant context for a model’s response quality. An important category of tools includes those that help augment the knowledge of your agent. Some of them have already been discussed: text retriever, image retriever, and SQL executor. Other potential tools include internal people search, an inventory API that returns the status of different products, Slack retrieval, an email reader, etc.
许多此类工具用你组织的私有流程和信息来增强模型。然而,工具也可以让模型访问公共信息,尤其是来自互联网的信息。
Many such tools augment a model with your organization’s private processes and information. However, tools can also give models access to public information, especially from the internet.
网页浏览是 ChatGPT 最早且最受期待的功能之一。网页浏览可以防止模型过时。当模型训练所用的数据变得过时时,模型就会过时。如果模型的训练数据截止于上周,它就无法回答需要本周信息的问题,除非这些信息被提供在上下文中。没有网页浏览,模型就无法告诉你天气、新闻、即将发生的事件、股票价格、航班状态等。
Web browsing was among the earliest and most anticipated capabilities to be incorporated into ChatGPT. Web browsing prevents a model from going stale. A model goes stale when the data it was trained on becomes outdated. If the model’s training data was cut off last week, it won’t be able to answer questions that require information from this week unless this information is provided in the context. Without web browsing, a model won’t be able to tell you about the weather, news, upcoming events, stock prices, flight status, etc.
我将网页浏览作为一个总称,涵盖所有访问互联网的工具,包括网页浏览器和 API,如搜索 API、新闻 API、GitHub API 或社交媒体 API。
I use web browsing as an umbrella term to cover all tools that access the internet, including web browsers and APIs such as search APIs, news APIs, GitHub APIs, or social media APIs.
虽然网页浏览允许你的智能体引用最新信息以生成更好的响应并减少幻觉,但它也可能让你的智能体暴露于互联网的污秽之中。请谨慎选择你的互联网 API。
While web browsing allows your agent to reference up-to-date information to generate better responses and reduce hallucinations, it can also open up your agent to the cesspools of the internet. Select your Internet APIs with care.
你还可以考虑使用工具来解决 AI 模型固有的局限性。这是提升模型性能的简便方法。例如,AI 模型以不擅长数学而闻名。如果你问模型 199,999 除以 292 等于多少,模型很可能会失败。然而,如果模型能使用计算器,这个计算就变得微不足道了。与其试图训练模型擅长算术,不如直接让模型使用工具,这样资源效率更高。
You might also consider tools that address the inherent limitations of AI models. They are easy ways to give your model a performance boost. For example, AI models are notorious for being bad at math. If you ask a model what is 199,999 divided by 292, the model will likely fail. However, this calculation would be trivial if the model had access to a calculator. Instead of trying to train the model to be good at arithmetic, it’s a lot more resource-efficient to just give the model access to a tool.
其他能显著提升模型能力的简单工具包括日历、时区转换器、单位转换器(例如,从磅到千克)以及翻译器,用于翻译模型不擅长的语言。
Other simple tools that can significantly boost a model’s capability include a calendar, timezone converter, unit converter (e.g., from lbs to kg), and translator that can translate to and from the languages that the model isn’t good at.
更复杂但强大的工具是代码解释器。与其训练模型理解代码,不如让它访问代码解释器来执行代码、返回结果或分析代码失败的原因。这种能力使你的智能体能够充当编码助手、数据分析师,甚至研究助手,可以编写代码来运行实验并报告结果。然而,自动代码执行存在代码注入攻击的风险,正如第 5 章“防御性提示工程”一节所讨论的。适当的安全措施对于保护你和用户的安全至关重要。
More complex but powerful tools are code interpreters. Instead of training a model to understand code, you can give it access to a code interpreter to execute a piece of code, return the results, or analyze the code’s failures. This capability lets your agents act as coding assistants, data analysts, and even research assistants that can write code to run experiments and report results. However, automated code execution comes with the risk of code injection attacks, as discussed in Chapter 5 in the section “Defensive Prompt Engineering“. Proper security measurements are crucial to keep you and your users safe.
工具可以将纯文本或纯图像模型转变为多模态模型。例如,只能生成文本的模型可以利用文本到图像模型作为工具,从而同时生成文本和图像。给定一个文本请求,智能体的 AI 规划器决定是调用文本生成、图像生成还是两者都调用。这就是 ChatGPT 能够生成文本和图像的方式——它使用 DALL-E 作为图像生成器。
Tools can turn a text-only or image-only model into a multimodal model. For example, a model that can generate only texts can leverage a text-to-image model as a tool, allowing it to generate both texts and images. Given a text request, the agent’s AI planner decides whether to invoke text generation, image generation, or both. This is how ChatGPT can generate both text and images—it uses DALL-E as its image generator.
智能体还可以使用代码解释器生成图表和图形,使用 LaTeX 编译器渲染数学公式,或使用浏览器从 HTML 代码渲染网页。
Agents can also use a code interpreter to generate charts and graphs, a LaTex compiler to render math equations, or a browser to render web pages from HTML code.
类似地,只能处理文本输入的模型可以使用图像描述工具处理图像,使用转录工具处理音频。它可以使用 OCR(光学字符识别)工具读取 PDF。
Similarly, a model that can process only text inputs can use an image captioning tool to process images and a transcription tool to process audio. It can use an OCR (optical character recognition) tool to read PDFs.
_与仅通过提示甚至微调相比,工具使用可以显著提升模型的性能_。Chameleon(Lu 等人,2023)表明,一个由 GPT-4 驱动的智能体,配备 13 个工具集,可以在多个基准测试上超越单独的 GPT-4。该智能体使用的工具示例包括知识检索、查询生成器、图像描述器、文本检测器和必应搜索。
_Tool use can significantly boost a model’s performance compared to just prompting or even finetuning_. Chameleon (Lu et al., 2023) shows that a GPT-4-powered agent, augmented with a set of 13 tools, can outperform GPT-4 alone on several benchmarks. Examples of tools this agent used are knowledge retrieval, a query generator, an image captioner, a text detector, and Bing search.
在 ScienceQA(一个科学问答基准测试)上,Chameleon 将最佳已发表的少样本结果提高了 11.37%。在 TabMWP(表格数学应用题)(Lu 等人,2022)上,一个涉及表格数学问题的基准测试,Chameleon 将准确率提高了 17%。
On ScienceQA, a science question answering benchmark, Chameleon improves the best published few-shot result by 11.37%. On TabMWP (Tabular Math Word Problems) (Lu et al., 2022), a benchmark involving tabular math questions, Chameleon improves the accuracy by 17%.
到目前为止,我们讨论了允许模型从其数据源读取的只读操作。但工具也可以执行写入操作,对数据源进行更改。SQL 执行器可以检索数据表(读取),也可以更改或删除表(写入)。电子邮件 API 可以读取邮件,但也可以回复邮件。银行 API 可以检索您的当前余额,但也可以发起银行转账。
So far, we’ve discussed read-only actions that allow a model to read from its data sources. But tools can also perform write actions, making changes to the data sources. An SQL executor can retrieve a data table (read) and change or delete the table (write). An email API can read an email but can also respond to it. A banking API can retrieve your current balance, but can also initiate a bank transfer.
写入操作使系统能够做更多事情。它们可以使您自动化整个客户外联工作流程:研究潜在客户、查找他们的联系方式、起草电子邮件、发送第一封邮件、阅读回复、跟进、提取订单、用新订单更新数据库等。
Write actions enable a system to do more. They can enable you to automate the whole customer outreach workflow: researching potential customers, finding their contacts, drafting emails, sending first emails, reading responses, following up, extracting orders, updating your databases with new orders, etc.
然而,赋予 AI 自动改变我们生活的能力这一前景令人恐惧。就像你不应该让实习生有权删除你的生产数据库一样,你也不应该允许不可靠的 AI 发起银行转账。对系统能力及其安全措施的信任至关重要。你需要确保系统免受可能试图操纵它执行有害行为的恶意行为者的侵害。
However, the prospect of giving AI the ability to automatically alter our lives is frightening. Just as you shouldn’t give an intern the authority to delete your production database, you shouldn’t allow an unreliable AI to initiate bank transfers. Trust in the system’s capabilities and its security measures is crucial. You need to ensure that the system is protected from bad actors who might try to manipulate it into performing harmful actions.
每当我与一群人谈论自主 AI 智能体时,总会有人提到自动驾驶汽车。“如果有人入侵汽车绑架你怎么办?”虽然自动驾驶汽车的例子因其物理性而显得直观,但 AI 系统无需物理存在就能造成伤害。它可以操纵股票市场、窃取版权、侵犯隐私、强化偏见、传播错误信息和宣传等,正如第 5 章“防御性提示工程”一节所讨论的。
Whenever I talk about autonomous AI agents to a group of people, there is often someone who brings up self-driving cars. “_What if someone hacks into the car to kidnap you?_” While the self-driving car example seems visceral because of its physicality, an AI system can cause harm without a presence in the physical world. It can manipulate the stock market, steal copyrights, violate privacy, reinforce biases, spread misinformation and propaganda, and more, as discussed in the section “Defensive Prompt Engineering” in Chapter 5.
这些都是合理的担忧,任何希望利用 AI 的组织都需要认真对待安全性和可靠性。然而,这并不意味着 AI 系统永远不应该被赋予在现实世界中行动的能力。如果我们能信任机器将我们送入太空,我希望有一天,安全措施足以让我们信任自主 AI 系统。此外,人类也可能犯错。就个人而言,我更信任自动驾驶汽车而不是普通陌生人载我一程。
These are all valid concerns, and any organization that wants to leverage AI needs to take safety and security seriously. However, this doesn’t mean that AI systems should never be given the ability to act in the real world. If we can trust a machine to take us into space, I hope that one day, security measures will be sufficient for us to trust autonomous AI systems. Besides, humans can fail, too. Personally, I would trust a self-driving car more than the average stranger to give me a lift.
正如合适的工具可以帮助人类大幅提高生产力——你能想象没有 Excel 做生意或没有起重机建造摩天大楼吗?——工具使模型能够完成更多任务。许多模型提供商已经支持其模型的工具使用,这一功能通常称为函数调用。展望未来,我预计大多数模型将普遍支持使用广泛工具的函数调用。
Just as the right tools can help humans be vastly more productive—can you imagine doing business without Excel or building a skyscraper without cranes?—tools enable models to accomplish many more tasks. Many model providers already support tool use with their models, a feature often called function calling. Going forward, I would expect function calling with a wide set of tools to be common with most models.
基础模型智能体的核心是负责解决用户提供任务的模型。任务由其目标和约束定义。例如,一个任务是规划一次从旧金山到印度的两周旅行,预算为 5000 美元。目标是两周旅行,约束是预算。
At the heart of a foundation model agent is the model responsible for solving user-provided tasks. A task is defined by its goal and constraints. For example, one task is to schedule a two-week trip from San Francisco to India with a budget of $5,000. The goal is the two-week trip. The constraint is the budget.
复杂任务需要规划。规划过程的输出是一个计划,它是概述完成任务所需步骤的路线图。有效的规划通常要求模型理解任务,考虑实现该任务的不同选项,并选择最有前景的一个。
Complex tasks require planning. The output of the planning process is a plan, which is a roadmap outlining the steps needed to accomplish a task. Effective planning typically requires the model to understand the task, consider different options to achieve this task, and choose the most promising one.
如果你参加过任何规划会议,你就会知道规划很难。作为一个重要的计算问题,规划已被深入研究,需要多卷著作才能覆盖。我在这里只能触及表面。
If you’ve ever been in any planning meeting, you know that planning is hard. As an important computational problem, planning is well studied and would require several volumes to cover. I’ll only be able to cover the surface here.
给定一个任务,有许多可能的方法来解决它,但并非所有方法都能带来成功的结果。在正确的解决方案中,有些比其他更高效。考虑查询:
Given a task, there are many possible ways to solve it, but not all of them will lead to a successful outcome. Among the correct solutions, some are more efficient than others. Consider the query,
“有多少家没有收入的公司筹集了至少 10 亿美元?”
"How many companies without revenue have raised at least $1 billion?"
1. 找出所有没有收入的公司,然后按筹集金额筛选。
1. Find all companies without revenue, then filter them by the amount raised.
2. 找出所有筹集了至少 10 亿美元的公司,然后按收入筛选。
2. Find all companies that have raised at least $1 billion, then filter them by revenue.
第二种方案更高效。没有收入的公司数量远远多于筹集了 10 亿美元的公司。仅在这两种方案中,一个智能体应该选择方案 2。
The second option is more efficient. There are vastly more companies without revenue than companies that have raised $1 billion. Given only these two options, an intelligent agent should choose option 2.
你可以将规划与执行耦合在同一个提示中。例如,你给模型一个提示,要求它逐步思考(例如使用思维链提示),然后在一个提示中执行这些步骤。但如果模型提出了一个 1000 步的计划,却连目标都无法实现呢?在没有监督的情况下,智能体可能会运行这些步骤数小时,浪费时间和 API 调用的费用,然后你才发现它毫无进展。
You can couple planning with execution in the same prompt. For example, you give the model a prompt, ask it to think step by step (such as with a chain-of-thought prompt), and then execute those steps all in one prompt. But what if the model comes up with a 1,000-step plan that doesn’t even accomplish the goal? Without oversight, an agent can run those steps for hours, wasting time and money on API calls, before you realize that it’s not going anywhere.
为了避免徒劳的执行,_规划_应该与_执行_解耦。你要求智能体首先生成一个计划,只有在这个计划被_验证_之后才执行。计划可以使用启发式方法进行验证。例如,一个简单的启发式方法是消除包含无效动作的计划。如果生成的计划需要谷歌搜索,而智能体没有谷歌搜索的访问权限,那么这个计划就是无效的。另一个简单的启发式方法可能是消除所有超过 X 步的计划。
To avoid fruitless execution, _planning_ should be decoupled from _execution_. You ask the agent to first generate a plan, and only after this plan is _validated_ is it executed. The plan can be validated using heuristics. For example, one simple heuristic is to eliminate plans with invalid actions. If the generated plan requires a Google search and the agent doesn’t have access to Google Search, this plan is invalid. Another simple heuristic might be eliminating all plans with more than X steps.
计划也可以使用 AI 评判器进行验证。你可以要求一个模型评估计划是否合理,或者如何改进它。
A plan can also be validated using AI judges. You can ask a model to evaluate whether the plan seems reasonable or how to improve it.
如果生成的计划被评估为糟糕,你可以要求规划器生成另一个计划。如果生成的计划很好,就执行它。
If the generated plan is evaluated to be bad, you can ask the planner to generate another plan. If the generated plan is good, execute it.
如果计划包含外部工具,将调用函数调用。执行该计划的输出将再次需要评估。注意,生成的计划不必是整个任务的端到端计划。它可以是一个子任务的小计划。整个过程如图 6-9 所示。
If the plan consists of external tools, function calling will be invoked. Outputs from executing this plan will then again need to be evaluated. Note that the generated plan doesn’t have to be an end-to-end plan for the whole task. It can be a small plan for a subtask. The whole process looks like Figure 6-9.
你的系统现在有三个组件:一个生成计划,一个验证计划,另一个执行计划。如果你将每个组件视为一个智能体,这可以被认为是一个多智能体系统。由于大多数智能体式工作流足够复杂,涉及多个组件,因此大多数智能体都是多智能体的。
Your system now has three components: one to generate plans, one to validate plans, and another to execute plans. If you consider each component an agent, this can be considered a multi-agent system. Because most agentic workflows are sufficiently complex to involve multiple components, most agents are multi-agent.
为了加速过程,你可以并行生成多个计划,而不是顺序生成,并要求评估者选择最有希望的一个。这是另一个延迟-成本权衡,因为同时生成多个计划会产生额外成本。
To speed up the process, instead of generating plans sequentially, you can generate several plans in parallel and ask the evaluator to pick the most promising one. This is another latency–cost tradeoff, as generating multiple plans simultaneously will incur extra costs.
规划需要理解任务背后的意图:用户通过这个查询试图做什么?意图分类器通常用于帮助智能体规划。如第 5 章“将复杂任务分解为更简单的子任务”一节所示,意图分类可以使用另一个提示或为此任务训练的分类模型来完成。意图分类机制可以被视为多智能体系统中的另一个智能体。
Planning requires understanding the intention behind a task: what’s the user trying to do with this query? An intent classifier is often used to help agents plan. As shown in Chapter 5 in the section “Break complex tasks into simpler subtasks“, intent classification can be done using another prompt or a classification model trained for this task. The intent classification mechanism can be considered another agent in your multi-agent system.
了解意图可以帮助智能体选择合适的工具。例如,对于客户支持,如果查询是关于账单的,智能体可能需要访问一个工具来检索用户最近的付款。但如果查询是关于如何重置密码,智能体可能需要访问文档检索。
Knowing the intent can help the agent pick the right tools. For example, for customer support, if the query is about billing, the agent might need access to a tool to retrieve a user’s recent payments. But if the query is about how to reset a password, the agent might need to access documentation retrieval.
到目前为止,我们假设智能体自动化了所有三个阶段:生成计划、验证计划和执行计划。实际上,人类可以在任何阶段参与,以帮助过程并减轻风险。
So far, we’ve assumed that the agent automates all three stages: generating plans, validating plans, and executing plans. In reality, humans can be involved at any stage to aid with the process and mitigate risks.
* 人类专家可以提供计划、验证计划或执行部分计划。例如,对于智能体难以生成完整计划的复杂任务,人类专家可以提供高级计划,智能体可以在此基础上扩展。
* A human expert can provide a plan, validate a plan, or execute parts of a plan. For example, for complex tasks for which an agent has trouble generating the whole plan, a human expert can provide a high-level plan that the agent can expand upon.
* 如果计划涉及风险操作,例如更新数据库或合并代码更改,系统可以在执行前请求明确的人类批准,或者将执行这些操作委托给人类。为了实现这一点,你需要明确定义智能体对每个动作的自动化程度。
* If a plan involves risky operations, such as updating a database or merging a code change, the system can ask for explicit human approval before executing or defer to humans to execute these operations. To make this possible, you need to clearly define the level of automation an agent can have for each action.
总结一下,解决一个任务通常涉及以下过程。注意,反思对于智能体不是强制性的,但它会显著提升智能体的性能。
To summarize, solving a task typically involves the following processes. Note that reflection isn’t mandatory for an agent, but it’ll significantly boost the agent’s performance.
1. _计划生成_:提出完成此任务的计划。计划是一系列可管理的动作,因此这个过程也称为任务分解。
1. _Plan generation_: come up with a plan for accomplishing this task. A plan is a sequence of manageable actions, so this process is also called task decomposition.
2. _反思与错误纠正_:评估生成的计划。如果计划不好,生成一个新的。
2. _Reflection and error correction_: evaluate the generated plan. If it’s a bad plan, generate a new one.
3. _执行_:执行生成计划中列出的动作。这通常涉及调用特定函数。
3. _Execution_: take actions outlined in the generated plan. This often involves calling specific functions.
4. _反思与错误纠正_:收到动作结果后,评估这些结果并确定目标是否已完成。识别并纠正错误。如果目标未完成,生成一个新计划。
4. _Reflection and error correction_: upon receiving the action outcomes, evaluate these outcomes and determine whether the goal has been accomplished. Identify and correct mistakes. If the goal is not completed, generate a new plan.
你已经在这本书中看到了一些计划生成和反思的技术。当你要求模型“逐步思考”时,你是在要求它分解任务。当你要求模型“验证你的答案是否正确”时,你是在要求它反思。
You’ve already seen some techniques for plan generation and reflection in this book. When you ask a model to “think step by step”, you’re asking it to decompose a task. When you ask a model to “verify if your answer is correct”, you’re asking it to reflect.
一个悬而未决的问题是基础模型在规划方面的能力如何。许多研究者认为,基础模型——至少那些基于自回归语言模型构建的——无法进行规划。Meta 的首席 AI 科学家 Yann LeCun 明确表示,自回归 LLM 无法规划(2023)。
An open question is how well foundation models can plan. Many researchers believe that foundation models, at least those built on top of autoregressive language models, cannot. Meta’s Chief AI Scientist Yann LeCun states unequivocably that autoregressive LLMs can’t plan (2023).
尽管有很多轶事证据表明 LLM 的规划能力很差,但尚不清楚这是因为我们不知道如何正确使用 LLM,还是因为 LLM 从根本上就无法规划。
While there is a lot of anecdotal evidence that LLMs are poor planners, it’s unclear whether it’s because we don’t know how to use LLMs the right way or because LLMs, fundamentally, can’t plan.
规划的核心是一个搜索问题。你在通往目标的不同路径中进行搜索,预测每条路径的结果(奖励),然后选择最有前景的路径。通常,你可能会发现不存在任何能带你到达目标的路径。
Planning, at its core, is a search problem. You search among different paths towards the goal, predict the outcome (reward) of each path, and pick the path with the most promising outcome. Often, you might determine that no path exists that can take you to the goal.
搜索通常需要_回溯_。例如,假设你在某一步有两个可能的动作:A 和 B。在执行动作 A 后,你进入了一个没有前景的状态,因此你需要回溯到之前的状态以执行动作 B。
Search often requires _backtracking_. For example, imagine you’re at a step where there are two possible actions: A and B. After taking action A, you enter a state that’s not promising, so you need to backtrack to the previous state to take action B.
有些人认为,自回归模型只能生成前向动作,无法回溯以生成替代动作。因此,他们得出结论:自回归模型无法规划。然而,这并不一定正确。在执行了动作 A 的路径后,如果模型确定这条路径不合理,它可以改用动作 B 来修正路径,从而实现回溯。模型也可以随时重新开始并选择另一条路径。
Some people argue that an autoregressive model can only generate forward actions. It can’t backtrack to generate alternate actions. Because of this, they conclude that autoregressive models can’t plan. However, this isn’t necessarily true. After executing a path with action A, if the model determines that this path doesn’t make sense, it can revise the path using action B instead, effectively backtracking. The model can also always start over and choose another path.
LLM 规划能力差也可能是因为它们没有被赋予规划所需的工具。要进行规划,不仅需要知道可用的动作,还需要知道_每个动作的潜在结果_。举个简单的例子,假设你想爬上一座山。你的潜在动作是右转、左转、掉头或直行。然而,如果右转会让你掉下悬崖,你可能就不会考虑这个动作。用技术术语来说,一个动作将你从一个状态带到另一个状态,而为了决定是否采取某个动作,必须知道结果状态。
It’s also possible that LLMs are poor planners because they aren’t given the toolings needed to plan. To plan, it’s necessary to know not only the available actions but also _the potential outcome of each action_. As a simple example, let’s say you want to walk up a mountain. Your potential actions are turn right, turn left, turn around, or go straight ahead. However, if turning right will cause you to fall off the cliff, you might not consider this action. In technical terms, an action takes you from one state to another, and it’s necessary to know the outcome state to determine whether to take an action.
这意味着,仅仅像流行的思维链提示技术那样提示模型生成一系列动作是不够的。论文《用语言模型推理即用世界模型规划》(Hao 等人,2023)认为,LLM 由于包含了大量关于世界的信息,能够预测每个动作的结果。这个 LLM 可以利用这种结果预测来生成连贯的计划。
This means that prompting a model to generate only a sequence of actions like what the popular chain-of-thought prompting technique does isn’t sufficient. The paper “Reasoning with Language Model is Planning with World Model” (Hao et al., 2023) argues that an LLM, by containing so much information about the world, is capable of predicting the outcome of each action. This LLM can incorporate this outcome prediction to generate coherent plans.
即使 AI 无法规划,它仍然可以成为规划器的一部分。有可能通过为 LLM 配备搜索工具和状态跟踪系统来帮助它进行规划。
Even if AI can’t plan, it can still be a part of a planner. It might be possible to augment an LLM with a search tool and state tracking system to help it plan.
_智能体_是强化学习中的一个核心概念,维基百科将其定义为一个领域,“_关注智能体如何在动态环境中采取行动以最大化累积奖励。_”
The _agent_ is a core concept in RL, which is defined in Wikipedia as a field “_concerned with how an intelligent agent ought to take actions in a dynamic environment in order to maximize the cumulative reward._”
RL 智能体和 FM 智能体在许多方面相似。它们都由其环境和可能的动作来表征。主要区别在于它们的规划器的工作方式。
RL agents and FM agents are similar in many ways. They are both characterized by their environments and possible actions. The main difference is in how their planners work.
* 在 RL 智能体中,规划器由 RL 算法训练。训练这个 RL 规划器可能需要大量的时间和资源。 * 在 FM 智能体中,模型本身就是规划器。这个模型可以通过提示或微调来提高其规划能力,通常需要较少的时间和资源。
* In an RL agent, the planner is trained by an RL algorithm. Training this RL planner can require a lot of time and resources.
然而,没有什么能阻止 FM 智能体结合 RL 算法来提高其性能。我怀疑从长远来看,FM 智能体和 RL 智能体会融合。
* In an FM agent, the model is the planner. This model can be prompted or finetuned to improve its planning capabilities, and generally requires less time and fewer resources.
However, there’s nothing to prevent an FM agent from incorporating RL algorithms to improve its performance. I suspect that in the long run, FM agents and RL agents will merge.
将模型转化为计划生成器的最简单方法是使用提示工程。假设你想创建一个智能体来帮助客户了解 Kitty Vogue 的产品。你给这个智能体提供三个外部工具:按价格检索产品、检索热门产品和检索产品信息。以下是一个用于计划生成的提示示例。此提示仅用于说明目的。生产环境中的提示可能更复杂。
The simplest way to turn a model into a plan generator is with prompt engineering. Imagine that you want to create an agent to help customers learn about products at Kitty Vogue. You give this agent access to three external tools: retrieve products by price, retrieve top products, and retrieve product information. Here’s an example of a prompt for plan generation. This prompt is for illustration purposes only. Production prompts are likely more complex.
提出一个解决任务的计划。你可以使用 5 个动作:
Propose a plan to solve the task. You have access to 5 actions:
* fetch_top_products(start_date, end_date, num_products)
* fetch_top_products(start_date, end_date, num_products)
* generate_query(task_history, tool_output)
* generate_query(task_history, tool_output)
计划必须是一系列有效的动作。
The plan must be a sequence of valid actions.
计划:[fetch_product_info, generate_query, generate_response]
Plan: [fetch_product_info, generate_query, generate_response]
任务:“上周最畅销的产品是什么?”
Task: "What was the best selling product last week?"
计划:[fetch_top_products, generate_query, generate_response]
Plan: [fetch_top_products, generate_query, generate_response]
关于这个示例,有两点需要注意:
There are two things to note about this example:
* 这里使用的计划格式——一个函数列表,其参数由智能体推断——只是构建智能体控制流的众多方式之一。
* The plan format used here—a list of functions whose parameters are inferred by the agent—is just one of many ways to structure the agent control flow.
函数接收任务的当前历史记录和最近的工具输出,以生成一个查询,供响应生成器使用。每一步的工具输出都会添加到任务的历史记录中。
function takes in the task’s current history and the most recent tool outputs to generate a query to be fed into the response generator. The tool output at each step is added to the task’s history.
给定用户输入“上周最畅销产品的价格是多少”,生成的计划可能如下所示:
Given the user input “What’s the price of the best-selling product last week”, a generated plan might look like this:
你可能会想,“每个函数所需的参数呢?”确切的参数很难提前预测,因为它们通常是从之前的工具输出中提取的。如果第一步,
You might wonder, “What about the parameters needed for each function?” The exact parameters are hard to predict in advance since they are often extracted from the previous tool outputs. If the first step,
输出“2030-09-13”,智能体可以推断下一步的参数应该使用以下参数调用:
, outputs “2030-09-13”, the agent can reason that the parameters for the next step should be called with the following parameters:
通常,没有足够的信息来确定函数的确切参数值。例如,如果用户问“最畅销产品的平均价格是多少?”,以下问题的答案并不明确:
Often, there’s insufficient information to determine the exact parameter values for a function. For example, if a user asks “What’s the average price of best-selling products?”, the answers to the following questions are unclear:
* 用户想查看多少种最畅销产品?
* How many best-selling products the user wants to look at?
* 用户想要上周、上个月还是所有时间的最畅销产品?
* Does the user want the best-selling products last week, last month, or of all time?
这意味着模型经常需要猜测,而猜测可能是错误的。
This means that models frequently have to make guesses, and guesses can be wrong.
由于动作序列和相关参数都是由 AI 模型生成的,因此它们可能产生幻觉。幻觉可能导致模型调用无效函数,或调用有效函数但参数错误。用于提高模型整体性能的技术也可用于改进模型的规划能力。
Because both the action sequence and the associated parameters are generated by AI models, they can be hallucinated. Hallucinations can cause the model to call an invalid function or call a valid function but with wrong parameters. Techniques for improving a model’s performance in general can be used to improve a model’s planning capabilities.
许多模型提供商为其模型提供工具使用功能,有效地将模型转变为智能体。工具是一个函数。因此,调用工具通常被称为_函数调用_。不同的模型 API 工作方式不同,但通常函数调用如下工作:
Many model providers offer tool use for their models, effectively turning their models into agents. A tool is a function. Invoking a tool is, therefore, often called _function calling_. Different model APIs work differently, but in general, function calling works as follows:
1. _创建工具清单。_ 声明你可能希望模型使用的所有工具。每个工具通过其执行入口点(例如,函数名)、参数及其文档(例如,函数的作用和所需参数)来描述。
1. _Create a tool inventory._ Declare all the tools that you might want a model to use. Each tool is described by its execution entry point (e.g., its function name), its parameters, and its documentation (e.g., what the function does and what parameters it needs).
2. _指定智能体可以为查询使用哪些工具。_
2. _Specify what tools the agent can use for a query._
由于不同的查询可能需要不同的工具,许多 API 允许你指定每个查询要使用的已声明工具列表。有些还允许你通过以下设置进一步控制工具使用:
Because different queries might need different tools, many APIs let you specify a list of declared tools to be used per query. Some let you control tool use further by the following settings:
函数调用如图 6-10 所示。这是用伪代码编写的,以使其代表多个 API。要使用特定的 API,请参考其文档。
Function calling is illustrated in Figure 6-10. This is written in pseudocode to make it representative of multiple APIs. To use a specific API, please refer to its documentation.
给定一个查询,如图 6-10 中定义的智能体将自动生成要使用的工具及其参数。一些函数调用 API 将确保只生成有效的函数,尽管它们无法保证正确的参数值。
Given a query, an agent defined as in Figure 6-10 will automatically generate what tools to use and their parameters. Some function calling APIs will make sure that only valid functions are generated, though they won’t be able to guarantee the correct parameter values.
例如,给定用户查询“40 磅是多少千克?”,智能体可能决定需要工具
For example, given the user query “How many kilograms are 40 pounds?”, the agent might decide that it needs the tool
参数值为 40。智能体的响应可能如下所示。
with one parameter value of 40. The agent’s response might look like this.
根据此响应,你可以调用函数
From this response, you can evoke the function
并使用其输出生成对用户的响应。
and use its output to generate a response to the users.
计划是概述完成任务所需步骤的路线图。路线图可以有不同粒度。例如,为一年做规划时,按季度划分的计划比按月划分的计划更宏观,而按月划分的计划又比按周划分的计划更宏观。
A plan is a roadmap outlining the steps needed to accomplish a task. A roadmap can be of different levels of granularity. To plan for a year, a quarter-by-quarter plan is higher-level than a month-by-month plan, which is, in turn, higher-level than a week-to-week plan.
存在规划与执行之间的权衡。详细的计划更难生成,但更容易执行;宏观的计划更容易生成,但更难执行。规避这种权衡的一种方法是分层规划。首先,使用规划器生成一个宏观计划,例如按季度划分的计划。然后,对于每个季度,使用相同或不同的规划器生成按月划分的计划。
There’s a planning/execution tradeoff. A detailed plan is harder to generate, but easier to execute. A higher-level plan is easier to generate, but harder to execute. An approach to circumvent this tradeoff is to plan hierarchically. First, use a planner to generate a high-level plan, such as a quarter-to-quarter plan. Then, for each quarter, use the same or a different planner to generate a month-to-month plan.
到目前为止,所有生成计划的示例都使用了精确的函数名称,这非常细粒度。这种方法的一个问题是,智能体的工具库会随时间变化。例如,获取当前日期的函数。当工具发生变化时,你需要更新提示词和所有示例。使用精确的函数名称还会使得在不同用例(具有不同工具 API)之间复用规划器变得更加困难。
So far, all examples of generated plans use the exact function names, which is very granular. A problem with this approach is that an agent’s tool inventory can change over time. For example, the function to get the current date
如果你之前微调了一个模型,使其基于旧工具库生成计划,那么你需要在新工具库上再次微调该模型。
. When a tool changes, you’ll need to update your prompt and all your examples. Using the exact function names also makes it harder to reuse a planner across different use cases with different tool APIs.
为了避免这个问题,计划也可以使用更自然的语言来生成,这比特定领域的函数名称更宏观。例如,对于查询“上周最畅销产品的价格是多少?”,可以指示智能体输出如下所示的计划:
If you’ve previously finetuned a model to generate plans based on the old tool inventory, you’ll need to finetune the model again on the new tool inventory.
检索上周最畅销的产品
To avoid this problem, plans can also be generated using a more natural language, which is higher-level than domain-specific function names. For example, given the query “What’s the price of the best-selling product last week”, an agent can be instructed to output a plan that looks like this:
使用更自然的语言有助于你的计划生成器对工具 API 的变化具有鲁棒性。如果你的模型主要是在自然语言上训练的,它可能更擅长理解和生成自然语言计划,并且更不容易产生幻觉。
retrieve the best-selling product last week
这种方法的缺点是你需要一个翻译器将每个自然语言动作转换为可执行的命令。Chameleon(Lu 等人,2023)将此翻译器称为程序生成器。然而,翻译比规划简单得多,可以由较弱的模型完成,且幻觉风险较低。
Using more natural language helps your plan generator become robust to changes in tool APIs. If your model was trained mostly on natural language, it’ll likely be better at understanding and generating plans in natural language and less likely to hallucinate.
The downside of this approach is that you need a translator to translate each natural language action into executable commands. Chameleon (Lu et al., 2023) calls this translator a program generator. However, translating is a much simpler task than planning and can be done by weaker models with a lower risk of hallucination.
到目前为止的计划示例都是顺序执行的:计划中的下一个动作总是在前一个动作完成后执行。动作可执行的顺序称为控制流。顺序形式只是控制流的一种类型。其他类型的控制流包括并行、if 语句和 for 循环。下面的列表概述了每种控制流,包括顺序形式以作比较:
The plan examples so far have been sequential: the next action in the plan is _always_ executed after the previous action is done. The order in which actions can be executed is called a _control flow_. The sequential form is just one type of control flow. Other types of control flows include the parallel, if statement, and for loop. The list below provides an overview of each control flow, including sequential for comparison:
在任务 A 完成后执行任务 B,可能是因为任务 B 依赖于任务 A。例如,SQL 查询只有在从自然语言输入翻译后才能执行。
Executing task B after task A is complete, possibly because task B depends on task A. For example, the SQL query can only be executed after it’s been translated from the natural language input.
同时执行任务 A 和任务 B。例如,对于查询“找出售价低于 100 美元的最畅销产品”,智能体可能首先检索前 100 个最畅销产品,然后对每个产品检索其价格。
Executing tasks A and B at the same time. For example, given the query “Find me best-selling products under $100”, an agent might first retrieve the top 100 best-selling products and, for each of these products, retrieve its price.
根据前一步的输出执行任务 B 或任务 C。例如,智能体首先检查 NVIDIA 的收益报告。根据这份报告,它可以决定卖出或买入 NVIDIA 股票。Anthropic 的博文将这种模式称为“路由”。
Executing task B or task C depending on the output from the previous step. For example, the agent first checks NVIDIA’s earnings report. Based on this report, it can then decide to sell or buy NVIDIA stocks. Anthropic’s post calls this pattern “routing”.
重复执行任务 A 直到满足特定条件。例如,不断生成随机数直到得到一个质数。
Repeat executing task A until a specific condition is met. For example, keep on generating random numbers until a prime number.
这些不同的控制流在图 6-11 中进行了可视化。
These different control flows are visualized in Figure 6-11.
在传统软件工程中,控制流的条件是精确的。而在基于 AI 的智能体中,AI 模型决定控制流。具有非顺序控制流的计划更难生成,也更难转换为可执行命令。
In traditional software engineering, conditions for control flows are exact. With AI-powered agents, AI models determine control flows. Plans with non-sequential control flows are more difficult to both generate and translate into executable commands.
即使是最佳计划也需要不断评估和调整,以最大化其成功概率。虽然反思对于智能体运行并非严格必要,但对于智能体取得成功却是必要的。
Even the best plans need to be constantly evaluated and adjusted to maximize their chance of success. While reflection isn’t strictly necessary for an agent to operate, it’s necessary for an agent to succeed.
在任务过程中,有许多地方可以进行反思:
There are many places during a task process where reflection can be useful:
* 在收到用户查询后,评估请求是否可行。
* After receiving a user query to evaluate if the request is feasible.
* 在初始计划生成后,评估计划是否合理。
* After the initial plan generation to evaluate whether the plan makes sense.
* 在每个执行步骤后,评估是否在正确的轨道上。
* After each execution step to evaluate if it’s on the right track.
* 在整个计划执行完毕后,确定任务是否已完成。
* After the whole plan has been executed to determine if the task has been accomplished.
反思和纠错是两种相辅相成的不同机制。反思产生洞察,有助于发现需要纠正的错误。
Reflection and error correction are two different mechanisms that go hand in hand. Reflection generates insights that help uncover errors to be corrected.
反思可以通过同一智能体使用自我批评提示来完成,也可以通过单独的组件(如专门的评分器:为每个结果输出具体分数的模型)来实现。
Reflection can be done with the same agent with self-critique prompts. It can also be done with a separate component, such as a specialized scorer: a model that outputs a concrete score for each outcome.
由 ReAct(Yao 等人,2022)首次提出,将推理和行动交错进行已成为智能体的常见模式。Yao 等人使用“推理”一词来涵盖规划和反思。在每个步骤中,智能体被要求解释其思考(规划)、采取行动,然后分析观察结果(反思),直到智能体认为任务完成。通常通过示例提示智能体按以下格式生成输出:
First proposed by ReAct (Yao et al., 2022), interleaving re asoning and ac tion has become a common pattern for agents. Yao et al. used the term “reasoning” to encompass both planning and reflection. At each step, the agent is asked to explain its thinking (planning), take actions, then analyze observations (reflection), until the task is considered finished by the agent. The agent is typically prompted, using examples, to generate outputs in the following format:
… [继续直到反思确定任务完成] …
… [continue until reflection determines that the task is finished] …
图 6-12 展示了一个遵循 ReAct 框架的智能体回答 HotpotQA(Yang 等人,2018)问题的示例,HotpotQA 是一个多跳问答基准。
Figure 6-12 shows an example of an agent following the ReAct framework responding to a question from HotpotQA (Yang et al., 2018), a benchmark for multi-hop question answering.
你可以在多智能体设置中实现反思:一个智能体规划和执行行动,另一个智能体在每个步骤或若干步骤后评估结果。
You can implement reflection in a multi-agent setting: one agent plans and takes actions and another agent evaluates the outcome after each step or after a number of steps.
如果智能体的响应未能完成任务,你可以提示智能体反思失败原因及改进方法。基于此建议,智能体生成新计划。这使得智能体能够从错误中学习。
If the agent’s response failed to accomplish the task, you can prompt the agent to reflect on why it failed and how to improve. Based on this suggestion, the agent generates a new plan. This allows agents to learn from their mistakes.
例如,在代码生成任务中,评估器可能评估生成的代码在三分之一的测试用例上失败。智能体随后反思失败是因为没有考虑所有数字均为负数的数组。执行者随后生成新代码,考虑了全负数数组。
For example, given a coding generation task, an evaluator might evaluate that the generated code fails ⅓ of the test cases. The agent then reflects that it failed because it didn’t take into account arrays where all numbers are negative. The actor then generates new code, taking into account all-negative arrays.
这是 Reflexion(Shinn 等人,2023)采用的方法。在该框架中,反思分为两个模块:一个评估结果的评估器和一个分析错误原因的自反思模块。图 6-13 展示了 Reflexion 智能体的运行示例。作者使用“轨迹”一词指代计划。在每个步骤中,经过评估和自反思后,智能体提出新轨迹。
This is the approach that Reflexion (Shinn et al., 2023) took. In this framework, reflection is separated into two modules: an evaluator that evaluates the outcome and a self-reflection module that analyzes what went wrong. Figure 6-13 shows examples of Reflexion agents in action. The authors used the term “trajectory” to refer to a plan. At each step, after evaluation and self-reflection, the agent proposes a new trajectory.
与计划生成相比,反思相对容易实现,并能带来令人惊讶的性能提升。这种方法的缺点是延迟和成本。思考、观察以及有时行动需要生成大量词元,这增加了成本和用户感知的延迟,尤其是对于具有许多中间步骤的任务。为了促使智能体遵循格式,ReAct 和 Reflexion 的作者在提示中使用了大量示例。这增加了计算输入词元的成本,并减少了可用于其他信息的上下文空间。
Compared to plan generation, reflection is relatively easy to implement and can bring surprisingly good performance improvement. The downside of this approach is latency and cost. Thoughts, observations, and sometimes actions can take a lot of tokens to generate, which increases cost and user-perceived latency, especially for tasks with many intermediate steps. To nudge their agents to follow the format, both ReAct and Reflexion authors used plenty of examples in their prompts. This increases the cost of computing input tokens and reduces the context space available for other information.
由于工具通常在任务成功中扮演关键角色,工具选择需要仔细考量。赋予智能体的工具取决于环境和任务,但也取决于驱动智能体的 AI 模型。
Because tools often play a crucial role in a task’s success, tool selection requires careful consideration. The tools to give your agent depend on the environment and the task, but also depends on the AI model that powers the agent.
没有万无一失的指南来选择最佳工具集。智能体文献包含各种各样的工具清单。例如:
There’s no foolproof guide on how to select the best set of tools. Agent literature consists of a wide range of tool inventories. For example:
* Toolformer (Schick et al., 2023) 微调了 GPT-J 以学习 5 种工具。
* Toolformer (Schick et al., 2023) finetuned GPT-J to learn 5 tools.
* Chameleon (Lu et al., 2023) 使用了 13 种工具。
* Chameleon (Lu et al., 2023) uses 13 tools.
* Gorilla (Patil et al., 2023) 尝试提示智能体在 1,645 个 API 中选择正确的 API 调用。
* Gorilla (Patil et al., 2023) attempted to prompt agents to select the right API call among 1,645 APIs.
更多的工具赋予智能体更强的能力。然而,工具越多,高效使用它们就越困难。这类似于人类掌握大量工具更难。添加工具也意味着增加工具描述,这可能超出模型的上下文窗口。
More tools give the agent more capabilities. However, the more tools there are, the harder it is to efficiently use them. It’s similar to how it’s harder for humans to master a large set of tools. Adding tools also means increasing tool descriptions, which might not fit into a model’s context.
与构建 AI 应用时的许多其他决策一样,工具选择需要实验和分析。以下是一些可以帮助你决策的方法:
Like many other decisions while building AI applications, tool selection requires experimentation and analysis. Here are a few things you can do to help you decide:
* 比较智能体在不同工具集下的表现。
* Compare how an agent performs with different sets of tools.
* 进行消融研究,观察如果从工具库中移除某个工具,智能体的性能下降多少。如果移除某个工具后性能没有下降,就移除它。
* Do an ablation study to see how much the agent’s performance drops if a tool is removed from its inventory. If a tool can be removed without a performance drop, remove it.
* 寻找智能体经常出错的工具。如果某个工具对智能体来说太难使用——例如,大量的提示甚至微调都无法让模型学会使用它——就更换该工具。
* Look for tools that the agent frequently makes mistakes on. If a tool proves too hard for the agent to use—for example, extensive prompting and even finetuning can’t get the model to learn to use it—change the tool.
* 绘制工具调用的分布图,查看哪些工具最常用,哪些最不常用。图 6-14 展示了 Chameleon (Lu et al., 2023)中 GPT-4 和 ChatGPT 工具使用模式的差异。
* Plot the distribution of tool calls to see what tools are most used and what tools are least used. Figure 6-14 shows the differences in tool use patterns of GPT-4 and ChatGPT in Chameleon (Lu et al., 2023).
Chameleon (Lu et al., 2023)的实验也证明了两点:
Experiments by Chameleon (Lu et al., 2023) also demonstrate two points:
1. 不同的任务需要不同的工具。科学问答任务 ScienceQA 比表格数学问题求解任务 TabMWP 更依赖知识检索工具。
1. Different tasks require different tools. ScienceQA, the science question answering task, relies much more on knowledge retrieval tools than TabMWP, a tabular math problem-solving task.
2. 不同的模型有不同的工具偏好。例如,GPT-4 似乎比 ChatGPT 选择更广泛的工具集。ChatGPT 似乎偏爱图像描述,而 GPT-4 似乎偏爱知识检索。
2. Different models have different tool preferences. For example, GPT-4 seems to select a wider set of tools than ChatGPT. ChatGPT seems to favor image captioning, while GPT-4 seems to favor knowledge retrieval.
作为人类,我们提高生产力不仅通过使用给定的工具,还通过从简单工具中逐步创造更强大的工具。AI 能否从其初始工具中创造新工具?
As humans, we become more productive not just by using the tools we’re given, but also by creating progressively more powerful tools from simpler ones. Can AI create new tools from its initial tools?
Chameleon (Lu et al., 2023)提出了工具转换的研究:在工具 X 之后,智能体调用工具 Y 的可能性有多大?图 6-15 展示了一个工具转换的例子。如果两个工具经常一起使用,它们可以合并成一个更大的工具。如果智能体知道这一信息,它本身就可以组合初始工具,持续构建更复杂的工具。
Chameleon (Lu et al., 2023) proposes the study of tool transition: after tool X, how likely is the agent to call tool Y? Figure 6-15 shows an example of tool transition. If two tools are frequently used together, they can be combined into a bigger tool. If an agent is aware of this information, the agent itself can combine initial tools to continually build more complex tools.
Voyager (Wang et al., 2023)提出了一个技能管理器,用于跟踪智能体获取的新技能(工具),以便后续重用。每个技能都是一个编码程序。当技能管理器确定一个新创建的技能有用时(例如,因为它成功帮助智能体完成了任务),它会将该技能添加到技能库(概念上类似于工具库)。该技能可以在以后被检索用于其他任务。
Vogager (Wang et al., 2023) proposes a skill manager to keep track of new skills (tools) that an agent acquires for later reuse. Each skill is a coding program. When the skill manager determines a newly created skill is to be useful (e.g., because it’s successfully helped an agent accomplish a task), it adds this skill to the skill library (conceptually similar to the tool inventory). This skill can be retrieved later to use for other tasks.
在本节前面,我们提到智能体在环境中的成功取决于其工具库和规划能力。任何一方面的失败都可能导致智能体失败。下一节将讨论智能体的不同失败模式以及如何评估它们。
Earlier in this section, we mentioned that the success of an agent in an environment depends on its tool inventory and its planning capabilities. Failures in either aspect can cause the agent to fail. The next section will discuss different failure modes of an agent and how to evaluate them.
评估关乎检测失败。智能体执行的任务越复杂,可能的失败点就越多。除了第 3 章和第 4 章讨论的所有 AI 应用共有的失败模式外,智能体还有由规划、工具执行和效率引起的独特失败。其中一些失败比其他更容易捕捉。
Evaluation is about detecting failures. The more complex a task an agent performs, the more possible failure points there are. Other than the failure modes common to all AI applications discussed in Chapters 3 and 4, agents also have unique failures caused by planning, tool execution, and efficiency. Some of the failures are easier to catch than others.
要评估一个智能体,需识别其失败模式并测量每种失败模式发生的频率。
To evaluate an agent, identify its failure modes and measure how often each of these failure modes happens.
规划是困难的,并且可能以多种方式失败。最常见的规划失败模式是工具使用失败。智能体可能生成包含一个或多个此类错误的计划。
Planning is hard and can fail in many ways. The most common mode of planning failure is tool use failure. The agent might generate a plan with one or more of these errors.
例如,它生成一个包含
For example, it generates a plan that contains
两个参数的计划,但该函数只需要一个参数,
with two parameters, but this function requires only one parameter,
* 有效工具,参数值不正确
* Valid tool, incorrect parameter values
,但将 lbs 的值设为 100,而实际应为 120。
, but uses the value 100 for lbs when it should be 120.
另一种规划失败模式是目标失败:智能体未能实现目标。这可能是因为计划没有解决任务,或者虽然解决了任务但没有遵循约束条件。为了说明这一点,假设你让模型规划一个从旧金山到印度的两周旅行,预算为 5000 美元。智能体可能规划了从旧金山到越南的旅行,或者规划了一个从旧金山到印度但花费远超预算的两周旅行。
Another mode of planning failure is goal failure: the agent fails to achieve the goal. This can be because the plan doesn’t solve a task, or it solves the task without following the constraints. To illustrate this, imagine you ask the model to plan a two-week trip from San Francisco to India with a budget of $5,000. The agent might plan a trip from San Francisco to Vietnam, or plan you a two-week trip from San Francisco to India that will cost you way over the budget.
智能体评估中常被忽略的一个常见约束是时间。在许多情况下,智能体花费的时间并不重要,因为你可以将任务分配给智能体,只需在任务完成时检查即可。然而,在许多情况下,智能体随着时间推移会变得不那么有用。例如,如果你让智能体准备一份资助提案,而智能体在资助截止日期后才完成,那么智能体就没有多大帮助。
A common constraint that is often overlooked by agent evaluation is time. In many cases, the time an agent takes matters less because you can assign a task to an agent and only need to check in when it’s done. However, in many cases, the agent becomes less useful with time. For example, if you ask an agent to prepare a grant proposal and the agent finishes it after the grant deadline, the agent isn’t very helpful.
一种有趣的规划失败模式是由反思错误引起的。智能体确信自己完成了任务,但实际上并未完成。例如,你让智能体将 50 人分配到 30 个酒店房间。智能体可能只分配了 40 人,并坚称任务已完成。
An interesting mode of planning failure is caused by errors in reflection. The agent is convinced that it’s accomplished a task when it hasn’t. For example, you ask the agent to assign 50 people to 30 hotel rooms. The agent might assign only 40 people and insist that the task has been accomplished.
为了评估智能体的规划失败,一种方法是创建一个规划数据集,其中每个示例是一个元组
To evaluate an agent for planning failures, one option is to create a planning dataset where each example is a tuple
。对于每个任务,使用智能体生成 K 个计划。计算以下指标:
. For each task, use the agent to generate a K number of plans. Compute the following metrics:
1. 在所有生成的计划中,有多少是有效的?
1. Out of all generated plans, how many are valid?
2. 对于给定任务,智能体需要生成多少个计划才能得到一个有效计划?
2. For a given task, how many plans does the agent have to generate to get a valid plan?
3. 在所有工具调用中,有多少是有效的?
3. Out of all tool calls, how many are valid?
5. 有效工具被调用时,参数无效的频率如何?
5. How often are valid tools called with invalid parameters?
6. 有效工具被调用时,参数值不正确的频率如何?
6. How often are valid tools called with incorrect parameter values?
分析智能体的输出以寻找模式。智能体在哪些类型的任务上更容易失败?你有什么假设来解释原因?模型经常在哪些工具上犯错?有些工具可能对智能体来说更难使用。你可以通过更好的提示、更多的示例或微调来提高智能体使用困难工具的能力。如果所有这些都失败了,你可以考虑将这个工具替换为更容易使用的工具。
Analyze the agent’s outputs for patterns. What types of tasks does the agent fail more on? Do you have a hypothesis why? What tools does the model frequently make mistakes with? Some tools might be harder for an agent to use. You can improve an agent’s ability to use a challenging tool by better prompting, more examples, or finetuning. If all fail, you might consider swapping out this tool for something easier to use.
工具故障发生在使用了正确的工具,但工具输出错误的情况下。一种故障模式是工具直接给出错误的输出。例如,图像描述器返回错误的描述,或 SQL 查询生成器返回错误的 SQL 查询。
Tool failures happen when the correct tool is used, but the tool output is wrong. One failure mode is when a tool just gives the wrong outputs. For example, an image captioner returns a wrong description, or an SQL query generator returns a wrong SQL query.
如果智能体仅生成高级计划,并且涉及一个翻译模块将每个计划动作转换为可执行命令,则可能因翻译错误而发生故障。
If the agent generates only high-level plans and a translation module is involved in translating from each planned action to executable commands, failures can happen because of translation errors.
工具故障是工具相关的。每个工具需要独立测试。始终打印每个工具调用及其输出,以便检查和评估。如果你有翻译器,创建基准来评估它。
Tool failures are tool-dependent. Each tool needs to be tested independently. Always print out each tool call and its output so that you can inspect and evaluate them. If you have a translator, create benchmarks to evaluate it.
检测缺失工具故障需要理解应该使用哪些工具。如果你的智能体在特定领域频繁失败,这可能是因为它缺少该领域的工具。与人类领域专家合作,观察他们会使用哪些工具。
Detecting missing tool failures requires an understanding of what tools should be used. If your agent frequently fails on a specific domain, this might be because it lacks tools for this domain. Work with human domain experts and observe what tools they would use.
智能体可能使用正确的工具生成有效的计划来完成任务,但可能效率低下。以下是一些你可能想要追踪以评估智能体效率的指标:
An agent might generate a valid plan using the right tools to accomplish a task, but it might be inefficient. Here are a few things you might want to track to evaluate an agent’s efficiency:
* 智能体平均需要多少步才能完成任务?
* How many steps does the agent need, on average, to complete a task?
* 智能体平均花费多少成本才能完成任务?
* How much does the agent cost, on average, to complete a task?
* 每个动作通常需要多长时间?是否有特别耗时或昂贵的动作?
* How long does each action typically take? Are there any actions that are especially time-consuming or expensive?
你可以将这些指标与你的基线进行比较,基线可以是另一个智能体或人类操作员。在比较 AI 智能体与人类智能体时,请记住人类和 AI 具有非常不同的操作模式,因此对人类来说高效的对于 AI 可能低效,反之亦然。例如,访问 100 个网页对于一次只能访问一个网页的人类智能体来说可能效率低下,但对于可以同时访问所有网页的 AI 智能体来说则微不足道。
You can compare these metrics with your baseline, which can be another agent or a human operator. When comparing AI agents to human agents, keep in mind that humans and AI have very different modes of operation, so what’s considered efficient for humans might be inefficient for AI and vice versa. For example, visiting 100 web pages might be inefficient for a human agent who can only visit one page at a time but trivial for an AI agent that can visit all the web pages at once.
其核心而言,智能体的概念相当简单。智能体由其运行的环境和可访问的工具集定义。在 AI 驱动的智能体中,AI 模型是大脑,利用其工具和来自环境的反馈来规划如何最好地完成任务。访问工具使模型能力大幅提升,因此智能体模式是不可避免的。
At its core, the concept of an agent is fairly simple. An agent is defined by the environment it operates in and the set of tools it has access to. In an AI-powered agent, the AI model is the brain that leverages its tools and feedback from the environment to plan how best to accomplish a task. Access to tools makes a model vastly more capable, so the agentic pattern is inevitable.
尽管“智能体”的概念听起来新颖,但它们建立在许多自 LLM 早期就已使用的概念之上,包括自我批评、思维链和结构化输出。
While the concept of “agents” sounds novel, they are built upon many concepts that have been used since the early days of LLMs, including self-critique, chain-of-thought, and structured outputs.
本文从概念上介绍了智能体的工作原理及其不同组成部分。在未来的文章中,我将讨论如何评估智能体框架。
This post covered conceptually how agents work and different components of an agent. In a future post, I’ll discuss how to evaluate agent frameworks.
智能体模式通常处理超出模型上下文限制的信息。一个补充模型上下文以处理信息的记忆系统可以显著增强智能体的能力。由于本文已经很长,我将在未来的博客文章中探讨记忆系统的工作原理。
The agentic pattern often deals with information that exceeds a model’s context limit. A memory system that supplements the model’s context in handling information can significantly enhance an agent’s capabilities. Since this post is already long, I’ll explore how a memory system works in a future blog post.
* [hi@[thiswebsite]](https://huyenchip.com/2025/01/07/agents.html)
* [hi@[thiswebsite]](https://huyenchip.com/2025/01/07/agents.html)
我致力于将 AI 投入生产。我撰写关于 AI 系统设计的文章。
I work to bring AI into production. I write about AI system design.
我目前还没有 Substack,但如果更多人推动我,我可能会开始写。
I don't have a Substack yet, but if more people nudge me, I might actually start it.