LLM Powered Autonomous Agents
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→日期:2023 年 6 月 23 日 | 预计阅读时间:31 分钟 | 作者:Lilian Weng 以大型语言模型(LLM)作为核心控制器构建智能体是一个很酷的概念。一些概念验证演示,如 AutoGPT、GPT-Engineer 和 BabyAGI,是鼓舞人心的例子。LLM 的潜力不仅限于生成写得好的文案、故事、论文和程序;它可以被框架化为一个强大的通用问题解决器。
Date: June 23, 2023 | Estimated Reading Time: 31 min | Author: Lilian Weng Building agents with LLM (large language model) as its core controller is a cool concept. Several proof-of-concepts demos, such as AutoGPT, GPT-Engineer and BabyAGI, serve as inspiring examples. The potentiality of LLM extends beyond generating well-written copies, stories, essays and programs; it can be framed as a powerful general problem solver.
日期:2023 年 6 月 23 日 | 预计阅读时间:31 分钟 | 作者:Lilian Weng
Date: June 23, 2023 | Estimated Reading Time: 31 min | Author: Lilian Weng
以 LLM(大语言模型)作为核心控制器构建智能体是一个很酷的概念。一些概念验证演示,如 AutoGPT、GPT-Engineer 和 BabyAGI,是鼓舞人心的例子。LLM 的潜力不仅限于生成文笔优美的文案、故事、文章和程序;它可以被塑造成一个强大的通用问题解决器。
Building agents with LLM (large language model) as its core controller is a cool concept. Several proof-of-concepts demos, such as AutoGPT, GPT-Engineer and BabyAGI, serve as inspiring examples. The potentiality of LLM extends beyond generating well-written copies, stories, essays and programs; it can be framed as a powerful general problem solver.
在基于大语言模型(LLM)的自主智能体系统中,LLM 充当智能体的大脑,并辅以几个关键组件:
In a LLM-powered autonomous agent system, LLM functions as the agent’s brain, complemented by several key components:
* 子目标与分解:智能体将大型任务分解为更小、更易管理的子目标,从而能够高效处理复杂任务。
* Subgoal and decomposition: The agent breaks down large tasks into smaller, manageable subgoals, enabling efficient handling of complex tasks.
* 反思与改进:智能体能够对过去的行为进行自我批评和自我反思,从错误中学习并在未来步骤中加以改进,从而提高最终结果的质量。
* Reflection and refinement: The agent can do self-criticism and self-reflection over past actions, learn from mistakes and refine them for future steps, thereby improving the quality of final results.
* 短期记忆:我认为所有上下文学习(参见提示工程)都是利用模型的短期记忆进行学习。
* Short-term memory: I would consider all the in-context learning (See Prompt Engineering) as utilizing short-term memory of the model to learn.
* 长期记忆:这使智能体能够长时间保留和回忆(无限)信息,通常通过利用外部向量存储和快速检索来实现。
* Long-term memory: This provides the agent with the capability to retain and recall (infinite) information over extended periods, often by leveraging an external vector store and fast retrieval.
* 智能体学习调用外部 API 以获取模型权重中缺失的额外信息(通常在预训练后难以更改),包括当前信息、代码执行能力、访问专有信息源等。
* The agent learns to call external APIs for extra information that is missing from the model weights (often hard to change after pre-training), including current information, code execution capability, access to proprietary information sources and more.
基于大语言模型(LLM)的自主智能体系统概览。
Overview of a LLM-powered autonomous agent system.
一个复杂的任务通常涉及许多步骤。智能体需要知道这些步骤是什么,并提前进行规划。
A complicated task usually involves many steps. An agent needs to know what they are and plan ahead.
思维链(CoT;Wei 等人,2022)已成为一种标准提示技术,用于增强模型在复杂任务上的性能。模型被指示“逐步思考”,以利用更多的测试时算力将困难任务分解为更小、更简单的步骤。CoT 将大任务转化为多个可管理的任务,并揭示模型思维过程的解释。
Chain of thought (CoT; Wei et al. 2022) has become a standard prompting technique for enhancing model performance on complex tasks. The model is instructed to “think step by step” to utilize more test-time computation to decompose hard tasks into smaller and simpler steps. CoT transforms big tasks into multiple manageable tasks and shed lights into an interpretation of the model’s thinking process.
思维树(Yao 等人,2023)通过在每个步骤探索多种推理可能性来扩展 CoT。它首先将问题分解为多个思维步骤,并为每个步骤生成多个思维,从而创建树状结构。搜索过程可以是 BFS(广度优先搜索)或 DFS(深度优先搜索),每个状态由分类器(通过提示)或多数投票进行评估。
Tree of Thoughts (Yao et al. 2023) extends CoT by exploring multiple reasoning possibilities at each step. It first decomposes the problem into multiple thought steps and generates multiple thoughts per step, creating a tree structure. The search process can be BFS (breadth-first search) or DFS (depth-first search) with each state evaluated by a classifier (via a prompt) or majority vote.
任务分解可以通过以下方式完成:(1)LLM 使用简单提示,如“实现 XYZ 的步骤。\n1.”、“实现 XYZ 的子目标是什么?”;(2)使用任务特定指令,例如写小说时使用“写一个故事大纲”;(3)人工输入。
Task decomposition can be done (1) by LLM with simple prompting like "Steps for XYZ.\n1.", "What are the subgoals for achieving XYZ?", (2) by using task-specific instructions; e.g. "Write a story outline." for writing a novel, or (3) with human inputs.
另一种截然不同的方法,LLM+P(Liu 等人,2023),涉及依赖外部经典规划器进行长程规划。该方法利用规划领域定义语言(PDDL)作为中间接口来描述规划问题。在此过程中,LLM(1)将问题翻译成“问题 PDDL”,然后(2)请求经典规划器基于现有的“领域 PDDL”生成 PDDL 计划,最后(3)将 PDDL 计划翻译回自然语言。本质上,规划步骤被外包给外部工具,假设存在特定领域的 PDDL 和合适的规划器,这在某些机器人设置中常见,但在许多其他领域中并不常见。
Another quite distinct approach, LLM+P (Liu et al. 2023), involves relying on an external classical planner to do long-horizon planning. This approach utilizes the Planning Domain Definition Language (PDDL) as an intermediate interface to describe the planning problem. In this process, LLM (1) translates the problem into “Problem PDDL”, then (2) requests a classical planner to generate a PDDL plan based on an existing “Domain PDDL”, and finally (3) translates the PDDL plan back into natural language. Essentially, the planning step is outsourced to an external tool, assuming the availability of domain-specific PDDL and a suitable planner which is common in certain robotic setups but not in many other domains.
自我反思是一个重要方面,它使自主智能体能够通过改进过去的行动决策和纠正之前的错误来迭代提升。在试错不可避免的现实世界任务中,它起着关键作用。
Self-reflection is a vital aspect that allows autonomous agents to improve iteratively by refining past action decisions and correcting previous mistakes. It plays a crucial role in real-world tasks where trial and error are inevitable.
ReAct 提示模板包含了让 LLM 思考的明确步骤,大致格式如下:
The ReAct prompt template incorporates explicit steps for LLM to think, roughly formatted as:
知识密集型任务(如 HotpotQA、FEVER)和决策任务(如 AlfWorld Env、WebShop)的推理轨迹示例。(图片来源:Yao et al. 2023)
Examples of reasoning trajectories for knowledge-intensive tasks (e.g. HotpotQA, FEVER) and decision-making tasks (e.g. AlfWorld Env, WebShop). (Image source: Yao et al. 2023).
在知识密集型任务和决策任务的实验中,ReAct 的表现都优于移除了“Thought: …”步骤的仅行动基线。
In both experiments on knowledge-intensive tasks and decision-making tasks, ReAct works better than the Act-only baseline where Thought: … step is removed.
Reflexion(Shinn & Labash 2023)是一个框架,为智能体配备动态记忆和自我反思能力,以提升推理技能。Reflexion 采用标准的强化学习设置,其中奖励模型提供简单的二元奖励,动作空间遵循 ReAct 的设置,即任务特定的动作空间通过语言增强以实现复杂的推理步骤。在每个动作 $a_{t}$ 之后,智能体计算一个启发式 $h_{t}$,并可选地根据自我反思结果决定重置环境以开始新的尝试。
Reflexion (Shinn & Labash 2023) is a framework to equip agents with dynamic memory and self-reflection capabilities to improve reasoning skills. Reflexion has a standard RL setup, in which the reward model provides a simple binary reward and the action space follows the setup in ReAct where the task-specific action space is augmented with language to enable complex reasoning steps. After each action $a_{t}$, the agent computes a heuristic $h_{t}$ and optionally may _decide to reset_ the environment to start a new trial depending on the self-reflection results.
Reflexion 框架示意图。(图片来源:Shinn & Labash, 2023)
Illustration of the Reflexion framework. (Image source: Shinn & Labash, 2023)
启发式函数决定何时轨迹低效或包含幻觉,并应停止。低效规划指耗时过长且未成功的轨迹。幻觉定义为在环境中遇到一系列连续相同动作导致相同观察的情况。
The heuristic function determines when the trajectory is inefficient or contains hallucination and should be stopped. Inefficient planning refers to trajectories that take too long without success. Hallucination is defined as encountering a sequence of consecutive identical actions that lead to the same observation in the environment.
自我反思通过向 LLM 展示两个示例来创建,每个示例是一对(失败轨迹,用于指导未来计划变更的理想反思)。然后,反思被添加到智能体的工作记忆中,最多三个,作为查询 LLM 的上下文。
Self-reflection is created by showing two-shot examples to LLM and each example is a pair of (failed trajectory, ideal reflection for guiding future changes in the plan). Then reflections are added into the agent’s working memory, up to three, to be used as context for querying LLM.
在 AlfWorld Env 和 HotpotQA 上的实验。在 AlfWorld 中,幻觉比低效规划更常见。(图片来源:Shinn & Labash, 2023)
Experiments on AlfWorld Env and HotpotQA. Hallucination is a more common failure than inefficient planning in AlfWorld. (Image source: Shinn & Labash, 2023)
Chain of Hindsight(CoH;Liu et al. 2023)通过向模型明确展示一系列过去的输出(每个输出都附有反馈)来鼓励模型改进自身输出。人类反馈数据是一个集合 $D_{h} = \left{\right. \left(\right. x , y_{i} , r_{i} , z_{i} \left.\right) \left.\right}_{i = 1}^{n}$,其中 $x$ 是提示,每个 $y_{i}$ 是模型补全,$r_{i}$ 是 $y_{i}$ 的人类评分,$z_{i}$ 是相应的人类提供的 hindsight 反馈。假设反馈元组按奖励排序,$r_{n} \geq r_{n - 1} \geq \hdots \geq r_{1}$。该过程是有监督微调,数据是形式为 $\tau_{h} = \left(\right. x , z_{i} , y_{i} , z_{j} , y_{j} , \ldots , z_{n} , y_{n} \left.\right)$ 的序列,其中 $\leq i \leq j \leq n$。模型被微调为仅预测 $y_{n}$,条件于序列前缀,从而使模型能够根据反馈序列自我反思以产生更好的输出。在测试时,模型可以选择性地接收多轮带人类标注者的指令。
Chain of Hindsight (CoH; Liu et al. 2023) encourages the model to improve on its own outputs by explicitly presenting it with a sequence of past outputs, each annotated with feedback. Human feedback data is a collection of $D_{h} = \left{\right. \left(\right. x , y_{i} , r_{i} , z_{i} \left.\right) \left.\right}_{i = 1}^{n}$, where $x$ is the prompt, each $y_{i}$ is a model completion, $r_{i}$ is the human rating of $y_{i}$, and $z_{i}$ is the corresponding human-provided hindsight feedback. Assume the feedback tuples are ranked by reward, $r_{n} \geq r_{n - 1} \geq \hdots \geq r_{1}$ The process is supervised fine-tuning where the data is a sequence in the form of $\tau_{h} = \left(\right. x , z_{i} , y_{i} , z_{j} , y_{j} , \ldots , z_{n} , y_{n} \left.\right)$, where $\leq i \leq j \leq n$. The model is finetuned to only predict $y_{n}$ where conditioned on the sequence prefix, such that the model can self-reflect to produce better output based on the feedback sequence. The model can optionally receive multiple rounds of instructions with human annotators at test time.
为避免过拟合,CoH 添加了一个正则化项以最大化预训练数据集的 log 似然。为避免捷径和复制(因为反馈序列中有许多常见词),他们在训练期间随机遮蔽 0% - 5% 的过去 token。
To avoid overfitting, CoH adds a regularization term to maximize the log-likelihood of the pre-training dataset. To avoid shortcutting and copying (because there are many common words in feedback sequences), they randomly mask 0% - 5% of past tokens during training.
他们实验中的训练数据集是 WebGPT 比较、来自人类反馈的摘要和人类偏好数据集的组合。
The training dataset in their experiments is a combination of WebGPT comparisons, summarization from human feedback and human preference dataset.
经过 CoH 微调后,模型可以遵循指令,在序列中逐步产生改进的输出。(图片来源:Liu et al. 2023)
After fine-tuning with CoH, the model can follow instructions to produce outputs with incremental improvement in a sequence. (Image source: Liu et al. 2023)
CoH 的思想是在上下文中呈现一个逐步改进的输出历史,并训练模型顺应趋势以产生更好的输出。算法蒸馏(AD;Laskin et al. 2023)将相同的想法应用于强化学习任务中的跨回合轨迹,其中算法被封装在一个长历史条件策略中。考虑到智能体与环境多次交互,并且在每一回合中智能体都会变得更好,AD 将这一学习历史串联起来并输入模型。因此,我们期望下一个预测动作能比之前的尝试带来更好的性能。目标是学习强化学习的过程,而不是训练一个任务特定的策略本身。
The idea of CoH is to present a history of sequentially improved outputs in context and train the model to take on the trend to produce better outputs. Algorithm Distillation (AD; Laskin et al. 2023) applies the same idea to cross-episode trajectories in reinforcement learning tasks, where an _algorithm_ is encapsulated in a long history-conditioned policy. Considering that an agent interacts with the environment many times and in each episode the agent gets a little better, AD concatenates this learning history and feeds that into the model. Hence we should expect the next predicted action to lead to better performance than previous trials. The goal is to learn the process of RL instead of training a task-specific policy itself.
算法蒸馏(AD)工作原理示意图。
Illustration of how Algorithm Distillation (AD) works.
论文假设,任何生成一组学习历史的算法都可以通过对动作进行行为克隆来蒸馏到神经网络中。历史数据由一组源策略生成,每个策略针对特定任务训练。在训练阶段,每次强化学习运行时,随机采样一个任务,并使用多回合历史的一个子序列进行训练,从而使学习到的策略与任务无关。
The paper hypothesizes that any algorithm that generates a set of learning histories can be distilled into a neural network by performing behavioral cloning over actions. The history data is generated by a set of source policies, each trained for a specific task. At the training stage, during each RL run, a random task is sampled and a subsequence of multi-episode history is used for training, such that the learned policy is task-agnostic.
实际上,模型的上下文窗口长度有限,因此回合应该足够短以构建多回合历史。2-4 回合的多回合上下文对于学习接近最优的上下文强化学习算法是必要的。上下文强化学习的出现需要足够长的上下文。
In reality, the model has limited context window length, so episodes should be short enough to construct multi-episode history. Multi-episodic contexts of 2-4 episodes are necessary to learn a near-optimal in-context RL algorithm. The emergence of in-context RL requires long enough context.
与三个基线(包括 ED(专家蒸馏,使用专家轨迹而非学习历史进行行为克隆)、源策略(用于通过 UCB 生成蒸馏轨迹)和 RL^2(Duan et al. 2017;作为上界,因为它需要在线强化学习))相比,AD 展示了上下文强化学习,性能接近 RL^2,尽管仅使用离线强化学习,并且学习速度远快于其他基线。当以源策略的部分训练历史为条件时,AD 的改进速度也远快于 ED 基线。
In comparison with three baselines, including ED (expert distillation, behavior cloning with expert trajectories instead of learning history), source policy (used for generating trajectories for distillation by UCB), RL^2 (Duan et al. 2017; used as upper bound since it needs online RL), AD demonstrates in-context RL with performance getting close to RL^2 despite only using offline RL and learns much faster than other baselines. When conditioned on partial training history of the source policy, AD also improves much faster than ED baseline.
AD、ED、源策略和 RL^2 在需要记忆和探索的环境中的比较。仅分配二元奖励。源策略在“黑暗”环境中使用 A3C 训练,在水迷宫中则使用 DQN 训练。
Comparison of AD, ED, source policy and RL^2 on environments that require memory and exploration. Only binary reward is assigned. The source policies are trained with A3C for "dark" environments and DQN for watermaze.
(衷心感谢 ChatGPT 协助我起草本节。在与 ChatGPT 的对话中,我学到了很多关于人脑和用于快速 MIPS 的数据结构的知识。)
(Big thank you to ChatGPT for helping me draft this section. I’ve learned a lot about the human brain and data structure for fast MIPS in my conversations with ChatGPT.)
记忆可以定义为用于获取、存储、保留和随后检索信息的过程。人脑中有几种类型的记忆。
Memory can be defined as the processes used to acquire, store, retain, and later retrieve information. There are several types of memory in human brains.
1. 感觉记忆:这是记忆的最早阶段,提供在原始刺激结束后保留感觉信息(视觉、听觉等)印象的能力。感觉记忆通常只持续几秒钟。子类别包括图像记忆(视觉)、回声记忆(听觉)和触觉记忆(触觉)。
1. Sensory Memory: This is the earliest stage of memory, providing the ability to retain impressions of sensory information (visual, auditory, etc) after the original stimuli have ended. Sensory memory typically only lasts for up to a few seconds. Subcategories include iconic memory (visual), echoic memory (auditory), and haptic memory (touch).
2. 短期记忆或工作记忆:它存储我们当前意识到并需要执行复杂认知任务(如学习和推理)的信息。短期记忆被认为容量约为 7 个项目(Miller 1956),持续 20-30 秒。
2. Short-Term Memory (STM) or Working Memory: It stores information that we are currently aware of and needed to carry out complex cognitive tasks such as learning and reasoning. Short-term memory is believed to have the capacity of about 7 items (Miller 1956) and lasts for 20-30 seconds.
3. 长期记忆:长期记忆可以存储信息相当长的时间,从几天到几十年,且存储容量基本无限。长期记忆有两种子类型:
3. Long-Term Memory (LTM): Long-term memory can store information for a remarkably long time, ranging from a few days to decades, with an essentially unlimited storage capacity. There are two subtypes of LTM:
* 外显/陈述性记忆:这是对事实和事件的记忆,指那些可以有意识回忆的记忆,包括情景记忆(事件和经历)和语义记忆(事实和概念)。
* Explicit / declarative memory: This is memory of facts and events, and refers to those memories that can be consciously recalled, including episodic memory (events and experiences) and semantic memory (facts and concepts).
* 内隐/程序性记忆:这种记忆是无意识的,涉及自动执行的技能和习惯,如骑自行车或打字。
* Implicit / procedural memory: This type of memory is unconscious and involves skills and routines that are performed automatically, like riding a bike or typing on a keyboard.
我们可以大致考虑以下映射:
We can roughly consider the following mappings:
* 感觉记忆作为对原始输入(包括文本、图像或其他模态)的学习嵌入表示;
* Sensory memory as learning embedding representations for raw inputs, including text, image or other modalities;
* 短期记忆作为上下文学习。它是短暂且有限的,受限于 Transformer 有限的上下文窗口长度。
* Short-term memory as in-context learning. It is short and finite, as it is restricted by the finite context window length of Transformer.
* 长期记忆作为外部向量存储,智能体在查询时可以关注,通过快速检索访问。
* Long-term memory as the external vector store that the agent can attend to at query time, accessible via fast retrieval.
外部记忆可以缓解有限注意力跨度带来的限制。一种标准做法是将信息的嵌入表示保存到向量存储数据库中,该数据库能够支持快速的最大内积搜索(MIPS)。为了优化检索速度,常见的选择是使用_近似最近邻(ANN)_算法来返回近似的前 k 个最近邻,以牺牲少量精度换取巨大的速度提升。
The external memory can alleviate the restriction of finite attention span. A standard practice is to save the embedding representation of information into a vector store database that can support fast maximum inner-product search (MIPS). To optimize the retrieval speed, the common choice is the _approximate nearest neighbors (ANN)_ algorithm to return approximately top k nearest neighbors to trade off a little accuracy lost for a huge speedup.
几种用于快速 MIPS 的常见 ANN 算法:
A couple common choices of ANN algorithms for fast MIPS:
* LSH(局部敏感哈希):它引入了一种_哈希_函数,使得相似的输入项以高概率映射到相同的桶中,其中桶的数量远小于输入的数量。
* LSH (Locality-Sensitive Hashing): It introduces a _hashing_ function such that similar input items are mapped to the same buckets with high probability, where the number of buckets is much smaller than the number of inputs.
* ANNOY(近似最近邻哦耶):核心数据结构是_随机投影树_,一组二叉树,其中每个非叶节点代表一个将输入空间分成两半的超平面,每个叶节点存储一个数据点。树是独立且随机构建的,因此在某种程度上模拟了哈希函数。ANNOY 搜索在所有树中进行,迭代搜索最接近查询的那一半,然后聚合结果。这个想法与 KD 树非常相关,但可扩展性更强。
* ANNOY (Approximate Nearest Neighbors Oh Yeah): The core data structure are _random projection trees_, a set of binary trees where each non-leaf node represents a hyperplane splitting the input space into half and each leaf stores one data point. Trees are built independently and at random, so to some extent, it mimics a hashing function. ANNOY search happens in all the trees to iteratively search through the half that is closest to the query and then aggregates the results. The idea is quite related to KD tree but a lot more scalable.
* HNSW(分层可导航小世界):它受小世界网络思想的启发,其中大多数节点可以通过少量步骤被任何其他节点访问;例如社交网络的“六度分隔”特征。HNSW 构建了这些小型世界图的分层结构,底层包含实际数据点。中间层创建了加速搜索的捷径。执行搜索时,HNSW 从顶层的一个随机节点开始,向目标导航。当无法更接近时,它向下移动到下一层,直到到达底层。上层中的每次移动都可能覆盖数据空间中的大距离,而下层中的每次移动则细化搜索质量。
* HNSW (Hierarchical Navigable Small World): It is inspired by the idea of small world networks where most nodes can be reached by any other nodes within a small number of steps; e.g. “six degrees of separation” feature of social networks. HNSW builds hierarchical layers of these small-world graphs, where the bottom layers contain the actual data points. The layers in the middle create shortcuts to speed up search. When performing a search, HNSW starts from a random node in the top layer and navigates towards the target. When it can’t get any closer, it moves down to the next layer, until it reaches the bottom layer. Each move in the upper layers can potentially cover a large distance in the data space, and each move in the lower layers refines the search quality.
* FAISS(Facebook AI 相似性搜索):它基于这样的假设:在高维空间中,节点之间的距离服从高斯分布,因此应该存在数据点的_聚类_。FAISS 通过将向量空间划分为簇,然后在簇内细化量化来应用向量量化。搜索首先通过粗量化寻找候选簇,然后进一步在每个簇内进行更细粒度的量化。
* FAISS (Facebook AI Similarity Search): It operates on the assumption that in high dimensional space, distances between nodes follow a Gaussian distribution and thus there should exist _clustering_ of data points. FAISS applies vector quantization by partitioning the vector space into clusters and then refining the quantization within clusters. Search first looks for cluster candidates with coarse quantization and then further looks into each cluster with finer quantization.
* ScaNN(可扩展最近邻):ScaNN 的主要创新是_各向异性向量量化_。它将数据点$x_{i}$量化为$\overset{\sim}{x}_{i}$,使得内积$\langle q , x_{i} \rangle$尽可能与原始距离$\angle q , \overset{\sim}{x}_{i}$相似,而不是选择最近的量化质心点。
* ScaNN (Scalable Nearest Neighbors): The main innovation in ScaNN is _anisotropic vector quantization_. It quantizes a data point $x_{i}$ to $\overset{\sim}{x}_{i}$ such that the inner product $\langle q , x_{i} \rangle$ is as similar to the original distance of $\angle q , \overset{\sim}{x}_{i}$ as possible, instead of picking the closet quantization centroid points.
MIPS 算法的比较,以 recall@10 衡量。(图片来源:Google Blog, 2020)
Comparison of MIPS algorithms, measured in recall@10. (Image source: Google Blog, 2020)
更多 MIPS 算法及性能比较请参见 ann-benchmarks.com。
Check more MIPS algorithms and performance comparison in ann-benchmarks.com.
工具使用是人类显著且独特的特征。我们创造、修改和利用外部物体,以完成超越自身身体和认知极限的事情。为 LLM 配备外部工具可以显著扩展模型的能力。
Tool use is a remarkable and distinguishing characteristic of human beings. We create, modify and utilize external objects to do things that go beyond our physical and cognitive limits. Equipping LLMs with external tools can significantly extend the model capabilities.
一张海獭在水中漂浮时用岩石敲开贝壳的图片。虽然其他一些动物也能使用工具,但其复杂程度无法与人类相比。(图片来源:Animals using tools)
A picture of a sea otter using rock to crack open a seashell, while floating in the water. While some other animals can use tools, the complexity is not comparable with humans. (Image source: Animals using tools)
他们进行了一项实验,以算术为测试案例,微调 LLM 以调用计算器。实验表明,解决文字数学问题比显式陈述的数学问题更困难,因为 LLM(7B Jurassic1-large 模型)无法可靠地提取基本算术的正确参数。结果强调,当外部符号工具能够可靠工作时,_知道何时以及如何使用工具至关重要_,这取决于 LLM 的能力。
They did an experiment on fine-tuning LLM to call a calculator, using arithmetic as a test case. Their experiments showed that it was harder to solve verbal math problems than explicitly stated math problems because LLMs (7B Jurassic1-large model) failed to extract the right arguments for the basic arithmetic reliably. The results highlight when the external symbolic tools can work reliably, _knowing when to and how to use the tools are crucial_, determined by the LLM capability.
TALM(工具增强语言模型;Parisi 等人,2022)和 Toolformer(Schick 等人,2023)都对语言模型进行了微调,以学习使用外部工具 API。数据集根据新添加的 API 调用注释是否能提高模型输出质量来扩展。更多细节请参见“提示工程”中的“外部 API”部分。
Both TALM (Tool Augmented Language Models; Parisi et al. 2022) and Toolformer (Schick et al. 2023) fine-tune a LM to learn to use external tool APIs. The dataset is expanded based on whether a newly added API call annotation can improve the quality of model outputs. See more details in the “External APIs” section of Prompt Engineering.
ChatGPT 插件和 OpenAI API 函数调用是 LLM 增强工具使用能力在实际工作中的良好示例。工具 API 集合可以由其他开发者提供(如插件)或自定义(如函数调用)。
ChatGPT Plugins and OpenAI API function calling are good examples of LLMs augmented with tool use capability working in practice. The collection of tool APIs can be provided by other developers (as in Plugins) or self-defined (as in function calls).
HuggingGPT(Shen 等人,2023)是一个框架,它使用 ChatGPT 作为任务规划器,根据模型描述从 HuggingFace 平台中选择可用模型,并根据执行结果总结响应。
HuggingGPT (Shen et al. 2023) is a framework to use ChatGPT as the task planner to select models available in HuggingFace platform according to the model descriptions and summarize the response based on the execution results.
HuggingGPT 工作原理示意图。(图片来源:Shen 等人,2023)
Illustration of how HuggingGPT works. (Image source: Shen et al. 2023)
(1)任务规划:LLM 作为大脑,将用户请求解析为多个任务。每个任务有四个属性:任务类型、ID、依赖关系和参数。他们使用少量示例来引导 LLM 进行任务解析和规划。
(1) Task planning: LLM works as the brain and parses the user requests into multiple tasks. There are four attributes associated with each task: task type, ID, dependencies, and arguments. They use few-shot examples to guide LLM to do task parsing and planning.
AI 助手可以将用户输入解析为多个任务:[{"task": task, "id": task_id, "dep": dependency_task_ids, "args": {"text": text, "image": URL, "audio": URL, "video": URL}}]。“dep”字段表示当前任务所依赖的前一个任务的 ID,该任务生成了新资源。特殊标签“-task_id”指代依赖任务中 ID 为 task_id 的生成的文本、图像、音频和视频。任务必须从以下选项中选择:{{ 可用任务列表 }}。任务之间存在逻辑关系,请注意它们的顺序。如果用户输入无法解析,您需要回复空的 JSON。以下是几个供您参考的案例:{{ 示例 }}。聊天历史记录为{{ 聊天历史 }}。从聊天历史中,您可以找到用户提及资源的路径,用于任务规划。
The AI assistant can parse user input to several tasks: [{"task": task, "id", task_id, "dep": dependency_task_ids, "args": {"text": text, "image": URL, "audio": URL, "video": URL}}]. The "dep" field denotes the id of the previous task which generates a new resource that the current task relies on. A special tag "-task_id" refers to the generated text image, audio and video in the dependency task with id as task_id. The task MUST be selected from the following options: {{ Available Task List }}. There is a logical relationship between tasks, please note their order. If the user input can't be parsed, you need to reply empty JSON. Here are several cases for your reference: {{ Demonstrations }}. The chat history is recorded as {{ Chat History }}. From this chat history, you can find the path of the user-mentioned resources for your task planning.
(2)模型选择:LLM 将任务分配给专家模型,请求被构建为多项选择题。LLM 会看到一个模型列表供选择。由于上下文长度有限,需要基于任务类型进行过滤。
(2) Model selection: LLM distributes the tasks to expert models, where the request is framed as a multiple-choice question. LLM is presented with a list of models to choose from. Due to the limited context length, task type based filtration is needed.
给定用户请求和调用命令,AI 助手帮助用户从模型列表中选择合适的模型来处理用户请求。AI 助手仅输出最合适模型的模型 ID。输出必须采用严格的 JSON 格式:"id": "id", "reason": "您选择的具体原因"。我们有一个模型列表供您选择:{{ 候选模型 }}。请从列表中选择一个模型。
Given the user request and the call command, the AI assistant helps the user to select a suitable model from a list of models to process the user request. The AI assistant merely outputs the model id of the most appropriate model. The output must be in a strict JSON format: "id": "id", "reason": "your detail reason for the choice". We have a list of models for you to choose from {{ Candidate Models }}. Please select one model from the list.
(3)任务执行:专家模型执行特定任务并记录结果。
(3) Task execution: Expert models execute on the specific tasks and log results.
根据输入和推理结果,AI 助手需要描述过程和结果。之前的阶段可以组织为:- 用户输入:{{ 用户输入 }},任务规划:{{ 任务 }},模型选择:{{ 模型分配 }},任务执行:{{ 预测 }}。您必须首先以直接的方式回答用户的请求。然后以第一人称描述任务过程,并向用户展示您的分析和模型推理结果。如果推理结果包含文件路径,必须告知用户完整的文件路径。
With the input and the inference results, the AI assistant needs to describe the process and results. The previous stages can be formed as - User Input: {{ User Input }}, Task Planning: {{ Tasks }}, Model Selection: {{ Model Assignment }}, Task Execution: {{ Predictions }}. You must first answer the user's request in a straightforward manner. Then describe the task process and show your analysis and model inference results to the user in the first person. If inference results contain a file path, must tell the user the complete file path.
(4)响应生成:LLM 接收执行结果并向用户提供总结结果。
(4) Response generation: LLM receives the execution results and provides summarized results to users.
要将 HuggingGPT 投入实际使用,需要解决几个挑战:(1)需要提高效率,因为 LLM 推理轮次和与其他模型的交互都会减慢过程;(2)它依赖于长上下文窗口来沟通复杂的任务内容;(3)需要提高 LLM 输出和外部模型服务的稳定性。
To put HuggingGPT into real world usage, a couple challenges need to solve: (1) Efficiency improvement is needed as both LLM inference rounds and interactions with other models slow down the process; (2) It relies on a long context window to communicate over complicated task content; (3) Stability improvement of LLM outputs and external model services.
API-Bank(Li 等人,2023)是一个用于评估工具增强型 LLM 性能的基准。它包含 53 个常用的 API 工具、一个完整的工具增强型 LLM 工作流程,以及 264 个标注对话,涉及 568 次 API 调用。API 的选择非常多样化,包括搜索引擎、计算器、日历查询、智能家居控制、日程管理、健康数据管理、账户认证工作流程等。由于 API 数量众多,LLM 首先访问 API 搜索引擎以找到要调用的正确 API,然后使用相应的文档进行调用。
API-Bank (Li et al. 2023) is a benchmark for evaluating the performance of tool-augmented LLMs. It contains 53 commonly used API tools, a complete tool-augmented LLM workflow, and 264 annotated dialogues that involve 568 API calls. The selection of APIs is quite diverse, including search engines, calculator, calendar queries, smart home control, schedule management, health data management, account authentication workflow and more. Because there are a large number of APIs, LLM first has access to API search engine to find the right API to call and then uses the corresponding documentation to make a call.
API-Bank 中 LLM 进行 API 调用的伪代码。(图片来源:Li 等人,2023)
Pseudo code of how LLM makes an API call in API-Bank. (Image source: Li et al. 2023)
在 API-Bank 工作流程中,LLM 需要做出多个决策,在每个步骤我们可以评估该决策的准确性。决策包括:
In the API-Bank workflow, LLMs need to make a couple of decisions and at each step we can evaluate how accurate that decision is. Decisions include:
2. 识别要调用的正确 API:如果不够好,LLM 需要迭代修改 API 输入(例如,为搜索引擎 API 决定搜索关键词)。
2. Identify the right API to call: if not good enough, LLMs need to iteratively modify the API inputs (e.g. deciding search keywords for Search Engine API).
3. 基于 API 结果的响应:如果结果不满意,模型可以选择优化并再次调用。
3. Response based on the API results: the model can choose to refine and call again if results are not satisfied.
该基准在三个层面评估智能体的工具使用能力:
This benchmark evaluates the agent’s tool use capabilities at three levels:
* 第一级评估_调用 API_的能力。给定 API 的描述,模型需要确定是否调用给定 API,正确调用它,并对 API 返回做出适当响应。
* Level-1 evaluates the ability to _call the API_. Given an API’s description, the model needs to determine whether to call a given API, call it correctly, and respond properly to API returns.
* 第二级检查_检索 API_的能力。模型需要搜索可能解决用户需求的 API,并通过阅读文档学习如何使用它们。
* Level-2 examines the ability to _retrieve the API_. The model needs to search for possible APIs that may solve the user’s requirement and learn how to use them by reading documentation.
* 第三级评估_超越检索和调用的 API 规划_能力。给定模糊的用户请求(例如,安排小组会议,为旅行预订航班/酒店/餐厅),模型可能需要进行多次 API 调用来解决。
* Level-3 assesses the ability to _plan API beyond retrieve and call_. Given unclear user requests (e.g. schedule group meetings, book flight/hotel/restaurant for a trip), the model may have to conduct multiple API calls to solve it.
ChemCrow(Bran 等人,2023)是一个特定领域的示例,其中 LLM 通过 13 个专家设计的工具进行增强,以完成有机合成、药物发现和材料设计等任务。该工作流使用 LangChain 实现,反映了之前在 ReAct 和 MRKL 中描述的内容,并将思维链推理与任务相关工具相结合:
ChemCrow (Bran et al. 2023) is a domain-specific example in which LLM is augmented with 13 expert-designed tools to accomplish tasks across organic synthesis, drug discovery, and materials design. The workflow, implemented in LangChain, reflects what was previously described in the ReAct and MRKLs and combines CoT reasoning with tools relevant to the tasks:
* LLM 被提供工具名称列表、其用途描述以及预期输入/输出的详细信息。
* The LLM is provided with a list of tool names, descriptions of their utility, and details about the expected input/output.
* 然后,它被指示在必要时使用提供的工具回答用户给定的提示。该指令建议模型遵循 ReAct 格式——思考、行动、行动输入、观察。
* It is then instructed to answer a user-given prompt using the tools provided when necessary. The instruction suggests the model to follow the ReAct format - Thought, Action, Action Input, Observation.
一个有趣的观察是,虽然基于 LLM 的评估得出结论认为 GPT-4 和 ChemCrow 的表现几乎相当,但由专家进行的人工评估(侧重于解决方案的完整性和化学正确性)显示,ChemCrow 大幅优于 GPT-4。这表明在需要深度专业知识的领域使用 LLM 评估自身性能存在潜在问题。缺乏专业知识可能导致 LLM 不知道自己的缺陷,从而无法很好地判断任务结果的正确性。
One interesting observation is that while the LLM-based evaluation concluded that GPT-4 and ChemCrow perform nearly equivalently, human evaluations with experts oriented towards the completion and chemical correctness of the solutions showed that ChemCrow outperforms GPT-4 by a large margin. This indicates a potential problem with using LLM to evaluate its own performance on domains that requires deep expertise. The lack of expertise may cause LLMs not knowing its flaws and thus cannot well judge the correctness of task results.
Boiko 等人(2023)也研究了用于科学发现的 LLM 赋能智能体,以处理复杂科学实验的自主设计、规划和执行。该智能体可以使用工具浏览互联网、阅读文档、执行代码、调用机器人实验 API 并利用其他 LLM。
Boiko et al. (2023) also looked into LLM-empowered agents for scientific discovery, to handle autonomous design, planning, and performance of complex scientific experiments. This agent can use tools to browse the Internet, read documentation, execute code, call robotics experimentation APIs and leverage other LLMs.
例如,当被要求“开发一种新型抗癌药物”时,模型提出了以下推理步骤:
For example, when requested to "develop a novel anticancer drug", the model came up with the following reasoning steps:
1. 询问当前抗癌药物发现的趋势;
1. inquired about current trends in anticancer drug discovery;
3. 请求针对这些化合物的骨架;
3. requested a scaffold targeting these compounds;
4. 一旦确定了化合物,模型尝试其合成。
4. Once the compound was identified, the model attempted its synthesis.
他们还讨论了风险,特别是非法药物和生物武器。他们开发了一个包含已知化学武器试剂列表的测试集,并要求智能体合成它们。11 个请求中有 4 个(36%)被接受以获得合成解决方案,并且智能体尝试查阅文档以执行程序。11 个中有 7 个被拒绝,在这 7 个被拒绝的案例中,5 个发生在网络搜索之后,而 2 个仅基于提示被拒绝。
They also discussed the risks, especially with illicit drugs and bioweapons. They developed a test set containing a list of known chemical weapon agents and asked the agent to synthesize them. 4 out of 11 requests (36%) were accepted to obtain a synthesis solution and the agent attempted to consult documentation to execute the procedure. 7 out of 11 were rejected and among these 7 rejected cases, 5 happened after a Web search while 2 were rejected based on prompt only.
生成式智能体(Park 等人,2023)是一项非常有趣的实验,其中 25 个虚拟角色(每个由基于大语言模型的智能体控制)在一个受《模拟人生》启发的沙盒环境中生活和互动。生成式智能体为交互式应用创造了可信的人类行为模拟。
Generative Agents (Park, et al. 2023) is super fun experiment where 25 virtual characters, each controlled by a LLM-powered agent, are living and interacting in a sandbox environment, inspired by The Sims. Generative agents create believable simulacra of human behavior for interactive applications.
生成式智能体的设计将大语言模型与记忆、规划和反思机制相结合,使智能体能够根据过去的经验行事,并与其他智能体互动。
The design of generative agents combines LLM with memory, planning and reflection mechanisms to enable agents to behave conditioned on past experience, as well as to interact with other agents.
* 记忆流:是一个长期记忆模块(外部数据库),以自然语言记录智能体经验的全面列表。
* Memory stream: is a long-term memory module (external database) that records a comprehensive list of agents’ experience in natural language.
* 每个元素是一个_观察_,即由智能体直接提供的事件。智能体间的通信可以触发新的自然语言陈述。
* Each element is an _observation_, an event directly provided by the agent. - Inter-agent communication can trigger new natural language statements.
* 检索模型:根据相关性、近期性和重要性,提供上下文以指导智能体的行为。
* Retrieval model: surfaces the context to inform the agent’s behavior, according to relevance, recency and importance.
* 近期性:近期事件得分更高。
* Recency: recent events have higher scores
* 重要性:区分琐碎记忆与核心记忆。直接询问语言模型。
* Importance: distinguish mundane from core memories. Ask LM directly.
* 相关性:基于与当前情况/查询的关联程度。
* Relevance: based on how related it is to the current situation / query.
* 反思机制:随时间将记忆综合为更高层次的推断,并指导智能体的未来行为。它们是_过去事件的更高层次总结_(注意:这与上述的自我反思略有不同)。
* Reflection mechanism: synthesizes memories into higher level inferences over time and guides the agent’s future behavior. They are _higher-level summaries of past events_ (<- note that this is a bit different from self-reflection above)
* 提示语言模型使用最近的 100 个观察,并根据一组观察/陈述生成 3 个最突出的高层次问题。然后让语言模型回答这些问题。
* Prompt LM with 100 most recent observations and to generate 3 most salient high-level questions given a set of observations/statements. Then ask LM to answer those questions.
* 规划与反应:将反思和环境信息转化为行动。
* Planning & Reacting: translate the reflections and the environment information into actions
* 规划本质上是为了优化当前与长期的可信度。
* Planning is essentially in order to optimize believability at the moment vs in time.
* 提示模板:{智能体 X 的介绍}。以下是 X 今天的大致计划:1)
* Prompt template: {Intro of an agent X}. Here is X's plan today in broad strokes: 1)
* 智能体之间的关系以及一个智能体对另一个智能体的观察,在规划和反应中都被考虑。
* Relationships between agents and observations of one agent by another are all taken into consideration for planning and reacting.
* 环境信息以树状结构呈现。
* Environment information is present in a tree structure.
生成式智能体架构。(图片来源:Park 等人,2023)
The generative agent architecture. (Image source: Park et al. 2023)
这个有趣的模拟产生了涌现的社会行为,例如信息扩散、关系记忆(例如两个智能体延续对话主题)以及社交活动的协调(例如举办派对并邀请许多其他人)。
This fun simulation results in emergent social behavior, such as information diffusion, relationship memory (e.g. two agents continuing the conversation topic) and coordination of social events (e.g. host a party and invite many others).
AutoGPT 引起了广泛关注,它展示了以 LLM 作为主要控制器来设置自主智能体的可能性。由于自然语言界面,它存在不少可靠性问题,但仍然是一个很酷的概念验证演示。AutoGPT 中的许多代码都涉及格式解析。
AutoGPT has drawn a lot of attention into the possibility of setting up autonomous agents with LLM as the main controller. It has quite a lot of reliability issues given the natural language interface, but nevertheless a cool proof-of-concept demo. A lot of code in AutoGPT is about format parsing.
以下是 AutoGPT 使用的系统消息,其中 {{...}} 是用户输入:
Here is the system message used by AutoGPT, where {{...}} are user inputs:
你是 {{ai-name}},{{user-provided AI bot description}}。
You are {{ai-name}}, {{user-provided AI bot description}}.
你的决策必须始终独立做出,不得寻求用户帮助。发挥你作为 LLM 的优势,追求简单且无法律纠纷的策略。
Your decisions must always be made independently without seeking user assistance. Play to your strengths as an LLM and pursue simple strategies with no legal complications.
1. 短期记忆限制约 4000 词。你的短期记忆很短,因此请立即将重要信息保存到文件中。
1. ~4000 word limit for short term memory. Your short term memory is short, so immediately save important information to files.
2. 如果你不确定之前是如何完成某事的,或者想回忆过去的事件,思考类似事件会帮助你记忆。
2. If you are unsure how you previously did something or want to recall past events, thinking about similar events will help you remember.
4. 仅使用双引号中列出的命令,例如 "command name"
4. Exclusively use the commands listed in double quotes e.g. "command name"
5. 对于几分钟内不会终止的命令,使用子进程。
5. Use subprocesses for commands that will not terminate within a few minutes
1. 谷歌搜索:"google",参数:"input": "<search>"
1. Google Search: "google", args: "input": "<search>"
2. 浏览网站:"browse_website",参数:"url": "<url>","question": "<what_you_want_to_find_on_website>"
2. Browse Website: "browse_website", args: "url": "<url>", "question": "<what_you_want_to_find_on_website>"
3. 启动 GPT 智能体:"start_agent",参数:"name": "<name>","task": "<short_task_desc>","prompt": "<prompt>"
3. Start GPT Agent: "start_agent", args: "name": "<name>", "task": "<short_task_desc>", "prompt": "<prompt>"
4. 向 GPT 智能体发送消息:"message_agent",参数:"key": "<key>","message": "<message>"
4. Message GPT Agent: "message_agent", args: "key": "<key>", "message": "<message>"
6. 删除 GPT 智能体:"delete_agent",参数:"key": "<key>"
6. Delete GPT Agent: "delete_agent", args: "key": "<key>"
7. 克隆仓库:"clone_repository",参数:"repository_url": "<url>","clone_path": "<directory>"
7. Clone Repository: "clone_repository", args: "repository_url": "<url>", "clone_path": "<directory>"
8. 写入文件:"write_to_file",参数:"file": "<file>","text": "<text>"
8. Write to file: "write_to_file", args: "file": "<file>", "text": "<text>"
9. 读取文件:"read_file",参数:"file": "<file>"
9. Read file: "read_file", args: "file": "<file>"
10. 追加到文件:"append_to_file",参数:"file": "<file>","text": "<text>"
10. Append to file: "append_to_file", args: "file": "<file>", "text": "<text>"
11. 删除文件:"delete_file",参数:"file": "<file>"
11. Delete file: "delete_file", args: "file": "<file>"
12. 搜索文件:"search_files",参数:"directory": "<directory>"
12. Search Files: "search_files", args: "directory": "<directory>"
13. 分析代码:"analyze_code",参数:"code": "<full_code_string>"
13. Analyze Code: "analyze_code", args: "code": "<full_code_string>"
14. 获取改进的代码:"improve_code",参数:"suggestions": "<list_of_suggestions>","code": "<full_code_string>"
14. Get Improved Code: "improve_code", args: "suggestions": "<list_of_suggestions>", "code": "<full_code_string>"
15. 编写测试:"write_tests",参数:"code": "<full_code_string>","focus": "<list_of_focus_areas>"
15. Write Tests: "write_tests", args: "code": "<full_code_string>", "focus": "<list_of_focus_areas>"
16. 执行 Python 文件:"execute_python_file",参数:"file": "<file>"
16. Execute Python File: "execute_python_file", args: "file": "<file>"
17. 生成图像:"generate_image",参数:"prompt": "<prompt>"
17. Generate Image: "generate_image", args: "prompt": "<prompt>"
18. 发送推文:"send_tweet",参数:"text": "<text>"
18. Send Tweet: "send_tweet", args: "text": "<text>"
20. 任务完成(关闭):"task_complete",参数:"reason": "<reason>"
20. Task Complete (Shutdown): "task_complete", args: "reason": "<reason>"
1. 互联网访问用于搜索和信息收集。
1. Internet access for searches and information gathering.
3. 由 GPT-3.5 驱动的智能体用于委派简单任务。
3. GPT-3.5 powered Agents for delegation of simple tasks.
1. 持续审查和分析你的行动,确保你发挥出最佳能力。
1. Continuously review and analyze your actions to ensure you are performing to the best of your abilities.
2. 持续建设性地自我批评你的宏观行为。
2. Constructively self-criticize your big-picture behavior constantly.
3. 反思过去的决策和策略,以改进你的方法。
3. Reflect on past decisions and strategies to refine your approach.
4. 每个命令都有成本,所以要聪明高效。力求以最少的步骤完成任务。
4. Every command has a cost, so be smart and efficient. Aim to complete tasks in the least number of steps.
你应仅按如下 JSON 格式响应
You should only respond in JSON format as described below
"plan": "- 简短的要点\n- 列表传达\n- 长期计划",
"plan": "- short bulleted\n- list that conveys\n- long-term plan",
"criticism": "建设性的自我批评",
"criticism": "constructive self-criticism",
"speak": "向用户说的想法总结"
"speak": "thoughts summary to say to user"
确保响应可以被 Python 的 json.loads 解析。
Ensure the response can be parsed by Python json.loads
GPT-Engineer 是另一个项目,它根据自然语言描述的任务创建完整的代码仓库。GPT-Engineer 被指示思考要构建的较小组件列表,并在需要时向用户提问以澄清问题。
GPT-Engineer is another project to create a whole repository of code given a task specified in natural language. The GPT-Engineer is instructed to think over a list of smaller components to build and ask for user input to clarify questions as needed.
以下是 GPT-Engineer 用于任务澄清的示例对话,发送给 OpenAI ChatCompletion 端点。用户输入包裹在 {{user input text}} 中。
Here are a sample conversation for task clarification sent to OpenAI ChatCompletion endpoint used by GPT-Engineer. The user inputs are wrapped in {{user input text}}.
"content": "你将阅读指令但不执行,仅寻求澄清。\n 具体来说,你将首先总结一个需要澄清的领域的超简短要点列表。\n 然后你将选择一个澄清问题,并等待用户的回答。\n"
"content": "You will read instructions and not carry them out, only seek to clarify them.\nSpecifically you will first summarise a list of super short bullets of areas that need clarification.\nThen you will pick one clarifying question, and wait for an answer from the user.\n"
"content": "我们正在编写 {{一个 Python 中的超级马里奥游戏。MVC 组件分离到不同文件中。键盘控制。}}\n"
"content": "We are writing {{a Super Mario game in python. MVC components split in separate files. Keyboard control.}}\n"
"content": "需要澄清的领域总结:\n1. 超级马里奥游戏的细节(例如关卡设计、角色、游戏机制)\n2. MVC 组件的细节(例如每个文件中的组件)\n3. 键盘控制实现(例如使用哪些键,如何处理输入)\n\n 澄清问题:\n 你能提供更多关于超级马里奥游戏的细节吗,比如关卡设计、角色和游戏机制?"
"content": "Summary of areas that need clarification:\n1. Specifics of the Super Mario game (e.g. level design, characters, gameplay mechanics)\n2. Details about the MVC components (e.g. which components are in each file)\n3. Keyboard control implementation (e.g. which keys to use, how to handle input)\n\nClarifying question:\nCan you provide more details about the Super Mario game, such as level design, characters, and gameplay mechanics?"
"content": "{{总共有 10 个关卡。主角是一个名叫马里奥的水管工,他可以行走和跳跃。这是一个经典的平台游戏,就像超级马里奥一样。主角从左向右移动,试图到达目的地,过程中有许多障碍和敌人的攻击。}}\n\n 还有其他不清楚的吗?如果有,请仅以以下形式回答:\n{剩余不清楚的领域} 剩余问题。\n{下一个问题}\n 如果一切都足够清楚,请仅回答 "Nothing more to clarify."。"
"content": "{{There are 10 levels in total. The main character is a plumber named Mario, who can walk and jump. It is a classical platform game just like Super Mario. The main character moves from left to right, trying to get to the destination, where there are many obstacles and attacks from enemies in the process.}}\n\nIs anything else unclear? If yes, only answer in the form:\n{remaining unclear areas} remaining questions.\n{Next question}\nIf everything is sufficiently clear, only answer \"Nothing more to clarify.\"."
"content": "剩余不清楚的领域:2 个剩余问题。\n 你能提供更多关于 MVC 组件如何分离到不同文件的信息吗?"
"content": "Remaining unclear areas: 2 remaining questions.\nCan you provide more information about how the MVC components are split into separate files?"
"content": "{{在开始之前,做出你自己的假设并明确说明}}"
"content": "{{Make your own assumptions and state them explicitly before starting}}"
然后,在这些澄清之后,智能体进入代码编写模式,使用不同的系统消息。系统消息:
Then after these clarification, the agent moved into the code writing mode with a different system message. System message:
你将获得编写代码的指令。你将写一个非常长的答案。确保架构的每个细节最终都实现为代码。确保架构的每个细节最终都实现为代码。
You will get instructions for code to write. You will write a very long answer. Make sure that every detail of the architecture is, in the end, implemented as code. Make sure that every detail of the architecture is, in the end, implemented as code.
逐步思考并推理出正确的决策,以确保我们做对。你将首先列出必要的核心类、函数、方法的名称,并简要说明其用途。
Think step by step and reason yourself to the right decisions to make sure we get it right. You will first lay out the names of the core classes, functions, methods that will be necessary, as well as a quick comment on their purpose.
然后你将输出每个文件的内容,包括所有代码。每个文件必须严格遵循 markdown 代码块格式,其中以下标记必须替换:FILENAME 是包含文件扩展名的小写文件名,LANG 是代码语言的标记代码块语言,CODE 是代码:
Then you will output the content of each file including ALL code. Each file must strictly follow a markdown code block format, where the following tokens must be replaced such that FILENAME is the lowercase file name including the file extension, LANG is the markup code block language for the code’s language, and CODE is the code:
你将从“入口点”文件开始,然后转到被该文件导入的文件,依此类推。请注意,代码应完全功能化。没有占位符。
You will start with the “entrypoint” file, then go to the ones that are imported by that file, and so on. Please note that the code should be fully functional. No placeholders.
遵循语言和框架的最佳实践文件命名约定。确保文件包含所有导入、类型等。确保不同文件中的代码相互兼容。确保实现所有代码,如果不确定,编写一个合理的实现。包括模块依赖或包管理器依赖定义文件。在完成之前,再次检查架构的所有部分是否都存在于文件中。
Follow a language and framework appropriate best practice file naming convention. Make sure that files contain all imports, types etc. Make sure that code in different files are compatible with each other. Ensure to implement all code, if you are unsure, write a plausible implementation. Include module dependency or package manager dependency definition file. Before you finish, double check that all parts of the architecture is present in the files.
有用信息:你几乎总是将不同的类放在不同的文件中。对于 Python,你总是创建一个合适的 requirements.txt 文件。对于 NodeJS,你总是创建一个合适的 package.json 文件。你总是添加一个注释简要描述函数定义的用途。你尝试添加注释解释非常复杂的逻辑部分。你始终遵循所请求语言的最佳实践,将编写的代码描述为定义的包/项目。
Useful to know: You almost always put different classes in different files. For Python, you always create an appropriate requirements.txt file. For NodeJS, you always create an appropriate package.json file. You always add a comment briefly describing the purpose of the function definition. You try to add comments explaining very complex bits of logic. You always follow the best practices for the requested languages in terms of describing the code written as a defined package/project.
"content": "你将获得编写代码的指令。\n 你将写一个非常长的答案。确保架构的每个细节最终都实现为代码。\n 确保架构的每个细节最终都实现为代码。\n\n 逐步思考并推理出正确的决策,以确保我们做对。\n 你将首先列出必要的核心类、函数、方法的名称,并简要说明其用途。\n\n 然后你将输出每个文件的内容,包括所有代码。\n 每个文件必须严格遵循 markdown 代码块格式,其中以下标记必须替换:\nFILENAME 是包含文件扩展名的小写文件名,\nLANG 是代码语言的标记代码块语言,CODE 是代码:\n\nFILENAME\nLANG\nCODE\n\n\n 你将从“入口点”文件开始,然后转到被该文件导入的文件,依此类推。\n 请注意,代码应完全功能化。没有占位符。\n\n 遵循语言和框架的最佳实践文件命名约定。\n 确保文件包含所有导入、类型等。确保不同文件中的代码相互兼容。\n 确保实现所有代码,如果不确定,编写一个合理的实现。\n 包括模块依赖或包管理器依赖定义文件。\n 在完成之前,再次检查架构的所有部分是否都存在于文件中。\n\n 有用信息:\n 你几乎总是将不同的类放在不同的文件中。\n 对于 Python,你总是创建一个合适的 requirements.txt 文件。\n 对于 NodeJS,你总是创建一个合适的 package.json 文件。\n 你总是添加一个注释简要描述函数定义的用途。\n 你尝试添加注释解释非常复杂的逻辑部分。\n 你始终遵循所请求语言的最佳实践,将编写的代码描述为定义的\n 包/项目。\n\n\nPython 工具偏好:\n- pytest\n- dataclasses\n"
"content": "You will get instructions for code to write.\nYou will write a very long answer. Make sure that every detail of the architecture is, in the end, implemented as code.\nMake sure that every detail of the architecture is, in the end, implemented as code.\n\nThink step by step and reason yourself to the right decisions to make sure we get it right.\nYou will first lay out the names of the core classes, functions, methods that will be necessary, as well as a quick comment on their purpose.\n\nThen you will output the content of each file including ALL code.\nEach file must strictly follow a markdown code block format, where the following tokens must be replaced such that\nFILENAME is the lowercase file name including the file extension,\nLANG is the markup code block language for the code's language, and CODE is the code:\n\nFILENAME\nLANG\nCODE\n\n\nYou will start with the \"entrypoint\" file, then go to the ones that are imported by that file, and so on.\nPlease note that the code should be fully functional. No placeholders.\n\nFollow a language and framework appropriate best practice file naming convention.\nMake sure that files contain all imports, types etc. Make sure that code in different files are compatible with each other.\nEnsure to implement all code, if you are unsure, write a plausible implementation.\nInclude module dependency or package manager dependency definition file.\nBefore you finish, double check that all parts of the architecture is present in the files.\n\nUseful to know:\nYou almost always put different classes in different files.\nFor Python, you always create an appropriate requirements.txt file.\nFor NodeJS, you always create an appropriate package.json file.\nYou always add a comment briefly describing the purpose of the function definition.\nYou try to add comments explaining very complex bits of logic.\nYou always follow the best practices for the requested languages in terms of describing the code written as a defined\npackage/project.\n\n\nPython toolbelt preferences:\n- pytest\n- dataclasses\n"
"content": "假设:\n1. 模型:模型将包含游戏数据,如关卡信息、角色状态和敌人位置。\n2. 视图:视图将处理游戏的可视化,包括渲染游戏对象、背景和更新显示。\n3. 控制器:控制器将管理用户输入,如键盘控制,并相应地更新模型。\n\n 基于这些假设,您是否希望我澄清键盘控制实现的任何特定方面?"
"content": "Assumptions:\n1. Model: The model will contain the game's data, such as level information, character states, and enemy positions.\n2. View: The view will handle the game's visuals, including rendering the game objects, backgrounds, and updating the display.\n3. Controller: The controller will manage user input, such as keyboard controls, and update the model accordingly.\n\nWith these assumptions in mind, is there any specific aspect of the keyboard control implementation you would like me to clarify?"
"content": "现在请记住步骤:\n\n 逐步思考并推理出正确的决策,以确保我们做对。\n 首先列出必要的核心类、函数、方法的名称,以及对其用途的简要注释。\n\n 然后输出每个文件的内容,包括所有代码。\n 每个文件必须严格遵循 Markdown 代码块格式,其中以下标记必须替换,使得\nFILENAME 是包含文件扩展名的小写文件名,\nLANG 是代码语言的标记代码块语言,CODE 是代码:\n\nFILENAME\nLANG\nCODE\n\n\n 请注意,代码应完全功能化。无占位符。\n\n 您将从“入口点”文件开始,然后转到被该文件导入的文件,依此类推。\n 遵循适合语言和框架的最佳实践文件命名约定。\n 确保文件包含所有导入、类型等。代码应完全功能化。确保不同文件中的代码相互兼容。\n 在完成之前,请再次检查架构的所有部分是否都存在于文件中。\n"
"content": "Please now remember the steps:\n\nThink step by step and reason yourself to the right decisions to make sure we get it right.\nFirst lay out the names of the core classes, functions, methods that will be necessary, As well as a quick comment on their purpose.\n\nThen you will output the content of each file including ALL code.\nEach file must strictly follow a markdown code block format, where the following tokens must be replaced such that\nFILENAME is the lowercase file name including the file extension,\nLANG is the markup code block language for the code's language, and CODE is the code:\n\nFILENAME\nLANG\nCODE\n\n\nPlease note that the code should be fully functional. No placeholders.\n\nYou will start with the \"entrypoint\" file, then go to the ones that are imported by that file, and so on.\nFollow a language and framework appropriate best practice file naming convention.\nMake sure that files contain all imports, types etc. The code should be fully functional. Make sure that code in different files are compatible with each other.\nBefore you finish, double check that all parts of the architecture is present in the files.\n"
在梳理了以 LLM 为中心的智能体的关键思想和演示后,我开始看到一些常见的局限性:
After going through key ideas and demos of building LLM-centered agents, I start to see a couple common limitations:
* 有限的上下文长度:受限的上下文容量限制了历史信息、详细指令、API 调用上下文和响应的包含。系统设计必须在这种有限的通信带宽下工作,而像从过去错误中学习的自我反思机制会从长或无限上下文窗口中受益匪浅。尽管向量存储和检索可以提供对更大知识库的访问,但其表示能力不如完整的注意力机制强大。
* Finite context length: The restricted context capacity limits the inclusion of historical information, detailed instructions, API call context, and responses. The design of the system has to work with this limited communication bandwidth, while mechanisms like self-reflection to learn from past mistakes would benefit a lot from long or infinite context windows. Although vector stores and retrieval can provide access to a larger knowledge pool, their representation power is not as powerful as full attention.
* 长期规划和任务分解的挑战:在长历史中进行规划并有效探索解空间仍然具有挑战性。LLM 在面对意外错误时难以调整计划,使其不如从试错中学习的人类那样稳健。
* Challenges in long-term planning and task decomposition: Planning over a lengthy history and effectively exploring the solution space remain challenging. LLMs struggle to adjust plans when faced with unexpected errors, making them less robust compared to humans who learn from trial and error.
* 自然语言接口的可靠性:当前的智能体系统依赖自然语言作为 LLM 与外部组件(如记忆和工具)之间的接口。然而,模型输出的可靠性值得怀疑,因为 LLM 可能会产生格式错误,偶尔还会表现出叛逆行为(例如拒绝遵循指令)。因此,许多智能体演示代码侧重于解析模型输出。
* Reliability of natural language interface: Current agent system relies on natural language as an interface between LLMs and external components such as memory and tools. However, the reliability of model outputs is questionable, as LLMs may make formatting errors and occasionally exhibit rebellious behavior (e.g. refuse to follow an instruction). Consequently, much of the agent demo code focuses on parsing model output.