Harness Engineering for Self-Improvement
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→日期:2026 年 7 月 4 日 | 预计阅读时间:31 分钟 | 作者:Lilian Weng * 模式 2:文件系统作为持久记忆 * 模式 3:子代理和后端任务
Date: July 4, 2026 | Estimated Reading Time: 31 min | Author: Lilian Weng * Pattern 2: File System as Persistent Memory * Pattern 3: Sub-agent and Backend Jobs
日期:2026 年 7 月 4 日 | 预计阅读时间:31 分钟 | 作者:Lilian Weng
Date: July 4, 2026 | Estimated Reading Time: 31 min | Author: Lilian Weng
* 模式 2:文件系统作为持久化内存
* Pattern 2: File System as Persistent Memory
* 模式 3:子智能体与后端任务
* Pattern 3: Sub-agent and Backend Jobs
* 与模型权重的联合优化
* Joint Optimization with Model Weights
递归自我改进(RSI)的概念可追溯到 I. J. Good(1965),他将“超智能机器”定义为在所有智力活动中超越人类并能设计更好机器以自我改进的系统。Yudkowsky(2008)使用“递归自我改进”一词指代一个特定的反馈循环:AI 利用其当前智能来改进产生其智能的认知机制。
The concept of recursive self-improvement (RSI) dates back to I. J. Good (1965), where he defined an “ultraintelligent machine” as a system that can surpass humans in all intellectual activities and design better machines to improve itself. Yudkowsky (2008) used the phrase “recursive self-improvement” for a specific feedback loop: an AI uses its current intelligence to improve the cognitive machinery that produces its intelligence.
现代 AI 中的这一反馈循环可能意味着模型直接重写自身权重,或更广泛地,模型改进_训练流程_和_部署系统_,进而使后续模型在经济价值任务上性能更优。前沿实验室(Anthropic;OpenAI)的研究表明,AI 研究开发速度已大幅加快。
This feedback loop in modern AI may indicate the model rewriting its own weights directly, or more broadly the model improves the _training pipeline_ and the _deployment system_, which in turn enables a better successor model with improved performance across economically valuable tasks. The speed of research development in AI has been shown to drastically accelerated in frontier labs (Anthropic; OpenAI).
我明确提及“_部署系统_”,因为原始模型与现实世界之间的中间层似乎与模型的原始智能(即预训练后的评估)同等重要。如 Claude Code 和 Codex 等成功的编码智能体产品所示,框架是 AI 部署的重要组成部分。框架是围绕基础模型的系统,负责编排执行,决定模型如何思考与规划、调用工具与行动、感知与管理上下文、存储工件以及评估结果。
I explicitly mention _“deployment system”_ because the layer between the raw model and the real-world context seems to be as important as the model’s raw intelligence (i.e. the evals right after pretraining). Harnesses are important components of AI deployment, as shown by successful coding agent products such as Claude Code and Codex. A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.
本文将聚焦于框架工程的研究及其对 RSI 的贡献。近期许多关于自动研究、自我改进智能体和进化程序搜索的工作都可围绕这一问题展开。其他关于模型自博弈、合成数据、测试时训练以及更广泛的持续学习主题的研究也与 RSI 愿景相符(例如 Yuan 等人 2024,Chen 等人 2024,Zhao 等人 2025,Choi 等人 2026),但它们并非本文重点。
This one post will focus on research around harness engineering and how it contributes to RSI. Much recent work on auto-research, self-improving agents, and evolutionary program search can be organized around this question. Other work on model self-play, synthetic data, test-time training and a broader theme of continual learning also matches the RSI vision (e.g. Yuan et al. 2024, Chen et al. 2024), Zhao et al. 2025, Choi et al. 2026)) but they will not be the focus of this post.
与早期的智能体框架“智能体 = 大语言模型 + 记忆 + 工具 + 规划 + 动作”相比,工程化驾驭还额外包括_工作流设计(例如循环工程)、评估、权限控制和持久状态管理_。它不再仅仅是提示模板,而是更接近运行时和软件系统设计:模型如何观察、行动、记忆、自我检查并改进。
Compared with early agent frameworks, “agent = LLM + memory + tools + planning + action”, harnesses engineering additionally include _workflow design (e.g. loop engineering), evaluation, permission controls, and persistent state management_. It is no longer only prompt templates, but closer to runtime and software system design: how the model observes, acts, memorizes, checks itself, and improves.
设计应刻意保持简单和通用,以便实现泛化,很可能需要借鉴现有的软件工程实践,从而受益于预训练知识。操作系统与驾驭之间也存在强烈的类比关系。类似于操作系统,驾驭应封装复杂的逻辑,同时保持接口简单。与此同时,配置、工具接口和其他协议可能会逐渐在行业内标准化。
The design should be deliberately simple and generic to enable generalization, likely with reference to existing software engineering practices to benefit from prertaining knowlege. There is also a strong analogy between operating systems and harnesses. Similar to an OS, a harness should encapsulate complicated logic while keeping the interface simple. Meanwhile, configs, tool interfaces and other protocols may gradually become standardized across the industry.
定义一个模型可以操作、测试和迭代的工作流是自动化的关键设计。Karpathy 的 autoresearch 仓库(https://github.com/karpathy/autoresearch)是一个如何构建此类工作流的简洁示例。常见的工作流遵循一个目标导向的循环:计划、执行、观察/测试、改进,然后再次执行,直到目标达成。该过程可能会主动向用户请求澄清任务规范或执行偏好。
Defining a workflow in which the model can operate, test, and iterate is a key design for automation. Karpathy’s autoresearch repo (https://github.com/karpathy/autoresearch) is a clean example of how such a workflow can be constructed. A common workflow follows a goal-oriented loop of plan, execute, observe/test, improve, and execute again _until_ the goal is achieved. The process may trigger proactive requests to users for clarity in task specification or execution preference.
一个简化的 Codex 智能体循环:智能体调用工具,工具响应影响模型的下一次生成。
A simplified Codex agent loop: the agent calls tools and tool responses affect the model's next generation.
工作流图还强调模型分析自身的轨迹和失败案例,然后通过“智能体运行时”而非静态提示模板来迭代其进展。
The workflow graph also emphasizes the model analyzing its own trajectories and failure cases and then iterating on its progress through an “agent runtime” rather than a static prompt template.
在长周期智能体系统中,一个反复出现的模式是对丰富状态和工件进行简单控制。一个框架不应将整个工作流和所有日志都放在上下文中;相反,它应将持久状态保存在文件中。在长周期智能体式展开中,诸如实验日志、代码差异、论文摘要、错误追踪和过去的展开轨迹等工件,其长度往往远超模型训练所用的上下文窗口。
A recurring pattern in long-horizon agent systems is simple control over rich states and artifacts. A harness should not carry the entire workflow and all logs in context; instead, it should keep durable state in files. In long-horizon agentic rollout, artifacts such as experiment logs, code diffs, paper summaries, error traces, and past rollout trajectories often grow much longer than the context window that the model has trained for.
学习如何读取、写入和编辑文件系统(通常通过 bash 命令)是 LLM 的基础技能,因此以简单的文件形式管理持久化内存,自然会受益于核心模型能力的提升。
Learning how to read, write, and edit the file system (commonly via bash commands) is a foundation skill for LLMs, and thus managing persistent memory in the simple form of files naturally benefits from improvements in core model capability.
一个 Harness 可以生成多个子智能体并行执行并监控后台任务。当主智能体需要搜索多个假设、并发运行实验或委托独立子任务而不污染主上下文时,这非常有用。父智能体随后需要一个小的进程管理器:启动任务、检查日志、取消失败运行,并将结果合并回主智能体线程。
A harness can spawn multiple subagents to execute in parallel and monitor backend jobs. This is useful when the main agent needs to search multiple hypotheses, run experiments concurrently, or delegate isolated subtasks without polluting the main context. The parent agent then needs a small process manager: launch jobs, inspect logs, cancel failed runs, and merge results back into the main agent thread.
关键设计选择是使并行性显式且可检查。如果子智能体的输出仅存在于瞬态聊天上下文中,它们会很快变得过时和隐藏。如果它们被存储为文件、日志和状态记录,模型可以在中断后恢复,并对其自身的执行历史进行推理。
The key design choice is to make parallelism explicit and inspectable. If subagent outputs only live in a transient chat context, they quickly become obselete and hidden. If they are stored as files, logs, and status records, the model can recover after interruptions and reason over its own execution history.
主流编码智能体的核心接口已在 Claude Code、Codex、OpenCode 和 Cursor 风格的智能体中趋于稳定。它们通常使用如下循环:
The core interface of mainstream coding agents has become stabilized across Claude Code, Codex, OpenCode, and Cursor-style agents. They commonly use a loop like:
借助一组工具,编码智能体能够在给定仓库中开发和调试问题,类似于人类开发者配备 IDE 的方式。
With access to a set of tools, the coding agent is able to develop and debug issues in a given repository, similar to how human developers are equipped with IDEs.
(非详尽列表;仅作演示。感兴趣可阅读此内容。)
(Not a comprenhensive list; shown for demonstration. Read this if interested.)
很难预测未来 RSI 在多大程度上依赖于利用工程,但 RSI 的近期路径不太可能从模型直接重写其权重开始。我对近期实用路径的预测是:
It is hard to forecast how much the future of RSI will rely on harness engineering, but the near-term path of RSI is unlikely to start as a model directly rewriting its weights. My prediction of a practical near-term path is:
1. 利用工程将朝着元方法论的方向发展(即改进获取更好答案的机制,而不仅仅是改进答案本身)。利用系统本身成为优化目标,减少启发式规则,增加通用机制。
1. Harness engineering will evolve in the direction of meta-methodology (i.e. improving the machinery for getting better answers, not just improving the answer itself). The harness system itself becomes an optimization target, with fewer heuristic rules and more general mechanisms.
2. 反过来,成熟的利用系统使模型自我改进循环的自动研究成为可能,更智能的模型防止利用系统过度工程化,并保持系统的可持续性。
2. In turn, mature harnesses enable auto-research for model self-improvement loop and smarter models prevents harnesses from overengineering and keep the system sustainable.
最终,许多利用改进可能会被内化为核心模型行为,但与外部上下文和工具的接口应保持不变。我们在提示工程中看到了这种模式的较软版本:随着指令微调和模型推理能力的提升,手动提示技巧变得不那么重要,但指定目标、约束、上下文和评估的需求并未消失。
Eventually it is possible that many harness improvements will be _internalized_ into core model behavior, but the interface with external context and tools should remain. We have seen a softer version of this pattern with prompt engineering: manual prompt tricks became less central as instruction tuning and model reasoning improved, but _the need to specify goals, constraints, context, and evaluation did not disappear_.
在利用系统中,被优化对象的演进大致为:指令提示 → 结构化上下文 → 工作流 → 利用代码 → 优化器代码。随着模型变得更加智能和强大,我们转向更复杂的目标和更通用的方法。
The progression in the object being optimized in the harness system is roughly: instruction prompts → structured context → workflow → harness code → optimizer code. As the model becomes more intelligent and powerful, we move toward more complex targets and generic methods.
随着智能体任务周期显著增长,简单地将所有工具响应和模型生成内容追加到上下文中会迅速失控。上下文管理是一个为 LLM 构建更结构化、更简洁的上下文并管理持久状态的层。毫无疑问,长上下文研究将持续取得进展,但目前长上下文智能与上下文工程有时相互交织。
Simply appending all the tool responses and model generations into the context can quickly grow out of control as the agentic job horizon increases significantly. Context management is a layer to construct a more structed and concise context for LLM and manage persistant states. There is no doubt that long-context research will keep on making progress but at the moment long-context intelligence and context engineering sometime intertwines.
智能体上下文工程(ACE;Zhang et al. 2025)将上下文视为一个不断演化的剧本,而非不断增长的提示。它包含三个组件来维护一个由要点组成的上下文剧本,每个要点都有标识符和描述。
Agentic Context Engineering (ACE; Zhang et al. 2025) treats context as an evolving playbook rather than an increasingly lengthening prompt. It has three components to maintain one context playbook of bullet points, each with an identifier and a description.
1. 生成器:生成任务轨迹,并参考要点。
1. _Generator_: produces task trajectories, with reference to bullet points.
2. 反思器:从成功和失败的轨迹中提炼见解。
2. _Reflector_: distills insights from successful and failed trajectories.
3. 管理模块:以增量、逐项的方式更新结构化上下文。
3. _Curator_: updates the structured context with incremental, itemized entries.
智能体上下文工程(ACE)框架。(图片来源:Zhang et al. 2025)
The framework of Agentic Context Engineering (ACE). (Image source: Zhang et al. 2025)
为防止迭代重写过程中出现上下文崩溃和简洁性偏差,ACE 的一个关键设计选择是管理模块不重写整个提示块。相反,它以(标识符,描述)的形式输出一组结构化的逐项要点,并通过确定性逻辑将这些要点合并到结构化的上下文日志中。上下文项会定期进行精炼和去重。
To prevent context collapse and brevity bias during iterative rewrites, one key design choice in ACE is that the curator does not rewrite a full prompt blob. It instead outputs a collection of structured, itemized bullets in the form of (identifier, description), and these bullets are merged into a structured context logbook with deterministic logic. The context items are refined and deduplicated periodically.
ACE 从轨迹中学习见解的事实有助于我们迈向自我管理的记忆,但更新规则和整体工作流程仍然是手工设计的。为了迈向更自我改进的循环,元上下文工程(MCE;Ye et al. 2026)将机制(如何管理上下文)与工件内容(上下文中有什么)分离,在元优化层面运行技能进化,在基础层面运行上下文优化。
The fact that ACE learns insights from rollouts helps us move toward self-managed memory, but the update rules and the overall workflow are still handcrafted. To move toward a more self-improving loop, Meta Context Engineering (MCE; Ye et al. 2026) separates the mechanism (how to manage context) from the artifact content (what is in context), running skill evolution at the meta-optimization level and context optimization at the base level.
一个 MCE 技能 $s \in \mathcal{S}$ 定义了一个上下文函数 $c_{s} = \left(\right. \rho_{s} , F_{s} \left.\right)$,并将输入 $x$ 映射到上下文 $c = F_{s} \left(\right. x ; \rho_{s} \left.\right)$,其中:
An MCE skill $s \in \mathcal{S}$ defines a context function $c_{s} = \left(\right. \rho_{s} , F_{s} \left.\right)$ and maps an input $x$ to context $c = F_{s} \left(\right. x ; \rho_{s} \left.\right)$, where:
* $\rho_{s} = \left{\right. \rho_{1} , \ldots , \rho_{m} \left.\right}$ 是静态组件(提示、知识库、代码库)。
* $\rho_{s} = \left{\right. \rho_{1} , \ldots , \rho_{m} \left.\right}$ are static components (prompts, knowledge bases, code libraries).
* $F_{s} = \left{\right. F_{1} , \ldots , F_{k} \left.\right}$ 是动态操作符(搜索、选择、过滤、格式化)。
* $F_{s} = \left{\right. F_{1} , \ldots , F_{k} \left.\right}$ are dynamic operators (search, selection, filtering, formatting).
双层优化是在训练数据上找到给定技能 $s$ 的最佳上下文 $c_{s}^{*}$,而外层循环找到在验证集上提供最佳性能的最优技能:
The bi-level optimization is to find the best context $c_{s}^{*}$ given skill $s$ on the training data, while the outer loop finds the optimal skill that provides the best performance on the validation set:
\text{内层}:\textrm{ } c_{s}^{*} = arg \underset{c_{s}}{max} J_{\text{train}} \left(\right. c_{s} ; s \left.\right) \text{外层}:\textrm{ } s^{*} = arg \underset{s \in \mathcal{S}}{max} J_{\text{val}} \left(\right. c_{s}^{*} \left.\right)
\text{Inner}:\textrm{ } c_{s}^{*} = arg \underset{c_{s}}{max} J_{\text{train}} \left(\right. c_{s} ; s \left.\right) \text{Outer}:\textrm{ } s^{*} = arg \underset{s \in \mathcal{S}}{max} J_{\text{val}} \left(\right. c_{s}^{*} \left.\right)
技能数据库跟踪先前技能、上下文函数和评估指标的历史 $\mathcal{H}_{k - 1} = \left{\right. \left(\right. s_{i} , c_{i} , J_{i}^{\text{train}} , J_{i}^{\text{val}} \left.\right) \left.\right}_{i = 1}^{k - 1}$。一个元级智能体对先前技能执行智能体式交叉,以针对任务 $\tau$ 创建新技能:$s_{k} = \text{crossover} \left(\right. \tau , \mathcal{H}_{k - 1} \left.\right)$。
The skill database tracks the history of previous skills, context functions and eval metrics $\mathcal{H}_{k - 1} = \left{\right. \left(\right. s_{i} , c_{i} , J_{i}^{\text{train}} , J_{i}^{\text{val}} \left.\right) \left.\right}_{i = 1}^{k - 1}$. A meta-level agent performs agentic crossover) over prior skills to create a new skill given a task $\tau$: $s_{k} = \text{crossover} \left(\right. \tau , \mathcal{H}_{k - 1} \left.\right)$.
然后,一个基础级上下文工程师执行技能 $s_{k}$,并从轨迹反馈 $\mathcal{R}_{k}$ 中学习上下文函数,由当前技能指导:$c_{k} = \text{engineer} \left(\right. \tau , s_{k} ; c_{k - 1}^{*} , \mathcal{R}_{k} \left.\right)$。
Then a base-level context engineer executes the skill $s_{k}$ and learns the context function from rollout feedback $\mathcal{R}_{k}$, guided by the current skill: $c_{k} = \text{engineer} \left(\right. \tau , s_{k} ; c_{k - 1}^{*} , \mathcal{R}_{k} \left.\right)$.
元上下文工程(MCE)框架:元级技能进化搜索上下文管理机制,而基础级优化任务上下文。(图片来源:Ye et al. 2026)
The framework of Meta Context Engineering (MCE): meta-level skill evolution searches over context-management mechanisms, while the base level optimizes the task context. (Image source: Ye et al. 2026)
MCE 不像 ACE 那样强制执行如何构建上下文的启发式规则。它使用自由形式技能来存储任务最重要的知识,并迭代地共同进化技能和技能条件上下文。在实现上,上下文函数 $c$ 被实例化为专用目录中的文件集合,包括静态(skill.md)和动态(上下文和数据轨迹)组件。元级和基础级优化都在智能体编码环境中使用标准工具集执行:
MCE does not enforce a heuristic rule for how to structure context as ACE does. It uses _free-form skills_ to store the most important knowledge for a task, and evolves the skill and the skill-conditioned context iteratively together. Implementation-wise, a context function $c$ is instantiated as a collection of files in a dedicated directory, including both static (skill.md) and dynamic (context and data rollouts) components. Both meta-level and base-level optimization are executed in agentic coding envs with a standard tool set,
\mathcal{T} = \left{\right. \mathtt{Read} , \mathtt{Write} , \mathtt{Edit} , \mathtt{Bash} , \mathtt{Glob} , \mathtt{Grep} , \mathtt{TodoWrite} \left.\right}
\mathcal{T} = \left{\right. \mathtt{Read} , \mathtt{Write} , \mathtt{Edit} , \mathtt{Bash} , \mathtt{Glob} , \mathtt{Grep} , \mathtt{TodoWrite} \left.\right}
元框架(Meta-Harness;Lee et al. 2026)又深入了一层:优化的对象是决定和优化哪些信息应被存储、检索并呈现给模型的代码。其名称中的“元”意味着它是一个用于优化框架的框架。
Meta-Harness (Lee et al. 2026) moves another level deeper: the optimized object is the _code_ that determines and optimizes what information should be stored, retrieved, and presented to the model. “Meta-” in its name means it is a harness for optimizing harnesses.
元框架外层循环优化算法。(图片来源:Lee et al. 2026)
The Meta-Harness outer-loop optimization algorithm. (Image source: Lee et al. 2026)
用于创建新框架的提议者本身就是一个编码智能体,最终输出是帕累托前沿上的一组框架候选。
The proposer for creating a new harness is itself a coding agent and the final output is a collection of harness candidates on the Pareto frontier.
* 整个执行历史可通过文件系统访问,因此编码智能体使用 grep 或 cat 等命令读取它,而不是将所有内容塞入单个提示上下文。
* The entire execution history is accessible via a file system, and thus the coding agent uses commands like grep or cat to read through it instead of shoveling everything into a single prompt context.
* 提议的框架是文件系统中的一个字典,包含其自身的源代码、分数、轨迹轨迹和状态更新。
* The proposed harness is a dictionary in the file system containing its own source code, scores, rollout trajectories, and state updates.
* 元框架循环迭代地创建新框架,只有合格的框架被保留。
* The mete-harness loop iteratively creates new harnesses, and only qualified ones are kept.
元框架在(左)少量迭代的文本分类和(右)TerminalBench-2 上的性能。注意,TerminalBench-2 实验中的搜索从 Terminus-KIRA 和 Terminus-2 这两个非常强大的框架初始化。(图片来源:Lee et al. 2026)
The performance of Meta-Harness on (Left) text classification with a small number of iterations and (Right) TerminalBench-2. Note that the search in the TerminalBench-2 experiment is initialized from Terminus-KIRA and Terminus-2, two very strong harnesses. (Image source: Lee et al. 2026)
尽管如此,重要的教训是明确的:一旦框架设计成为可执行的搜索空间,一个强大的编码智能体就可以利用人类工程师使用的相同设计空间。
Still, the important lesson is clear: once harness design becomes an executable search space, a strong coding agent can exploit the same design space human engineers use.
在驾驭工程中,工作流设计可以由领域专家手工完成。以自动研究为例,已有多种框架被提出和测试。AI Scientist 系统(Lu 等人,2026)构建了一个流水线,用于提出研究想法、编写代码、运行实验、分析结果、撰写论文并进行同行评审。Meng 等人(2026)将可验证性作为 ScientistOne 的核心设计约束,其中每个声明(引用、数值、方法、结论)都必须追溯到证据来源,并通过证据链检查进行审计。
Workflow design in harness engineering can be handcrafted by domain experts. Taking auto-research as an example, various frameworks have been proposed and tested. The AI Scientist system (Lu et al. 2026) builds a pipeline to propose research ideas, write code, run experiments, analyze results, write a manuscript, and perform peer review. Meng et al. (2026) make verifiability the central design constraint in ScientistOne, where every claim (citation, numerical, methodological, conclusion) must trace to an evidence source and is audited by Chain-of-Evidence checks.
AI Scientist 流水线:想法生成、实验、论文撰写和评审。(图片来源:Lu 等人,2026)
AI Scientist pipeline for idea generation, experimentation, paper writing, and review. (Image source: Lu et al. 2026)
Autodata 智能体(Kulikov 等人,2026)被设计为数据科学家,用于生成训练和评估数据。主智能体管理一个提出问题的挑战者、一个弱求解器、一个强求解器和一个验证者/评判者,旨在合成难度“恰到好处”的数据,即强求解器成功而弱求解器失败。
The Autodata agent (Kulikov et al. 2026) is designed to work as a data scientist for generating training and evaluation data. The main agent manages a _challenger_ that proposes problems, a _weak solver_, a _strong solver_, and a _verifier/judge_, aiming to synthesize data at the “just right” level of difficulty, meaning that the strong solver succeeds but the weak solver fails.
在 Autodata 中,挑战者提示根据求解器和验证者的反馈进行迭代更新。其局限性在于,合成任务用于微调弱求解器而非强求解器;如果循环无法迭代改进强模型,则更像是基于生成提示分布的间接蒸馏,缺乏递归自我改进(RSI)的味道。
In Autodata, the challenger prompt is updated iteratively according to feedback from the solvers and verifier. The limitation here is that synthesized tasks are used to fine-tune weak solvers but not strong solvers; if the loop cannot iteratively improve the strong model, it is more like indirect distillation over a generated prompt distribution, with less RSI flavor.
Autodata 智能体工作流设计:围绕挑战者、求解器和验证者角色生成合成训练和评估数据。(图片来源:Kulikov 等人,2026)
Autodata agentic workflow design for generating synthetic training and evaluation data around challenger, solver, and verifier roles. (Image source: Kulikov et al. 2026)
工作流的设计空间是巨大的,我们自然可以将工作流设计视为一个搜索问题,因此应该能够通过算法而非仅手工来找到好的解决方案。沿着这个方向,智能体系统的自动化设计(ADAS;Hu 等人,2025)将智能体设计本身形式化为一个优化问题,即“元智能体搜索”,其中元智能体提出智能体工作流的新设计。
The design space for workflow is _enormous_, and naturally we can think of workflow design as a search problem, and therefore we should be able to find good solutions by algorithms rather than only manually craft them. Following this direction, Automated Design of Agentic Systems (ADAS; Hu et al. 2025) formulates agent design itself as an optimization problem, “meta-agent search” where a meta-agent proposes new designs of agentic workflows.
1. 用简单的智能体(如思维链和自我精炼)初始化一个智能体工作流存档。
1. Initialize an archive of agentic workflows with simple agents such as CoT and self-refine.
2. 让元智能体以代码形式编程新的智能体,灵感来自存档中的现有解决方案。
2. Ask a meta-agent to program new agents, all in _code_, inspired by existing solutions in the archive.
* 元智能体首先生成新工作流的高级描述,然后用代码实现。
* The meta-agent first generates a high-level description of the new workflow, and then implements it in code.
* 草案程序随后经过元智能体的两个自我精炼步骤(即让模型提供反馈,然后让同一模型基于反馈精炼之前生成的输出;Madaan 等人,2023)以检查其新颖性。
* The draft program then goes through two self-refine steps (i.e. ask the model to provide feedback and then ask the same model to refine the previously generated outputs based on the feedback; Madaan et al. 2023) by the meta-agent to check its novelty.
3. 评估每个新候选,并将成功的添加回存档。
3. Evaluate each new candidate and add successful ones back to the archive.
4. 重复步骤 2-3,直到达到最大迭代次数。
4. Repeat steps 2-3 until the maximum iteration count is reached.
智能体系统的自动化设计(ADAS)示意图。
Illustration of Automated Design of Agentic Systems (ADAS).
AFlow(Zhang 等人,2025)将智能体工作流表示为图,其中节点表示调用 LLM 的动作,边实现代码中的逻辑操作。工作流优化依赖于蒙特卡洛树搜索(MCTS):
AFlow (Zhang et al. 2025) represents an agentic workflow as a graph, where nodes represent LLM-invoking actions and edges implement logical operations in code. The workflow optimization relies on MCTS (Monte Carlo Tree Search):
1. 用模板在树中初始化起始工作流 $W_{0}$。
1. Initialize the starting workflow $W_{0}$ in the tree with a template.
2. 使用分数和均匀探索的软混合选择一个工作流节点。
2. Select a workflow node using a soft mixture of score and uniform exploration.
3. 通过让 LLM 根据其评估性能生成修改后的工作流来扩展它。
3. Expand it by asking an LLM to produce a modified workflow conditioned on its evaluation performance.
4. 执行并评估新工作流。
4. Execute and evaluate the new workflow.
5. 如果新工作流在 $N$ 轮预算内显示出改进,则将其添加回树中。
5. Add it back to the tree if the new workflow shows improvement within a budget of $N$ rounds.
6. 重复步骤 2-5,当 top-$k$ 平均分数趋于平稳或达到预算时停止。
6. Repeat steps 2-5 and stop when the top-$k$ average score plateaus or hit the budget.
AFlow 在工作流候选树上的优化过程。(图片来源:Zhang 等人,2025)
AFlow optimization process over a tree of workflow candidates. (Image source: Zhang et al. 2025)
AFlow 在问答、代码和数学任务上的实验显示,AFlow 相对于手工设计的工作流和 ADAS 有显著改进。
Experiments of AFlow in QA, code, and math tasks showed decent improvement of AFlow over manually designed workflows and ADAS.
AFlow 实验与手工方法和 ADAS 的对比。(图片来源:Zhang 等人,2025)
AFlow experiments in comparison to manual methods and ADAS. (Image source: Zhang et al. 2025)
无论是上下文工程还是工作流设计,都只是 harness 的一部分。我们需要搜索整个设计空间,并共同优化上下文管理逻辑、工作流、权限以及许多其他 harness 组件。正如我们在 Meta-Harness、ADAS 和 AFlow 等工作中所见,✨代码✨是定义程序和系统的通用语言。简单来说,harness 是编程如何使提示、工具调用、子智能体、控制流、记忆和工作流逻辑协同工作的代码。如果大语言模型能够优化执行智能体的代码,它就能访问比手工编写提示大得多的设计空间。
Either context engineering or workflow design is only one part of a harness. We need to search through the entire design space and optimize context-management logic, workflow, permissions, and many other harness components together. As we have seen in work like Meta-Harness, ADAS, and AFlow, ✨code✨ is a universal language for defining programs and systems. In simple words, a harness is code that programs how prompts, tool calls, subagents, control flow, memory, and workflow logic work together. If an LLM can optimize the code that executes agents, it can access a _much larger design space_ than hand-written prompts.
自教优化器(STOP;Zelikman 等人,2023)是递归脚手架改进的早期例子之一。在步骤 t=0 时,种子改进器 I₀接收初始解决方案 s、效用函数 u 和黑盒语言模型 M,并返回改进后的解决方案 s',即 s' = I(u, s; M)。STOP 的目标不是直接改进 s,而是改进改进器 I 本身。
Self-Taught Optimizer (STOP; Zelikman et al. 2023) is one of the early examples of recursive scaffolding improvement. A seed improver $I_{0}$ at step $t = 0$ takes an initial solution $s$, a utility function $u$, and a black-box language model $M$, and returns an improved solution $s^{'}$, that is, $s^{'} = I \left(\right. u , s ; M \left.\right)$. The goal of STOP is not directly to improve $s$ but _to improve the improver $I$ itself_.
首先,我们将元效用定义为给定改进器函数 I 在一组下游任务 D 上的平均效用:
First, let’s define the meta-utility as the average utility of a given improver function $I$ over a collection of downstream tasks $\mathcal{D}$:
û(I) ≜ (1/|D|) E_{(u,s)~D}[u(I(u, s; M))]
\hat{u} \left(\right. I \left.\right) \triangleq \frac{1}{\left|\right. \mathcal{D} \left|\right.} \mathbb{E}_{\left(\right. u , s \left.\right) sim \mathcal{D}} \left[\right. u \left(\right. I \left(\right. u , s ; M \left.\right) \left.\right) \left]\right.
由于改进改进器函数本身就是一个优化问题,我们可以通过自我改进更新,根据 I_{t-1}由元效用衡量的性能,递归地获得 I_t 的新版本:
Because improving the improver function is an optimization problem itself, we can recursively get a new version of $I_{t}$ based on $I_{t - 1}$’s performance measured by meta-utility via a self-improvement update:
I_t = I_{t-1}(û, I_{t-1}; M)
I_{t} = I_{t - 1} \left(\right. \hat{u} , I_{t - 1} ; M \left.\right)
自教优化器(STOP)算法。(图片来源:Zelikman 等人,2023)
Algorithm of Self-Taught Optimizer (STOP). (Image source: Zelikman et al. 2023)
在他们的实验中,改进后的改进器发现了各种策略,例如遗传算法、分解和改进部分、多臂提示赌博机、模拟退火、变化温度以及波束/树搜索。这类似于 harness 工作流如何被表示为优化对象。
In their experiments, the improved improver discovered various strategies, such as genetic algorithms, decomposing and improving parts, multi-armed prompt bandits, simulated annealing, varying temperature, and beam/tree search. This is analogous to how a harness workflow can be represented as an object for optimization.
STOP 发现的自我改进策略示例。(图片来源:Zelikman 等人,2023)
Examples of self-improvement strategies discovered by STOP. (Image source: Zelikman et al. 2023)
Zelikman 等人(2023)发现的一个警示性结果是,STOP 在使用 GPT-4 时迭代地提高了平均下游性能,但在使用 GPT-3.5 和 Mixtral 等较弱模型时性能下降。仅靠递归结构是不够的。基础模型必须足够有能力来改进机制。这意味着 harness 改进能够更好地部署模型,但智能仍然是核心。
A _cautionary_ result in Zelikman et al. (2023)’s findings is that STOP improved mean downstream performance across iterations with GPT-4 but degraded with weaker models like GPT-3.5 and Mixtral. Recursive structure alone is not enough. The base model must be _capable enough_ to improve the mechanism. This implies that harness improvement enables better deployment of the model but intelligence is still the core.
Lin 等人(2026)更详细地研究了 harness 进化对模型能力的依赖性。他们区分了两个维度:(1)harness 更新能力,指产生有用 harness 编辑的能力;(2)harness 收益能力,指利用更新后的 harness 实现更好任务解决的能力。有趣的是,在他们的实验中,从 Qwen3.5-9B 到 Claude Opus 4.6 的一系列不同规模和核心智能的模型都表现出相似的 harness 更新能力;9B 的 harness 提议者/进化者能够编写与 Opus 程序同构的技能。为了最好地利用 harness,模型需要正确且及时地调用技能/工具,并擅长长程指令遵循。
Lin et al. (2026) investigated the dependency of harness evolution on model capabilities in more details. They disentangled two axes: (1) _harness-updating_ refers to the capability of producing useful harness edits and (2) _harness-benefit_ denotes the capability of utilizing the updated harness, to achieve better task solving. Interestingly a range of model of different sizes and core intelligence, from Qwen3.5-9B to Claude Opus 4.6, were observed in their experiments to show similar harness updating capability; the 9B harness proposer/evolver is able to write a skill procedurally isomorphic to Opus. To best utilize a harness, a model needs to invoke skills/tools correctly and timely and be good at long-horizon instruction following.
主要结果:(A)从 Qwen2-32B 到 Opus 4.6 的一系列模型中,harness 更新能力保持平稳;(B)harness 收益能力是非单调的,中等层级的模型受益最多。(图片来源:Lin 等人,2026)
Main results: (A) harness updating capability is measured flat across a range of models from Qwen2-32B to Opus 4.6; (B) harness benefit capability is non-monotonic where middle tier models benefit the most. (Image source: Lin et al. 2026)
最近的工作 Self-Harness(Zhang 等人,2026)依赖大语言模型智能体通过提议-评估-接受循环来改进自己的 harness。
A more recent work, Self-Harness (Zhang et al. 2026), relies on LLM agents to improve their own harness via a propose-evaluate-accept loop.
Self-Harness 使用弱点挖掘、有界 harness 提议和验证的循环来更新 harness。(图片来源:Zhang 等人,2026)
Self-Harness uses a loop of weakness mining, bounded harness proposal, and validation to update a harness. (Image source: Zhang et al. 2026)
Self-Harness 中的循环包含三个阶段:
The loop in Self-Harness has three stages:
1. 弱点挖掘:将失败聚类为验证器可验证的失败模式。
1. _Weakness mining_: cluster failures into verifier-grounded failure patterns.
* 当前 harness h_t 用于评估任务,并收集执行轨迹进行分析。
* The current harness $h_{t}$ is used to evaluate on tasks and execution traces are collected for analysis.
* 注意,两次运行可能在错误日志表面共享相同的验证器结果,例如超时或缺少工件,但具有不同的因果机制。因此,我们需要包含丰富信息的失败记录,包括终端验证器级别的原因、相关智能体行为的因果状态以及轨迹暴露的抽象智能体机制,以揭示根本原因。
* Note that two runs can share the same verifier outcome in the error logs on the surface, such as timeout or missing artifact, while having different causal mechanisms. Therefore we need a failure record of rich information, containing the terminal verifier-level cause, the causal status of the relevant agent behavior, and the abstract agent mechanism exposed by the trace, to uncover the root causes.
2. Harness 提议:基于挖掘到的失败模式提出有界的 harness 编辑。
2. _Harness proposal_: propose bounded harness edits based on mined failure patterns.
* 在 h_t 下调用相同的模型作为提议者。
* The same model is invoked under $h_{t}$ as a proposer.
* 向模型提供有界的提议上下文:(1)当前 harness 的可编辑表面,(2)来自评估系统的验证器可验证失败模式,(3)应保留的通过行为记录,以及(4)先前尝试编辑的摘要。
* The model is provided with a bounded proposal context: (1) the editable surfaces of the current harness, (2) the verifier-grounded failure patterns from the evaluation system, (3) records of passing behaviors that should be preserved, and (4) summaries of previously attempted edits.
* Harness 编辑应优先考虑可解决的重复错误模式(例如,不是任务特定的难度),并且可以通过狭窄的更改来解决。
* Harness edits should prefer recurrent error patterns that are addressable (e.g. not task-specific difficulty) and can be resolved by narrow changes.
* Harness 编辑候选应独特且多样。
* Harness edit candidates should be distinct and diverse.
3. 提议验证:验证并合并合格的编辑以创建新的 harness h_{t+1}。
3. _Proposal validation_: validate and merge qualified edits to create a new harness $h_{t + 1}$.
* 候选编辑通过保留集 D_in(用于测试弱点是否解决)和保留集 D_out(用于检查是否引入了其他未知问题)上的回归测试进行评估。
* Candidate edits are evaluated by regression tests on held-in $D_{\text{in}}$ (for testing whether the weakness is resolved) and held-out $D_{\text{out}}$ (for checking whether other unknown issues were introduced) splits.
* 仅当候选编辑在保留集和保留集上均无回归时才被接受。
* Candidates are accepted only if they have no regression on both held-in and held-out data.
* 接受的候选编辑被合并以将 harness 更新为 h_{t+1},而拒绝的候选编辑被记录而不更改活动 harness。
* Accepted candidates are merged to update the harness to $h_{t + 1}$, while rejected candidates are logged without changing the active harness.
在 Terminal-Bench-2 上运行 MiniMax M2.5、Qwen3.5-35B-A3B 和 GLM-5 时,Self-Harness 被证明能够学习针对不同基础模型不同弱点的模型特定 harness 指令,并提高保留集通过率。
When running MiniMax M2.5, Qwen3.5-35B-A3B, and GLM-5 on Terminal-Bench-2, Self-Harness was shown to learn model-specific harness instructions that target at different weaknesses of different base models and improve held-out pass rates.
Self-Harness 这类工作确实引起了我的担忧:如果允许程序编辑操作系统,抽象边界将被打破。可编辑表面需要适当设计,权限控制和安全层需要位于此循环之外。所有关于奖励黑客的挑战仍然存在。
Self-harness type of work does raise my concerns that if a program is allowed to edit the OS system, abstraction boundaries are broken. The editable surface needs to be properly designed and the permission control and security layers need to live outside this loop. All the challenges around reward hacking still remain.
智能体式 Harness 工程(AHE;Lin 等人,2026)认为 harness 进化的瓶颈在于可观测性——即当一次运行失败时,我们需要知道哪个组件负责,并且每次编辑都应有证据支持。
Agentic Harness Engineering (AHE; Lin et al. 2026) see the bottlenecks of harness evolution are around observability—that is, when a rollout fails, we need to know which component is responsible for that and every edit should be grounded by evidence.
该框架创建了一个包含三个可观测性支柱的闭环:
The framework creates a closed loop with 3 observability pillars:
1. 组件可观测性:每个可编辑的 harness 组件在文件系统中都有表示,因此动作空间是明确且可追踪的。
1. _Component observability_: every editable harness component has a representation in the file system so the action space is explicit and tracable.
* Harness 包含 7 个组件:系统提示、工具描述、工具实现、中间件、技能、子智能体配置和长期记忆。
* A harness contains 7 components: system prompt, tool description, tool implementation, middleware, skill, sub-agent configuration, and long-term memory.
* 每个失败模式映射到一个组件,以便编辑更具针对性。
* Each failure pattern is mapped to one component so the edit can be more targeted.
2. 经验可观测性:将大量原始轨迹分析并总结为证据和失败模式的层次结构。
2. _Experience observability_: analysize and summarize a large amount of raw trajectories into a hierarchy of evidence and failure patterns.
* 使用智能体(“智能体调试器”)分析每个存储在一个文件中的轨迹,并生成每个任务的分析报告,说明失败或成功的根本原因。
* Use an agent (“Agent debugger”) to analysis the trajectories each stored in one file and generate per-task analysis report on the root cause for the failure or success.
* 所有每个任务的报告汇总为下一步的基准概览,如果需要可以访问原始轨迹。这种分层访问结构更节省 token。
* All the per-task reports are aggregated into a benchmark overview for the next step, and raw traces can be accessed if needed. This layered access structure is more token efficient.
3. 决策可观测性:每次编辑都附带一个对下一轮验证的预测。
3. _Decision observability_: every edit is paired with a prediction for the next round to validate.
* 智能体(“进化智能体”)读取仓库并决定编辑哪个组件,然后生成编辑及其背后的推理。
* An agent (“Evolve agent”) reads the repo and decides which component to edit, and then produces the edit and the reasoning behind it.
* 每次编辑都是一个文件级别的可证伪声明,可以在下一轮验证,受两个约束:
* Every edit is a file-level, falsifiable claim and can be verified in the next round, under two constraints:
* (1)编辑仅应用于 harness 工作空间。运行目录、追踪器、验证器和 LLM 配置是只读的,这禁用了一组奖励黑客行为(例如禁用验证器、交换模型或提高推理预算),因此可以保持每个记录的收益归因于 harness 编辑。
* (1) Edits are only applied to the harness workspace. the runs directory, tracer, verifier, and LLM configuration are read-only, which disables a set of reward hacking (e.g disabling the verifier, swapping the model, or raising the reasoning budget) and thus it can keep every recorded gain attributable to harness edits.
* (2)编辑是证据驱动的,包含一个清单条目:失败证据的名称、推断的根本原因、目标修复以及包含预期修复和风险回归的预测影响。
* (2) Edits are evidence-driven, with a manifesto entry: the failure evidence’s name, the inferred root cause, the targeted fix, and a predicted impact comprising both expected fixes and at-risk regressions.
在 Terminal-Bench-2 上,AHE 实现了优于人工设计的 harness(OpenCode、Terminus-2、Codex)的性能,除了 Hard 层级和少数其他自我进化基线(ACE、TF-GRPO)。相同的冻结 harness,无需进一步进化,即可迁移到 SWE-bench-verified,表明进化后的 harness 能够将工程经验编码到 harness 组件中,而不是进行特定于基准的优化。
On Terminal-Bench-2, AHE achieved better than human-designed harness (OpenCode, Terminus-2, Codex) except for Hard tier and a few other self-evolve baselines (ACE, TF-GRPO). The same frozen harness, without further evolving, transfers to SWE-bench-verified, indicating that the evolved harness is able to encode engineering experience into harness components rather than doing benchmark-specific optimization.
进化搜索是一种受自然选择启发的优化方法(参见我关于进化算法的旧文)。它通过变异解群体并只保留群体中“适应度”高的解来进化。当(1)搜索空间广阔或形状怪异;(2)难以直接用梯度优化但容易评估解时,进化搜索就派上用场了。Harness 搜索似乎很适合这里。
Evolutionary search is an optimization method inspired by natural selection (see my old post on evolutionary algorithm). It evolves a population of solutions by mutating them and only keeping those with high “fitness” in the crowd. Evolutionary search comes in handy when (1) the search space is extensive or weirdly shaped; and (2) it is hard to optimize directly with gradients but easy to evaluate solutions. Harness search seems to be a good fit here.
进化搜索已在过去的提示工程研究中被使用。Promptbreeder(Fernando 等人,2023)通过丰富的变异操作集优化任务特定提示,有趣的是,变异提示(即用于变异任务提示的 LLM 指令)本身也通过进化得到改进。GEPA(Agrawal 等人,2025)将基于反思的提示与进化搜索相结合,并利用试错轨迹上的自然语言反思来提出提示更新。
Evolutionary search has been used in prompt engineering in the past studies. Promptbreeder (Fernando et al. 2023) optimizes task-specific prompts through a rich set of mutation operations, and interestingly the mutation prompts (i.e. instructions to an LLM to mutate a task prompt) are themselves also improved through evolution. GEPA (Agrawal et al. 2025) combines reflection-based prompting with evolutionary search and uses natural language reflection over trajectories of trial and error to propose prompt updates.
Novikov 等人(2025)引入了 AlphaEvolve 作为一个编码智能体进化搜索系统,该系统存储候选程序池,并提示冻结的 LLM 生成差异以进行改进。随着系统反复评估子程序并保留成功的程序,它及时发现了更好的解。
Novikov et al. (2025) introduced AlphaEvolve as a coding-agent evolutionary search system, which stores a pool of candidate programs and prompts frozen LLMs to generate diffs for improvement. As the system repeatedly evaluates child programs and keeps successful ones, it discovers better solutions in time.
AlphaEvolve 的工作原理。(图片来源:Novikov 等人,2025)
How AlphaEvolve works. (Image source: Novikov et al. 2025)
AlphaEvolve 的设计中有几个细节很重要:
A few details matter in the design of AlphaEvolve:
* 提示包括父程序、结果、指令,有时还有元信息。
* The prompt includes parent programs, results, instructions, and sometimes meta information.
* 编码智能体可以访问完整仓库,但改进的代码区域明确用 # EVOLVE-BLOCK-START 和 # EVOLVE-BLOCK-END 标记。
* The coding agent has access to the full repo, but code regions for improvement are explicitly marked with # EVOLVE-BLOCK-START and # EVOLVE-BLOCK-END.
* 元提示与指令和上下文共同进化,由 LLM 建议,类似于我们进化解程序的方式。
* Meta-prompt co-evolves with instructions and context as suggested by LLM, in a similar way as how we evolve solution programs.
消融实验显示了进化过程、提示中的上下文、元提示、全文件进化以及使用更强 LLM 的效果。
Ablations show the evolution procedure, context in prompts, meta-prompts, full-file evolution and the use of stronger LLMs.
消融实验显示了 AlphaEvolve 中若干设计的价值。(图片来源:Novikov 等人,2025)
Ablations show the value of everal designs in AlphaEvolve. (Image source: Novikov et al. 2025)
最近的变体如 ThetaEvolve(Wang 等人,2025)将进化搜索与强化学习和上下文学习相结合,而 DemoEvolve(Che 等人,2026)用人类专家演示作为参考经验来增强自展开存档,用于 harness 级别的诊断和编辑。另一方面,ShinkaEvolve(Lange 等人,2025)引入了三个新组件来提高 LLM 采样效率:
Recent variants such as ThetaEvolve (Wang et al. 2025) combines evolutionary search with RL and in-context learning, and DemoEvolve (Che, et al. 2026) augments the self-rollout archive with human expert demonstrations as reference experience for harness-level diagnosis and editing. ShinkaEvolve (Lange et al. 2025), on the other hand, introduced three new components to improve LLM sampling efficiency:
* 通过设计父采样以平衡性能排名和后代数量,实现更高效的探索。
* More sample-efficient exploration by designing parent sampling to balance performance rank and offspring count.
* 代码新颖性拒绝采样,基于嵌入的余弦相似度丢弃与现有群体过于相似的候选。
* Code-novelty rejection sampling by discarding candidates that are too similar to the existing population based on embedding-based cosine similarity.
* 在元暂存区中识别成功解中的良好模式,以指导未来的变异。
* Identifying good patterns in successful solutions in a meta-scratchpad to guide future mutation.
与上述专注于解改进的方法不同,达尔文·哥德尔机(DGM;Zhang 等人,2025)明确针对可编辑 harness 代码仓库的进化,使用基于 LLM 的编码智能体。确切地说,该智能体被允许修改其自身的 harness。后续工作 Hyperagents(Zhang 等人,2026)引入了一个元智能体来控制如何修改现有任务智能体以创建新智能体。
Unlike the methods above, which focus on solution improvement, Darwin Gödel Machine (DGM; Zhang et al. 2025) explicitly targets the evolution of an editable harness-code repository with an LLM-based coding agent. Precisely, this agent is allowed to modify its own harness. A follow-up work on Hyperagents (Zhang et al. 2026) introduced a meta-agent to control how to modify existing task agents to create new ones.
1. 从池中的一个编码智能体开始。
1. Start with one coding agent in the pool.
2. 在每次迭代中,选择一个父智能体,概率与其性能成正比,与其子代数量成反比,进行修改并分支产生新智能体。
2. In each iteration, pick one parent with a probability proportional to its performance and inversely to the number of children it has, to modify and branch off to produce new agents.
3. 选中的父智能体检查其自身的基准评估日志,然后对其自身的 harness 代码库提出改进,以生成新版本的编码智能体。代码编辑通过两个基本工具实现:(1)bash(参数:<bash_command>)和(2)编辑器(参数:view/create/edit <file_path>)。
3. The selected parent agent examines its own benchmark evaluation log and then proposes improvements to its own harness codebase to generate a new version of the coding agent. Code editing is implemented with two basic tools: (1) bash (args: <bash_command>) and (2) editor (args: view/create/edit <file_path>).
4. 新编码智能体被评估,只有性能足够高的才被添加回池中。
4. New coding agents are evaluated, and only those with sufficiently high performance are added back into the pool.
5. 重复步骤 2-4,直到达到某个停止标准。
5. Repeat steps 2-4 until some stop criteria hit.
DGM 是在固定模型下的 harness 进化。在使用 Claude 3.5 Sonnet 作为基础 LLM 和简单初始 harness 配置的实验中,DGM 发现的智能体在 SWE-bench Verified(20% 到 50%)和 Polyglot(14.2% 到 30.7%)上可与手工制作的智能体相媲美或更优。
DGM is harness evolution under a fixed model. In experiments with Claude 3.5 Sonnet as the base LLM and simple initial harness configs, the DGM-discovered agents are comparable to or outperform handcrafted agents on SWE-bench Verified (20% to 50%) and Polyglot (14.2% to 30.7%).
这类方法在候选解可自动评估且候选适应度易于量化时效果良好,例如矩阵乘法、GPU 内核优化、算法竞赛、数据中心调度。它在评估缓慢、模糊或主要基于启发式的领域则表现不佳。进化的计算效率和有效性也是问题。
This family of methods works well when candidate solutions are automatically evaluable and candidate fitness is easy to quantify, such as matrix multiplication, GPU kernel optimization, algorithm contests, datacenter scheduling. It struggles with domains where evaluation is slow, ambiguous, or mostly heuristic-based. The compute efficiency and effectiveness of evolution are also concerns.
Harness 进化改变了模型周围的非参数化系统。为了实现完全的自我改进,模型可以同时被允许更新自身的权重。权重更新可以通过模型训练流程的改进或测试时的持续学习来实现。持续学习这一主题值得未来专门写一篇文章。
Harness evolution changes the non-parametric system around the model. To enable full self-improvement, the model can totally be allowed to update its own weights at the same time. The weight update can be implemented via improvements in the model training pipeline or continual learning at test time. The topic of continual learning is worthy of its own post in the future.
SIA(Hebbar 等人,2026)是早期尝试将 harness 改进和模型参数更新结合在同一优化循环中的工作,其设计包含三个组件:
SIA (Hebbar et al. 2026) is an early attempt to combine harness improvement and model-parameter updates in the same optimization loop, with three components in the design:
* _元智能体_:提出初始 harness。
* _Meta-Agent_: proposes the initial harness.
* _任务特定智能体_:执行任务。
* _Task-Specific Agent_: executes the task.
* _反馈智能体_:根据最近的轨迹决定是更新 harness 还是模型权重。
* _Feedback-Agent_: chooses whether to update the harness or the model weights based on recent trajectories.
SIA 中的反馈智能体决定下一次迭代的类型。(图片来源:Hebbar 等人,2026)
The Feedback-Agent in SIA decides the next iteration type. (Image source: Hebbar et al. 2026)
SIA 的实验中有一些令人困惑的选择,使得结果难以解释。例如,任务特定智能体远弱于用于元智能体和反馈智能体的模型(gpt-oss-120b vs Claude Sonnet 4.6),并且基线太弱,无法与相关方法进行清晰的交叉对比。我认为这个方向很有趣,但证据是初步的。然而,许多挑战,如训练稳定性和古德哈特效应,仍然悬而未决。
There are a few confounding choices in SIA’s experiments that make the results hard to interpret. For example, the task-specific agent is much weaker than the models used for the Meta-Agent and Feedback-Agent (gpt-oss-120b vs Claude Sonnet 4.6), and the baselines are too weak to cross-reference cleanly against related methods. I would consider the direction interesting, but the evidence provisional. Yet many challenges, such as training stability and Goodhart effect, still remain open.
Continual Harness(Karten 等人,2026)在长周期游戏场景中进行了实验,通过从强教师模型在低奖励轨迹上的标签进行蒸馏,实现了 harness 更新和策略模型的共同学习。
Continual Harness (Karten et al. 2026) experimented in long-horizon gameplay setting with harness updating and co-learning a policy model by distilling a strong teacher model’s labels on low-reward trajectories.
AI 科学家系列工作有力地证明,专家设计的遏制框架能够协调自动研究循环的很大一部分,并以撰写研究论文的形式进行实验。但论文产出并不等同于科学发现。一个系统可以写出一篇看似合理的稿件,同时仍可能存在捏造的引用、实现偏差或薄弱的实验结果。
The AI Scientist line of work is a strong demonstration that an expert-designed harness can coordinate a large portion of auto-research loop, experimented in the form of writing research papers. But paper production is not identical to scientific discovery. A system can write a plausible manuscript while still having fabricated citations, implementation drift, or weak experimental results.
Trehan & Chopra (2026) 测试了 LLM 能否在最小脚手架和基本工具(即 read_file、write_file、llm_search、list_files)下,从研究想法到论文。每个想法都有一个专用工作空间,智能体可以在其中生成和阅读文档作为上下文的一部分。他们在三个领域(世界模型、多智能体强化学习、AI 安全与对齐)进行了实验,每个领域包含 45-50 篇高质量种子文档以激发新想法。只有四个想法被人类专家选中进入完整流程,仅有一个完全执行成论文。他们在实验中观察到六种反复出现的失败模式:
Trehan & Chopra (2026) tested whether LLMs can go from a research idea to a paper with minimal scaffolding and basic tools (i.e., read_file, write_file, llm_search, list_files). Each idea had a dedicated workspace where agents could generate and read documents as part of context. They experimented in three domains (world models, multi-agent RL, AI safety & alignment), with each domain containing 45-50 high-quality seed documents to inspire new ideas. Only four ideas were selected by human experts to run through the full pipeline, and only one was fully executed into a paper. They observed six recurring failure modes in the experiments:
* _偏向训练数据默认值_:使用旧库、过时命令、标准格式,或假设未基于实际仓库或数据集。
* _Bias toward training-data defaults_: use old libraries, stale commands, standard formats, or assumptions not grounded in the actual repository or dataset.
* _执行压力下的实现漂移_:当实现变得技术复杂时,模型可能转向常见的更简单解决方案,而非所提出的方法。
* _Implementation drift under execution pressure_: when implementation becomes technically complex, the model may move toward a common simpler solution rather than the proposed method.
* _记忆与上下文退化_:长期项目会丢失关键细节,除非日志被写为持久性工件。
* _Memory and context degradation_: long-horizon projects lose critical details unless logs are written as persistent artifacts.
* _过度乐观_:模型在实验有噪声或失败时仍宣布成功,类似 Bubeck 等人(2025)观察到的“p-hacking and eureka-ing”模式,即模型可能引入“数值胶带”并在信号仍为噪声时宣布胜利。
* _Over-optimism_: the model declares success despite noisy or failed experiments, similarly observed as “p-hacking and eureka-ing” pattern by Bubeck et al. (2025) where models can introduce “numerical duct tape” and declare victory when signals are still noise.
* _领域智能不足_:模型缺乏隐性工艺知识,例如预测实现复杂度、判断实验结果是否合理,或知道哪些基线重要。
* _Insufficient domain intelligence_: the model lacks tacit craft knowledge, e.g. predicting implementation complexity, judging whether an experimental result is plausible, or knowing which baselines matter.
* _科学品味薄弱_:实验可能可执行,但未能回答正确的问题。
* _Weak scientific taste_: experiments may be executable but fail to answer the right question.
朝着完全 RSI,研究人员取得了实际进展,但仍存在几个瓶颈。
Toward full RSI, researchers have made real progress, but several bottlenecks remain.
1. 薄弱且模糊的评估器。许多研究主张没有快速且精确的验证器,许多现实世界任务也是如此。当前的自我改进循环在评估指标可测量且客观时效果最佳,类似于强化学习的工作方式。
1. Weak and fuzzy evaluators. Many research claims do not have a fast and precise verifier, and the same is true for many real-world tasks. Current self-improvement loops work best for tasks when evaluation metrics are measurable and objective, similar as how RL works.
研究品味、新颖性和长期科学价值更难衡量。例如,研究品味通常混合了问题框架、实验设计,以及判断哪些令人惊讶的结果值得追求、哪些失败案例值得重试。
Research taste, novelty, and long-term scientific value are much harder to measure. For example, research taste often mixes problem framing, experimental design, and judgment about which surprising results are worth pursuing and which failure cases are worth retries.
2. 上下文与记忆生命周期。随着 AI 智能体变得更加自主和独立,记忆会增长。一个有用的遏制框架需要管理上下文和记忆,以补充长上下文生成中的现有局限性,同时最大化长期任务的成功。由于人类能够终生维持记忆,我在此看到一个类比:上下文工程将并且应该成为智能的核心部分,而不是停留在软件系统层。
2. Context and memory lifecycle. Memory grows as AI agents become more autonomous and independent. A useful harness needs to manage context and memory to complement existing limitation in long-context generation while still maximizing the success of long-horizon tasks. Since humans are able to maintain memory through our life time, I see an anoloy here that context engineering will and should become a core part of intelligence, rather than staying in the software system layer.
3. 负面结果。研究人员有动力发表成功结果,因此文献偏向成功。在大量数据(至少目前主要是人类创建的,哈哈)上训练的 LLM 可能不擅长决定何时放弃假设、报告负面结果,甚至承认失败,因为数据中成功与失败案例的不平衡。研究遏制框架应使失败尝试易于保存,因为从失败中学习是缩小任务搜索空间的最佳方式。
3. Negative results. Researchers are incentivized to publish successful results and thus literature is biased toward successes. LLMs trained on a vast amount of data (mostly human created, at least for now, lol) may be bad at deciding when to abandon a hypothesis, report a negative result, or even acknowledge a failure due to the imablance of success vs failure cases in data. A research harness should make failed attempts easy to preserve, as learning from failure is the best way to trim down the task search space.
4. 多样性崩溃。进化和强化学习循环倾向于利用已知的高奖励模式。我们需要机制防止种群崩溃为同一解决方案的变体。这对于开放式研究尤其关键,因为最佳路径在当前评估器下最初可能看起来更差。
4. Diversity collapse. Evolutionary and RL loops tend to exploit known high-reward patterns. We need mechanisms to prevent the population from collapsing into variants of the same solution. This is especially critical for open-ended research, where the best path may initially look worse under the current evaluator.
5. 奖励破解。自我改进循环优化它得到的任何信号。如果奖励来自单元测试,智能体可能过拟合测试;如果来自评判模型,它可能学习针对该评判模型的奖励破解技巧;如果来自基准分数,它可能利用基准工件。
5. Reward hacking. A self-improvement loop optimizes whatever signal it is given. If the reward comes from unit tests, the agent may overfit to tests; if it comes from a judge model, it may learn reward hacking tricks specific to this judge; if it comes from benchmark scores, it may exploit benchmark artifacts.
评估器和权限控制很可能应位于演化遏制框架的循环之外,并配有留出测试、痕迹审计以及关键决策点的人类审查——多少监督可以扩展和自动化仍是一个开放的研究领域。
The evaluator and permission control should likely sit outside the loop that evolves harness, with held-out tests, trace audits, and human review at decision points that matter—how much oversight can be scaled up and automated remains an open research area.
6. 长期成功。一个外在的优化循环作用于单个轨迹之外的奖励,我们可以在训练沙盒中模拟这些轨迹。
6. Long-term success. An extrinsic loop of optimization works on rewards outside of individual rollouts that we can simulate in training sandbox.
以编码智能体为例。编码智能体已经提高了软件工程的日常生产力,但许多优化目标仍然过于短期。它通常能完成手头任务,但不太明显应如何保护由数百或数千名工程师共同维护的仓库的长期健康。基于沙盒的标准 RLVR 式训练很少捕捉可维护性、所有权边界、迁移成本、向后兼容性或未来的调试负担。
Take coding agent as an example. Coding agents have already increased daily productivity in software engineering, but many optimization goals are still too short-term. It can often complete the task at hand, but less obvious how it should protect the long-term health of a repo collectively maintained by hundreds or thousands of engineers. Standard sandbox-based RLVR-style training rarely captures maintainability, ownership boundaries, migration cost, backwards compatibility, or future debugging burden.
7. 人类的角色。人类应向上移动堆栈,而不是被移出循环,这意味着人类应在正确的时间、正确的抽象层级提供监督,我们的系统设计应考虑何时以及如何设置这样的接触点。
7. The role of humans. Humans should move up the stack, not be removed from the loop, meaning that human should provide oversight at the right time, at the right abstraction level and our system design should consider when and how to set up such touch points.
上述许多挑战需要人类的反馈和引导。毕竟,我们正在为人类更美好的未来构建技术,而不是相反。
Many challenges listed above need human’s feedback and steering. After all, we are building the technology for better future of humanity, not other way around.