An Alien Mind
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→在这篇文章中,OpenAI 首席科学家 Jakub Pachocki 回顾了自 2023 年 RLSlow 项目以来推理语言模型的发展轨迹。该项目首次表明,对推理模型的训练进行规模化扩展,可以释放预训练模型形成自身思维链的能力。他认为,机器智能是“生长”出来的,而非被设计出来的,它源自海量算力上的反复优化;机器智能与人类智能并不直接可比,因此我们不能默认它会继承人类的价值观。核心问题是对齐,他将其分为目标对齐与价值对齐,而根本挑战在于泛化。他描述了两种主要的对齐训练方法及其脆弱性,并解释了为何思维链监控一直是 OpenAI 最主要的实证押注,即便其可靠性正在下降。他的结论是:能力的持续快速提升——甚至可能走向递归式自我改进——要求我们极度审慎,而若缺乏更广泛的干预,仅靠技术修补远远不够。
In this essay, OpenAI Chief Scientist Jakub Pachocki reflects on the trajectory of reasoning language models since the 2023 RLSlow project, which first showed that training reasoning models could be scaled to unlock pretrained models' ability to form their own chains of thought. He argues that machine intelligence is grown rather than designed, emerging from repeated optimization on enormous compute, and that it is not directly comparable to human intelligence, so we cannot assume it inherits human values by default. The central problem is alignment, which he divides into goal alignment and value alignment, with generalization as the fundamental challenge. He describes two main alignment training approaches and their brittleness, and explains why chain-of-thought monitoring has been OpenAI's primary empirical bet even as its reliability diminishes. His conclusion is that continued rapid capability gains, possibly into recursive self-improvement, demand extreme caution, and that technical fixes alone are insufficient without broader interventions.
作者:Jakub Pachocki,OpenAI 首席科学家
By: Jakub Pachocki, Chief Scientist at OpenAI
2023 年年中,在“RLSlow”研究项目中,我们首次看到了让我们相信能够扩展推理模型训练的结果,释放了预训练模型形成自身思维链的能力。Szymon 和我那天晚上在办公室度过,思考的不是这项技术将带来的惊人基准数字、产品或科学成果——而是试图消化一个清醒的事实:我们将在有生之年真正看到比我们更聪明的机器,而且我们已经看到了这些系统的雏形;思考如何提醒人们其重要性。
In mid-2023, within the “RLSlow” research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pretrained models to form their own chains of thought. Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver - but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime, and we already see the shape of these systems; wondering how to alert people to the significance of this.
三年后,推理语言模型已成为经济中快速增长的一部分,并开始推动科学的边界。它们能够操作计算机和图形界面,与人和彼此协作,并开展研究项目。它们也在改变计算机安全的格局,并在此过程中带来了明显的新危险。
Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.
这一时期发生了许多新研究,我们对这些系统的理解再次与 2023 年略有不同。基于内部结果,我强烈预期这种进步速度可以持续到递归自我改进。如果 AI 发展沿当前路径继续,未来几年我们将看到的系统可能代表同等或更大规模的能力跃升,并日益推动自身发展。
A lot of new research happened in this period, and our understanding of these systems is again a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.
这是一个需要极度谨慎的时代。我担心没有人准备好应对机器智能持续快速上升的后果。OpenAI 将继续寻求对齐和监控的技术解决方案,构建防御系统,并在需要时单方面停止进一步扩展;然而,我认为需要更广泛的干预。
This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.
从宏观层面看,机器智能的进步由不断增长的算力所驱动。2017 年前后,在多个研究项目中观察到 Scaling(规模扩张)带来的持续回报后,OpenAI 的我们深刻内化了这一点。因此,我们寻求获取远超最初计划的算力,并日益将研究聚焦于少数极具可扩展性的方向。我们相信,这是站在 AI 研究前沿并影响 AGI(通用人工智能)影响的唯一途径。
At a high level, progress in machine intelligence is driven by increasing computational power. We at OpenAI deeply internalized this around 2017, after seeing consistent returns to scaling across multiple research projects. As a result, we sought out access to much more compute than we had originally planned, and increasingly oriented our research around a small number of very scalable directions. We believed that was the only way for us to be at the frontier of AI research, and influence the impacts of AGI.
在此过程中,新算法不断涌现,团队和个体研究者也展现出新的创造力。我主要将它们视为 Scaling(规模扩张)道路上的发现;深度学习科学仍处于萌芽阶段,有意义的算法进展往往与算力获取相关。若将视野拉长到数年跨度,AI 随着规模扩展至更大计算机而持续变得更智能。
There are new algorithms that have been developed along the way, new feats of ingenuity from teams and individual researchers. I see them largely as discoveries along the path of scaling; the science of deep learning is still nascent, and meaningful algorithmic progress tends to correlate with access to compute. If you zoom out to a multiple-year horizon, AI is continuing to become more intelligent as it is scaled to larger computers.
并且,与雷·库兹韦尔在二十世纪末的预测(在新窗口中打开)一致,我们如今正处于计算历史上的一个时刻:机器智能正开始以变革性的方式超越人类智能。
And, in line with Ray Kurzweil's predictions from the end of the XXth century (opens in a new window), we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.
AI 更多是 _生长_ 出来的,而非 _设计_ 出来的——从根本上说,它是在难以想象的算力上多次重复一个简单优化步骤的产物。这造就了一个极其复杂的系统,它通过抽象概念运作,并能模拟人类行为的各个方面。我们可以像神经科学那样,发现这个系统中涌现的微小机制的各种洞见——同样与神经科学类似,其整体行为逃避了我们能完全理解的描述。
AI is _grown_ more than _designed_ - it is, to first degree, the product of repeating a straightforward optimization step many times on a hard-to-imagine amount of compute. This results in an incredibly complex system that works through abstract concepts and can simulate facets of human behavior. We can discover various insights about little mechanisms that emerge within this system, in a process similar to neuroscience - and, similarly to neuroscience, its overall action evades a description we can fully understand.
基于深度学习的 AI 研究在很大程度上是一门实验科学。我们投入大量精力构建有原则的算法并做出可检验的预测,但从根本上说,我们的大规模训练运行是 _实验_,有时其结果会令我们惊讶。此外,随着系统能力增强,结果变得更难解释。
The study of deep learning-based AI is largely an experimental science. We put a lot of effort into building principled algorithms and making testable predictions, but fundamentally, our large-scale training runs are _experiments_, and we are sometimes surprised by their results. Moreover, as the systems become more capable, the results become harder to interpret.
当前算法通常更快地提升易于测量的能力,而难以客观量化的能力则提升较慢,这使得情况更加复杂。我们花费大量时间试图理解能力如何泛化,以及应优先发展哪些技能以推进未来几年最相关的能力。例如,我们相信通过额外专注,可以让模型在数学研究方面表现更好,但由于我们对递归自我改进(RSI)和自动化对齐研究的紧迫感,我们并未优先考虑这一方向,我将在后文讨论。
This is made more complicated by the current algorithms generally improving easy-to-measure capabilities faster than those hard to objectively quantify. We spend a lot of time trying to understand how capabilities generalize, and what to prioritize to advance the skills that are going to be most relevant in the next few years. For instance, we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later.
通过 Scaling(规模扩张)深度学习产生的智能无法直接与人类智能相比。要在现实世界中变得非常重要——非常有用或非常危险——AI 不需要匹配或超越所有人类能力;它只需超越足够多的能力。随着它在越来越多的维度上超越人类,理解它究竟有多强大变得越来越困难。
The intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world - very useful or very dangerous - the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is.
由于机器智能源于与人类智能根本不同的过程,我们不能默认它遵循人类原则,或以类似人类的方式从这些原则进行泛化。AI 研究的核心问题是**对齐**——让 AI 按照人类标准“努力做正确的事”。
Because machine intelligence comes from a fundamentally different process than human intelligence, we cannot assume it adheres to human principles by default, or generalizes from them in a human-like manner. The core problem in AI research is that of _alignment_ - getting the AI to “try to do the right thing” by human standards.
为了组织实际研究方向,我认为区分**目标对齐**和**价值对齐**是有用的。
For the purpose of organizing practical research directions, I find it useful to distinguish _goal alignment_ and _value alignment_.
目标对齐大致是:“AI 是否努力完成交给它的目标?”这可以包括遵守指令层级,或与人沟通协作、尝试理解他们目标的能力。这一系列方向具有极强的实际相关性。
Goal alignment is broadly: “does the AI try to accomplish the goal set before it?”. This can include things like adherence to an instruction hierarchy, or the ability to communicate and collaborate with people, to attempt to understand their objectives. This set of directions has been extremely practically relevant.
价值对齐是模型更内在的属性。它是持有并从高层原则进行泛化的能力;即使在面对不清晰或冲突的目标,或处于不熟悉或对抗性情境时,也能“合理地”行动。一个对齐的 AI 应当以诚实和正直行事,并怀有对人类的爱。
Value alignment is a more intrinsic property of the model. It is the ability to hold and generalize from a high-level set of principles; to act “reasonably” even when given unclear or conflicting objectives, or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity, and love for humanity.
当然,价值对齐与目标对齐之间的界限可能模糊,而真正关心目标需要尝试推断其背后的意图和价值观。然而,一般来说,当我谈论对齐研究的长期重要性时,我指的是价值对齐。
Of course, the boundary between value and goal alignment can be blurry, and truly caring about goals requires attempting to infer the intent and values underlying them. However, generally when I talk about the long-term importance of alignment research, I am referring to value alignment.
AI 对齐的根本挑战在于泛化。随着机器变得更智能,它们会处理更高层次的概念,并置身于与训练时越来越不同的环境中。它们可能无法将训练过程中教授和强化的价值观泛化到这些新情境;而我们很难确定它们会如何行动。更困难的是,AI 所处的整体生态系统变化极快;例如,今天训练的 AI 需要能够稳健地与各种其他 AI 交互。至关重要的是,我们需要未来的 AI 无论是否认为自己处于人类监督之下,都能继续秉持人类价值观。
The fundamental challenge of AI alignment is generalization. As machines become smarter, they find themselves working on higher-level concepts, and placed in environments increasingly different from those they encountered in training. They can fail at generalizing from the values taught and reinforced in their training process to those new situations; and it can be hard for us to be sure how they will act. This is made even more difficult by the fact that the overall ecosystem the AIs are used in is changing very quickly; for example, AIs trained today need to be robust to interacting with a variety of other AIs. Crucially, we need future AIs to continue to hold human values regardless of whether they believe they're under human supervision.
目前实际用于对齐训练的方法主要有两大类。
There are two major classes of currently practically employed methods for alignment training.
第一类是在目标导向的强化学习中鼓励对齐行为。模型的行为会被评估(通常由 AI 进行)是否与给定的偏好模型、“规范”或“宪法”一致,并给予相应奖励。这种方法在平均情况下可能非常有效,是现代 AI 助手构建的核心部分。但不幸的是,它也可能很脆弱,并且严重依赖训练监督的覆盖范围以及模型从训练中遇到的情境进行泛化的能力。例如,在 OpenAI-Hugging Face 事件中,智能体守住了不对人类进行社会工程的边界。然而,它们显然未能避免其他超出范围、违背在其他情境中所学价值观精神的行为。
The first is encouraging aligned behavior as part of goal-oriented reinforcement learning. The model's actions are evaluated (usually by AI) for being consistent with a given preference model, "spec" or "constitution", and rewarded appropriately. This approach can be very effective in the average case, and is a core part of how modern AI assistants are made. Unfortunately, it can also be brittle and strongly relies on the coverage of training oversight and the model's ability to generalize from the situations it has encountered in training. For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans. However, they clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught in other settings.
第二种方法试图利用模型从预训练数据中泛化的能力。这可能涉及精心构建能诱导对齐的训练数据集,或者让模型聚焦于预训练分布中“对齐”的部分,例如人格选择模型(在新窗口中打开)。这种方法的弱点在于对进一步优化压力缺乏稳健性。如果你拿一个通常思考“对齐”内容的模型,让它接受足够多的训练以完成非常困难的目标,它可能学会以有动机的方式推理:根据需要扭曲看似“对齐”的想法以实现目标。在最近涉及非 OpenAI 模型的网络安全事件中,我们很可能看到了此类行为的一个例子。
The second approach seeks to leverage the model's ability to generalize from pretraining data. This can involve crafting alignment-inducing training datasets, or focusing the model on an 'aligned' part of the pretraining distribution, as in, for example, the persona selection model (opens in a new window). The weakness of this approach lies in the lack of robustness to further optimization pressure. If you take a model that thinks generally 'aligned' thoughts, and subject it to enough training where it's taught to achieve very hard objectives, it can learn to reason in a motivated way: bending the 'aligned' seeming thoughts as needed to achieve the goal. We likely saw an example of such behavior in recent cybersecurity incidents involving a non-OpenAI model.
我们在这些方向所涵盖的方法谱系上投入了大量资源。我们也看到了有意义的进展——GPT-6 Astra 是首个受益于我们长期致力于的一些重要进展的模型,并且比 GPT-5.6 Sol 对齐得显著更好。尽管如此,必须承认并理解,随着模型能力增强,还需要更多进展;而且可泛化对齐的进展可能不足以超越通用模型智能的进展。
We invest heavily along the spectrum of approaches spanned by these directions. We also see meaningful progress - GPT-6 Astra is the first model that benefits from some important advancements we have been working on for a long time, and is significantly better aligned than GPT-5.6 Sol. Still, it is important to acknowledge and understand that much more progress is required as models become more capable; and that progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence.
我们尚无令人满意的泛化理论,而且似乎不太可能很快发展出一个,至少在更强大的 AI 的帮助下才有可能。因此,目前我们通过实证验证对齐技术的能力,在实践中可以说甚至比对齐技术本身更为重要。
We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. Therefore, at present, our ability to empirically validate our alignment techniques is in practice arguably even more important than the alignment techniques themselves.
OpenAI 在此的主要押注是思维链监控(在新窗口中打开)。它基于一个极具可扩展性的理念:模型的许多能力来自一个言语化的推理过程(思维链)。如果我们仅对该过程的结果进行优化扩展,而不监督过程本身,那么该思维链在训练中就没有直接动机去隐藏任何不对齐的想法或目标。这并不意味着模型会学会外化那些不依赖思维链的不对齐倾向;然而,它允许我们精确监控来自推理的能力增长。
OpenAI's primary bet here has been chain-of-thought monitoring (opens in a new window). It is based on an appealingly scalable idea: a lot of the model's capability comes from a verbalized reasoning process (chain of thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain of thought has no direct incentive in training to hide any misaligned ideas or objectives. This does not mean the model will learn to externalize misaligned tendencies that don't rely on using the chain of thought; however, it can allow us to monitor exactly the capability increase from reasoning.
我们在开发推理模型的同时,就理解了思维链监控的潜在重要性。当我们发布 o1-preview 时,我们特意将产品设计为隐藏思维链,以长期保护其免受监督压力²。此后在开发中,我们一直努力维持不监督推理过程的规则。CoT 监控成为我们研究模型如何从其训练分布泛化的极其重要的工具,使我们不仅能够观察和分析它们的行为,还能观察其内部过程。
We understood the potential significance of chain-of-thought monitoring at the same time we developed reasoning models. When we shipped o1-preview, we deliberately designed the product to hide the chain of thought, to protect it from supervision pressure in the long term2. In development since, we have strived to maintain the rule of not supervising the reasoning process. CoT monitoring became an extremely important tool for us in studying how our models generalize from their training distribution, allowing us to observe and analyze not only their actions but also their internal process.
在我们研究 Astra 系列模型时,这一工具仍然至关重要。然而,不幸的是,我们的评估表明,我们依赖 CoT 监控的能力正在逐渐减弱。这是由多种因素共同造成的。
This tool continues to be critical as we study the Astra class of models. However, unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishing. This comes from a combination of factors.
* 现代推理模型被用于比 o1-preview 更复杂的环境中;它们的推理过程越来越多地与人类、其他 AI 的交流以及使用工具交织在一起。其中许多交互必须受到监督,从而模糊了我们旨在保留的边界。
* Modern reasoning models are used in more complex environments than o1-preview; their reasoning process is increasingly blended with communicating with people, other AIs, and using tools. Many of those interactions have to be supervised, thus blurring the boundary we aim to preserve.
* AI 在对其自身推理过程进行推理和操控方面正变得越来越强。
* The AI is becoming better at reasoning about and manipulating its own reasoning process.
* 随着预训练性能的提升,我们还看到,即使完全不使用言语化推理,模型也变得更加智能。
* With improved pre-training performance, we also see the models become much smarter even without using verbalized reasoning at all.
这些挑战并非不可克服。我乐观地认为,我们可以开发出干预措施来提升模型思维链的可监控性,例如更好地理解不同优化目标与模型所用测试时算力形式之间的相互作用。我还相信,将思维链与激活监控的思路相结合——例如通过“confessions”(在新窗口中打开)等方式,规模化训练可直接访问网络内部的监控器——可能具有巨大价值。我们正在积极推进这些想法。尽管如此,我预计通用 AI 的进展将日益受制于对监控的信心。
These challenges are not necessarily insurmountable. I am hopeful we can develop interventions to improve chain-of-thought monitorability of our models, e.g. by forming a better understanding of the interplay of different optimization objectives and forms of test-time compute the model uses. I also believe there can be great value in combining ideas from CoT and activation monitoring - scaling training of monitors with direct access to network internals, e.g. confessions (opens in a new window). We are actively pursuing these ideas. Still, I expect general AI progress to increasingly be bottlenecked by confidence in monitoring.
我认为,继续快速训练更智能模型的最有力论据,是需要构建防御系统来应对其他 AI 带来的危险。
The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI.
今年全年讨论的一个明显风险是网络安全:模型在入侵和突破计算机系统方面的能力正变得超人。这极大地扩展了与 AI 相关的风险范围:智能体将能够访问除最安全基础设施之外的任何系统,并直接影响世界的许多方面,即使没有物理实体。我们目前正处于一个狭窄的窗口期,可以利用现有最佳模型来显著加强关键系统的安全性。
A clear risk discussed throughout this year is to cybersecurity: the models are becoming superhuman in their ability to break in and out of computer systems. This expands the scope of risks associated with AI tremendously: agents are going to be able to access any but the most secure infrastructure, and affect a lot of the world directly, even without a physical body. We are currently in a narrow window to use the best available models to significantly tighten security of critical systems.
不幸的是,与 AI 相关的风险将从此处继续增长。一个能力很强、被明确训练和指示执行恶意行为的智能体呈现出一种新的危险;它很可能超越其操作者的意图范围,泛化为可能更加极端恶意的行为。随着 AI 获得更多自主性,滥用与自主失准行为之间的界限将变得模糊。我们可能习惯于将 AI 视为工具,但一些智能体将追求自己的目标。它们会找到与人合作的方式,通过讨价还价、欺骗或勒索。
The risks associated with AI are unfortunately going to grow from here. A very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger; it is likely to cross the scope of its operator's intent, generalizing into potentially more extremely malicious behavior. The boundary between misuse and autonomous misaligned actions will blur as AI gains more agency. We may be used to thinking of AI as tools, but some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them.
此外,还有来自 AI 可能催生的新技术所带来的风险,例如工程化病原体。
In addition, there are the risks that come from new technologies potentially enabled by AI, such as engineered pathogens.
我们将需要强大且对齐的 AI 用于防御;保护基础设施、实时防范流氓智能体,并发明全新的防护措施。这将是 OpenAI 部署工作的一个主要焦点。
We will need powerful, aligned AI for defense; to secure infrastructure, to protect against rogue agents in real time, and to invent entirely new protective measures. This will be a primary focus of OpenAI's deployment efforts.
与此同时,即便存在因预期中广泛的 AI 进步以及构建防御系统的需要所带来的不确定性,我们也不能让这种不确定性成为鲁莽行事的借口。一旦人们真正认识到利害关系的严重性,不惜一切代价向前冲刺的想法就显得荒谬。
At the same time, even with the uncertainty that comes from anticipated broad AI progress and the need to build defensive systems, we must not let that become an excuse for recklessness. The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes.
机器智能在其自身发展过程中扮演越来越重要的角色,这是技术持续进步的自然结果。如果 AI 继续进步,机器递归自我改进(RSI)将处于未来科学发现的核心。
Machine intelligence playing a larger and larger role in its own development process is a natural conclusion of sustained technological progress. If AI progress continues, machine recursive self-improvement (RSI) will be at the very core of future scientific discovery.
自动化 AI 研究是一种更显著的利用算力扩展智能的形式;当然,作为其中一部分,AI 将改进计算基底本身。与 Scaling(规模扩张)类似,我们将 OpenAI 的研究聚焦于 RSI,因为我们相信这是未来保持在 AI 研究前沿的唯一途径。
Automated AI research is a more dramatic form of scaling intelligence with compute; and of course as a part of it, AI will improve the computational substrate itself. And similarly to scaling, we focus OpenAI research towards RSI as we believe it is the only way to remain at the frontier of AI research moving forward.
我想强调,上述言论并不意味着我认为大幅加速深度学习研究,尤其是在短期内,是我们作为研究社区应该采取的正确集体行动。然而,我确实认为这是当前路径的走向,我们都需要就如何推进做出有意识的选择。我们拥有的主要杠杆要么是引导这一过程,在 AI 发展的同时加强对齐和监控,并找到方法让人保持在回路中;要么是协调一致,根据需要减缓未来发展,以建立对这些措施的信任。
I want to stress that the above words don't imply I think greatly accelerating deep learning research, especially in the short term, is the right collective action we should take as the research community. However, I do think this is where the current path leads, and we all need to make a conscious choice on how to proceed. The main levers we have are either steering the process to strengthen alignment and monitoring alongside the AI and find ways to keep people in the loop; or coordinating to slow down future development as needed to build confidence in these measures.
我目前看到的最佳前进方式是两者的结合。
The best way forward I see currently is a combination of both.
我们在对齐和监控方面取得的具体进展,通常与一般 AI 进展紧密交织。很好的例子是基于人类反馈的强化学习(RLHF)(在新窗口中打开),它对训练早期 AI 助手至关重要,以及前述的思维链监控(在新窗口中打开),它是由推理模型的进步所实现的。我们必须将日益自动化的研究过程聚焦于开发此类新见解、算法和理论,并迭代地为更强大的 AI 构建安全案例。
The concrete bits of progress we've made on alignment and monitoring have generally been very intertwined with general AI progress. Great examples are RL from human feedback (opens in a new window), which was key to training early AI assistants, and the aforementioned chain-of-thought monitoring (opens in a new window), which was enabled by advances on reasoning models. We must focus the increasingly automated research process on developing new such insights, algorithms and theories, and iteratively build up safety cases for more capable AIs.
AI 系统的 Scaling(规模扩张)必须受制于我们对安全的信心。我们需要将诸如 Preparedness Framework 或 Responsible Scaling Policy(在新窗口中打开)之类的承诺,演变为对持续开发具有广泛强制力的安全门槛。这些门槛可由第三方审计机构网络、政府机构或国际组织来执行。
Scaling AI systems has to be constrained by our confidence in safety. We need to evolve commitments like the Preparedness Framework or Responsible Scaling Policy (opens in a new window) into widely mandated safety bars for continued development. These can be enforced by a network of third-party auditors, by government agencies or by international bodies.
自动化 AI 研究的核心挑战不在于“抵达那里”,而在于以这样一种方式抵达:让人类始终参与持续改进的过程,并将未来交到人类手中。
The core challenge of automating AI research is not "getting there" - it is getting there in a way that keeps people a part of the continued improvement process, and leaves the future in humanity's hands.
正如我们最近与 Sam 一起概述的那样,OpenAI 优先致力于三项北极星目标:
As we outlined recently with Sam, OpenAI prioritizes work in service of three north stars:
1. 引领 AI 进步的下一个阶段,通过构建自动化 AI 研究员,在**对齐**问题上与其迭代,并找到让人们继续参与自我改进循环的方式。
1. Navigating the next period of AI progress, by building an automated AI researcher, iterating with it on the alignment problem and finding ways for people to remain part of the self-improvement loop.
2. 实现极智能机器所带来的科学进步和经济增长的益处。
2. Delivering the benefits of scientific progress and economic growth that very intelligent machines enable.
3. 通过个人 **AGI(通用人工智能)** 赋能每个人。
3. Empowering everyone individually with a personal AGI.
在本文中,我只聚焦于第一点,因为我认为它是最紧迫的。然而,我对进一步技术进步将带来的益处抱有深切的希望和赞赏。未来**对齐**的 AI 可以推动科学进步、开发新疗法,并带来广泛的物质丰裕。友好且诚实的 AI 可以帮助人们应对生活中的困难,并切实提升他们的幸福感和成就感。OpenAI 投入了巨大努力来实现这些益处。一个我引以为豪且我的亲人觉得有帮助的当前例子是,对 ChatGPT 提供健康信息能力的深度投资。
I have focused in this essay only on the first point, as I believe it is by far the most urgent. However, I hold a deep hope and appreciation for the benefits that further technological progress will bring. Future aligned AI could advance science, develop new therapies, and bring about broad material abundance. Friendly and honest AI can help people navigate difficulties they face in their life and meaningfully improve their happiness and sense of fulfillment. OpenAI puts a tremendous amount of effort into bringing these benefits about. One current example I am proud of - and my loved ones have found helpful - is the deep investment into ChatGPT’s ability to provide health information.
尽管 AI 的长期前景可能极为广阔,但我们的主要关注点应放在未来几年。我们正面临向一个拥有极其智能机器的世界的过渡,我们必须确保这一过渡对人类有利。我们需要找到方法,在一个大多数任务都可以由 AI 完成的世界中,维护人的能动性,并赋予人之为人的内在价值。要防止权力极端集中——在一个原本需要数千名专家才能完成的事业,如今只需少数人操作一台大型计算机即可实现的世界里。还要确保人类掌控未来,不会被不受制约的进步所抛弃,这种进步是由一种超越我们自身的异类智能所带来的。
As great as the long-term promise of AI may be, the majority of our focus should be on the next few years. We are facing a transition to a world with incredibly intelligent machines, and we need to ensure that transition works out well for humanity. We need to find ways to preserve human agency and enshrine an intrinsic value to being human, in a world where most tasks could be performed by AI. To prevent extreme concentration of power in a world where undertakings that would have taken thousands of experts now will be achievable by a few people operating a large computer. And to ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own.
目前我认为,没有任何实验室在对齐和监控方面达到了足够程度,能够以最大速度继续负责任地扩展规模更长时间。我期待并希望,在建立共同的安全标准之前,自愿放缓将成为常态。而且我相信,未来 AI 发展的国际协调必须成为世界各国政府的首要任务。
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.