下半场

The Second Half

姚顺雨 Shunyu Yao · OpenAI · 2025-04-01 · Shunyu Yao Blog ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

几十年来,人工智能主要致力于开发新的训练方法和模型。这确实奏效了:从击败国际象棋和围棋世界冠军,在 SAT 和律师资格考试中超越大多数人类,到获得 IMO 和 IOI 金牌。这些历史里程碑——深蓝、AlphaGo、GPT-4 和 o 系列——背后是 AI 方法的根本创新:搜索、深度强化学习、扩展和推理。事情随着时间的推移越来越好。用三个词概括:强化学习终于奏效了。更准确地说:强化学习终于泛化了。在经历了几次重大曲折和一系列里程碑的积累之后,我们找到了一种有效的配方,可以利用语言和推理解决广泛的强化学习任务。即使是一年前,如果你告诉大多数 AI 研究人员,一个单一的配方可以解决软件工程、创意写作、IMO 级别的数学、鼠标键盘操作和长形式问答——他们会嘲笑你的幻觉。这些任务中的每一个都极其困难,许多研究人员花费整个博士生涯只专注于其中一个狭窄的领域。那么接下来呢?AI 的下半场——从现在开始——将从解决问题转向定义问题。在这个新时代,评估变得更加重要。

For decades, AI has largely been about developing new training methods and models. And it worked: from beating world champions at chess and Go, surpassing most humans on the SAT and bar exams, to earning IMO and IOI gold medals. Behind these milestones in the history book — DeepBlue, AlphaGo, GPT-4, and the o-series — are fundamental innovations in AI methods: search, deep RL, scaling, and reasoning. Things just get better over time. In three words: RL finally works. More precisely: RL finally generalizes. After several major detours and a culmination of milestones, we’ve landed on a working recipe to solve a wide range of RL tasks using language and reasoning. Even a year ago, if you told most AI researchers that a single recipe could tackle software engineering, creative writing, IMO-level math, mouse-and-keyboard manipulation, and long-form question answering — they’d laugh at your hallucinations. Each of these tasks is incredibly difficult and many researchers spend their entire PhDs focused on just one narrow slice. So what comes next? The second half of AI — starting now — will shift focus from solving problems to defining problems. In this new era, evaluation becomes more im

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 4)

阅读逐段中英对照全文 →