几十年来,人工智能主要致力于开发新的训练方法和模型。这确实奏效了:从击败国际象棋和围棋世界冠军,在 SAT 和律师资格考试中超越大多数人类,到获得 IMO 和 IOI 金牌。这些历史里程碑——深蓝、AlphaGo、GPT-4 和 o 系列——背后是 AI 方法的根本创新:搜索、深度强化学习、扩展和推理。事情随着时间的推移越来越好。用三个词概括:强化学习终于奏效了。更准确地说:强化学习终于泛化了。在经历了几次重大曲折和一系列里程碑的积累之后,我们找到了一种有效的配方,可以利用语言和推理解决广泛的强化学习任务。即使是一年前,如果你告诉大多数 AI 研究人员,一个单一的配方可以解决软件工程、创意写作、IMO 级别的数学、鼠标键盘操作和长形式问答——他们会嘲笑你的幻觉。这些任务中的每一个都极其困难,许多研究人员花费整个博士生涯只专注于其中一个狭窄的领域。那么接下来呢?AI 的下半场——从现在开始——将从解决问题转向定义问题。在这个新时代,评估变得更加重要。
For decades, AI has largely been about developing new training methods and models. And it worked: from beating world champions at chess and Go, surpassing most humans on the SAT and bar exams, to earning IMO and IOI gold medals. Behind these milestones in the history book — DeepBlue, AlphaGo, GPT-4, and the o-series — are fundamental innovations in AI methods: search, deep RL, scaling, and reasoning. Things just get better over time. In three words: RL finally works. More precisely: RL finally generalizes. After several major detours and a culmination of milestones, we’ve landed on a working recipe to solve a wide range of RL tasks using language and reasoning. Even a year ago, if you told most AI researchers that a single recipe could tackle software engineering, creative writing, IMO-level math, mouse-and-keyboard manipulation, and long-form question answering — they’d laugh at your hallucinations. Each of these tasks is incredibly difficult and many researchers spend their entire PhDs focused on just one narrow slice. So what comes next? The second half of AI — starting now — will shift focus from solving problems to defining problems. In this new era, evaluation becomes more im
核心贡献 · Key contributions
识别出 AI 从方法驱动转向问题定义驱动的转变,称为‘下半场’。 Identifies a shift from method-driven AI to problem-definition-driven AI, termed the 'second half'.
论证强化学习因语言预训练先验和推理作为动作而最终实现泛化。 Argues that RL finally generalizes due to language pre-training priors and reasoning as actions.
提出在新时代评估比训练更重要。 Proposes that evaluation becomes more important than training in the new era.
强调‘效用问题’:AI 在基准测试中表现出色,但对现实经济影响有限。 Highlights the 'utility problem': AI excels at benchmarks but has limited real-world economic impact.
呼吁重新思考评估设置,挑战如独立同分布和自主性等假设。 Calls for rethinking evaluation setups, challenging assumptions like i.i.d. and autonomy.
指出通用配方(预训练、Scaling、推理)将主导增量方法。 Suggests that the general recipe (pre-training, scaling, reasoning) will dominate incremental methods.
局限 · Limitations
‘下半场’论点具有推测性,缺乏实证验证。 The 'second half' thesis is speculative and lacks empirical validation.
声称强化学习算法微不足道可能忽视新领域所需的算法创新。 The claim that RL algorithm is trivial may overlook algorithmic innovations needed for new domains.
效用问题可能是暂时的;AI 的经济影响可能随部署快速增长。 The utility problem may be temporary; AI's economic impact could grow rapidly with deployment.
配方的通用性限于预训练数据分布内的任务。 The recipe's generality is limited to tasks within distribution of pre-training data.
重新思考评估可能因机构惯性和缺乏明确指标而面临阻力。 Rethinking evaluation may face resistance due to institutional inertia and lack of clear metrics.
论文章节 · Sections(共 4)
Shunyu Yao (https://ysymyth.github.io/)Shunyu Yao(https://ysymyth.github.io/)