控制大语言模型中的推理努力程度

Controlling Reasoning Effort in LLMs

塞巴斯蒂安·拉施卡 Sebastian Raschka · · 2026-07-18 · Ahead of AI ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

自 OpenAI 发布 o1 以来,已经过去了近两年,o1 是一个普及了基于 LLM 的推理模型理念的模型。大约四个月后,DeepSeek-R1 紧随其后,并提供了使用可验证奖励的强化学习(RLVR)配方来训练此类推理模型的细节。上周,OpenAI 发布了 GPT-5.6 模型系列。该系列包含三种尺寸,每种尺寸大约有五六种推理努力设置。图 1:GPT 5.6 Sol 模型在不同推理努力设置下的表现。(Ultra 的基准数字目前尚不可用,但应与 Max 相对相似,因为它使用类似的努力水平,但通过四个子代理加速工作。)

It has been almost two years since OpenAI released o1, a model that popularized the idea of LLM-based reasoning models. DeepSeek-R1 followed about four months later, together with details of a reinforcement learning with verifiable rewards (RLVR) recipe to train such reasoning models. Last week, OpenAI released the GPT-5.6 model family. It comes in three sizes, each with roughly five or six reasoning-effort settings. Figure 1: The GPT 5.6 Sol model with different reasoning effort settings. (Benchmark numbers for Ultra are currently not available but should be relatively similar to Max, since it uses a similar effort level but accelerates the work with four subagents.)

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 28)

阅读逐段中英对照全文 →