Seed1.5-Thinking:通过强化学习推进卓越推理模型

Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning

于琪颖 Qiying Yu · ByteDance Seed · 2025-04-10 · arXiv:2504.13914 ↗ · 被引 133

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们推出了 Seed1.5-Thinking,该模型能够在回答前进行思考推理,从而在广泛基准测试中表现出色。Seed1.5-Thinking 在 AIME 2024 上达到 86.7,Codeforces 上 55.0,GPQA 上 77.3,展现了在 STEM 和编程领域的卓越推理能力。除推理任务外,该方法在多个领域展现出显著的泛化能力。例如,在非推理任务上,其胜率比 DeepSeek R1 高出 8%,表明其更广泛的适用性。与其他最先进的推理模型相比,Seed1.5-Thinking 是一个相对较小的混合专家(MoE)模型,拥有 200 亿激活参数和 2000 亿总参数。作为评估泛化推理能力的一部分,我们开发了两个内部基准测试 BeyondAIME 和 Codeforces,这两个基准测试将公开发布以支持未来研究。模型试用链接:https://www.volcengine.com/experience/ark。

We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 86.7 on AIME 2024, 55.0 on Codeforces and 77.3 on GPQA, demonstrating excellent reasoning abilities in STEM and coding. Beyond reasoning tasks, the method demonstrates notable generalization across diverse domains. For instance, it surpasses DeepSeek R1 by 8% in win rate on non-reasoning tasks, indicating its broader applicability. Compared to other state-of-the-art reasoning models, Seed1.5-Thinking is a Mixture-of-Experts (MoE) model with a relatively small size, featuring 20B activated and 200B total parameters. As part of our effort to assess generalized reasoning, we develop two internal benchmarks, BeyondAIME and Codeforces, both of which will be publicly released to support future research. Model trial link: https://www.volcengine.com/experience/ark.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 20)

阅读逐段中英对照全文 →