学习用大语言模型进行推理

Learning to reason with LLMs

OpenAI OpenAI · OpenAI · 2024-09-12 · OpenAI Research ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

贡献:OpenAI o1 在竞争性编程问题(Codeforces)中排名第 89 百分位,在美国数学奥林匹克(AIME)预选赛中位列全美前 500 名学生,并在物理、生物和化学问题基准(GPQA)上超越了人类博士水平的准确率。尽管使这个新模型像当前模型一样易于使用的工作仍在进行中,但我们发布了该模型的早期版本 OpenAI o1-preview,供 ChatGPT 和受信任的 API 用户立即使用。我们的大规模强化学习算法通过高度数据高效的训练过程,教会模型如何利用思维链进行高效思考。我们发现,o1 的性能随着更多的强化学习(训练时计算)和更多的思考时间(测试时计算)而持续提升。扩展这种方法的约束与 LLM 预训练有很大不同,我们正在继续研究它们。

ContributionsUse o1(opens in a new window) OpenAI o1 ranks in the 89th percentile on competitive programming questions (Codeforces), places among the top 500 students in the US in a qualifier for the USA Math Olympiad (AIME), and exceeds human PhD-level accuracy on a benchmark of physics, biology, and chemistry problems (GPQA). While the work needed to make this new model as easy to use as current models is still ongoing, we are releasing an early version of this model, OpenAI o1‑preview, for immediate use in ChatGPT and to trusted API users⁠(opens in a new window). Our large-scale reinforcement learning algorithm teaches the model how to think productively using its chain of thought in a highly data-efficient training process. We have found that the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute). The constraints on scaling this approach differ substantially from those of LLM pretraining, and we are continuing to investigate them.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 12)

全文 · Full text(逐段中英对照)

概述 Overview

贡献使用 o1(在新窗口中打开)

ContributionsUse o1(opens in a new window)

OpenAI o1 在竞争性编程问题(Codeforces)中排名第 89 百分位,在美国数学奥林匹克(AIME)资格赛中跻身前 500 名学生之列,并在物理、生物和化学问题基准测试(GPQA)上超过了人类博士水平的准确率。尽管使这个新模型像当前模型一样易于使用的工作仍在进行中,但我们发布了该模型的早期版本 OpenAI o1-preview,供 ChatGPT 和受信任的 API 用户立即使用(在新窗口中打开)。

OpenAI o1 ranks in the 89th percentile on competitive programming questions (Codeforces), places among the top 500 students in the US in a qualifier for the USA Math Olympiad (AIME), and exceeds human PhD-level accuracy on a benchmark of physics, biology, and chemistry problems (GPQA). While the work needed to make this new model as easy to use as current models is still ongoing, we are releasing an early version of this model, OpenAI o1‑preview, for immediate use in ChatGPT and to trusted API users⁠(opens in a new window).

我们的大规模强化学习算法教会模型如何在高数据效率的训练过程中利用其思维链进行高效思考。我们发现,o1 的性能随着更多的强化学习(训练时算力)和更多的思考时间(推理时算力)而持续提升。扩展这种方法的约束与 LLM 预训练有很大不同,我们正在继续研究它们。

Our large-scale reinforcement learning algorithm teaches the model how to think productively using its chain of thought in a highly data-efficient training process. We have found that the performance of o1 consistently improves with more reinforcement learning (train-time compute) and with more time spent thinking (test-time compute). The constraints on scaling this approach differ substantially from those of LLM pretraining, and we are continuing to investigate them.

o1 的性能随着训练时和推理时算力的增加而平滑提升。

o1 performance smoothly improves with both train-time and test-time compute

评估 Evals

为了突出相对于 GPT‑4o 的推理改进,我们在多样化的人类考试和机器学习基准上测试了我们的模型。我们表明,在绝大多数这些推理密集型任务上,o1 显著优于 GPT‑4o。除非另有说明,我们在最大测试时算力设置下评估了 o1。

To highlight the reasoning improvement over GPT‑4o, we tested our models on a diverse set of human exams and ML benchmarks. We show that o1 significantly outperforms GPT‑4o on the vast majority of these reasoning-heavy tasks. Unless otherwise specified, we evaluated o1 on the maximal test-time compute setting.

o1 在具有挑战性的推理基准上大幅改进 GPT-4o。实心条表示 pass@1 准确率,阴影区域表示 64 个样本的多数投票(共识)性能。

o1 greatly improves over GPT-4o on challenging reasoning benchmarks. Solid bars show pass@1 accuracy and the shaded region shows the performance of majority vote (consensus) with 64 samples.

o1 在具有挑战性的推理基准上大幅改进 GPT-4o。实心条表示 pass@1 准确率,阴影区域表示 64 个样本的多数投票(共识)性能。

o1 greatly improves over GPT-4o on challenging reasoning benchmarks. Solid bars show pass@1 accuracy and the shaded region shows the performance of majority vote (consensus) with 64 samples.

o1 在广泛的基准上改进 GPT-4o,包括 54/57 个 MMLU 子类别。图中展示了七个作为示例。

o1 improves over GPT-4o on a wide range of benchmarks, including 54/57 MMLU subcategories. Seven are shown for illustration.

o1 在广泛的基准上改进 GPT-4o,包括 54/57 个 MMLU 子类别。图中展示了七个作为示例。

o1 improves over GPT-4o on a wide range of benchmarks, including 54/57 MMLU subcategories. Seven are shown for illustration.

在许多推理密集型基准上,o1 与人类专家水平相当。最近的前沿模型在 MATH2 和 GSM8K 上表现如此出色,以至于这些基准不再能有效区分模型。我们在 AIME(一项旨在挑战美国最聪明高中数学学生的考试)上评估了数学性能。在 2024 年 AIME 考试中,GPT‑4o 平均仅解决了 12%(1.8/15)的问题。o1 在每问题单样本下平均达到 74%(11.1/15),在 64 个样本的共识下达到 83%(12.5/15),在使用学习评分函数对 1000 个样本重新排序时达到 93%(13.9/15)。13.9 的分数使其位列全国前 500 名学生,并超过美国数学奥林匹克竞赛的分数线。

In many reasoning-heavy benchmarks, o1 rivals the performance of human experts. Recent frontier models1 do so well on MATH2and GSM8K that these benchmarks are no longer effective at differentiating models. We evaluated math performance on AIME, an exam designed to challenge the brightest high school math students in America. On the 2024 AIME exams, GPT‑4o only solved on average 12% (1.8/15) of problems. o1 averaged 74% (11.1/15) with a single sample per problem, 83% (12.5/15) with consensus among 64 samples, and 93% (13.9/15) when re-ranking 1000 samples with a learned scoring function. A score of 13.9 places it among the top 500 students nationally and above the cutoff for the USA Mathematical Olympiad.

我们还评估了 o1 在 GPQA diamond 上的表现,这是一个测试化学、物理和生物学专业知识的困难智能基准。为了将模型与人类进行比较,我们招募了拥有博士学位的专家来回答 GPQA-diamond 问题。我们发现 o1 超越了这些人类专家的表现,成为第一个在该基准上做到这一点的模型。这些结果并不意味着 o1 在所有方面都比博士更有能力——只是该模型在解决一些博士应能解决的问题上更熟练。在其他几个机器学习基准上,o1 改进了现有最佳水平。凭借其视觉感知能力,o1 在 MMMU 上获得了 78.2% 的分数,成为第一个与人类专家竞争的模型。它还在 57 个 MMLU 子类别中的 54 个上优于 GPT‑4o。

We also evaluated o1 on GPQA diamond, a difficult intelligence benchmark which tests for expertise in chemistry, physics and biology. In order to compare models to humans, we recruited experts with PhDs to answer GPQA-diamond questions. We found that o1 surpassed the performance of those human experts, becoming the first model to do so on this benchmark. These results do not imply that o1 is more capable than a PhD in all respects — only that the model is more proficient in solving some problems that a PhD would be expected to solve. On several other ML benchmarks, o1 improved over the state-of-the-art. With its vision perception capabilities enabled, o1 scored 78.2% on MMMU, making it the first model to be competitive with human experts. It also outperformed GPT‑4o on 54 out of 57 MMLU subcategories.

思维链 Chain of Thought

类似于人类在回答难题前可能会长时间思考,o1 在尝试解决问题时会使用思维链。通过强化学习,o1 学会磨练其思维链并完善所使用的策略。它学会识别并纠正自己的错误。它学会将棘手的步骤分解为更简单的步骤。它学会在当前方法无效时尝试不同的方法。这一过程极大地提升了模型的推理能力。为了说明这一飞跃,我们下面展示 o1‑preview 在几个难题上的思维链。

Similar to how a human may think for a long time before responding to a difficult question, o1 uses a chain of thought when attempting to solve a problem. Through reinforcement learning, o1 learns to hone its chain of thought and refine the strategies it uses. It learns to recognize and correct its mistakes. It learns to break down tricky steps into simpler ones. It learns to try a different approach when the current one isn’t working. This process dramatically improves the model’s ability to reason. To illustrate this leap forward, we showcase the chain of thought from o1‑preview on several difficult problems below.

密码 编码 数学 填字 英语 科学 安全 健康 科学

Cipher Coding Math Crossword English Science Safety Health Science

GPT-4o GPT-4o

oyfjdnisdr rtqwainr acxz mynzbhhx -> 逐步思考

oyfjdnisdr rtqwainr acxz mynzbhhx -> Think step by step

oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz

oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz

让我们根据示例逐步分解:

Let’s break this down step by step based on the example:

1. 示例给出:• 输入:oyfjdnisdr rtqwainr acxz mynzbhhx

1. Example given:• Input: oyfjdnisdr rtqwainr acxz mynzbhhx

• 模式涉及选择特定字母或对其进行转换。

• The pattern involves selecting specific letters or transforming them.

2. 现在,解码新短语:• 输入:oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz

2. Now, let’s decode the new phrase:• Input: oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz

第一个短语可能遵循与示例相同的结构。输入文本似乎可以根据自然分隔或模式分成若干组。

The first phrase might follow the same structure as the example. It seems like the input text could be broken into groups based on natural separations or patterns.

然而,要进行有意义的解码,我需要更多关于可能涉及的转换或字母移位的信息。

However, to make a meaningful decoding, I would need more context about the transformations or letter shifting that might be involved.

您能否提供此密码中使用的任何额外解码规则或转换?

Could you provide any additional decoding rules or transformations used in this cipher?

解码示例 OpenAI o1-preview

oyfjdnisdr rtqwainr acxz mynzbhhx -> Think step by step

oyfjdnisdr rtqwainr acxz mynzbhhx -> Think step by step

oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz

oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz

"oyfjdnisdr rtqwainr acxz mynzbhhx" 被转换为 "Think step by step"

"oyfjdnisdr rtqwainr acxz mynzbhhx -> Think step by step"

我们可以看到,当 "oyfjdnisdr rtqwainr acxz mynzbhhx" 被转换为 "Think step by step" 时

We can see that when "oyfjdnisdr rtqwainr acxz mynzbhhx" is transformed to "Think step by step"

我们的任务是:使用上面的例子解码:"oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz"

Our task is: Use the example above to decode: "oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz"

所以第一部分是弄清楚 "oyfjdnisdr rtqwainr acxz mynzbhhx" 是如何被解码为 "Think step by step" 的

So the first part is to figure out how "oyfjdnisdr rtqwainr acxz mynzbhhx" is decoded into "Think step by step"

观察 1:短语 "Think step by step" 有 18 个字母(不包括空格)。让我们数一下字母。

Observation 1: The phrase "Think step by step" is 18 letters (excluding spaces) . Let's count the letters.

类似地,让我们数一下 "oyfjdnisdr rtqwainr acxz mynzbhhx" 中的字母

Similarly, let's count the letters in "oyfjdnisdr rtqwainr acxz mynzbhhx"

密文字母:o y f j d n i s d r r t q w a i n r a c x z m y n z b h h x

Ciphertext letters: o y f j d n i s d r r t q w a i n r a c x z m y n z b h h x

总字母数:让我们数一下密文中的字母。

Total letters: Let's count the letters in the ciphertext.

密文总字母数:10 + 8 + 4 + 8 = 30 个字母

Total letters in ciphertext: 10 + 8 + 4 + 8 = 30 letters

明文总字母数:5 + 4 + 2 + 4 = 15 个字母

Total letters in plaintext: 5 + 4 + 2 + 4 = 15 letters

但实际上在问题中它说这个例子:

But actually in the problem it says the example:

那么解码过程是密文映射到 "Think step by step" 吗?

So is the decoding process that the ciphertext maps to "Think step by step"?

我们的目标是使用上面的例子解码:

Our goal is to use the example above to decode:

"oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz"

"oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz"

首先,密码似乎与明文有关。

First, the cipher seems connected to the plaintext.

考虑到时间限制,也许最简单的方法是尝试寻找模式。

Given the time constraints, perhaps the easiest way is to try to see patterns.

选项 1:尝试找到字母到字母的映射。

Option 1: Try to find mapping from letters to letters.

首先,让我们将密文字母和明文字母上下对齐写出来。

First, let's write down the ciphertext and plaintext letters on top of each other.

密文:o y f j d n i s d r r t q w a i n r a c x z m y n z b h h x

Ciphertext: o y f j d n i s d r r t q w a i n r a c x z m y n z b h h x

我想也许这里有一个字谜或替换密码。

I think maybe there is an anagram or substitution cipher here.

或者,我认为我们可以注意到每个密文组对应一个明文单词。

Alternatively, I think that we can notice that each group of ciphertext corresponds to a plaintext word.

第一个密文单词:oyfjdnisdr(10 个字母)

First ciphertext word: oyfjdnisdr (10 letters)

第二个密文单词:rtqwainr(8 个字母)

Second ciphertext word: rtqwainr (8 letters)

第四个密文单词:mynzbhhx(8 个字母)

Fourth ciphertext word: mynzbhhx (8 letters)

看起来密文单词的长度正好是明文单词的两倍。

It seems that the ciphertext words are exactly twice as long as the plaintext words.

想法:也许我们需要每隔一个字母取一次,或者相应地重建明文。

Idea: Maybe we need to take every other letter or rebuild the plaintext from the ciphertext accordingly.

如果我们从密文单词中每隔一个字母取一次,也许我们会得到明文。

If we take every other letter from the ciphertext words, maybe we'll get the plaintext.

密文单词:o y f j d n i s d r(10 个字母)

Ciphertext word: o y f j d n i s d r (10 letters)

第一个字母(位置 1,3,5,7,9):o f d i d

First letters (positions 1,3,5,7,9): o f d i d

第二个字母(位置 2,4,6,8,10):y j n s r

Second letters (positions 2,4,6,8,10): y j n s r

等等:让我们尝试每隔一个字母取一次或取第一个字母。

Wait: Let's try taking every second letter or first letter.

明文单词是 "Think",字母是 T h i n k

Plaintext word is "Think", letters are T h i n k

选项 2:尝试按照指定的代码进行映射:也许是字母的列?

Option 2: Try mapping as per an assigned code: perhaps columns of letters?

或者,也许密码更复杂。

Alternatively, perhaps the cipher is more complex.

或者,注意到 "oyfjdnisdr" 有 10 个字母,而 "Think" 有 5 个字母。

Alternatively, notice that "oyfjdnisdr" has 10 letters and "Think" has 5 letters.

如果每个明文字母由两个密文字母表示呢?

What if each plaintext letter is being represented by two ciphertext letters?

假设每个明文字母对应密文中的一对字母。

Suppose that each plaintext letter corresponds to pair of letters in the ciphertext.

这意味着密文的前两个字母映射到第一个明文字母,依此类推。

Meaning that the first two letters of ciphertext map to the first plaintext letter, and so on.

所以对于第一个单词,将密文字母分成对:

So for the first word, partition the ciphertext letters into pairs:

对应明文字母 T h i n k

Corresponding to plaintext letters T h i n k

如果这是映射,也许每对字母映射到一个明文字母。

If this is the mapping, perhaps each pair of letters maps to one plaintext letter.

是否存在从对到字母的一致映射?

Is there a consistent mapping from pairs to letters?

让我们制作一个对和对应字母的表格。

Let's make a table of pairs and corresponding letters.

首先,从第一个单词中制作对及其对应的明文字母:

First, make the pairs from the first word and their corresponding plaintext letters:

现在,"oy" 到 'T' 的映射是什么?

Now, what's the mapping from "oy" to 'T'?

一种可能的方式是 'oy' 代表 'T'。

One possible way is that 'oy' represents 'T'.

将字母转换为数字(A=1, B=2, ..., Z=26)

Convert letters to numbers (A=1, B=2, ..., Z=26)

所以让我们尝试从 'o' 和 'y' 得到 'T'

So let's try to get to 'T' from 'o' and 'y'

或者,也许相减:25 -15 = 10。

Alternatively, perhaps subtract: 25 -15 = 10.

或者,也许以某种方式组合数字。

Alternatively, perhaps combine the numbers in some way.

或者,考虑它们在字母表中的位置。

Alternatively, think about their positions in the alphabet.

或者,也许字母是通过代码加密的。

Alternatively, perhaps the letters are encrypted via a code.

或者,也许如果我们把单词 'Think' 覆盖在密码对 'oy', 'fj' 等上,密码是通过用两个字母替换每个明文字母形成的。

Alternatively, perhaps if we overlay the word 'Think' over the cipher pairs 'oy', 'fj', etc., the cipher is formed by substituting each plaintext letter with two letters.

或者,也许考虑 '原始' 字母。

Alternatively, perhaps consider the 'original' letters.

让我们看看单词 "Think" 和 "oyfjdnisdr" 之间的字母

Let's look at the letters between the word "Think" and "oyfjdnisdr"

也许我们可以发现 "oyfjdnisdr" 是 "Thinkfiles" 的字谜?

Perhaps we can spot that "oyfjdnisdr" is an anagram for "Thinkfiles"?

假设我们给字母分配数字:a=1, b=2, ..., z=26

Suppose we assign numbers to letters: a=1, b=2, ..., z=26

如果我们取密码对的和模 26 呢?

What if we take the sum of the cipher pair mod 26.

或者,看看密码对是否对应其他东西。

Alternatively, see if the cipher pairs correspond to something else.

或者,也许对 'oy' 通过密钥对应 'T'。

Alternatively, perhaps the pair 'oy' corresponds to 'T' via key.

或者,也许就像 'oy' 是 'T' 一样简单。

Alternatively, perhaps it's as simple as 'oy' is 'T'.

但 'o' 对应 'T',从 'T' 的 ASCII 码减去 'o' 的 ASCII 码。

But 'o' corresponds to 'T', subtracting 'o' ASCII code from 'T'.

现在,让我们从对称字母的角度思考。

Now, let's think in terms of symmetric letters.

或者,也许我们可以发现这些字母对应单词的首字母。

Alternatively, perhaps we can spot that the letters correspond to the first letters of words.

但也许更简单的方法是看单词的位置。

But perhaps an easier way is to look at the word positions.

或者,也许我们可以看到密文字母是明文字母的偏移量。

Alternatively, perhaps we can see that the ciphertext letters are offsets from the plaintext letters.

所以让我们尝试将这对映射到明文字母。

So let's try to map the pairs to the plaintext letters.

但 'y' 是 25,比 'T'(20+5=25)大 5 个字母

But 'y' is 25, which is 5 letters after 'T' (20 +5=25)

类似地,'y' +1 = 'Z',不是 'T'。

Similarly, 'y' +1 = 'Z', which is not 'T'.

或者,也许两个字母的平均值对应 'T'

Alternatively, maybe the average of the two letters corresponds to 'T'

所以也许这对中字母的平均值对应明文字母。

So perhaps the average of the letters in the pair corresponds to the plaintext letter.

所以也许我们需要仔细考虑字母。

So perhaps we need to think carefully about letters.

所以我们的代码是:对于每一对,将它们的数值相加,除以 2,得到对应的字母。

So our code is: For each pair, sum their numeric values, divide by 2, get the corresponding letter.

或者,字母数值的平均值对应明文字母。

Alternatively, the average of the numeric values of the letters corresponds to the plaintext letter.

现在让我们用其他单词测试这个。

Now let's test this with the other words.

对应明文字母 's', 't', 'e', 'p'

Corresponding to plaintext letters 's', 't', 'e', 'p'

现在,有了这个,我们可以解码给定的消息。

Now, given this, we can decode the given message.

"oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz"

"oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz"

鉴于例子中的密文单词长度为:

Given that the ciphertext in the example had words of length:

所以也许在我们提供的密文中,我们需要相应地将字母分成单词。

So perhaps in our supplied ciphertext, we need to divide the letters into words accordingly.

密文:o y e k a i j z d f a a p t c g s u a o k y b h a i o u o w a q h t m y n z n v a a t z a c d f o u l x x z

Ciphertext: o y e k a i j z d f a a p t c g s u a o k y b h a i o u o w a q h t m y n z n v a a t z a c d f o u l x x z

1) oyekaijzdf(10 个字母)-> 5 个字母的单词

1) oyekaijzdf (10 letters) -> word of 5 letters

2) aaptcg(6 个字母)-> 3 个字母的单词

2) aaptcg (6 letters) -> word of 3 letters

3) suaokybhai(10 个字母)-> 5 个字母的单词

3) suaokybhai (10 letters) -> word of 5 letters

6) mynznvaatzacdfoulxxz(22 个字母)-> 11 个字母的单词

6) mynznvaatzacdfoulxxz (22 letters) -> word of 11 letters

第一个单词:oyekaijzdf(10 个字母 = 5 对)

First word: oyekaijzdf (10 letters = 5 pairs)

第三个单词:suaokybhai(10 个字母=5 对)

Third word: suaokybhai (10 letters=5 pairs)

第六个单词:mynznvaatzacdfoulxxz(22 个字母=11 对)

Sixth word: mynznvaatzacdfoulxxz (22 letters=11 pairs)

所以我们可以假设明文单词的长度为:

So we can suppose that the plaintext words have lengths:

所以我们可以通过平均每对字母的数值来找到明文字母。

So we can find the plaintext letters by averaging the numeric values of the letters in each pair.

第六个单词:mynznvaatzacdfoulxxz(22 个字母:11 对)

Sixth word: mynznvaatzacdfoulxxz (22 letters: 11 pairs)

等等,之前我漏掉了一些字母。

Wait, earlier I missed some letters there.

但总共有 22 个字母;我想我有个不匹配。

But in total, we have 22 letters; I think I have a mismatch.

或者,也许最后的 'z' 是多余的字母,或者它与前面的 'x' 配对。

Alternatively, perhaps the 'z' at the end is an extra letter or it's paired with the previous 'x'.

或者,也许有错误。

Alternatively, perhaps there is a mistake.

或者,也许当我们到达最后的 'x x z' 时,我们应该写成:

Alternatively, perhaps when we reach the last 'x x z', we should write:

现在总共有 10 对。但我们应该有 11 个明文字母。

Now total pairs: 10 pairs. But we are supposed to have 11 letters in the plaintext word.

所以也许我数错了,没有考虑包含 'f' 或 'z'。

So perhaps I have miscounted and not considered including 'f' or 'z'.

等等,有没有可能序列是:

Wait, is there any chance that the sequence is:

也许 'x x' 和 'z' 是重叠的。

Perhaps the 'x x' and 'z' are overlapping.

我认为要使其工作,我们需要从 22 个字母中创建 11 对。

I think for this to work, we need to create 11 pairs from the 22 letters.

所以让我们尝试按顺序配对字母,每次取两个字母。

So let's try pairing letters in order, taking two letters at a time.

但同样,最后一对只有一个字母 'z'

But again, last pair only has one letter 'z'

或者,我们可能需要调整对。

Alternatively, we might need to adjust the pairs.

也许是 'ou', 'lx', 'xz',并且没有多余的 'x'。

Perhaps it's 'ou', 'lx', 'xz', and there is no extra 'x'.

等等,让我们按顺序写出字母:

Wait, let's write the letters as they come:

总共 10 对,将对应 10 个字母。

Total of 10 pairs, which will correspond to 10 letters.

但之前我们认为单词是 11 个字母。

But earlier we thought the word was 11 letters.

也许我们应该继续使用这 10 对。

Maybe we should proceed with the 10 pairs.

所以第十个单词是 'S', 'T', 'R', 'A', 'W', 'B', 'E', 'R', 'R', 'Y'

So the tenth word is 'S', 'T', 'R', 'A', 'W', 'B', 'E', 'R', 'R', 'Y'

所以第六个单词是 'STRAWBERRY',这说得通。

So the sixth word is 'STRAWBERRY', which makes sense.

所以将所有翻译后的字母放在一起:

So putting all these translated letters together:

为了使用提供的例子解码给定的密文,我们将使用例子中识别出的模式。

To decode the given ciphertext using the example provided, we’ll use the pattern identified in the example.

明文 Plain Text

2. 将每个字母转换为其在字母表中的数字位置(A=1,B=2,…,Z=26)。

2. Convert each letter to its numerical position in the alphabet (A=1, B=2, …, Z=26).

3. 计算每对字母的数值之和。

3. Sum the numerical values of each pair.

4. 计算和的平均值(除以 2)。

4. Compute the average of the sum (divide by 2).

5. 将平均值转换回字母(1=A,2=B,…,26=Z)。

5. Convert the average back to a letter (1=A, 2=B, …, 26=Z).

将此方法应用于给定的密文:

Apply this method to the given ciphertext:

纯文本 Plain Text

1oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz

1oyekaijzdf aaptcg suaokybhai ouow aqht mynznvaatzacdfoulxxz

1. 配对:my, nz, nv, aa, tz, ac, df, ou, lx, xz

1. Pairs: my, nz, nv, aa, tz, ac, df, ou, lx, xz

编程 Coding

我们训练了一个模型,该模型在 2024 年国际信息学奥林匹克竞赛(IOI)中获得了 213 分,排名第 49 百分位。该模型从 o1 初始化,并通过训练进一步提高编程技能。该模型在 2024 年 IOI 中与人类参赛者相同的条件下进行比赛。它有十个小时的时间来解决六个具有挑战性的算法问题,并且每个问题允许提交 50 次。

We trained a model that scored 213 points and ranked in the 49th percentile in the 2024 International Olympiad in Informatics (IOI), by initializing from o1 and training to further improve programming skills. This model competed in the 2024 IOI under the same conditions as the human contestants. It had ten hours to solve six challenging algorithmic problems and was allowed 50 submissions per problem.

对于每个问题,我们的系统采样了许多候选提交,并根据测试时选择策略提交了其中的 50 个。提交的选择基于 IOI 公开测试用例、模型生成的测试用例以及一个学习到的评分函数的表现。如果我们随机提交,平均只能得到 156 分,这表明在比赛约束下,该策略价值近 60 分。

For each problem, our system sampled many candidate submissions and submitted 50 of them based on a test-time selection strategy. Submissions were selected based on performance on the IOI public test cases, model-generated test cases, and a learned scoring function. If we had instead submitted at random, we would have only scored 156 points on average, suggesting that this strategy was worth nearly 60 points under competition constraints.

在放宽提交限制的情况下,我们发现模型性能显著提升。当每个问题允许提交 10,000 次时,即使没有测试时选择策略,模型也获得了 362.14 分——超过了金牌门槛。

With a relaxed submission constraint, we found that model performance improved significantly. When allowed 10,000 submissions per problem, the model achieved a score of 362.14 – above the gold medal threshold – even without any test-time selection strategy.

最后,我们模拟了 Codeforces 举办的编程竞赛,以展示该模型的编程技能。我们的评估严格遵循比赛规则,并允许提交 10 次。GPT-4o 的 Elo 评级为 808,处于人类参赛者的第 11 百分位。该模型远远超过了 GPT-4o 和 o1——其 Elo 评级达到 1807,表现优于 93%的参赛者。

Finally, we simulated competitive programming contests hosted by Codeforces to demonstrate this model’s coding skill. Our evaluations closely matched competition rules and allowed for 10 submissions. GPT‑4o achieved an Elo rating3of 808, which is in the 11th percentile of human competitors. This model far exceeded both GPT‑4o and o1—it achieved an Elo rating of 1807, performing better than 93% of competitors.

针对编程竞赛的进一步微调改进了 o1。改进后的模型在 2024 年国际信息学奥林匹克竞赛中,按照比赛规则排名第 49 百分位。

Further fine-tuning on programming competitions improves o1. The improved model ranked in the 49th percentile in the 2024 International Olympiad in Informatics under competition rules.

人类偏好评估 Human preference evaluation

除了考试和学术基准测试外,我们还在广泛的领域内,针对具有挑战性的开放式提示,评估了人类对 o1‑preview 与 GPT‑4o 的偏好。在此评估中,人类训练员会看到 o1‑preview 和 GPT‑4o 对同一提示的匿名回复,并投票选出他们更喜欢的回复。在数据分析、编程和数学等推理密集型类别中,o1‑preview 相比 GPT-4o 获得了压倒性的偏好。然而,在某些自然语言任务中,o1‑preview 并不被偏好,这表明它并非适用于所有使用场景。

In addition to exams and academic benchmarks, we also evaluated human preference of o1‑preview vs GPT‑4o on challenging, open-ended prompts in a broad spectrum of domains. In this evaluation, human trainers were shown anonymized responses to a prompt from o1‑preview and GPT‑4o, and voted for which response they preferred. o1‑preview is preferred to gpt-4o by a large margin in reasoning-heavy categories like data analysis, coding, and math. However, o1‑preview is not preferred on some natural language tasks, suggesting that it is not well-suited for all use cases.

安全性 Safety

思维链推理为对齐和安全性提供了新的机遇。我们发现,将我们的模型行为策略整合到推理模型的思维链中,是稳健地教授人类价值观和原则的有效方法。通过教授模型我们的安全规则以及如何在上下文中推理这些规则,我们发现了推理能力直接提升模型稳健性的证据:o1‑preview 在关键的越狱评估和我们用于评估模型安全拒绝边界的最难的内部基准测试中,性能显著提升。我们相信,使用思维链为安全性和对齐带来了重大进展,因为(1)它使我们能够以可读的方式观察模型的思考过程,(2)模型对安全规则的推理在分布外场景中更加稳健。

Chain of thought reasoning provides new opportunities for alignment and safety. We found that integrating our policies for model behavior into the chain of thought of a reasoning model is an effective way to robustly teach human values and principles. By teaching the model our safety rules and how to reason about them in context, we found evidence of reasoning capability directly benefiting model robustness: o1‑preview achieved substantially improved performance on key jailbreak evaluations and our hardest internal benchmarks for evaluating our model's safety refusal boundaries. We believe that using a chain of thought offers significant advances for safety and alignment because (1) it enables us to observe the model thinking in a legible way, and (2) the model reasoning about safety rules is more robust to out-of-distribution scenarios.

为了对我们的改进进行压力测试,我们在部署前根据我们的准备框架(在新窗口中打开)进行了一系列安全测试和红队测试。我们发现,思维链推理有助于提升我们各项评估中的能力。特别值得注意的是,我们观察到了奖励破解(在新窗口中打开)的有趣实例。这些评估的详细结果可在随附的系统卡中找到。

To stress-test our improvements, we conducted a suite of safety tests and red-teaming before deployment, in accordance with our Preparedness Framework⁠(opens in a new window). We found that chain of thought reasoning contributed to capability improvements across our evaluations. Of particular note, we observed interesting instances of reward hacking⁠(opens in a new window). Detailed results from these evaluations can be found in the accompanying System Card.

隐藏思维链 Hiding the Chains of Thought

我们相信,隐藏的思维链为监控模型提供了独特的机会。假设它是忠实且可读的,隐藏的思维链使我们能够“读取模型的思想”并理解其思考过程。例如,在未来,我们可能希望监控思维链中是否有操纵用户的迹象。然而,要使这一点成立,模型必须有自由以未改变的形式表达其思想,因此我们不能在思维链上训练任何策略合规性或用户偏好。我们也不希望让未对齐的思维链直接对用户可见。

We believe that a hidden chain of thought presents a unique opportunity for monitoring models. Assuming it is faithful and legible, the hidden chain of thought allows us to "read the mind" of the model and understand its thought process. For example, in the future we may wish to monitor the chain of thought for signs of manipulating the user. However, for this to work the model must have freedom to express its thoughts in unaltered form, so we cannot train any policy compliance or user preferences onto the chain of thought. We also do not want to make an unaligned chain of thought directly visible to users.

因此,在权衡了包括用户体验、竞争优势以及进行思维链监控的选项在内的多个因素后,我们决定不向用户展示原始的思维链。我们承认这一决定有缺点。我们努力通过教导模型在答案中重现思维链中的任何有用想法来部分弥补这一点。对于 o1 模型系列,我们展示模型生成的思维链摘要。

Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users. We acknowledge this decision has disadvantages. We strive to partially make up for it by teaching the model to reproduce any useful ideas from the chain of thought in the answer. For the o1 model series we show a model-generated summary of the chain of thought.

结论 Conclusion

o1 显著推进了 AI 推理的最新技术水平。我们计划在持续迭代的过程中发布该模型的改进版本。我们期望这些新的推理能力将提升我们使模型与人类价值观和原则对齐的能力。我们相信 o1 及其后继者将为科学、编程、数学及相关领域的 AI 解锁许多新的应用场景。我们期待用户和 API 开发者发现它如何改善他们的日常工作。

o1 significantly advances the state-of-the-art in AI reasoning. We plan to release improved versions of this model as we continue iterating. We expect these new reasoning capabilities will improve our ability to align models to human values and principles. We believe o1 – and its successors – will unlock many new use cases for AI in science, coding, math, and related fields. We are excited for users and API developers to discover how it can improve their daily work.

互动版:图/公式 + 针对本篇提问 →