GPT 5.6 有 72 种可能配置。什么是好的默认设置?

GPT 5.6 Has 72 Possible Configurations. What's A Good Default?

塞巴斯蒂安·拉施卡 Sebastian Raschka · · 2026-07-09 · Ahead of AI ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

是的,GPT 5.6 版本中可能有过多的选项可供选择。然而,在推理模型的背景下,我发现这些选择如何映射到训练时和推理时的扩展上很有趣。如果我们粗略地将这些选项映射到经典的 o1 图上,Sol、Terra 和 Luna 代表三种模型规模和训练预算,位于训练计算轴上。而 effort 设置则位于推理时计算轴上。图 1 展示了 OpenAI 经典 o1 扩展图与 GPT 5.6 选项的粗略映射。三种模型选项代表不同的模型规模和训练预算,而 effort 设置代表推理时计算。

Yes, there are probably too many options to choose from in the GPT 5.6 release. In the context of reasoning models, though, I find it interesting how these choices map onto training-time and inference-time scaling. If we loosely map the options onto the classic o1 plots, Sol, Terra, and Luna stand in for three model sizes and training budgets along the training-compute axis. The effort settings then sit on the inference-time-compute axis. Figure 1. A loose mapping from OpenAI's classic o1 scaling plots to the GPT 5.6 choices. The three model options represent different model sizes and training budgets, while the effort setting represents inference-time compute.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 1)

全文 · Full text(逐段中英对照)

概述 Overview

是的,在 GPT 5.6 发布中,可供选择的选项可能太多了。

Yes, there are probably too many options to choose from in the GPT 5.6 release.

然而,在推理模型的背景下,我觉得有趣的是这些选择如何映射到训练时间和推理时间的 Scaling(规模扩张)。如果我们粗略地将这些选项映射到经典的 o1 图上,Sol、Terra 和 Luna 代表三种模型规模和训练预算,位于训练算力轴上。而 effort 设置则位于推理时间算力轴上。

In the context of reasoning models, though, I find it interesting how these choices map onto training-time and inference-time scaling. If we loosely map the options onto the classic o1 plots, Sol, Terra, and Luna stand in for three model sizes and training budgets along the training-compute axis. The effort settings then sit on the inference-time-compute axis.

图 1:OpenAI 经典 o1 缩放图到 GPT 5.6 选项的粗略映射。三个模型选项代表不同的模型规模和训练预算,而努力程度设置代表推理时算力。

Figure 1. A loose mapping from OpenAI's classic o1 scaling plots to the GPT 5.6 choices. The three model options represent different model sizes and training budgets, while the effort setting represents inference-time compute.

完整列表包含三个模型选项和六个推理努力级别(Light、Medium、High、Extra High、Max 和 Ultra)。一旦加入 Work 与 Codex、Standard 与 Fast 的区分,完整配置矩阵变为:

The full list has three model choices and six reasoning-effort levels (Light, Medium, High, Extra High, Max, and Ultra). Once Work versus Codex and Standard versus Fast are included, the full configuration matrix becomes

Work/Codex × Sol/Terra/Luna × Light/Medium/High/Extra High/Max/Ultra × Standard/Fast

Work/Codex × Sol/Terra/Luna × Light/Medium/High/Extra High/Max/Ultra × Standard/Fast

那么,现在什么是好的默认配置?Luna 搭配 High 努力?Sol 搭配 Light 努力?还是 Terra 搭配 Medium 努力?

So, what is a good default now? Luna with High effort? Sol with Light effort? Terra with Medium effort?

当然,性能与成本对比图有助于识别性价比高的组合。例如,Luna 搭配 Extra High 努力可能比 Sol 搭配 Medium 努力更好且更便宜。

Sure, a performance-versus-cost chart can help identify the good bang-for-the-buck combinations. For instance, Luna with Extra High effort may be better and cheaper than Sol with Medium effort.

图 2. 人工分析编码智能体指数 v1.1 得分与 API 成本的关系图,涵盖 GPT 5.6 及多个对比模型。该图基于 OpenAI 在 X 上发布的 GPT 5.6 发布帖。

Figure 2. Artificial Analysis Coding Agent Index v1.1 scores plotted against API cost across GPT 5.6 and several comparison models. Figure based on OpenAI's GPT 5.6 release post on X.

但确实,72 种可能的配置给我们留下了很多选择 🤯。

But yeah, 72 possible configurations leave us with a lot of choices 🤯.

互动版:图/公式 + 针对本篇提问 →