The Second Half
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→几十年来,人工智能主要致力于开发新的训练方法和模型。这确实奏效了:从击败国际象棋和围棋世界冠军,在 SAT 和律师资格考试中超越大多数人类,到获得 IMO 和 IOI 金牌。这些历史里程碑——深蓝、AlphaGo、GPT-4 和 o 系列——背后是 AI 方法的根本创新:搜索、深度强化学习、扩展和推理。事情随着时间的推移越来越好。用三个词概括:强化学习终于奏效了。更准确地说:强化学习终于泛化了。在经历了几次重大曲折和一系列里程碑的积累之后,我们找到了一种有效的配方,可以利用语言和推理解决广泛的强化学习任务。即使是一年前,如果你告诉大多数 AI 研究人员,一个单一的配方可以解决软件工程、创意写作、IMO 级别的数学、鼠标键盘操作和长形式问答——他们会嘲笑你的幻觉。这些任务中的每一个都极其困难,许多研究人员花费整个博士生涯只专注于其中一个狭窄的领域。那么接下来呢?AI 的下半场——从现在开始——将从解决问题转向定义问题。在这个新时代,评估变得更加重要。
For decades, AI has largely been about developing new training methods and models. And it worked: from beating world champions at chess and Go, surpassing most humans on the SAT and bar exams, to earning IMO and IOI gold medals. Behind these milestones in the history book — DeepBlue, AlphaGo, GPT-4, and the o-series — are fundamental innovations in AI methods: search, deep RL, scaling, and reasoning. Things just get better over time. In three words: RL finally works. More precisely: RL finally generalizes. After several major detours and a culmination of milestones, we’ve landed on a working recipe to solve a wide range of RL tasks using language and reasoning. Even a year ago, if you told most AI researchers that a single recipe could tackle software engineering, creative writing, IMO-level math, mouse-and-keyboard manipulation, and long-form question answering — they’d laugh at your hallucinations. Each of these tasks is incredibly difficult and many researchers spend their entire PhDs focused on just one narrow slice. So what comes next? The second half of AI — starting now — will shift focus from solving problems to defining problems. In this new era, evaluation becomes more im
几十年来,人工智能主要围绕开发新的训练方法和模型展开。这确实奏效了:从在国际象棋和围棋中击败世界冠军,在 SAT 和律师资格考试中超越大多数人类,到获得 IMO 和 IOI 金牌。这些历史里程碑——深蓝、AlphaGo、GPT-4 以及 o 系列——背后是 AI 方法的根本性创新:搜索、深度强化学习、Scaling(规模扩张)和推理。事情只会随着时间的推移变得更好。
For decades, AI has largely been about developing new training methods and models. And it worked: from beating world champions at chess and Go, surpassing most humans on the SAT and bar exams, to earning IMO and IOI gold medals. Behind these milestones in the history book — DeepBlue, AlphaGo, GPT-4, and the o-series — are fundamental innovations in AI methods: search, deep RL, scaling, and reasoning. Things just get better over time.
用三个词概括:强化学习终于奏效了。更准确地说:强化学习终于实现了泛化。在经历了几次重大曲折和一系列里程碑的积累之后,我们找到了一种可行的方案,利用语言和推理来解决广泛的强化学习任务。即使在一年前,如果你告诉大多数 AI 研究人员,一个单一的方案可以解决软件工程、创意写作、IMO 级别的数学、鼠标键盘操作以及长文本问答——他们会嘲笑你的幻觉。这些任务中的每一个都极其困难,许多研究人员花费整个博士生涯只专注于其中一个狭窄的领域。
In three words: RL finally works. More precisely: RL finally generalizes. After several major detours and a culmination of milestones, we’ve landed on a working recipe to solve a wide range of RL tasks using language and reasoning. Even a year ago, if you told most AI researchers that a single recipe could tackle software engineering, creative writing, IMO-level math, mouse-and-keyboard manipulation, and long-form question answering — they’d laugh at your hallucinations. Each of these tasks is incredibly difficult and many researchers spend their entire PhDs focused on just one narrow slice.
那么接下来呢?AI 的下半场——从现在开始——将把焦点从解决问题转向定义问题。在这个新时代,评估变得比训练更重要。我们不再仅仅问“我们能训练一个模型来解决 X 吗?”,而是问“我们应该训练 AI 做什么,以及如何衡量真正的进步?”为了在下半场蓬勃发展,我们需要及时转变思维方式和技能组合,或许更接近产品经理的角色。
So what comes next? The second half of AI — starting now — will shift focus from solving problems to defining problems. In this new era, evaluation becomes more important than training. Instead of just asking, “Can we train a model to solve X?”, we’re asking, “What should we be training AI to do, and how do we measure real progress?” To thrive in this second half, we’ll need a timely shift in mindset and skill set, ones perhaps closer to a product manager.
要理解上半场,看看它的赢家。你认为迄今为止最具影响力的 AI 论文有哪些?
To make sense of the first half, look at its winners. What do you consider to be the most impactful AI papers so far?
我试了斯坦福 224N 的测验,答案并不意外:Transformer、AlexNet、GPT-3 等。这些论文有什么共同点?它们提出了训练更好模型的基本突破。但同样,它们通过在某些基准上展示(显著的)改进而成功发表了论文。
I tried the quiz in Stanford 224N, and the answers were not surprising: Transformer, AlexNet, GPT-3, etc. What’s common about these papers? They propose some fundamental breakthroughs to train better models. But also, they managed to publish their papers by showing some (significant) improvements on some benchmarks.
然而有一个潜在共性:这些“赢家”都是训练方法或模型,而不是基准或任务。即使是有史以来最具影响力的基准 ImageNet,其引用量也不到 AlexNet 的三分之一。方法 vs 基准的对比在其他地方更为悬殊——例如,Transformer 的主要基准是 WMT'14,其研讨会报告约有 1300 次引用,而 Transformer 超过 16 万次。
There is a latent commonality though: these “winners” are all training methods or models, not benchmarks or tasks. Even arguably the most impactful benchmark of all, ImageNet, has less than one third of the citation of AlexNet. The contrast of method vs benchmark is even more drastic anywhere else —- for example, the main benchmark of Transformer is WMT’14, whose workshop report has ~1,300 citations, while Transformer had >160,000.
这说明了上半场的游戏规则:专注于构建新模型和新方法,评估和基准是次要的(尽管为了让论文体系运作是必要的)。
That illustrates the game of the first half: focus on building new models and methods, and evaluation and benchmark are secondary (although necessary to make the paper system work).
为什么?一个主要原因是,在 AI 的上半场,方法比任务更难、更令人兴奋。从头创建新的算法或模型架构——想想反向传播算法、卷积网络(AlexNet)或 GPT-3 中使用的 Transformer 等突破——需要非凡的洞察力和工程能力。相比之下,为 AI 定义任务往往更直接:我们只是把人类已经做的任务(如翻译、图像识别或国际象棋)变成基准。不需要太多洞察力甚至工程。
Why? A big reason is that, in the first half of AI, methods were harder and more exciting than tasks. Creating a new algorithm or model architecture from scratch – think of breakthroughs like the backpropagation algorithm, convolutional networks (AlexNet), or the Transformer used in GPT-3 – required remarkable insight and engineering. In contrast, defining tasks for AI often felt more straightforward: we simply took tasks humans already do (like translation, image recognition, or chess) and turned them into benchmarks. Not much insight or even engineering.
方法也往往比单个任务更通用、更广泛适用,因此特别有价值。例如,Transformer 架构最终推动了 CV、NLP、RL 和许多其他领域的进步——远远超出了它最初证明自己的单一数据集(WMT'14 翻译)。一个伟大的新方法可以攀登许多不同的基准,因为它简单且通用,因此影响往往超越单个任务。
Methods also tended to be more general and widely applicable than individual tasks, making them especially valuable. For example, the Transformer architecture ended up powering progress in CV, NLP, RL, and many other domains – far beyond the single dataset (WMT’14 translation) where it first proved itself. A great new method can hillclimb many different benchmarks because it’s simple and general, thus the impact tends to go beyond an individual task.
这个游戏已经运行了几十年,激发了改变世界的想法和突破,这些体现在各个领域不断增长的基准性能上。为什么游戏会改变?因为这些想法和突破的积累在创建解决任务的有效配方方面产生了质的差异。
This game has worked for decades and sparked world-changing ideas and breakthroughs, which manifested themselves by ever-increasing benchmark performances in various domains. Why would the game change at all? Because the cumulation of these ideas and breakthroughs have made a qualitative difference in creating a working recipe in solving tasks.
配方是什么?其成分不出所料包括大规模语言预训练、规模(数据和算力)以及推理和行动的理念。这些听起来像是你在旧金山每天听到的流行词,但为什么称之为配方呢?
What’s the recipe? Its ingredients, not surprisingly, include massive language pre-training, scale (in data and compute), and the idea of reasoning and acting. These might sound like buzzwords that you hear daily in SF, but why call them a recipe??
我们可以通过强化学习的视角来理解这一点,强化学习通常被认为是人工智能的“终局”——毕竟,强化学习在理论上保证能赢得游戏,并且从经验上看,很难想象没有强化学习的超人类系统(例如 AlphaGo)。
We can understand this by looking through the lens of reinforcement learning (RL), which is often thought of as the “end game” of AI — after all, RL is theoretically guaranteed to win games, and empirically it’s hard to imagine any superhuman systems (e.g. AlphaGo) without RL.
在强化学习中,有三个关键组成部分:算法、环境和先验。长期以来,强化学习研究者主要关注算法(例如 REINFORCE、DQN、TD 学习、演员-评论家、PPO、TRPO……)——这是智能体如何学习的智力核心——而将环境和先验视为固定或最小化。例如,Sutton 和 Barto 的经典教科书全是关于算法,几乎不涉及环境或先验。
In RL, there are three key components: algorithm, environment, and priors. For a long time, RL researchers focused mostly on the algorithm (e.g. REINFORCE, DQN, TD-learning, actor-critic, PPO, TRPO…) – the intellectual core of how an agent learns – while treating the environment and priors as fixed or minimal. For example, Sutton and Barto’s classical textbook is all about algorithms and almost nothing about environments or priors.
然而,在深度强化学习时代,环境在经验上显然非常重要:算法的性能通常高度依赖于其开发和测试的环境。如果你忽略环境,你可能会构建一个仅在玩具设置中表现出色的“最优”算法。那么,我们为什么不先弄清楚我们真正想解决的环境,然后找到最适合它的算法呢?
However, in the era of deep RL, it became clear that environments matter a lot empirically: an algorithm’s performance is often highly specific to the environment it was developed and tested in. If you ignore the environment, you risk building an “optimal” algorithm that only excels in toy settings. So why don’t we first figure out the environment we actually want to solve, then find the algorithm best suited for it?
这正是 OpenAI 最初的计划。它构建了 gym,一个用于各种游戏的标准强化学习环境,然后是 World of Bits 和 Universe 项目,试图将互联网或计算机变成一个游戏。这是个好计划,不是吗?一旦我们将所有数字世界转化为环境,用智能的强化学习算法解决它,我们就拥有了数字 AGI。
That’s exactly OpenAI’s initial plan. It built gym, a standard RL environment for various games, then the World of Bits and Universe projects, trying to turn the Internet or computer into a game. A good plan, isn’t it? Once we turn all digital worlds into an environment, solve it with smart RL algorithms, we have digital AGI.
好计划,但并未完全奏效。OpenAI 在这条道路上取得了巨大进展,使用强化学习解决了 Dota、机器人手等问题。但它从未接近解决计算机使用或网页导航问题,而且在一个领域工作的强化学习智能体无法迁移到另一个领域。缺少了一些东西。
A good plan, but not entirely working. OpenAI made tremendous progress down the path, using RL to solve Dota, robotic hands, etc. But it never came close to solving computer use or web navigation, and the RL agents working in one domain do not transfer to another. Something is missing.
直到 GPT-2 或 GPT-3 之后,才发现缺失的部分是先验。你需要强大的语言预训练来将通用常识和语言知识提炼到模型中,然后可以微调这些模型成为网页(WebGPT)或聊天(ChatGPT)智能体(并改变世界)。事实证明,强化学习最重要的部分可能甚至不是强化学习算法或环境,而是先验,这些先验可以通过与强化学习完全无关的方式获得。
Only after GPT-2 or GPT-3, it turned out that the missing piece is priors. You need powerful language pre-training to distill general commonsense and language knowledge into models, which then can be fine-tuned to become web (WebGPT) or chat (ChatGPT) agents (and change the world). It turned out the most important part of RL might not even be the RL algorithm or environment, but the priors, which can be obtained in a way totally unrelated from RL.
语言预训练为聊天创造了良好的先验,但对于控制计算机或玩视频游戏并不同样有效。为什么?这些领域与互联网文本的分布相距更远,并且在这些领域上简单地进行 SFT/RL 泛化效果很差。我在 2019 年注意到了这个问题,当时 GPT-2 刚刚发布,我在其基础上进行 SFT/RL 来解决基于文本的游戏——CALM 是世界上第一个通过预训练语言模型构建的智能体。但它花费了数百万次强化学习步骤才能让智能体在一个游戏中爬坡,并且无法迁移到新游戏。尽管这正是强化学习的特征,对强化学习研究者来说并不奇怪,但我发现这很奇怪,因为我们人类可以轻松地玩一个新游戏,并且零样本表现显著更好。然后我迎来了人生中第一个顿悟时刻——我们之所以能泛化,是因为我们可以选择做更多的事情,而不仅仅是“去柜子 2”或“用钥匙 1 打开箱子 3”或“用剑杀死地牢”,我们还可以选择思考诸如“地牢很危险,我需要武器与之战斗。没有可见的武器,所以也许我需要在锁着的盒子或箱子里找一把。箱子 3 在柜子 2 里,让我先去那里解锁它”之类的事情。
Language pre-training created good priors for chatting, but not equally good for controlling computers or playing video games. Why? These domains are further from the distribution of Internet text, and naively doing SFT / RL on these domains generalizes poorly. I noticed the problem in 2019, when GPT-2 just came out and I did SFT / RL on top of it to solve text-based games - CALM was the first agent in the world built via pre-trained language models. But it took millions of RL steps for the agent to hillclimb a single game, and it doesn’t transfer to new games. Though that’s exactly the characteristic of RL and nothing strange to RL researchers, I found it weird because we humans can easily play a new game and be significantly better zero-shot. Then I hit one of the first eureka moment in my life - we generalize because we can choose to do more than “go to cabinet 2” or “open chest 3 with key 1” or “kill dungeon with sword”, we can also choose to think about things like “The dungeon is dangerous and I need a weapon to fight with it. There is no visible weapon so maybe I need to find one in locked boxes or chests. Chest 3 is in Cabinet 2, let me first go there and unlock it”.
思考,或推理,是一种奇怪的行为——它不直接影响外部世界,然而推理的空间是开放且组合无限的——你可以思考一个词、一个句子、一整段话,或 10000 个随机英语单词,但你周围的世界不会立即改变。在经典强化学习理论中,这是一个糟糕的交易,使得决策变得不可能。想象一下,你需要从两个盒子中选择一个,只有一个盒子有 100 万美元,另一个是空的。你预计能赚 50 万美元。现在想象我添加了无限个空盒子。你预计什么也赚不到。但是通过将推理添加到任何强化学习环境的行为空间中,我们利用语言预训练先验进行泛化,并且我们能够为不同的决策提供灵活的测试时算力。这真的很神奇,我抱歉在这里没有完全解释清楚,我可能需要再写一篇博客专门讨论它。欢迎阅读 ReAct 了解智能体推理的原始故事,并感受我当时的心境。现在,我的直观解释是:即使你添加了无限个空盒子,你在生活中各种游戏里都见过它们,选择这些盒子能让你更好地为任何给定游戏选择有钱的盒子。我的抽象解释是:语言通过智能体中的推理进行泛化。
Thinking, or reasoning, is a strange kind of action - it does not directly affect the external world, yet the space of reasoning is open-ended and combintocially infinite — you can think about a word, a sentence, a whole passage, or 10000 random English words, but the world around you doesn’t immediate change. In the classical RL theory, it is a terrible deal and makes decision-making impossible. Imagine you need to choose one out of two boxes, and there’s only one box with $1M and the other one empty. You’re expected to earn $500k. Now imagine I add infinite empty boxes. You’re expected to earn nothing. But by adding reasoning into the action space of any RL environment, we make use of the language pre-training priors to generalize, and we afford to have flexible test-time compute for different decisions. It is a really magical thing and I apologize for not fully making sense of it here, I might need to write another blog post just for it. You’re welcome to read ReAct for the original story of reasoning for agents and read my vibes at the time. For now, my intuitive explanation is: even though you add infinite empty boxes, you have seen them throughout your life in all kinds of games, and choosing these boxes prepare you to better choose the box with money for any given game. My abstract explanation would be: language generalizes through reasoning in agents.
一旦我们有了正确的强化学习先验(语言预训练)和强化学习环境(将语言推理作为行为添加),强化学习算法可能是最微不足道的部分。于是我们有了 o 系列、R1、深度研究、计算机使用智能体,以及更多即将到来的东西。多么讽刺的转折!长期以来,强化学习研究者对算法的关注远多于环境,而且没有人关注先验——所有强化学习实验基本上都是从零开始。但我们花了数十年的弯路才意识到,也许我们的优先级应该完全颠倒过来。
Once we have the right RL priors (language pre-training) and RL environment (adding language reasoning as actions), it turns out RL algorithm might be the most trivial part. Thus we have o-series, R1, deep research, computer-using agent, and so much more to come. What a sarcastic turn of events! For so long RL researchers cared about algorithms way more than environments, and no one paid any attention to priors — all RL experiments essentially start from scratch. But it took us decades of detours to realize maybe our prioritization should have be completely reversed.
但正如史蒂夫·乔布斯所说:你无法向前看将点连接起来;你只能向后看将它们连接起来。
But just like Steve Jobs said: You can’t connect the dots looking forward; you can only connect them looking backward.
这个配方正在彻底改变游戏规则。回顾上半场的游戏:
This recipe is completely changing the game. To recap the game of the first half:
* 我们开发新颖的训练方法或模型,以攀登基准测试。
* We develop novel training methods or models that hillclimb benchmarks.
* 我们创建更难的基准测试,并继续循环。
* We create harder benchmarks and continue the loop.
* 这个配方基本上已经标准化和工业化地攀登基准测试,而不需要太多新想法。随着配方的扩展和泛化良好,你针对特定任务的新方法可能提升 5%,而下一个 o 系列模型在没有明确针对的情况下提升了 30%。
* The recipe has essentially standardized and industried benchmark hillclimbing without requiring much more new ideas. As the recipe scales and generalizes well, your novel method for a particular task might improve it by 5%, while the next o-series model improve it by 30% without explicitly targeting it.
* 即使我们创建更难的基准测试,很快(而且越来越快)它们也会被这个配方解决。我的同事 Jason Wei 制作了一张漂亮的图来很好地可视化这一趋势:
* Even if we create harder benchmarks, pretty soon (and increasingly soon) they get solved by the recipe. My colleague Jason Wei made a beautiful figure to visualize the trend well:
那么下半场还有什么可玩的?如果不再需要新方法,而更难的基准测试只会越来越快地被解决,我们该怎么办?
Then what’s left to play in the second half? If novel methods are no longer needed and harder benchmarks will just get solved increasingly soon, what should we do?
我认为我们应该从根本上重新思考评估。这意味着不仅要创建新的、更难的基准测试,还要从根本上质疑现有的评估设置并创建新的设置,这样我们才能被迫发明超越现有配方的新方法。这很难,因为人类有惯性,很少质疑基本假设——你只是理所当然地接受它们,而没有意识到它们是假设,而不是定律。
I think we should fundamentally re-think evaluation. It means not just to create new and harder benchmarks, but to fundamentally question existing evaluation setups and create new ones, so that we are forced to invent new methods beyond the working recipe. It is hard because humans have inertia and seldom question basic assumptions - you just take them for granted without realizing they are assumptions, not laws.
为了解释惯性,假设你发明了历史上最成功的评估之一,基于人类考试。这在 2021 年是一个极其大胆的想法,但 3 年后它已经饱和了。你会怎么做?很可能创建一个更难的考试。或者假设你解决了简单的编码任务。你会怎么做?很可能找到更难的编码任务来解决,直到你达到 IOI 金牌水平。
To explain inertia, suppose you invented one of the most successful evals in history based on human exams. It was an extremely bold idea in 2021, but 3 years later it’s saturated. What would you do? Most likely create a much harder exam. Or suppose you solved simply coding tasks. What would you do? Most likely find harder coding tasks to solve until you have reached IOI gold level.
惯性是自然的,但问题在于:AI 已经在国际象棋和围棋中击败了世界冠军,在 SAT 和律师考试中超越了大多数人类,并在 IOI 和 IMO 上达到了金牌水平。但世界并没有太大改变,至少从经济和 GDP 来看是这样。
Inertia is natural, but here is the problem. AI has beat world champions at chess and Go, surpassed most humans on SAT and bar exams, and reached gold medal level on IOI and IMO. But the world hasn’t changed much, at least judged by economics and GDP.
我称之为效用问题,并认为它是 AI 最重要的问题。
I call this the utility problem, and deem it the most important problem for AI.
也许我们很快就会解决效用问题,也许不会。无论哪种情况,这个问题的根本原因可能出奇地简单:我们的评估设置在许多基本方面与现实世界的设置不同。举两个例子:
Perhaps we will solve the utility problem pretty soon, perhaps not. Either way, the root cause of this problem might be deceptively simple: our evaluation setups are different from real-world setups in many basic ways. To name two examples:
* 评估“应该”自动运行,所以通常智能体接收任务输入,自主执行,然后获得任务奖励。但在现实中,智能体必须在整个任务过程中与人类互动——你不会给客服发一条超长消息,等 10 分钟,然后期望一个详细的回复来解决所有问题。通过质疑这一设置,新的基准测试被发明出来,要么让真实人类参与(例如 Chatbot Arena),要么在循环中使用用户模拟(例如 tau-bench)。
* Evaluation “should” run automatically, so typically an agent receives a task input, do things autonomously, then receive a task reward. But in reality, an agent has to engage with a human throughout the task — you don’t just text customer service a super long message, wait for 10 minutes, then expect a detailed response to settle everything. By questioning this setup, new benchmarks are invented to either engage real humans (e.g. Chatbot Arena) or user simulation (e.g. tau-bench) in the loop.
* 评估“应该”独立同分布运行。如果你有一个包含 500 个任务的测试集,你独立运行每个任务,平均任务指标,得到总体指标。但在现实中,你是顺序解决任务,而不是并行。一个谷歌软件工程师在越来越熟悉代码库后,能越来越好地解决 google3 的问题,但一个软件工程智能体在解决同一代码库中的多个问题时,并没有获得这种熟悉度。我们显然需要长期记忆方法(并且确实存在),但学术界没有合适的基准测试来证明这种需求,甚至没有足够的勇气质疑作为机器学习基础的独立同分布假设。
* Evaluation “should” run i.i.d. If you have a test set with 500 tasks, you run each task independently, average the task metrics, and get an overall metric. But in reality, you solve tasks sequentially rather than in parallel. A Google SWE solves google3 issues increasingly better as she gets more familiar with the repo, but a SWE agent solves many issues in the same repo without gaining such familiarity. We obviously need long-term memory methods (and thereare), but academia does not have the proper benchmarks to justify the need, or even the proper courage to question i.i.d. assumption that has been the foundation of machine learning.
这些假设“一直”都是这样,在 AI 的上半场,在这些假设下开发基准测试是没问题的,因为当智能水平低时,提高智能通常能提高效用。但现在,通用配方在这些假设下保证有效。所以玩下半场新游戏的方式是:
These assumptions have “always” been like this, and developing benchmarks in these assumptions were fine in the first half of AI, because when the intelligence is low, improving intelligence generally improves utility. But now, the general recipe is guaranteed to work under these assumptions. So the way to play the new game of the second half is
* 我们为现实世界的效用开发新颖的评估设置或任务。
* We develop novel evaluation setups or tasks for real-world utility.
* 我们用配方解决它们,或者用新组件增强配方。继续循环。
* We solve them with the recipe or augment the recipe with novel components. Continue the loop.
这个游戏很难,因为它不熟悉。但它令人兴奋。上半场的玩家解决视频游戏和考试,而下半场的玩家则通过利用智能构建有用的产品来建立价值数十亿或数万亿美元的公司。上半场充满了增量方法和模型,而下半场在某种程度上会过滤它们。通用配方会碾压你的增量方法,除非你创建新的假设来打破这个配方。这样你就能做真正改变游戏规则的研究。
This game is hard because it is unfamiliar. But it is exciting. While players in the first half solve video games and exams, players in the second half get to build billion or trillion dollar companies by building useful products out of intelligence. While the first half is filled with incremental methods and models, the second half filters them to some degree. The general recipe would just crush your incremental methods, unless you create new assumptions that break the recipe. Then you get to do truly game-changing research.