数学与理论计算机科学的十项进展(转载)几天前,Anthropic 使用 Claude 的 Mythos Preview 版本发现了密码学弱点,花费了 10 万美元的 token,并提示“我们不是在寻找容易的成果,我们想要真正的研究来发现真正困难的发现”。现在轮到 OpenAI 展示实力了。他们让“Astra 的内部版本,我们的下一个主要模型”来解决十个数学问题,这些问题“主要结果至少十年没有进展”。他们声称每个问题在 GPT-5.6 Sol 代币价格上花费不到 2000 美元。(没有关于他们在多少个问题上花费了 2000 美元却没有得到解决方案的消息。)
Ten advances in mathematics and theoretical computer science (via) A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinely hard findings." Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade". They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one. (No news on how many problems they spent $2,000 on _without_ reaching a solution though.)
核心贡献 · Key contributions
OpenAI 的内部模型 Astra 解决了十个十多年未有进展的数学难题。 OpenAI's internal model Astra solved ten long-standing mathematical problems with no progress for over a decade.
每个解决方案的代币成本不到 2000 美元,展示了 AI 驱动研究的高效率。 Each solution cost less than $2,000 in token prices, demonstrating high efficiency in AI-driven research.
结果在 Lean 4 中形式化,确保证明的严格验证。 The results are formalized in Lean 4, ensuring rigorous verification of the proofs.
OpenAI 发布了一篇论文和一份由 LLM 生成的 PDF,从推理轨迹中重建证明过程。 OpenAI released a paper and an LLM-generated PDF reconstructing the proof process from reasoning traces.
这项工作凸显了前沿模型在推进数学和理论计算机科学方面的潜力。 The work highlights the potential of frontier models in advancing mathematics and theoretical computer science.
局限 · Limitations
未披露尝试但未成功的问题数量,成功率未知。 The number of problems attempted without success is undisclosed, leaving the success rate unknown.
用于指导模型的提示词未公开,限制了可复现性。 The prompts used to guide the model are not released, limiting reproducibility.
由于所选问题具有特定性,结果可能无法推广到所有数学领域。 The results may not generalize to all mathematical domains, as the selected problems were specific.
成本估算未包括潜在的隐藏计算开销,如开发和训练。 The cost estimate excludes potential hidden computational expenses, such as development and training.
该方法依赖模型的内部推理,可能不透明且难以审计。 The approach relies on the model's internal reasoning, which may be opaque and hard to audit.
论文章节 · Sections(共 1)
Simon Willison 的博客(https://simonwillison.net/)Overview