Ten advances in mathematics and theoretical computer science
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→数学与理论计算机科学的十项进展(转载)几天前,Anthropic 使用 Claude 的 Mythos Preview 版本发现了密码学弱点,花费了 10 万美元的 token,并提示“我们不是在寻找容易的成果,我们想要真正的研究来发现真正困难的发现”。现在轮到 OpenAI 展示实力了。他们让“Astra 的内部版本,我们的下一个主要模型”来解决十个数学问题,这些问题“主要结果至少十年没有进展”。他们声称每个问题在 GPT-5.6 Sol 代币价格上花费不到 2000 美元。(没有关于他们在多少个问题上花费了 2000 美元却没有得到解决方案的消息。)
Ten advances in mathematics and theoretical computer science (via) A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinely hard findings." Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade". They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one. (No news on how many problems they spent $2,000 on _without_ reaching a solution though.)
数学与理论计算机科学领域的十项进展(via)几天前,Anthropic 使用 Claude 和 Mythos Preview 发现了密码学弱点,花费了 10 万美元的 token,提示词中包括“我们不是在寻找唾手可得的果实,我们想要真正的研究,以发现真正困难的发现。”
Ten advances in mathematics and theoretical computer science (via) A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinely hard findings."
现在轮到 OpenAI 展示实力了。他们让“Astra 的内部版本,我们的下一个主要模型”去解决十个数学问题,这些问题“主要结果至少十年没有进展”。他们声称每个问题在 GPT-5.6 Sol 的 token 价格下花费不到 2000 美元。
Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade". They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one.
(不过,关于他们花了 2000 美元却未解决问题的情况,没有消息。)
(No news on how many problems they spent $2,000 on _without_ reaching a solution though.)
openai/ten-proofs 仓库包含他们结果的 Lean 4 形式化,还有一篇描述解决方案的论文,以及一个额外的 LLM 生成的 PDF,其中模型基于未发表的推理轨迹“重构了证明是如何形成的”。
The openai/ten-proofs repository has Lean 4 formalizations of their results, and there's also a paper describing the solutions and an additional LLM-generated PDF where the model "reconstructs how the proof came together" based on the unpublished reasoning traces.
这种透明度还算不错,但我想看到他们使用的提示词!
That's a decent level of transparency, but I want to see the prompts they used!
网上许多数学家正经历一场集体性的“深蓝”爆发。数学家基尔温·汉普希尔上周发表了一篇充满激情的文章《数学的暗夜》,描述了此前(且不那么重要的)结果所引发的“深刻的精神危机”。
A lot of mathematicians online are experiencing a collective burst of Deep Blue. Mathematician Kirwin Hampshire published an impassioned essay last week, The Dark Night of Mathematics, describing "a profound spiritual crisis" brought on by previous (and less significant) results.
OpenAI 的结果让我想起特伦斯·陶在六月《IEEE Spectrum》中所描述的“大数学”:
OpenAI's results remind me of what Terence Tao described as "big mathematics" in IEEE Spectrum in June: