从《Don't Worry About the Vase》发现更多内容——一个由齿轮构成的世界。既提供快速的优质短期更新,也进行长期的世界模型构建。目前专注于每周 AI 更新。探索领域包括 AI、政策、理性、医学与生育、教育及游戏。订阅即表示你同意 Substack 的使用条款,并确认已阅读其信息收集通知和隐私政策。
Discover more from Don't Worry About the Vase A world made of gears. Doing both speed premium short term updates and long term world model building. Currently focused on weekly AI updates. Explorations include AI, policy, rationality, medicine and fertility, education and games. By subscribing, you agree to Substack's Terms of Use, and acknowledge its Information Collection Notice and Privacy Policy.
核心贡献 · Key contributions
Kimi K3 是一个 2.8T 参数的开源权重模型,在原始能力上是最强的开源模型,综合基准排名第三。 Kimi K3 is a 2.8T-parameter open-weight model, making it the strongest open model in raw capability and third on overall benchmarks.
该模型正好位于中国 ECI 趋势线上,估计落后封闭前沿六个月;预训练落后程度大于后训练。 It sits exactly on the Chinese ECI trend line, estimated six months behind the closed frontier; pre-training lags more than post-training.
从 Claude 蒸馏显然是部分原因,但远非全部;Moonshot 将快速跟进与真正的创新相结合。 Distillation from Claude is clearly part of the story, but far from the whole story; Moonshot combines fast-following with genuine innovation.
Kimi K3 在智能体式编码、前端生成和 3D 任务上表现出色,但在网络基准上表现不佳,能力呈锯齿状。 Kimi K3 excels at agentic coding, frontend generation, and 3D tasks, while underperforming on cyber benchmarks and showing jagged capability.
基准测试以最大努力运行,可能高估实际表现;该模型更慢、更耗 token,且服务成本高。 Benchmarks are run at maximum effort and likely overstate practical performance; the model is slower, token-hungry, and expensive to serve.
该发布对网络和生物滥用构成非零尾部风险,但不是 Mythos 级别的开源模型时刻。 The release poses non-zero tail risks for cyber and bio misuse, but is not a Mythos-level open-model moment.
局限 · Limitations
所有官方基准均采用最大努力设置,使相对能力看起来比实际表现更好;评估期间访问不稳定。 All official benchmarks used maximum effort, making relative capabilities look better than real-world performance; access was spotty during evaluation.
该模型又大又慢,消耗大量 token 和算力,与较小的开源模型相比限制了实际部署。 The model is large and slow, consuming many tokens and requiring substantial compute, limiting its practical deployment compared with smaller open models.
网络能力相对于编码显得异常薄弱,且未报告针对生物滥用的实质性测试。 Cyber capability appears strangely weak relative to coding, and no substantive testing against bio misuse was reported.
由于权重将公开释放,任何防护措施都可能被剥离,增加了滥用和下游危害的可能性。 Because weights will be released openly, any safeguards can be stripped, increasing the potential for misuse and downstream harm.
不确定性仍然很高:发布周访问受限,蒸馏引发原创性问题,监管审查可能阻碍商业采用。 Uncertainty remains high: release-week access was limited, distillation raises questions about originality, and regulatory scrutiny may hamper commercial adoption.
论文章节 · Sections(共 22)
别担心花瓶Don't Worry About the Vase(https://thezvi.substack.com/)
目录Table of Contents
DeepSeek 时刻:又来了DeepSeek Moments: Here We Go Again
我们曾有过一个时刻(2025 年 6 月重演)We Had a Moment (Reprise from June 2025(https://thezvi.substack.com/i/165339410/we-had-a-moment))
此后的故事The Story Since Then
Kimi K3 发布、宣传与基本事实The Kimi K3 Announcement, Pitch and Basic Facts
论现代刷榜On Modern Benchmaxxing
他人的基准Other People’s Benchmarks
基准测试并非真实世界Benchmarks Are Not The Real World
技术防护措施?那是什么?Technical Safeguards? What Are Those?
Kimi 能做的事Things Kimi Can Do
Kimi 无法做到的事Things Kimi Cannot Do
让 Kimi 难以做到的事情Things It Is Not Easy To Get Kimi To Do
开放权重模型是不安全的,且无法修复Open Weight Models Are Unsafe And Nothing Can Fix This
Dean Ball 试图提出建设性意见Dean Ball Attempts To Be Constructive
特朗普政府考虑发布行政命令,在美国境内禁止中国的开放模型Trump Administration Considering Executive Order Banning Chinese Open Models Within the United States
OpenAI 员工对此相对乐观OpenAI Employees Are Relatively Bullish On This One
Kimi K3 在典型智能体式编码、前端工作和 3D 方面相对最强Kimi K3 Is Relatively Strongest At Typical Agentic Coding, Front End Work and 3D