AI Paper › 西蒙·威利森 › 论文
2025:大语言模型之年 2025: The year in LLMs
西蒙·威利森 Simon Willison · · 2025-12-31 · Simon Willison's Weblog ↗
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→
摘要 · Abstract 由 Atlassian 赞助——给你的智能体一个计划,而不是一条提示。新的 Jira 功能为 AI 原生软件开发解锁了完整上下文。现在可以直接从 Jira 将任务分配给 Claude、Cursor 或 GitHub Copilot。了解更多。
Sponsored by: Atlassian — Give your agents a plan. Not a prompt. New Jira capabilities unlock full-context for AI-native software development. Assign tasks to Claude, Cursor, or GitHub Copilot, now directly from Jira. Learn more
核心贡献 · Key contributions 推理模型(基于可验证奖励的强化学习)成为标配,实现了多步骤工具使用和编码智能体。 Reasoning models (RLVR) became standard, enabling multi-step tool use and coding agents. Claude Code 等编码智能体年化收入达 10 亿美元,证明命令行工具可行。 Coding agents like Claude Code reached $1B run-rate revenue, proving CLI tools are viable. 中国开源权重模型(GLM-4.7、DeepSeek V3.2)排名领先,媲美前沿模型。 Chinese open-weight models (GLM-4.7, DeepSeek V3.2) topped rankings, rivaling frontier models. 提示驱动图像编辑(GPT-image-1、Nano Banana)一周内带来 1 亿 ChatGPT 注册。 Prompt-driven image editing (GPT-image-1, Nano Banana) drove 100M ChatGPT signups in a week. 模型在国际数学奥林匹克和国际大学生程序设计竞赛中获得金牌,展现了无需训练数据解决新问题的能力。 Models achieved gold in IMO and ICPC, demonstrating novel problem-solving without training data. 一致性测试套件使编码智能体仅通过提示就能移植复杂库。 Conformance suites enable coding agents to port complex libraries via prompting alone.
局限 · Limitations 推理模型在简单任务(如数字母)上仍表现不佳,实用性依赖工具。 Reasoning models still struggle with simple tasks like counting letters; utility is tool-driven. 编码智能体的 YOLO 模式存在安全风险;提示注入问题仍未解决。 YOLO mode in coding agents poses security risks; prompt injection remains unsolved. 中国模型未公开完整训练数据和代码,限制了可复现性。 Chinese models lack full training data and code, limiting reproducibility. 本地模型在编码智能体的可靠工具调用方面仍无法与云端模型匹敌。 Local models cannot yet match cloud models for reliable tool calling in coding agents. AI 数据中心能耗面临日益增长的公众反对;杰文斯悖论削弱了效率提升的效果。 AI data center energy use faces growing public opposition; Jevons paradox undermines efficiency gains.
论文章节 · Sections(共 31) Simon Willison 的博客 Simon Willison’s Weblog(https://simonwillison.net/) 2025:大语言模型年度回顾 2025: The year in LLMs “推理”之年 The year of “reasoning” 智能体之年 The year of agents 编程智能体与 Claude Code 之年 The year of coding agents and Claude Code 命令行上的 LLM 之年 The year of LLMs on the command-line YOLO 之年与偏差常态化 The year of YOLO and the Normalization of Deviance 每月 200 美元订阅的一年 The year of $200/month subscriptions 中国顶级开源权重模型之年 The year of top-ranked Chinese open weight models 长任务之年 The year of long tasks 提示驱动图像编辑之年 The year of prompt-driven image editing 模型在学术竞赛中夺金之年 The year models won gold in academic competitions Llama 迷失方向的一年 The year that Llama lost its way OpenAI 失去领先地位的一年 The year that OpenAI lost their lead 双子座之年 The year of Gemini 鹈鹕骑自行车之年 The year of pelicans riding bicycles 我构建了 110 个工具的那一年 The year I built 110 tools 告密之年! The year of the snitch! 氛围编程之年 The year of vibe coding MCP 的(唯一?)一年 The (only?) year of MCP 令人警惕的 AI 赋能浏览器之年 The year of alarmingly AI-enabled browsers 致命三重奏之年 The year of the lethal trifecta 在手机上编程的一年 The year of programming on my phone 合规套件之年 The year of conformance suites 本地模型变好的一年,但云端模型变得更好 The year local models got good, but cloud models got even better 垃圾内容之年 The year of slop 数据中心变得极不受欢迎的一年 The year that data centers got extremely unpopular 我的年度词汇 My own words of the year 2025 年总结 That’s a wrap for 2025 近期文章 More recent articles 月度简报 Monthly briefing
阅读逐段中英对照全文 →
© AI Paper · aipaper.jasonlin.tech — 著名 AI 学者的代表论文,逐段中英对照。论文正文/摘要版权归原作者与 arXiv,译文 AI 生成仅供参考,应权利人要求即下架(linzheng3535@gmail.com)。