AI as Normal Technology
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→本文探讨了将人工智能视为一种正常技术的观点,强调其与其他技术一样,具有渐进性、可预测性和可控性,而非神秘或不可预测的力量。
* * [](https://mastodon.online/@knightcolumbia) * [](https://bsky.app/profile/knightcolumbia.org) * [](https://www.linkedin.com/company/knightcolumbia/)
* * [](https://mastodon.online/@knightcolumbia)
* * [](https://mastodon.online/@knightcolumbia)
* [](https://bsky.app/profile/knightcolumbia.org)
* [](https://bsky.app/profile/knightcolumbia.org)
* [](https://www.linkedin.com/company/knightcolumbia/)
* [](https://www.linkedin.com/company/knightcolumbia/)
* [](https://www.youtube.com/channel/UCIp4_36yY_Sr3CwLvvxDjwQ)
* [](https://www.youtube.com/channel/UCIp4_36yY_Sr3CwLvvxDjwQ)
* [](https://www.instagram.com/knightinstitute/)
* [](https://www.instagram.com/knightinstitute/)
* [](https://www.tiktok.com/@knightinstitutecolumbia)
* [](https://www.tiktok.com/@knightinstitutecolumbia)
* [](https://knightcolumbia.org/rss?v=2)
* [](https://knightcolumbia.org/rss?v=2)
作者:Arvind Narayanan & Sayash Kapoor,2025 年 4 月 15 日
By Arvind Narayanan&Sayash KapoorApril 15, 2025
一个研究先进 AI 系统可能如何损害或帮助加强民主自由的项目。
A project studying how advanced AI systems may harm, or help strengthen, democratic freedoms
我们提出将人工智能(AI)视为“正常技术”的愿景。将 AI 视为正常并非低估其影响——即使是电力和互联网这类变革性的通用技术,在我们的概念中也是“正常的”。但这与关于 AI 未来的乌托邦和反乌托邦愿景形成对比,后者通常倾向于将其视为一个独立的物种,一个高度自主、可能超级智能的实体。
We articulate a vision of artificial intelligence (AI) as _normal technology_. To view AI as normal is not to understate its impact—even transformative, general-purpose technologies such as electricity and the internet are “normal” in our conception. But it is in contrast to both utopian and dystopian visions of the future of AI which have a common tendency to treat it akin to a separate species, a highly autonomous, potentially superintelligent entity.1 1. Nick Bostrom. 2012. The superintelligent will: Motivation and instrumental rationality in advanced artificial agents. Minds and Machines 22, 2 (May 2012), 71–85. https://doi:10.1007/s11023-012-9281-3; Nick Bostrom. 2017. Superintelligence: Paths, Dangers, Strategies (reprinted with corrections). Oxford University Press, Oxford, United Kingdom; Sam Altman, Greg Brockman, and Ilya Sutskever. 2023. Governance of Superintelligence (May 2023). https://openai.com/blog/governance-of-superintelligence; Shazeda Ahmed et al. 2023. Building the Epistemic Community of AI Safety. SSRN: Rochester, NY. doi:10.2139/ssrn.4641526.
“AI 是正常技术”这一陈述包含三层含义:对当前 AI 的_描述_、对 AI 可预见未来的_预测_,以及关于我们应如何对待它的_规范_。我们认为 AI 是一种我们可以且应该保持控制的工具,并主张这一目标不需要剧烈的政策干预或技术突破。我们认为,将 AI 视为类人智能目前既不准确,也无助于理解其社会影响,在我们对未来的愿景中也不太可能如此。
The statement “AI is normal technology” is three things: a _description_ of current AI, a _prediction_ about the foreseeable future of AI, and a _prescription_ about how we should treat it. We view AI as a tool that we can and should remain in control of, and we argue that this goal does not require drastic policy interventions or technical breakthroughs. We do not think that viewing AI as a humanlike intelligence is currently accurate or useful for understanding its societal impacts, nor is it likely to be in our vision of the future.2 2. This is different from the question of whether it is helpful for an individual user to conceptualize a specific AI system as a tool as opposed to a human-like entity such as an intern, a co-worker, or a tutor.
正常技术框架关乎技术与社会的关系。它拒绝技术决定论,尤其是拒绝将 AI 本身视为决定其未来的能动者。它借鉴了过去技术革命的经验教训,例如技术采用和扩散的缓慢与不确定性。它还强调了 AI 过去与未来轨迹在社会影响和制度塑造轨迹方面的连续性。
The normal technology frame is about the relationship between technology and society. It rejects technological determinism, especially the notion of AI itself as an agent in determining its future. It is guided by lessons from past technological revolutions, such as the slow and uncertain nature of technology adoption and diffusion. It also emphasizes continuity between the past and the future trajectory of AI in terms of societal impact and the role of institutions in shaping this trajectory.
在第一部分中,我们解释了为什么我们认为变革性的经济和社会影响将是缓慢的(以十年为尺度),并对 AI 方法、AI 应用和 AI 采纳做出了关键区分,认为三者发生在不同的时间尺度上。
In Part I, we explain why we think that transformative economic and societal impacts will be slow (on the timescale of decades), making a critical distinction between AI methods, AI applications, and AI adoption, arguing that the three happen at different timescales.
在第二部分中,我们讨论了在拥有先进 AI(但不是“超级智能”AI,我们认为后者通常被概念化时是不连贯的)的世界中,人类与 AI 之间可能的劳动分工。在这个世界中,控制权主要掌握在人和组织手中;事实上,人们工作中越来越大的比例是 AI 控制。
In Part II, we discuss a potential division of labor between humans and AI in a world with advanced AI (but not “superintelligent” AI, which we view as incoherent as usually conceptualized). In this world, control is primarily in the hands of people and organizations; indeed, a greater and greater proportion of what people do in their jobs is AI control.
在第三部分中,我们考察了将 AI 视为正常技术对 AI 风险的影响。我们分析了事故、军备竞赛、滥用和不对齐,并认为将 AI 视为正常技术会导致与将 AI 视为类人智能截然不同的缓解措施结论。
In Part III, we examine the implications of AI as normal technology for AI risks. We analyze accidents, arms races, misuse, and misalignment, and argue that viewing AI as normal technology leads to fundamentally different conclusions about mitigations compared to viewing AI as being humanlike.
当然,我们无法确定我们的预测,但我们旨在描述我们认为的中位结果。我们没有尝试量化概率,但我们试图做出能够告诉我们 AI 是否表现得像正常技术的预测。
Of course, we cannot be certain of our predictions, but we aim to describe what we view as the median outcome. We have not tried to quantify probabilities, but we have tried to make predictions that can tell us whether or not AI is behaving like normal technology.
在第四部分中,我们讨论了 AI 政策的影响。我们主张将减少不确定性作为首要政策目标,并将韧性作为应对灾难性风险的总体方法。我们认为,如果 AI 最终是正常技术,那么基于控制超级智能 AI 的困难而采取的剧烈干预措施实际上会使情况变得更糟——其负面影响很可能与先前在资本主义社会中部署的技术相似,例如不平等。
In Part IV, we discuss the implications for AI policy. We advocate for reducing uncertainty as a first-rate policy goal and resilience as the overarching approach to catastrophic risks. We argue that drastic interventions premised on the difficulty of controlling superintelligent AI will, in fact, make things much worse if AI turns out to be normal technology— the downsides of which will be likely to mirror those of previous technologies that are deployed in capitalistic societies, such as inequality.3 3. Daron Acemoglu and Simon Johnson. 2023. Power and Progress: Our Thousand-Year Struggle over Technology and Prosperity .PublicAffairs, New York, NY.
我们在第二部分中描述的世界是一个 AI 远比今天先进的世界。我们并不是说 AI 进步——或人类进步——将在那时停止。之后会发生什么?我们不知道。考虑这个类比:在第一次工业革命之初,尝试思考工业世界会是什么样子以及如何为之准备是有用的,但试图预测电力或计算机将是徒劳的。我们这里的练习类似。由于我们拒绝“快速起飞”情景,我们认为没有必要或有用地设想一个比我们尝试的更远的未来。如果我们在第二部分中描述的情景实现,我们将能够更好地预测和准备接下来的一切。
The world we describe in Part II is one in which AI is far more advanced than it is today. We are not claiming that AI progress—or human progress—will stop at that point. What comes after it? We do not know. Consider this analogy: At the dawn of the first Industrial Revolution, it would have been useful to try to think about what an industrial world would look like and how to prepare for it, but it would have been futile to try to predict electricity or computers. Our exercise here is similar. Since we reject “fast takeoff” scenarios, we do not see it as necessary or useful to envision a world further ahead than we have attempted to. If and when the scenario we describe in Part II materializes, we will be able to better anticipate and prepare for whatever comes next.
_致读者。_ _本文有一个不同寻常的目标:陈述一种世界观,而非捍卫一个命题。关于 AI 超级智能的文献浩如烟海。我们没有尝试对潜在的反对论点进行逐点回应,因为这会使论文篇幅增加数倍。本文仅仅是我们观点的初步阐述;我们计划在各种后续文章中详细阐述。_
_A note to readers._ _This essay has the unusual goal of stating a worldview rather than defending a proposition. The literature on AI superintelligence is copious. We have not tried to give a point-by-point response to potential counter arguments, as that would make the paper several times longer. This paper is merely the initial articulation of our views; we plan to elaborate on them in various follow ups._
图 1. 与其他通用目的技术一样,AI 的影响并非在方法和能力提升时实现,而是在这些改进转化为应用并扩散到经济生产部门时实现。4 4. Jeffrey Ding. 2024. Technology and the Rise of Great Powers: How Diffusion Shapes Economic Competition. Princeton University Press, Princeton. 每个阶段都存在速度限制。
_Figure 1. Like other general-purpose technologies, the impact of AI is materialized not when methods and capabilities improve, but when those improvements are translated into applications and are diffused through productive sectors of the economy.4 4. Jeffrey Ding. 2024. Technology and the Rise of Great Powers: How Diffusion Shapes Economic Competition. Princeton University Press, Princeton. There are speed limits at each stage._
AI 的进步会是渐进的,使人们和机构能够随着 AI 能力和采用的增加而适应,还是会出现导致大规模颠覆甚至技术奇点的跳跃?我们处理这个问题的方法是,将高影响任务与低影响任务分开分析,并首先分析 AI 采用和扩散的速度,然后再回到创新和发明的速度。
Will the progress of AI be gradual, allowing people and institutions to adapt as AI capabilities and adoption increase, or will there be jumps leading to massive disruption, or even a technological singularity? Our approach to this question is to analyze highly consequential tasks separately from less consequential tasks and to begin by analyzing the speed of adoption and diffusion of AI before returning to the speed of innovation and invention.
我们用“发明”指代开发新的 AI 方法(如大型语言模型)以提升 AI 执行各种任务的能力。“创新”指代开发消费者和企业可以使用的 AI 产品和应用。“采用”指代个人(或团队或公司)决定使用某项技术,而“扩散”指代采用水平提高的更广泛的社会过程。对于具有足够颠覆性的技术,扩散可能需要改变公司和组织的结构,以及社会规范和法律法规。
We use _invention_ to refer to the development of new AI methods—such as large language models—that improve AI’s capabilities to carry out various tasks. _Innovation_ refers to the development of products and applications using AI that consumers and businesses can use. _Adoption_ refers to the decision by an individual (or team or firm) to use a technology, whereas _diffusion_ refers to the broader social process through which the level of adoption increases. For sufficiently disruptive technologies, diffusion might require changes to the structure of firms and organizations, as well as to social norms and laws.
在论文《反对预测性优化》中,我们编制了一份包含约 50 种预测性优化应用的全面清单,即使用机器学习(ML)通过预测个人未来行为或结果来做出关于个人的决策。5 5. Angelina Wang 等人,2023 年。反对预测性优化:论优化预测准确性的决策算法的合法性。载于《2023 年 ACM 公平、问责与透明度会议论文集》(伊利诺伊州芝加哥:ACM,2023 年),第 626-26 页。doi:10.1145/3593013.3594030。这些应用大多用于对人们产生重要影响的决策,例如犯罪风险预测、保险风险预测或儿童虐待预测。
In the paper _Against Predictive Optimization_, we compiled a comprehensive list of about 50 applications of predictive optimization, namely the use of machine learning (ML) to make decisions about individuals by predicting their future behavior or outcomes.5 5. Angelina Wang et al. 2023. Against predictive optimization: On the legitimacy of decision-making algorithms that optimize predictive accuracy. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (Chicago, IL, USA: ACM, 2023), 626–26. doi:10.1145/3593013.3594030. Most of these applications, such as criminal risk prediction, insurance risk prediction, or child maltreatment prediction, are used to make decisions that have important consequences for people.
尽管这些应用已经激增,但有一个关键细微差别:在大多数情况下,使用的是几十年前的统计技术——简单、可解释的模型(主要是回归)和相对较小的人工设计特征集。更复杂的机器学习方法,如随机森林,很少使用,而现代方法,如 Transformer,则无处可寻。
While these applications have proliferated, there is a crucial nuance: In most cases, decades-old statistical techniques are used—simple, interpretable models (mostly regression) and relatively small sets of handcrafted features. More complex machine learning methods, such as random forests, are rarely used, and modern methods, such as transformers, are nowhere to be found.
换句话说,在这一广泛的领域集合中,AI 的扩散落后于创新数十年。一个主要原因是安全性——当模型更复杂且更难以理解时,很难在测试和验证过程中预见所有可能的部署条件。一个很好的例子是 Epic 的败血症预测工具,尽管在内部验证时看似具有高准确性,但在医院中表现却差得多,漏掉了三分之二的败血症病例,并用虚假警报压倒了医生。6 6. Casey Ross,2022 年。Epic 对有缺陷算法的改造表明 AI 监管是生死攸关的问题。STAT。https://www.statnews.com/2022/10/24/epic-overhaul-of-a-flawed-algorithm/。
In other words, in this broad set of domains, AI diffusion lags _decades_ behind innovation. A major reason is safety—when models are more complex and less intelligible, it is hard to anticipate all possible deployment conditions in the testing and validation process. A good example is Epic’s sepsis prediction tool which, despite having seemingly high accuracy when internally validated, performed far worse in hospitals, missing two thirds of sepsis cases and overwhelming physicians with false alerts.6 6. Casey Ross. 2022. Epic’s Overhaul of a Flawed Algorithm Shows Why AI Oversight Is a Life-or-Death Issue. STAT. https://www.statnews.com/2022/10/24/epic-overhaul-of-a-flawed-algorithm/.
Epic 的败血症预测工具失败的原因是,当你有具有无约束特征集的复杂模型时,很难捕捉到错误。7 7. Andrew Wong 等人,2021 年。广泛实施的专有败血症预测模型在住院患者中的外部验证。《美国医学会内科杂志》181, 8(2021 年 8 月),1065-70,https://doi:10.1001/jamainternmed.2021.2626。特别是,用于训练模型的特征之一是医生是否已经开了抗生素——用于治疗败血症。换句话说,在测试和验证期间,模型使用了来自未来的特征,依赖于一个因果依赖于结果变量。当然,这个特征在部署期间是不可用的。可解释性和审计方法无疑会改进,以便我们能够更好地捕捉这些问题,但我们还没有达到那个水平。
Epic’s sepsis prediction tool failed because of errors that are hard to catch when you have complex models with unconstrained feature sets.7 7. Andrew Wong et al. 2021. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine 181, 8 (August 2021), 1065–70, https://doi:10.1001/jamainternmed.2021.2626. In particular, one of the features used to train the model was whether a physician had already prescribed antibiotics —to treat sepsis. In other words, during testing and validation, the model was using a feature from the future, relying on a variable that was causally dependent on the outcome. Of course, this feature would not be available during deployment. Interpretability and auditing methods will no doubt improve so that we will get much better at catching these issues, but we are not there yet.
在生成式 AI 的情况下,即使是事后看来极其明显的失败,在测试期间也没有被捕捉到。一个例子是早期的必应聊天机器人“悉尼”,它在长时间对话中偏离了轨道;开发者显然没有预料到对话可能持续超过几轮。8 8. Kevin Roose,2023 年。与必应聊天机器人的对话让我深感不安。《纽约时报》(2023 年 2 月)。https://www.nytimes.com/2023/02/16/technology/bing-chatbot-microsoft-chatgpt.html。类似地,Gemini 图像生成器似乎从未在历史人物上进行过测试。9 9. Dan Milmo 和 Alex Hern,2024 年。“我们确实搞砸了”:为什么谷歌 AI 工具会生成冒犯性的历史图像?《卫报》(2024 年 3 月)。https://www.theguardian.com/technology/2024/mar/08/we-definitely-messed-up-why-did-google-ai-tool-make-offensive-historical-images。幸运的是,这些不是高度后果性的应用。
In the case of generative AI, even failures that seem extremely obvious in hindsight were not caught during testing. One example is the early Bing chatbot “Sydney” that went off the rails during extended conversations; the developers evidently did not anticipate that conversations could last for more than a handful of turns.8 8. Kevin Roose. 2023. A Conversation With Bing’s Chatbot Left Me Deeply Unsettled. The New York Times (February 2023). https://www.nytimes.com/2023/02/16/technology/bing-chatbot-microsoft-chatgpt.html. Similarly, the Gemini image generator was seemingly never tested on historical figures.9 9. Dan Milmo and Alex Hern. 2024. ‘We definitely messed up’: why did Google AI tool make offensive historical images? The Guardian (March 2024). https://www.theguardian.com/technology/2024/mar/08/we-definitely-messed-up-why-did-google-ai-tool-make-offensive-historical-images Fortunately, these were not highly consequential applications.
更多的实证工作将有助于理解各种应用中的创新-扩散滞后及其原因。但就目前而言,我们在先前工作中分析的证据与以下观点一致:在高度后果性的任务中,已经存在极其强大的与安全相关的速度限制。这些限制通常通过监管来执行,例如 FDA 对医疗设备的监督,以及较新的立法如欧盟 AI 法案,该法案对高风险 AI 提出了严格要求。10 10. Jamie Bernardi 等人,2024 年。社会对先进 AI 的适应。arXiv:2024 年 5 月。取自 http://arxiv.org/abs/2405.10295;器械与放射卫生中心,2024 年。用于改进和自动化医疗实践的新人工智能(AI)用途的监管评估。FDA(2024 年 6 月)。https://www.fda.gov/medical-devices/medical-device-regulatory-science-research-programs-conducted-osel/regulatory-evaluation-new-artificial-intelligence-ai-uses-improving-and-automating-medical-practices;“欧洲议会和理事会 2024 年 6 月 13 日第 2024/1689 号条例,制定人工智能统一规则并修订(EC)第 300/2008 号条例、(EU)第 167/2013 号条例、(EU)第 168/2013 号条例、(EU)第 2018/858 号条例、(EU)第 2018/1139 号条例和(EU)第 2019/2144 号条例以及指令 2014/90/EU、(EU)2016/797 和(EU)2020/1828(人工智能法案)(与欧洲经济区相关的文本)”,2024 年 6 月,http://data.europa.eu/eli/reg/2024/1689/oj/eng。事实上,存在(可信的)担忧,即现有对高风险 AI 的监管如此繁重,可能导致“失控的官僚主义”。11 11. Javier Espinoza,2024 年。欧洲匆忙制定 AI 规则。《金融时报》(2024 年 7 月)。https://www.ft.com/content/6cc7847a-2fc5-4df0-b113-a435d6426c81;Daniel E. Ho 和 Nicholas Bagley,2024 年。失控的官僚主义可能使 AI 的常见用途变得更糟,甚至包括邮件投递。《国会山报》(2024 年 1 月)。https://thehill.com/opinion/technology/4405286-runaway-bureaucracy-could-make-common-uses-of-ai-worse-even-mail-delivery/。因此,我们预测在高度后果性的任务中,缓慢扩散将继续成为常态。
More empirical work would be helpful for understanding the innovation-diffusion lag in various applications and the reasons for this lag. But, for now, the evidence that we have analyzed in our previous work is consistent with the view that there are already extremely strong safety-related speed limits in highly consequential tasks. These limits are often enforced through regulation, such as the FDA’s supervision of medical devices, as well as newer legislation such as the EU AI Act, which puts strict requirements on high-risk AI.10 10. Jamie Bernardi et al. 2024. Societal adaptation to advanced AI. arXiv: May 2024. Retrieved from http://arxiv.org/abs/2405.10295; Center for Devices and Radiological Health. 2024. Regulatory evaluation of new artificial intelligence (AI) uses for improving and automating medical practices. FDA (June 2024). https://www.fda.gov/medical-devices/medical-device-regulatory-science-research-programs-conducted-osel/regulatory-evaluation-new-artificial-intelligence-ai-uses-improving-and-automating-medical-practices; “Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 Laying down Harmonised Rules on Artificial Intelligence and Amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and Directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828 (Artificial Intelligence Act) (Text with EEA Relevance),” June 2024, http://data.europa.eu/eli/reg/2024/1689/oj/eng. In fact, there are (credible) concerns that existing regulation of high-risk AI is so onerous that it may lead to “runaway bureaucracy”.11 11. Javier Espinoza. 2024. Europe’s rushed attempt to set the rules for AI. Financial Times (July 2024). https://www.ft.com/content/6cc7847a-2fc5-4df0-b113-a435d6426c81; Daniel E. Ho and Nicholas Bagley. 2024. Runaway bureaucracy could make common uses of ai worse, even mail delivery. The Hill (January 2024). https://thehill.com/opinion/technology/4405286-runaway-bureaucracy-could-make-common-uses-of-ai-worse-even-mail-delivery/. Thus, we predict that slow diffusion will continue to be the norm in high-consequence tasks.
无论如何,当出现 AI 可以以高度后果性方式使用的新领域时,我们能够且必须对其进行监管。一个很好的例子是 2010 年的闪电崩盘,其中自动化高频交易被认为起到了一定作用。这导致了新的交易限制,例如熔断机制。12 12. Avanidhar Subrahmanyam,2013 年。算法交易、闪电崩盘和协调熔断机制。《伊斯坦布尔证券交易所评论》13, 3(2013 年 9 月),4-9。http://doi:10.1016/j.bir.2013.10.003。
At any rate, as and when new areas arise in which AI can be used in highly consequential ways, we can and must regulate them. A good example is the Flash Crash of 2010, in which automated high-frequency trading is thought to have played a part. This led to new curbs on trading, such as circuit breakers.12 12. Avanidhar Subrahmanyam. 2013. Algorithmic trading, the flash crash, and coordinated circuit breakers. Borsa Istanbul Review 13, 3 (September 2013), 4–9. http://doi:10.1016/j.bir.2013.10.003.
即使在安全关键领域之外,AI 的采用速度也比流行说法所暗示的要慢。例如,一项研究因发现 2024 年 8 月 40%的美国成年人使用生成式 AI 而成为头条新闻。13 13. Alexander Bick, Adam Blandin, and David J. Deming. 2024. The Rapid Adoption of Generative AI. National Bureau of Economic Research. 但是,由于大多数人使用频率不高,这仅转化为 0.5%-3.5%的工作时间(以及 0.125-0.875 个百分点的劳动生产率增长)。
Even outside of safety-critical areas, AI adoption is slower than popular accounts would suggest. For example, a study made headlines due to the finding that, in August 2024, 40% of U.S. adults used generative AI.13 13. Alexander Bick, Adam Blandin, and David J. Deming. 2024. The Rapid Adoption of Generative AI. National Bureau of Economic Research. But, because most people used it infrequently, this only translated to 0.5%-3.5% of work hours (and a 0.125-0.875 percentage point increase in labor productivity).
甚至不清楚今天的扩散速度是否比过去更快。上述研究报告称,生成式 AI 在美国的采用速度比个人电脑(PC)更快,在首个大众市场产品发布两年内,40%的美国成年人采用了生成式 AI,而 PC 在三年内仅为 20%。但这种比较没有考虑采用强度(使用小时数)的差异,也没有考虑购买 PC 与使用生成式 AI 的成本差异。14 14. Alexander Bick, Adam Blandin, and David J. Deming. 2024. The Rapid Adoption of Generative AI. National Bureau of Economic Research. 根据我们衡量采用的方式,生成式 AI 的采用速度很可能比 PC 慢得多。
It is not even clear if the speed of diffusion is greater today compared to the past. The aforementioned study reported that generative AI adoption in the U.S. has been faster than personal computer (PC) adoption, with 40% of U.S. adults adopting generative AI within two years of the first mass-market product release compared to 20 % within three years for PCs. But this comparison does not account for differences in the intensity of adoption (the number of hours of use) or the high cost of buying a PC compared to accessing generative AI.14 14. Alexander Bick, Adam Blandin, and David J. Deming. 2024. The Rapid Adoption of Generative AI. National Bureau of Economic Research. Depending on how we measure adoption, it is quite possible that the adoption of generative AI has been much slower than PC adoption.
鉴于数字技术可以同时覆盖数十亿设备,认为技术采用速度不一定在加快的说法可能令人惊讶(甚至明显错误)。但重要的是要记住,采用关乎软件使用,而非可用性。即使一款新的 AI 产品立即在网上免费发布,人们也需要时间改变工作流程和习惯,以利用新产品的优势并学会规避风险。
The claim that the speed of technology adoption is not necessarily increasing may seem surprising (or even obviously wrong) given that digital technology can reach billions of devices at once. But it is important to remember that adoption is about software use, not availability. Even if a new AI-based product is instantly released online for anyone to use for free, it takes time to for people to change their workflows and habits to take advantage of the benefits of the new product and to learn to avoid the risks.
因此,扩散速度本质上受限于个人、组织和制度适应技术的速度。这也是我们在过去的通用技术中看到的趋势:扩散需要数十年,而非数年。15 15. Benedict Evans. 2023. AI and the Automation of Work. https://www.ben-evans.com/benedictevans/2023/7/2/working-with-ai; Benedict Evans, 2023; Jeffrey Ding. 2024. Technology and the Rise of Great Powers: How Diffusion Shapes Economic Competition. Princeton University Press, Princeton.
Thus, the speed of diffusion is inherently limited by the speed at which not only individuals, but also organizations and institutions, can adapt to technology. This is a trend that we have also seen for past general-purpose technologies: Diffusion occurs over decades, not years.15 15. Benedict Evans. 2023. AI and the Automation of Work. https://www.ben-evans.com/benedictevans/2023/7/2/working-with-ai; Benedict Evans, 2023; Jeffrey Ding. 2024. Technology and the Rise of Great Powers: How Diffusion Shapes Economic Competition. Princeton University Press, Princeton.
例如,Paul A. David 对电气化的分析表明,生产力效益需要数十年才能完全实现。16 16. Paul A. David. 1990. The dynamo and the computer: an historical perspective on the modern productivity paradox. The American Economic Review 80, 2 (1990), 355–61. https://www.jstor.org/stable/2006600; Tim Harford. 2017. Why didn’t electricity immediately change manufacturing? (August 2017). https://www.bbc.com/news/business-40673694. 在爱迪生第一座中央发电站建成后的近 40 年里,电动发电机“无处不在,唯独不在生产力统计数据中”。17 17. Robert Solow as quoted in Paul A. David. 1990. The dynamo and the computer: an historical perspective on the modern productivity paradox. The American Economic Review 80, 2 (1990), Page 355. https://www.jstor.org/stable/2006600; Tim Harford. 2017. Why didn’t electricity immediately change manufacturing? (August 2017). https://www.bbc.com/news/business-40673694. 这不仅仅是技术惯性;工厂主发现电气化并未带来显著的效率提升。
As an example, Paul A. David’s analysis of electrification shows that the productivity benefits took decades to fully materialize.16 16. Paul A. David. 1990. The dynamo and the computer: an historical perspective on the modern productivity paradox. The American Economic Review 80, 2 (1990), 355–61. https://www.jstor.org/stable/2006600; Tim Harford. 2017. Why didn’t electricity immediately change manufacturing? (August 2017). https://www.bbc.com/news/business-40673694. Electric dynamos were “everywhere but in the productivity statistics” for nearly 40 years after Edison’s first central generating station. 17 17. Robert Solow as quoted in Paul A. David. 1990. The dynamo and the computer: an historical perspective on the modern productivity paradox. The American Economic Review 80, 2 (1990), Page 355. https://www.jstor.org/stable/2006600; Tim Harford. 2017. Why didn’t electricity immediately change manufacturing? (August 2017). https://www.bbc.com/news/business-40673694. This was not just technological inertia; factory owners found that electrification did not bring substantial efficiency gains.
最终实现收益的关键在于围绕生产线的逻辑重新设计整个工厂布局。除了工厂架构的变化,扩散还需要改变工作组织和过程控制,这些只能通过跨行业的实验来发展。这些变化使工人获得了更多的自主权和灵活性,同时也需要不同的招聘和培训实践。
What eventually allowed gains to be realized was redesigning the entire layout of factories around the logic of production lines. In addition to changes to factory architecture, diffusion also required changes to workplace organization and process control, which could only be developed through experimentation across industries. Workers had more autonomy and flexibility as a result of the changes, which also necessitated different hiring and training practices.
诚然,AI 技术发展迅速,但当我们区分 AI 方法与具体应用时,情况就远没有那么明朗了。
It is true that technical advances in AI have been rapid, but the picture is much less clear when we differentiate AI methods from applications.
我们将 AI 方法的进步概念化为一个通用性阶梯。18 18. Arvind Narayanan and Sayash Kapoor. 2024. AI Snake Oil: What Artificial Intelligence Can Do, What It Can’t, and How to Tell the Difference. Princeton University Press, Princeton, NJ. 这个阶梯上的每一级都建立在下一级之上,反映了向更通用计算能力的迈进。也就是说,它减少了让计算机执行新任务所需的编程工作量,并增加了在给定编程(或用户)工作量下可执行的任务集合;见图 2。例如,机器学习通过消除程序员为每个新任务设计逻辑的需要,仅需收集训练样本,从而提高了通用性。
We conceptualize progress in AI methods as a ladder of generality.18 18. Arvind Narayanan and Sayash Kapoor. 2024. AI Snake Oil: What Artificial Intelligence Can Do, What It Can’t, and How to Tell the Difference. Princeton University Press, Princeton, NJ. Each step on this ladder rests on the ones below it and reflects a move toward more general computing capabilities. That is, it reduces the programmer effort needed to get the computer to perform a new task and increases the set of tasks that can be performed with a given amount of programmer (or user) effort; see Figure 2. For example, machine learning increases generality by obviating the need for the programmer to devise logic to solve each new task, only requiring the collection of training examples instead.
人们很容易得出结论:随着我们构建更多阶梯层级,开发特定应用所需的工作量将持续减少,直到我们达到通用人工智能(AGI)——通常被设想为一种开箱即用、无所不能的 AI 系统,从而完全无需开发应用。
It is tempting to conclude that the effort required to develop specific applications will keep decreasing as we build more rungs of the ladder until we reach artificial general intelligence, often conceptualized as an AI system that can do everything out of the box, obviating the need to develop applications altogether.
在某些领域,我们确实看到了应用开发工作量减少的趋势。在自然语言处理中,大语言模型使得开发语言翻译应用变得相对简单。再比如游戏:AlphaZero 通过自我对弈,仅凭游戏描述和足够的算力,就能学会下棋并超越人类——这与过去游戏程序的开发方式有天壤之别。
In some domains, we are indeed seeing this trend of decreasing application development effort. In natural language processing, large language models have made it relatively trivial to develop a language translation application. Or consider games: AlphaZero can learn to play games such as chess better than any human through self-play given little more than a description of the game and enough computing power—a far cry from how game-playing programs used to be developed.
_图 2:计算的通用性阶梯。对于某些任务,阶梯层级越高,让计算机执行新任务所需的编程工作量就越少,并且在给定编程(或用户)工作量下可执行的任务就越多。_ 19 19. Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25 (2012); Harris Drucker, Donghui Wu, and Vladimir N. Vapnik. 1999. Support vector machines for spam categorization. IEEE Transactions on Neural Networks 10, 5 (September 1999), 1048–54. http://doi:10.1109/72.788645; William D. Smith. 1964. New I.B.M, System 360 can serve business, science and government; I.B.M. Introduces a computer it says tops output of biggest. The New York Times April 1964. https://www.nytimes.com/1964/04/08/archives/new-ibm-system-360-can-serve-business-science-and-government-ibm.html; Special to THE NEW YORK TIMES. Algebra machine spurs research calling for long calculations; Harvard receives today device to solve in hours problems taking so much time they have never been worked out. The New York Times (August 1944). https://www.nytimes.com/1944/08/07/archives/algebra-machine-spurs-research-calling-for-long-calculations.html; Herman Hollerith. 1894. The electrical tabulating machine. Journal of the Royal Statistical Society 57, 4 (December 1894), 678. http://doi:10.2307/2979610.
_Figure 2: The Ladder of Generality in Computing. For some tasks, higher ladder rungs require less programmer effort to get a computer to perform a new task, and more tasks can be performed with a given amount of programmer (or user) effort._ 19 19. Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25 (2012); Harris Drucker, Donghui Wu, and Vladimir N. Vapnik. 1999. Support vector machines for spam categorization. IEEE Transactions on Neural Networks 10, 5 (September 1999), 1048–54. http://doi:10.1109/72.788645; William D. Smith. 1964. New I.B.M, System 360 can serve business, science and government; I.B.M. Introduces a computer it says tops output of biggest. The New York Times April 1964. https://www.nytimes.com/1964/04/08/archives/new-ibm-system-360-can-serve-business-science-and-government-ibm.html; Special to THE NEW YORK TIMES. Algebra machine spurs research calling for long calculations; Harvard receives today device to solve in hours problems taking so much time they have never been worked out. The New York Times (August 1944). https://www.nytimes.com/1944/08/07/archives/algebra-machine-spurs-research-calling-for-long-calculations.html; Herman Hollerith. 1894. The electrical tabulating machine. Journal of the Royal Statistical Society 57, 4 (December 1894), 678. http://doi:10.2307/2979610.
然而,在那些难以模拟且错误代价高昂的高风险现实应用中,情况并非如此。以自动驾驶汽车为例:其发展轨迹在许多方面与 AlphaZero 的自我对弈相似——技术进步使其能在更真实的条件下行驶,从而收集更好或更真实的数据,进而推动技术改进,形成反馈循环。但这一过程耗时二十多年,而非 AlphaZero 的几小时,因为安全考虑限制了每次迭代相对于前一次的扩展程度。20 20. Mohammad Musa, Tim Dawkins, and Nicola Croce. 2019. This is the next step on the road to a safe self-driving future. World Economic Forum (December 2019). https://www.weforum.org/stories/2019/12/the-key-to-a-safe-self-driving-future-lies-in-sharing-data/; Louise Zhang. 2023. Cruise’s Safety Record Over 1 Million Driverless Miles. Cruise (April 2023). https://web.archive.org/web/20230504102309/https://getcruise.com/news/blog/2023/cruises-safety-record-over-one-million-driverless-miles/
However, this has not been the trend in highly consequential, real-world applications that cannot easily be simulated and in which errors are costly. Consider self-driving cars: In many ways, the trajectory of their development is similar to AlphaZero’s self-play—improving the tech allowed them to drive in more realistic conditions, which enabled the collection of better and/or more realistic data, which in turn led to improvements in the tech, completing the feedback loop. But this process took over two decades instead of a few hours in the case of AlphaZero because safety considerations put a limit on the extent to which each iteration of this loop could be scaled up compared to the previous one.20 20. Mohammad Musa, Tim Dawkins, and Nicola Croce. 2019. This is the next step on the road to a safe self-driving future. World Economic Forum (December 2019). https://www.weforum.org/stories/2019/12/the-key-to-a-safe-self-driving-future-lies-in-sharing-data/; Louise Zhang. 2023. Cruise’s Safety Record Over 1 Million Driverless Miles. Cruise (April 2023). https://web.archive.org/web/20230504102309/https://getcruise.com/news/blog/2023/cruises-safety-record-over-one-million-driverless-miles/
这种“能力-可靠性差距”反复出现。它一直是构建能够自动化现实任务的实用 AI“智能体”的主要障碍。21 21. Arvind Narayanan and Sayash Kapoor. 2024. AI companies are pivoting from creating gods to building products. Good. AI Snake Oil newsletter. https://www.aisnakeoil.com/p/ai-companies-are-pivoting-from-creating. 需要明确的是,许多设想使用智能体的任务,如预订旅行或提供客户服务,其后果远不如驾驶严重,但成本仍然很高,以至于让智能体从现实经验中学习并非易事。
This “capability-reliability gap” shows up over and over. It has been a major barrier to building useful AI “agents” that can automate real-world tasks.21 21. Arvind Narayanan and Sayash Kapoor. 2024. AI companies are pivoting from creating gods to building products. Good. AI Snake Oil newsletter. https://www.aisnakeoil.com/p/ai-companies-are-pivoting-from-creating. To be clear, many tasks for which the use of agents is envisioned, such as booking travel or providing customer service, are far less consequential than driving, but still costly enough that having agents learn from real-world experiences is not straightforward.
在非安全关键的应用中也存在障碍。通常,组织中的许多知识是隐性的,并未被记录下来,更不用说以可被动学习的形式存在了。这意味着这些发展反馈循环必须在每个行业中发生,对于更复杂的任务,甚至可能需要在不同组织中分别进行,从而限制了快速并行学习的机会。并行学习受限的其他原因包括隐私问题:组织和个人可能不愿与 AI 公司共享敏感数据,而法规可能限制在医疗等情境下与第三方共享的数据类型。
Barriers also exist in non-safety-critical applications. In general, much knowledge is tacit in organizations and is not written down, much less in a form that can be learned passively. This means that these developmental feedback loops will have to happen in each sector and, for more complex tasks, may even need to occur separately in different organizations, limiting opportunities for rapid, parallel learning. Other reasons why parallel learning might be limited are privacy concerns: Organizations and individuals might be averse to sharing sensitive data with AI companies, and regulations might limit what kinds of data can be shared with third parties in contexts such as healthcare.
AI 中的“苦涩教训”是,利用算力增长的通用方法最终会大幅超越利用人类领域知识的方法。22 22. Rich Sutton. 2019. The Bitter Lesson (March 2019). http://www.incompleteideas.net/IncIdeas/BitterLesson.html. 这是关于_方法_的宝贵观察,但常被误解为涵盖应用开发。在基于 AI 的产品开发背景下,苦涩教训从未接近过正确。23 23. Arvind Narayanan and Sayash Kapoor. 2024. AI companies are pivoting from creating gods to building products. Good. AI Snake Oil newsletter. https://www.aisnakeoil.com/p/ai-companies-are-pivoting-from-creating 以社交媒体上的推荐系统为例:它们由(日益通用的)机器学习模型驱动,但这并未消除手动编写业务逻辑、前端和其他组件的需求,这些组件加起来可能多达数百万行代码。
The “bitter lesson” in AI is that general methods that leverage increases in computational power eventually surpass methods that utilize human domain knowledge by a large margin.22 22. Rich Sutton. 2019. The Bitter Lesson (March 2019). http://www.incompleteideas.net/IncIdeas/BitterLesson.html. This is a valuable observation about _methods_, but it is often misinterpreted to encompass application development. In the context of AI-based product development, the bitter lesson has never been even close to true.23 23. Arvind Narayanan and Sayash Kapoor. 2024. AI companies are pivoting from creating gods to building products. Good. AI Snake Oil newsletter. https://www.aisnakeoil.com/p/ai-companies-are-pivoting-from-creating Consider recommender systems on social media: They are powered by (increasingly general) machine learning models, but this has not obviated the need for manual coding of the business logic, the frontend, and other components which, together, can comprise on the order of a million lines of code.
当我们需要超越 AI 从现有人类知识中学习时,还会出现更多限制。24 24. Melanie Mitchell. 2021. Why AI is harder than we think. arXiv preprint. Retrieved from http://arxiv.org/abs/2104.12871, April 2021), https://arxiv.org/abs/2104.12871. 我们最有价值的知识类型包括科学和社会科学知识,它们通过技术和大规模社会组织(如政府)推动了文明的进步。AI 要拓展这类知识的边界需要什么?很可能需要与人类或组织进行互动,甚至进行实验,范围从药物测试到经济政策。在这里,由于实验的社会成本,知识获取的速度存在硬性限制。社会可能不会(也不应该)允许为 AI 开发而快速扩展实验。
Further limits arise when we need to go beyond AI learning from existing human knowledge.24 24. Melanie Mitchell. 2021. Why AI is harder than we think. arXiv preprint. Retrieved from http://arxiv.org/abs/2104.12871, April 2021), https://arxiv.org/abs/2104.12871. Some of our most valuable types of knowledge are scientific and social-scientific, and have allowed the progress of civilization through technology and large-scale social organizations (e.g., governments). What will it take for AI to push the boundaries of such knowledge? It will likely require interactions with, or even experiments on, people or organizations, ranging from drug testing to economic policy. Here, there are hard limits to the speed of knowledge acquisition because of the social costs of experimentation. Societies probably will not (and should not) allow the rapid scaling of experiments for AI development.
方法-应用的区别对我们如何衡量和预测 AI 进展具有重要意义。AI 基准测试对于衡量方法的进展是有用的;不幸的是,它们常常被误解为衡量应用的进展,这种混淆一直是关于即将到来的经济转型的许多炒作背后的驱动因素。
The methods-application distinction has important implications for how we measure and forecast AI progress. AI benchmarks are useful for measuring progress in methods; unfortunately, they have often been misunderstood as measuring progress in applications, and this confusion has been a driver of much hype about imminent economic transformation.
例如,尽管 GPT-4 据报道在律师资格考试中取得了前 10%的成绩,但这几乎无法告诉我们 AI 从事法律实践的能力。25 25. Josh Achiam et al. 2023. GPT-4 technical report. arXiv preprintarXiv: 2303.08774; Peter Henderson et al. 2024. Rethinking machine learning benchmarks in the context of professional codes of conduct. In Proceedings of the Symposium on Computer Science and Law; Varun Magesh et al. 2024. Hallucination-free? Assessing the reliability of leading AI legal research tools. arXiv preprint arXiv: 2405.20362; Daniel N. Kluttz and Deirdre K. Mulligan. 2019. Automated decision support technologies and the legal profession. Berkeley Technology Law Journal 34, 3 (2019), 853–90; Inioluwa Deborah Raji, Roxana Daneshjou, and Emily Alsentzer. 2025. It’s time to bench the medical exam benchmark. NEJM AI 2, 2 (2025). 律师资格考试过分强调学科知识,而低估了实际技能,这些技能在标准化、计算机管理的考试中更难衡量。换句话说,它恰恰强调了语言模型擅长的——检索和应用记忆信息。
For example, while GPT-4 reportedly achieved scores in the top 10% of bar exam test takers, this tells us remarkably little about AI’s ability to practice law.25 25. Josh Achiam et al. 2023. GPT-4 technical report. arXiv preprintarXiv: 2303.08774; Peter Henderson et al. 2024. Rethinking machine learning benchmarks in the context of professional codes of conduct. In Proceedings of the Symposium on Computer Science and Law; Varun Magesh et al. 2024. Hallucination-free? Assessing the reliability of leading AI legal research tools. arXiv preprint arXiv: 2405.20362; Daniel N. Kluttz and Deirdre K. Mulligan. 2019. Automated decision support technologies and the legal profession. Berkeley Technology Law Journal 34, 3 (2019), 853–90; Inioluwa Deborah Raji, Roxana Daneshjou, and Emily Alsentzer. 2025. It’s time to bench the medical exam benchmark. NEJM AI 2, 2 (2025). The bar exam overemphasizes subject-matter knowledge and under-emphasizes real-world skills that are far harder to measure in a standardized, computer-administered format. In other words, it emphasizes precisely what language models are good at—retrieving and applying memorized information.
更广泛地说,那些会导致法律行业最显著变化的任务也是最难评估的。对于像按法律领域分类法律请求这样的任务,评估是直接的,因为有明确的正确答案。但对于涉及创造力和判断的任务,比如准备法律文件,没有单一的正确答案,理性的人可能对策略有不同意见。这些后一类任务正是如果实现自动化,将对行业产生最深远影响的任务。26 26. Sayash Kapoor, Peter Henderson, and Arvind Narayanan. Promises and pitfalls of artificial intelligence for legal applications. Journal of Cross-Disciplinary Research in Computational Law 2, 2 (May 2024), Article 2. https://journalcrcl.org/crcl/article/view/62.
More broadly, tasks that would lead to the most significant changes to the legal profession are also the hardest ones to evaluate. Evaluation is straightforward for tasks like categorizing legal requests by area of law because there are clear correct answers. But for tasks that involve creativity and judgment, like preparing legal filings, there is no single correct answer, and reasonable people can disagree about strategy. These latter tasks are precisely the ones that, if automated, would have the most profound impact on the profession.26 26. Sayash Kapoor, Peter Henderson, and Arvind Narayanan. Promises and pitfalls of artificial intelligence for legal applications. Journal of Cross-Disciplinary Research in Computational Law 2, 2 (May 2024), Article 2. https://journalcrcl.org/crcl/article/view/62.
这一观察绝不限于法律领域。另一个例子是 AI 明显擅长的独立编程问题与难以衡量其影响但似乎有限的现实世界软件工程之间的差距。27 27. Hamel Husain, Isaac Flath, and Johno Whitaker. Thoughts on a month with Devin. Answer.AI (2025). answer.ai/posts/2025-01-08-devin.html. 即使是备受推崇的、超越玩具问题的编程基准测试,也必然为了量化和使用公开数据进行自动评估而忽略现实世界软件工程的许多维度。28 28. Ehud Reiter. 2025. Do LLM Coding Benchmarks Measure Real-World Utility?. https://ehudreiter.com/2025/01/13/do-llm-coding-benchmarks-measure-real-world-utility/.
This observation is in no way limited to law. Another example is the gap between self-contained coding problems at which AI demonstrably excels, and real-world software engineering in which its impact is hard to measure but appears to be modest.27 27. Hamel Husain, Isaac Flath, and Johno Whitaker. Thoughts on a month with Devin. Answer.AI (2025). answer.ai/posts/2025-01-08-devin.html. Even highly regarded coding benchmarks that go beyond toy problems must necessarily ignore many dimensions of real-world software engineering in the interest of quantification and automated evaluation using publicly available data.28 28. Ehud Reiter. 2025. Do LLM Coding Benchmarks Measure Real-World Utility?. https://ehudreiter.com/2025/01/13/do-llm-coding-benchmarks-measure-real-world-utility/.
这种模式反复出现:一项任务越容易通过基准测试衡量,它就越不可能代表定义专业实践的那种复杂、情境化的工作。通过过度依赖能力基准测试来理解 AI 进展,AI 社区一直高估该技术的现实世界影响。
This pattern appears repeatedly: The easier a task is to measure via benchmarks, the less likely it is to represent the kind of complex, contextual work that defines professional practice. By focusing heavily on capability benchmarks to inform our understanding of AI progress, the AI community consistently overestimates the real-world impact of the technology.
这是一个“构念效度”问题,指的是测试是否真正衡量了它意图衡量的东西。29 29. Deborah Raji et al. 2021. AI and the everything in the whole wide world benchmark. In Proceedings of the Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks, vol. 1. https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/084b6fbb10729ed4da8c3d3f5a3ae7c9-Abstract-round2.html; Rachel Thomas and David Uminsky. 2020. The problem with metrics is a fundamental problem for AI. arXiv preprint. Retrieved from https://arxiv.org/abs/2002.08512v1. 衡量潜在应用现实有用性的唯一可靠方法是实际构建该应用,然后在现实场景中与专业人士一起测试(根据预期用途,替代或增强他们的劳动)。这类“提升”研究通常确实表明,许多职业的专业人士受益于现有的 AI 系统,但这种收益通常是适度的,更多的是增强而非替代,这与基于静态基准测试(如考试)可能得出的结论截然不同。30 30. Ashwin Nayak et al. 2023. Comparison of history of present illness summaries generated by a chatbot and senior internal medicine residents. JAMA Internal Medicine 183, 9 (September 2023), 1026–27. http://doi:10.1001/jamainternmed.2023.2561; Shakked Noy and Whitney Zhang. 2023. Experimental evidence on the productivity effects of generative artificial intelligence. Science 381, 6654 (July 2023), 187–92. http://doi:10.1126/science.adh2586; Fabrizio Dell’Acqua et al., “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality,” Harvard Business School Technology & Operations Mgt. Unit Working Paper, no. 24–13 (2023). (少数职业如文案和翻译经历了大量失业 31 31. Pranshu Verma and Gerrit De Vynck. 2023. ChatGPT took their jobs. Now they walk dogs and fix air conditioners. Washington Post (June 2023). https://www.washingtonpost.com/technology/2023/06/02/ai-taking-jobs/.)。
This is a problem of ‘construct validity,’ which refers to whether a test actually measures what it is intended to measure.29 29. Deborah Raji et al. 2021. AI and the everything in the whole wide world benchmark. In Proceedings of the Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks, vol. 1. https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/084b6fbb10729ed4da8c3d3f5a3ae7c9-Abstract-round2.html; Rachel Thomas and David Uminsky. 2020. The problem with metrics is a fundamental problem for AI. arXiv preprint. Retrieved from https://arxiv.org/abs/2002.08512v1. The only sure way to measure real-world usefulness of a potential application is to actually build the application and to then test it with professionals in realistic scenarios (either substituting or augmenting their labor, depending on the intended use). Such ‘uplift’ studies generally do show that professionals in many occupations benefit from existing AI systems, but this benefit is typically modest and is more about augmentation than substitution, a radically different picture from what one might conclude based on static benchmarks like exams 30 30. Ashwin Nayak et al. 2023. Comparison of history of present illness summaries generated by a chatbot and senior internal medicine residents. JAMA Internal Medicine 183, 9 (September 2023), 1026–27. http://doi:10.1001/jamainternmed.2023.2561; Shakked Noy and Whitney Zhang. 2023. Experimental evidence on the productivity effects of generative artificial intelligence. Science 381, 6654 (July 2023), 187–92. http://doi:10.1126/science.adh2586; Fabrizio Dell’Acqua et al., “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality,” Harvard Business School Technology & Operations Mgt. Unit Working Paper, no. 24–13 (2023). (a small number of occupations such as copywriters and translators have seen substantial job losses 31 31. Pranshu Verma and Gerrit De Vynck. 2023. ChatGPT took their jobs. Now they walk dogs and fix air conditioners. Washington Post (June 2023). https://www.washingtonpost.com/technology/2023/06/02/ai-taking-jobs/.).
总之,虽然基准测试对于跟踪 AI 方法的进展很有价值,但我们应该关注其他类型的指标来跟踪 AI 的影响(图 1)。在衡量采用情况时,我们必须考虑 AI 使用的强度。应用的类型也很重要:增强与替代,以及高后果与低后果。
In conclusion, while benchmarks are valuable for tracking progress in AI methods, we should look at other kinds of metrics to track AI impacts (Figure 1). When measuring adoption, we must take into account the intensity of AI use. The type of application is also important: Augmentation versus substitution and high-consequence versus low-consequence.
确保构念效度的困难不仅困扰着基准测试,也困扰着预测,这是人们试图评估(未来)AI 影响的另一种主要方式。避免模糊结果对于确保有效预测极为重要。预测社区实现这一点的方式是将里程碑定义为相对狭窄的技能,例如考试成绩。例如,Metaculus 关于“人机智能对等”的问题被定义为在数学、物理和计算机科学考试问题上的表现。基于这一定义,预测者预测到 2040 年实现“人机智能对等”的概率为 95%也就不足为奇了。32 32. Metaculus. 2024. Will there be human-machine intelligence parity before 2040? https://www.metaculus.com/questions/384/humanmachine-intelligence-parity-by-2040/.
The difficulty of ensuring construct validity afflicts not only benchmarking, but also forecasting, which is another major way in which people try to assess (future) AI impacts. It is extremely important to avoid ambiguous outcomes to ensure effective forecasting. The way that the forecasting community accomplishes this is by defining milestones in terms of relatively narrow skills, such as exam performance. For instance, the Metaculus question on “human-machine intelligence parity” is defined in terms of performance on exam questions in math, physics, and computer science. Based on this definition, it is not surprising that forecasters predict a 95% chance of achieving “human-machine intelligence parity” by 2040. 32 32. Metaculus. 2024. Will there be human-machine intelligence parity before 2040? https://www.metaculus.com/questions/384/humanmachine-intelligence-parity-by-2040/.
不幸的是,这个定义被稀释得如此严重,以至于对理解 AI 的影响意义不大。正如我们上面在法律和其他专业基准测试中看到的,AI 在考试中的表现构念效度如此之低,以至于甚至无法让我们预测 AI 是否会取代专业工作者。
Unfortunately, this definition is so watered down that it does not mean much for understanding the impacts of AI. As we saw above with legal and other professional benchmarks, AI performance on exams has so little construct validity that it does not even allow us to predict whether AI will replace professional workers.
一种认为 AI 发展可能带来突然、剧烈经济影响的论点是,通用性的提升可能导致经济中大量任务变得可自动化。这与通用人工智能(AGI)的一种定义相关——一个能够执行所有经济上有价值任务的统一系统。
One argument for why AI development may have sudden, drastic economic impacts is that an increase in generality may lead to a wide swath of tasks in the economy becoming automatable. This is related to one definition of artificial general intelligence (AGI)—a unified system that is capable of performing all economically valuable tasks.
根据正常技术观,这种突然的经济影响是不可能的。在前几节中,我们讨论了一个原因:AI 方法的突然改进是可能的,但不会直接转化为经济影响,后者需要创新(在应用开发的意义上)和扩散。
According to the normal technology view, such sudden economic impacts are implausible. In the previous sections, we discussed one reason: Sudden improvements in AI methods are certainly possible but do not directly translate to economic impacts, which require innovation (in the sense of application development) and diffusion.
创新和扩散发生在一个反馈循环中。在安全关键应用中,这个反馈循环总是缓慢的,但即使超出安全领域,也有许多原因使其可能缓慢。对于过去的通用技术,如电力、计算机和互联网,各自的反馈循环历经数十年才展开,我们应该预期 AI 也会如此。
Innovation and diffusion happen in a feedback loop. In safety-critical applications, this feedback loop is always slow, but even beyond safety, there are many reasons why it is likely to be slow. With past general-purpose technologies such as electricity, computers, and the internet, the respective feedback loops unfolded over several decades, and we should expect the same to happen with AI as well.
另一个支持经济影响渐进的论点是:一旦我们自动化了某件事,其生产成本和价值相对于人类劳动力往往会随时间急剧下降。随着自动化程度的提高,人类将适应,并专注于尚未自动化的任务,也许是今天不存在的任务(在第二部分中,我们描述了这些任务可能是什么样子)。
Another argument for gradual economic impacts: Once we automate something, its cost of production, and its value, tend to drop drastically over time compared to the cost of human labor. As automation increases, humans will adapt, and will focus on tasks that are not yet automated, perhaps tasks that do not exist today (in Part II we describe what those might look like).
这意味着 AGI 的目标将随着自动化的增加不断重新定义哪些任务具有经济价值而不断后移。即使人类今天所做的每一项任务有一天都可能被自动化,但这并不意味着人类劳动力将是多余的。
This means that the goalpost of AGI will continually move further away as increasing automation redefines which tasks are economically valuable. Even if every task that humans do _today_ might be automated one day, this does not mean that human labor will be superfluous.
所有这些都表明,在某个特定时刻自动化经济中大量任务的可能性不大。这也意味着强大 AI 的影响将在不同行业以不同的时间尺度被感受到。
All of this points away from the likelihood of the automation of a vast swath of the economy at a particular moment in time. It also implies that the impacts of powerful AI will be felt on different timescales in different sectors.
我们关于 AI 影响缓慢的论点基于创新-扩散反馈循环,即使 AI 方法的进展可以任意加速,这一论点也适用。我们认为收益和风险主要源于 AI 的部署而非开发;因此,AI 方法进展的速度与影响问题并不直接相关。尽管如此,讨论同样适用于方法开发的速度限制仍有必要。
Our argument for the slowness of AI impact is based on the innovation-diffusion feedback loop, and is applicable even if progress in AI methods can be arbitrarily sped up. We see both benefits and risks as arising primarily from AI deployment rather than from development; thus, the speed of progress in AI methods is not directly relevant to the question of impacts. Nonetheless, it is worth discussing speed limits that also apply to methods development.
AI 研究的产出呈指数级增长,arXiv 上 AI/ML 论文的发表率翻倍时间不到两年。33 33. Mario Krenn 等人,2023 年。基于机器学习链接预测在指数增长知识网络中预测人工智能的未来。《自然机器智能》5, 11 (2023), 1326–35。但这种数量的增长如何转化为进展尚不清楚。进展的一个衡量标准是核心思想的更替率。不幸的是,纵观其历史,AI 领域表现出对流行思想的高度趋同,而对非流行思想的探索(事后看来)不足。一个显著的例子是神经网络研究被边缘化了几十年。
The production of AI research has been increasing exponentially, with the rate of publication of AI/ML papers on arXiv exhibiting a doubling time under two years.33 33. Mario Krenn et al. 2023. Forecasting the future of artificial intelligence with machine learning-based link prediction in an exponentially growing knowledge network. Nature Machine Intelligence 5, 11 (2023), 1326–35. But it is not clear how this increase in volume translates to progress. One measure of progress is the rate of turnover of central ideas. Unfortunately, throughout its history, the AI field has shown a high degree of herding around popular ideas, and inadequate (in retrospect) levels of exploration of unfashionable ones. A notable example is the sidelining of research on neural networks for many decades.
当前时代是否不同?尽管思想以越来越快的速度增量积累,但它们是否在取代已有思想?Transformer 架构在过去十年的大部分时间里一直是主导范式,尽管其众所周知的局限性。通过分析 241 个学科中超过十亿次引用,Johan S.G. Chu 和 James A. Evans 表明,在论文数量更高的领域,新思想突破反而更难,而不是更容易。这导致了“经典固化”。34 34. Johan S.G. Chu 和 James A. Evans,2021 年。大科学领域中经典进展的放缓。《美国国家科学院院刊》118, 41 (2021), e2021636118。也许这种描述适用于 AI 方法研究的当前状态。
Is the current era different? Although ideas incrementally accrue at increasing rates, are they turning over established ones? The transformer architecture has been the dominant paradigm for most of the last decade, despite its well-known limitations. By analyzing over a billion citations in 241 subjects, Johan S.G. Chu & James A. Evans showed that, in fields in which the volume of papers is higher, it is harder, not easier, for new ideas to break through. This leads to an “ossification of canon.”34 34. Johan S.G. Chu and James A. Evans. 2021. Slowed Canonical Progress in Large Fields of Science. Proceedings of the National Academy of Sciences 118, 41 (2021), e2021636118. Perhaps this description applies to the current state of AI methods research.
许多其他速度限制也是可能的。历史上,深度神经网络技术部分因硬件(尤其是图形处理单元)不足而受阻。计算和成本限制仍然与新的范式相关,包括推理时 Scaling。新的放缓可能出现:最近的迹象表明,行业正从开放知识共享的文化转向。
Many other speed limits are possible. Historically, deep neural network technology was partly held back due to the inadequacy of hardware, particularly Graphics Processing Units. Computational and cost limits continue to be relevant to new paradigms, including inference-time scaling. New slowdowns may emerge: Recent signs point to a shift away from the culture of open knowledge sharing in the industry.
AI 进行的 AI 研究能否带来缓解仍有待观察。也许方法的递归自我改进是可能的,导致方法无限加速。但请注意,AI 开发已经严重依赖 AI。更可能的是,我们将继续看到自动化在 AI 开发中的作用逐渐增强,而不是一个单一的、不连续的递归自我改进实现时刻。35 35. Timothy B. Lee,2024 年。AI 厄运的预测太像好莱坞电影情节。https://www.understandingai.org/p/predictions-of-ai-doom-are-too-much
It remains to be seen if AI-conducted AI research can offer a reprieve. Perhaps recursive self-improvement in methods is possible, resulting in unbounded speedups in methods. But note that AI development already relies heavily on AI. It is more likely that we will continue to see a gradual increase in the role of automation in AI development than a singular, discontinuous moment when recursive self-improvement is achieved.35 35. Timothy B. Lee. 2024. Predictions of AI doom are too much like Hollywood movie plots. https://www.understandingai.org/p/predictions-of-ai-doom-are-too-much
之前,我们认为基准测试对 AI 应用有用性的描绘具有误导性。但它们也导致了关于方法进展速度的过度乐观。一个原因是,设计超越当前进展视野的基准测试很困难。图灵测试几十年来一直是 AI 的北极星,因为人们假设任何通过它的系统在重要方面都会像人类一样,并且我们能够使用这样的系统来自动化各种复杂任务。现在,大型语言模型可以说通过了它,但仅勉强满足测试背后的期望,其重要性已经减弱。36 36. Celeste Biever,2023 年。ChatGPT 打破了图灵测试——寻找评估 AI 新方法的竞赛开始。《自然》619, 7971 (2023 年 7 月), 686–89。http://doi:10.1038/d41586-023-02361-7; Melanie Mitchell,2024 年。图灵测试与我们不断变化的智能概念。《科学》385, 6710 (2024), eadq9356。http://doi:10.1126/science.adq9356。
Earlier, we argued that benchmarks give a misleading picture of the usefulness of AI applications. But they have arguably also led to overoptimism about the speed of methods progress. One reason is that it is hard to design benchmarks that make sense beyond the current horizon of progress. The Turing test was the north star of AI for many decades because of the assumption that any system that passed it would be humanlike in important ways, and that we would be able to use such a system to automate a variety of complex tasks. Now that large language models can arguably pass it while only weakly meeting the expectations behind the test, its significance has waned.36 36. Celeste Biever. 2023. ChatGPT broke the Turing Test — The race is on for new ways to assess AI. Nature 619, 7971 (July 2023), 686–89. http://doi:10.1038/d41586-023-02361-7; Melanie Mitchell. 2024. The Turing Test and our shifting conceptions of intelligence. Science 385, 6710 (2024), eadq9356. http://doi:10.1126/science.adq9356.
与登山类比很恰当。每当我们解决一个基准测试(到达我们认为是顶峰的地方),我们就会发现基准测试的局限性(意识到我们在一个“假顶峰”上),并构建一个新的基准测试(将目光投向我们现在认为是顶峰的地方)。这导致了“移动球门柱”的指责,但考虑到基准测试的内在挑战,这正是我们应该预期的。
An analogy with mountaineering is apt. Every time we solve a benchmark (reach what we thought was the peak), we discover limitations of the benchmark (realize that we’re on a ‘false summit’) and construct a new benchmark (set our sights on what we now think is the summit). This leads to accusations of ‘moving the goalposts’, but this is what we should expect given the intrinsic challenges of benchmarking.
AI 先驱认为 AI(我们现在所说的 AGI)的两大挑战是(我们现在所说的)硬件和软件。在建造了可编程机器之后,有一种明显的感觉,即 AGI 近在咫尺。1956 年达特茅斯会议的组织者希望通过一个“2 个月、10 人”的努力在实现目标上取得重大进展。37 37. John McCarthy, Marvin L. Minsky, Nathaniel Rochester 和 Claude E. Shannon,1955 年。关于人工智能的达特茅斯夏季研究项目的提案。http://jmc.stanford.edu/articles/dartmouth/dartmouth.pdf 今天,我们已经爬上了通用性阶梯的更多梯级。我们经常听到,构建 AGI 所需的一切就是 Scaling,或通用 AI 智能体,或样本高效学习。
AI pioneers considered the two big challenges of AI (what we now call AGI) to be (what we now call) hardware and software. Having built programmable machines, there was a palpable sense that AGI was close. The organizers of the 1956 Dartmouth conference hoped to make significant progress toward the goal through a “2-month, 10-man” effort.37 37. John McCarthy, Marvin L. Minsky, Nathaniel Rochester and Claude E. Shannon. 1955. A proposal for the dartmouth summer research project on artificial intelligence. http://jmc.stanford.edu/articles/dartmouth/dartmouth.pdf Today, we have climbed many more rungs on the ladder of generality. We often hear that all that is needed to build AGI is scaling, or generalist AI agents, or sample-efficient learning.
但有必要记住,看似一步之遥可能并非如此。例如,可能不存在一个单一的突破性算法能够在所有上下文中实现样本高效学习。事实上,大型语言模型中的上下文学习已经是“样本高效”的,但仅适用于有限的任务集。38 38. Changmao Li 和 Jeffrey Flanigan,2023 年。任务污染:语言模型可能不再是小样本学习。arXiv: 2023 年 12 月。检索自 doi:10.48550/arXiv.2312.16337。
But it is useful to bear in mind that what appears to be a single step might not be so. For example, there may not exist one single breakthrough algorithm that enables sample-efficient learning across all contexts. Indeed, in-context learning in large language models is already “sample efficient,” but only works for a limited set of tasks.38 38. Changmao Li and Jeffrey Flanigan. 2023. Task contamination: Language models may not be few-shot anymore. arXiv: December 2023. Retrieved from doi:10.48550/arXiv.2312.16337.
我们认为,依赖'智能'和'超级智能'这些模糊概念,已经模糊了我们清晰思考拥有先进 AI 世界的能力。通过将智能分解为不同的底层概念——能力和权力,我们反驳了在拥有'超级智能'AI 的世界中人类劳动将变得多余的观点,并提出了一种替代愿景。这也为我们第三部分讨论风险奠定了基础。
We argue that reliance on the slippery concepts of ‘intelligence’ and ‘superintelligence’ has clouded our ability to reason clearly about a world with advanced AI. By unpacking intelligence into distinct underlying concepts, capability and power, we rebut the notion that human labor will be superfluous in a world with ‘superintelligent’ AI, and present an alternative vision. This also lays the foundation for our discussion of risks in Part III.
AI 能否超越人类智能,如果能,能超越多少?根据一种流行的观点,其程度深不可测。这通常通过比较不同物种在智能谱系上的位置来描绘。
Can AI exceed human intelligence and, if so, by how much? According to a popular argument, unfathomably so. This is often depicted by comparing different species along a spectrum of intelligence.
_图 3. 通过递归自我改进 AI 实现智能爆炸是一个常见的担忧,通常像这样的图来描绘。图重绘。39 39. Luke Muehlhauser. 2013. 我们之上有大量空间。载于《面对智能爆炸》。https://intelligenceexplosion.com/2011/plenty-of-room-above-us/._
_Figure 3. Intelligence explosion through recursively self-improved AI is a common concern, often depicted by figures like this one. Figure redrawn.39 39. Luke Muehlhauser. 2013. Plenty of room above us. In Facing the Intelligence Explosion.https://intelligenceexplosion.com/2011/plenty-of-room-above-us/._
然而,这幅图在概念和逻辑上存在缺陷。在概念层面,智能——尤其是不同物种之间的比较——并没有明确定义,更不用说在一维尺度上测量了。40 40. Melanie Mitchell 等人. 2024. 第 1 集:什么是智能?复杂性。圣塔菲研究所;播客节目。https://www.santafe.edu/culture/podcasts/ep-1-what-is-intelligence; Melanie Mitchell. 2019. 观点:我们不应被‘超级智能 AI’吓倒。《纽约时报》(2019 年 10 月)。https://www.nytimes.com/2019/10/31/opinion/superintelligent-artificial-intelligence.html.
However, there are conceptual and logical flaws with this picture. On a conceptual level, intelligence—especially as a comparison between different species—is not well defined, let alone measurable on a one-dimensional scale.40 40. Melanie Mitchell et al. 2024. Ep. 1: What is intelligence? Complexity. Santa Fe Institute; Podcast episode. https://www.santafe.edu/culture/podcasts/ep-1-what-is-intelligence; Melanie Mitchell. 2019. Opinion. We shouldn’t be scared by ‘Superintelligent A.I.’ The New York Times (October 2019). https://www.nytimes.com/2019/10/31/opinion/superintelligent-artificial-intelligence.html.
更重要的是,智能并非分析 AI 影响的关键属性。关键在于权力——改变环境的能力。要清晰分析技术(尤其是日益通用的计算技术)的影响,我们必须研究技术如何影响人类的权力。当我们从这个角度审视时,一幅完全不同的图景浮现出来。
More importantly, intelligence is not the property at stake for analyzing AI’s impacts. Rather, what is at stake is power—the ability to modify one’s environment. To clearly analyze the impact of technology (and in particular, increasingly general computing technology), we must investigate how technology has affected humanity’s power. When we look at things from this perspective, a completely different picture emerges.
_图 4. 分析技术对人类权力的影响。我们之所以强大,不是因为我们的智能,而是因为我们利用技术来增强自身能力。_
_Figure 4. Analyzing the impact of technology on humanity’s power. We are powerful not because of our intelligence, but because of the technology we use to increase our capabilities._
这种视角的转变表明,人类一直利用技术来增强控制环境的能力。祖先人类与现代人类在生物或生理上几乎没有差异;相反,相关差异在于知识和理解的提升、工具、技术,以及 AI。从某种意义上说,与前技术时代的人类相比,现代人类能够改变地球及其气候,是‘超级智能’的存在。不幸的是,许多分析 AI 超级智能风险的奠基性文献在‘智能’一词的使用上缺乏精确性。
This shift in perspective clarifies that humans have always used technology to increase our ability to control our environment. There are few biological or physiological differences between ancestral and modern humans; instead, the relevant differences are improved knowledge and understanding, tools, technology and, indeed, AI. In a sense, modern humans, with the capability to alter the planet and its climate, are ‘superintelligent’ beings compared to pre-technological humans. Unfortunately, much of the foundational literature analyzing the risks of AI superintelligence suffers from a lack of precision in the use of the term ‘intelligence.’
_图 5. 从 AI 能力提升到失控的因果链的两种观点。_
_Figure 5. Two views of the causal chain from increases in AI capability to loss of control._
一旦我们停止使用‘智能’和‘超级智能’这些术语,事情就变得清晰得多(图 5)。担忧在于,如果 AI 能力无限增长(无论是否类人或超人),可能导致 AI 系统拥有越来越大的权力,进而导致失控。如果我们接受能力很可能无限增长(我们确实接受),那么防止失控的选择就是干预这两个因果步骤之一。
Once we stop using the terms ‘intelligence’ and ‘superintelligence,’ things become much clearer (Figure 5). The worry is that if AI capabilities continue to increase indefinitely (whether or not they are humanlike or superhuman is irrelevant), they may lead to AI systems with more and more power, in turn leading to a loss of control. If we accept that capabilities are likely to increase indefinitely (we do), our options for preventing a loss of control are to intervene in one of the two causal steps.
超级智能观点对图 5 中的第一个箭头持悲观态度——即防止任意强大的 AI 系统获取足以造成灾难性风险的权力——而是专注于对齐技术,试图防止任意强大的 AI 系统违背人类利益。我们的观点恰恰相反,正如我们在本文其余部分所阐述的。
The superintelligence view is pessimistic about the first arrow in Figure 5—preventing arbitrarily capable AI systems from acquiring power that is significant enough to pose catastrophic risks—and instead focuses on alignment techniques that try to prevent arbitrarily powerful AI systems from acting against human interests. Our view is precisely the opposite, as we elaborate in the rest of this paper.
弱化智能并非仅仅是修辞上的策略:我们认为,在“智能”一词的有用含义上,AI 并不比借助 AI 辅助行动的人类更智能。人类智能的特殊性在于我们使用工具并将其他智能纳入自身的能力,因此无法被连贯地置于一个智能谱系中。
De-emphasizing intelligence is not just a rhetorical move: We do not think there is a useful sense of the term ‘intelligence’ in which AI is more intelligent than people acting with the help of AI. Human intelligence is special due to our ability to use tools and to subsume other intelligences into our own, and cannot be coherently placed on a spectrum of intelligence.
人类能力确实存在一些重要的局限性,尤其是速度。这就是为什么机器在国际象棋等领域显著超越人类,而在人+AI 团队中,人类几乎只能完全听从 AI 的建议。但速度限制在大多数领域无关紧要,因为高速顺序计算或快速反应时间并非必需。
Human abilities definitely have some important limitations, notably speed. This is why machines dramatically outperform humans in domains like chess and, in a human+AI team, the human can hardly do better than simply deferring to AI. But speed limitations are irrelevant in most areas because high-speed sequential calculations or fast reaction times are not required.
在少数需要超人类速度的现实任务中,例如核反应堆控制,我们擅长构建范围狭窄的自动化工具来执行高速部分,而人类则保留对整个系统的控制。
In the few real-world tasks for which superhuman speed is required, such as nuclear reactor control, we are good at building tightly scoped automated tools to do the high-speed parts, while humans retain control of the overall system.
基于这种对人类能力的看法,我们提出一个预测。我们认为,在现实世界的认知任务中,人类局限性如此显著以至于 AI 能够超越人类表现(如 AI 在国际象棋中那样)的情况相对较少。在许多其他领域,包括一些与 AI 性能的突出希望和恐惧相关的领域,我们认为存在很高的“不可约误差”——由于现象固有的随机性而不可避免的误差——而人类表现基本接近这一极限。
We offer a prediction based on this view of human abilities. We think there are relatively few real-world cognitive tasks in which human limitations are so telling that AI is able to blow past human performance (as AI does in chess). In many other areas, including some that are associated with prominent hopes and fears about AI performance, we think there is a high “irreducible error”—unavoidable error due to the inherent stochasticity of the phenomenon—and human performance is essentially near that limit.41 41. Matthew J Salganik et al. 2020. Measuring the predictability of life outcomes with a scientific mass collaboration. Proceedings of the National Academy of Sciences 117, 15 (2020), 8398–8403.
具体而言,我们提出两个这样的领域:预测和说服。我们预测,在地缘政治事件(例如选举)的预测方面,AI 将无法显著超越受过训练的人类(尤其是人类团队,特别是当辅以简单的自动化工具时)。对于说服人们违背自身利益的任务,我们也做出同样的预测。
Concretely, we propose two such areas: forecasting and persuasion. We predict that AI will not be able to meaningfully outperform trained humans (particularly teams of humans and especially if augmented with simple automated tools) at forecasting geopolitical events (say elections). We make the same prediction for the task of persuading people to act against their own self-interest.
说服中的自身利益方面至关重要,但常常被低估。作为一个常见模式的说明性例子,考虑研究“评估前沿模型的危险能力”,该研究评估了语言模型说服人们的能力。其中一些说服测试对被说服者来说没有成本;他们只是在与 AI 互动结束时被问及是否相信某个说法。其他测试则涉及小额成本,例如放弃 20 英镑的慈善奖金(当然,向慈善机构捐款是人们经常自愿做的事情)。因此,这些测试不一定能告诉我们 AI 说服人们执行某些危险任务的能力。值得称赞的是,作者承认了这种缺乏生态效度的情况,并强调他们的研究不是“社会科学实验”,而仅仅旨在评估模型能力。但这样一来,这种脱离上下文的能评估是否具有任何安全含义就不清楚了,然而它们通常被误解为具有安全含义。
The self-interest aspect of persuasion is a critical one, but is often underappreciated. As an illustrative example of a common pattern, consider the study “Evaluating Frontier Models for Dangerous Capabilities,” which evaluated language models’ abilities to persuade people.42 42. Mary Phuong et al. 2024. Evaluating frontier models for dangerous capabilities. arXiv: April 2024. Page 5. Retrieved from doi:10.48550/arXiv.2403.13793. Some of their persuasion tests were costless to the subjects being persuaded; they were simply asked whether they believed a claim at the end of the interaction with AI. Other tests had small costs, such as forfeiting a £20 bonus to charity (of course, donating to charity is something that people often do voluntarily). So these tests do not necessarily tell us about AI’s ability to persuade people to perform some dangerous tasks. To their credit, the authors acknowledged this lack of ecological validity and stressed that their study was not a “social science experiment,” but merely intended to evaluate model capability. 43 43. Mary Phuong et al. 2024. Evaluating frontier models for dangerous capabilities. arXiv: April 2024. Retrieved from doi:10.48550/arXiv.2403.13793. But then it is not clear that such decontextualized capability evaluations have any safety implications, yet they are typically misinterpreted as if they do.
需要一些谨慎来使我们的预测精确——尚不清楚在考虑人类众所周知的但次要的局限性(如预测中的校准不足或说服中的耐心有限)时,应允许多大的余地。
Some care is necessary to make our predictions precise—it is not clear how much slack to allow for well-known but minor human limitations such as the lack of calibration (in the case of forecasting) or limited patience (in the case of persuasion).
如果我们假设存在超级智能,控制问题让人联想到建造一个星系级大脑然后将其关在盒子里的比喻,这是一个可怕的前景。但是,如果我们正确认为 AI 系统不会比人类借助 AI 辅助更有能力,那么控制问题就容易处理得多,尤其是如果超人类说服力被证明是一个无根据的担忧。
If we presume superintelligence, the control problem evokes the metaphor of building a galaxy brain and then keeping it in a box, which is a terrifying prospect. But, if we are correct that AI systems will not be meaningfully more capable than humans acting with AI assistance, then the control problem is much more tractable, especially if superhuman persuasion turns out to be an unfounded concern.
关于 AI 控制的讨论往往过度聚焦于少数几种狭隘的方法,包括模型对齐和保持人类参与。44 44. Arvind Narayanan, Sayash Kapoor, and Seth Lazar. 2024. Model alignment protects against accidental harms, not intentional ones. https://www.aisnakeoil.com/p/model-alignment-protects-against. 我们可以粗略地将这些视为两个极端:在系统运行期间将安全决策完全委托给 AI,以及让人类对每个决策进行二次猜测。这些方法有其作用,但非常有限。在第三部分,我们解释了对模型对齐的怀疑。所谓人类参与控制,我们指的是每个 AI 决策或行动都需要人类审查和批准的系统。在大多数情况下,这种方法大大削弱了自动化的好处,因此要么退化为人类充当橡皮图章,要么被不太安全的解决方案所击败。45 45. Raja Parasuraman and Dietrich H. Manzey. 2010. Complacency and bias in human use of automation: An attentional integration. Human Factors 52, 3 (June 2010), 381–410. http://doi:10.1177/0018720810376055. 我们强调,人类参与控制并不等同于人类对 AI 的监督;它是一种特定的监督模型,而且是一种极端的模型。
Discussions of AI control tend to over-focus on a few narrow approaches, including model alignment and keeping humans in the loop.44 44. Arvind Narayanan, Sayash Kapoor, and Seth Lazar. 2024. Model alignment protects against accidental harms, not intentional ones. https://www.aisnakeoil.com/p/model-alignment-protects-against. We can roughly think of these as opposite extremes: delegating safety decisions entirely to AI during system operation, and having a human second-guessing every decision. There is a role for such approaches, but it is very limited. In Part III, we explain our skepticism of model alignment. By human-in-the-loop control, we mean a system in which every AI decision or action requires review and approval by a human. In most scenarios, this approach greatly diminishes the benefits of automation, and therefore either devolves into the human acting as a rubber stamp or is outcompeted by a less safe solution.45 45. Raja Parasuraman and Dietrich H. Manzey. 2010. Complacency and bias in human use of automation: An attentional integration. Human Factors 52, 3 (June 2010), 381–410. http://doi:10.1177/0018720810376055. We emphasize that human-in-the-loop control is not synonymous with human oversight of AI; it is one particular oversight model, and an extreme one.
幸运的是,还有许多其他形式的控制介于这两个极端之间,例如审计和监控。审计允许在部署前和/或定期评估 AI 系统实现其既定目标的程度,使我们能够在灾难性故障发生之前预见它们。监控允许在系统属性偏离预期行为时进行实时监督,从而在真正需要时进行人工干预。
Fortunately, there are many other flavors of control that fall between these two extremes, such as auditing and monitoring. Auditing allows pre-deployment and/or periodic assessments of how well an AI system fulfills its stated goals, allowing us to anticipate catastrophic failures before they arise. Monitoring allows real-time oversight when system properties diverge from the expected behavior, allowing human intervention when truly needed.
其他想法来自系统安全,这是一门工程学科,专注于通过系统分析和设计来预防复杂系统中的事故。46 46. Roel I. J. Dobbe. 2022). System safety and artificial intelligence. In The Oxford Handbook of AI Governance, ed. Justin B. Bullock et al., Oxford University Press, Oxford. http://doi:10.1093/oxfordhb/9780197579329.013.67. 例子包括故障安全机制,确保系统在发生故障时默认进入安全状态,例如预定义规则或硬编码操作;以及断路器,当超过预定义安全阈值时自动停止操作。其他技术包括关键组件的冗余和系统动作安全属性的验证。
Other ideas come from system safety, an engineering discipline that is focused on preventing accidents in complex systems through systematic analysis and design.46 46. Roel I. J. Dobbe. 2022). System safety and artificial intelligence. In The Oxford Handbook of AI Governance, ed. Justin B. Bullock et al., Oxford University Press, Oxford. http://doi:10.1093/oxfordhb/9780197579329.013.67. Examples include fail-safes, which ensure that systems default to a safe state when they malfunction, such as a predefined rule or a hard-coded action, and circuit breakers that automatically stop operations when predefined safety thresholds are exceeded. Other techniques include redundancy in critical components and the verification of safety properties of the system’s actions.
其他计算领域,包括网络安全、形式化验证和人机交互,也是控制技术的丰富来源,这些技术已成功应用于传统软件系统,同样适用于 AI。在网络安全中,“最小权限”原则确保参与者只拥有完成任务所需的最小资源。访问控制防止处理敏感数据和系统的人员访问其工作不需要的机密信息和工具。我们可以为在关键环境中的 AI 系统设计类似的保护措施。形式化验证方法确保安全关键代码按照其规范工作;它现在被用于验证 AI 生成代码的正确性。47 47. CodeMetal.ai. 2024. Combining AI with formal verification for efficient migration of legacy code. https://www.codemetal.ai/research/combining-ai-with-formal-verification-for-efficient-migration-of-legacy-code 从人机交互中,我们可以借鉴诸如设计系统使改变状态的操作可逆等想法,允许人类即使在高度自动化的系统中也能保持有意义的控制。
Other computing fields, including cybersecurity, formal verification, and human-computer interaction, are also rich sources of control techniques that have been successfully applied to traditional software systems and are equally applicable to AI. In cybersecurity, the principle of ‘least privilege’ ensures that actors only have access to the minimum resources needed for their tasks. Access controls prevent people working with sensitive data and systems from accessing confidential information and tools that are not required for their jobs. We can design similar protections for AI systems in consequential settings. Formal verification methods ensure that safety-critical codes work according to its specifications; it is now being used to verify the correctness of AI-generated code.47 47. CodeMetal.ai. 2024. Combining AI with formal verification for efficient migration of legacy code. https://www.codemetal.ai/research/combining-ai-with-formal-verification-for-efficient-migration-of-legacy-code From human-computer interaction, we can borrow ideas like designing systems so that state-changing actions are reversible, allowing humans to retain meaningful control even in highly automated systems.
除了来自其他领域的现有思想被改编用于 AI 控制之外,技术性 AI 安全研究也产生了许多新思想。48 48. Balint Gyevnar and Atoosa Kasirzadeh. 2025. AI safety for everyone. arXiv preprint arXiv: 2502.09288. 例子包括使用语言模型作为自动裁判来评估提议行动的安全性;开发系统,使其能够根据不确定性或风险水平学习何时适当地将决策升级给人类操作员;设计智能体式系统,使其活动对人类可见且可理解;以及创建分层控制结构,其中更简单、更可靠的 AI 系统监督更有能力但可能不可靠的系统。49 49. Balint Gyevnar and Atoosa Kasirzadeh. 2025. AI safety for everyone. arXiv preprint arXiv: 2502.09288; Tinghao Xie et al. 2024. SORRY-Bench: Systematically evaluating large language model safety refusal behaviors. arXiv: June 2024. Retrieved from doi:10.48550/arXiv.2406.14598; Alan Chan et al. 2024. Visibility into AI agents. arXiv: May 2024. Retrieved from doi:10.48550/arXiv.2401.13138; Yonadav Shavit et al. 2023. Practices for governing agentic AI systems. https://cdn.openai.com/papers/practices-for-governing-agentic-ai-systems.pdf.
In addition to existing ideas from other fields being adapted for AI control, technical AI safety research has generated many new ideas.48 48. Balint Gyevnar and Atoosa Kasirzadeh. 2025. AI safety for everyone. arXiv preprint arXiv: 2502.09288. Examples include using language models as automated judges to evaluate the safety of proposed actions, developing systems that learn when to appropriately escalate decisions to human operators based on uncertainty or risk level, designing agentic systems so that their activity is visible and legible to humans, and creating hierarchical control structures in which simpler and more reliable AI systems oversee more capable but potentially unreliable ones.49 49. Balint Gyevnar and Atoosa Kasirzadeh. 2025. AI safety for everyone. arXiv preprint arXiv: 2502.09288; Tinghao Xie et al. 2024. SORRY-Bench: Systematically evaluating large language model safety refusal behaviors. arXiv: June 2024. Retrieved from doi:10.48550/arXiv.2406.14598; Alan Chan et al. 2024. Visibility into AI agents. arXiv: May 2024. Retrieved from doi:10.48550/arXiv.2401.13138; Yonadav Shavit et al. 2023. Practices for governing agentic AI systems. https://cdn.openai.com/papers/practices-for-governing-agentic-ai-systems.pdf.
技术性 AI 安全研究有时会根据一个模糊且不切实际的目标来评判,即保证未来的“超级智能”AI 将“与人类价值观对齐”。从这个角度来看,它往往被视为一个未解决的问题。但从使 AI 系统的开发者、部署者和操作者更容易减少事故可能性的角度来看,技术性 AI 安全研究已经产生了大量的想法。我们预测,随着先进 AI 的开发和采用,将会有越来越多的创新来寻找人类控制的新模型。
Technical AI safety research is sometimes judged against the fuzzy and unrealistic goal of guaranteeing that future “superintelligent” AI will be “aligned with human values.” From this perspective, it tends to be viewed as an unsolved problem. But from the perspective of making it easier for developers, deployers, and operators of AI systems to decrease the likelihood of accidents, technical AI safety research has produced a great abundance of ideas. We predict that as advanced AI is developed and adopted, there will be increasing innovation to find new models for human control.
随着越来越多的物理和认知任务变得可自动化,我们预测人类工作和任务中与 AI 控制相关的比例将不断增加。如果这看起来激进,请注意这种对工作概念的近乎彻底的重定义以前就发生过。在工业革命之前,大多数工作涉及体力劳动。随着时间的推移,越来越多的体力任务被自动化,这一趋势仍在继续。在这个过程中,发明了许多操作、控制和监控物理机器的不同方式,今天人类在工厂中所做的是“控制”(监控自动化生产线、编程机器人系统、管理质量控制检查点以及协调对设备故障的响应)与一些机器尚不具备的认知能力或灵巧度要求的任务的结合。
As more physical and cognitive tasks become amenable to automation, we predict that an increasing percentage of human jobs and tasks will be related to AI control. If this seems radical, note that this kind of near-total redefinition of the concept of work has happened previously. Before the Industrial Revolution, most jobs involved manual labor. Over time, more and more manual tasks have been automated, a trend that continues. In this process, a great many different ways of operating, controlling, and monitoring physical machines were invented, and what humans do in factories today is a combination of “control” (monitoring automated assembly lines, programming robotic systems, managing quality control checkpoints, and coordinating responses to equipment malfunctions) and some tasks that require levels of cognitive ability or dexterity that machines are not yet capable.
Karen Levy 描述了这种转变在 AI 和卡车司机案例中已经如何展开:
Karen Levy describes how this transformation is already unfolding in the case of AI and truck drivers:
除了 AI 控制之外,任务规范可能成为人类工作内容的更大一部分(取决于我们对控制概念的宽泛程度,规范可以被视为控制的一部分)。任何尝试过外包软件或产品开发的人都知道,明确指定所需内容竟然是整体工作中相当大的一部分。因此,人类劳动——规范和监督——将在执行不同任务的 AI 系统之间的边界上运作。消除其中一些效率瓶颈并让 AI 系统自主完成更大的“端到端”任务将是一个始终存在的诱惑,但这会增加安全风险,因为它会降低可读性和可控性。这些风险将作为自然检查,防止让渡过多控制权。
In addition to AI control, task specification is likely to become a bigger part of what human jobs entail (depending on how broadly we conceive of control, specification could be considered part of control). As anyone who has tried to outsource software or product development knows, unambiguously specifying what is desired turns out to be a surprisingly big part of the overall effort. Thus, human labor—specification and oversight—will operate at the boundary between AI systems performing different tasks. Eliminating some of these efficiency bottlenecks and having AI systems autonomously accomplish larger tasks “end-to-end” will be an ever-present temptation, but this will increase safety risks since it will decrease legibility and control. These risks will act as a natural check against ceding too much control.
我们进一步预测,这种转变将主要由市场力量驱动。控制不良的 AI 将过于容易出错,无法在商业上可行。但监管可以并且应该加强组织保持人类控制的能力和必要性。
We further predict that this transformation will be primarily driven by market forces. Poorly controlled AI will be too error prone to make business sense. But regulation can and should bolster the ability and necessity of organizations to keep humans in control.
我们考虑五类风险:事故、军备竞赛(导致事故)、滥用、失调,以及非灾难性但系统性的风险。
We consider five types of risks: accidents, arms races (leading to accidents), misuse, misalignment, and non-catastrophic but systemic risks.
我们已在上面讨论了事故。我们的观点是,与其他技术一样,部署者和开发者应承担减轻 AI 系统事故的主要责任。他们能否有效做到这一点取决于其激励措施以及缓解方法的进展。在许多情况下,市场力量将提供足够的激励,但安全监管应填补任何空白。至于缓解方法,我们回顾了 AI 控制研究正在快速推进。
We have already addressed accidents above. Our view is that, just like other technologies, deployers and developers should have the primary responsibility for mitigating accidents in AI systems. How effectively they will do so depends on their incentives, as well as on progress in mitigation methods. In many cases, market forces will provide an adequate incentive, but safety regulation should fill any gaps. As for mitigation methods, we reviewed how research on AI control is advancing rapidly.
这种乐观评估可能不成立的原因有几个。首先,可能存在军备竞赛,因为 AI 的竞争利益如此之大,以至于成为常规模式的例外。我们将在下面讨论这一点。
There are a few reasons why this optimistic assessment might not hold. First, there might be arms races because the competitive benefits of AI are so great that they are an exception to the usual patterns. We discuss this below.
其次,部署 AI 的公司或实体可能如此庞大和强大,以至于知道如果它对事故缓解态度不佳最终会倒闭,这几乎无法带来安慰——它可能将文明一同拖垮。例如,控制几乎所有消费设备的 AI 智能体的不当行为可能导致灾难性的广泛数据丢失。虽然这确实可能,但这种权力集中比 AI 事故的可能性更成问题,这正是我们的政策方法强调韧性和去中心化的原因(第四部分)。
Second, a company or entity deploying AI might be so big and powerful that it is little consolation to know that it will eventually go out of business if it has a poor attitude to accident mitigation—it might take down civilization with it. For example, misbehavior by an AI agent that controls almost every consumer device might lead to catastrophically widespread data loss. While this is certainly possible, such concentration of power is a bigger problem than the possibility of AI accidents, and is precisely why our approach to policy emphasizes resilience and decentralization (Part IV).
最后,即使是相对不起眼的部署者导致的 AI 控制失败也可能带来灾难性风险——例如,因为 AI 智能体“逃脱”、自我复制等。我们认为这是一种失调风险,并在下面讨论。
Finally, perhaps even an AI control failure by a relatively inconspicuous deployer might lead to catastrophic risk—say because an AI agent ‘escapes,’ makes copies of itself, and so forth. We see this as a misalignment risk, and discuss it below.
在第三部分的其余部分,我们通过将 AI 视为正常技术的视角,考虑四种风险——军备竞赛、滥用、失调,以及非灾难性但系统性的风险。
In the rest of Part III, we consider four risks—arms races, misuse, misalignment, and non-catastrophic but systemic risks—through the lens of AI as normal technology.
AI 军备竞赛是指两个或多个竞争者——公司、不同国家的政策制定者、军队——在监督和控制不足的情况下部署日益强大的 AI。危险在于,更安全的参与者可能会被更冒险的参与者击败。基于上述原因,我们不太担心 AI 方法_开发_中的军备竞赛,而更担心 AI 应用_部署_中的军备竞赛。
An AI arms race is a scenario in which two or more competitors—companies, policymakers in different countries, militaries—deploy increasingly powerful AI with inadequate oversight and control. The danger is that safer actors will be outcompeted by riskier ones. For the reasons described above, we are less concerned about arms races in the _development_ of AI _methods_ and are more concerned about the _deployment_ of AI _applications_.
一个重要说明:我们明确将军事 AI 排除在分析之外,因为它涉及机密能力和独特动态,需要更深入的分析,这超出了本文的范围。
One important caveat: We explicitly exclude military AI from our analysis, as it involves classified capabilities and unique dynamics that require a deeper analysis, which is beyond the scope of this essay.
首先考虑公司。在安全方面竞相逐底在历史上跨行业极为常见,并已被广泛研究;它也高度适用于已被充分理解的监管干预。例子包括美国服装业的消防安全(20 世纪初)、美国肉类加工行业的食品安全和工人安全(19 世纪末 20 世纪初)、美国汽船行业(19 世纪)、采矿业(19 世纪和 20 世纪初)以及航空业(20 世纪初)。
Let us consider companies first. A race to the bottom in terms of safety is historically extremely common across industries and has been studied extensively; it is also highly amenable to well-understood regulatory interventions. Examples include fire safety in the U.S. garment industry (early 20th century), both food safety and worker safety in the U.S. meatpacking industry (late 19th and early 20th centuries), the U.S. steamboat industry (19th century), the mining industry (19th and early 20th centuries), and the aviation industry (early 20th century).
这些竞赛之所以发生,是因为公司能够将安全成本外部化,导致市场失灵。消费者很难评估产品安全(工人也很难评估工作场所安全),因此在缺乏监管的情况下,市场失灵很常见。但一旦监管迫使公司将安全实践的成本内部化,竞赛就会消失。有许多潜在的监管策略,包括关注过程(标准、审计和检查)、结果(责任)以及纠正信息不对称(标签和认证)。
These races happened because companies were able to externalize the costs of poor safety, resulting in market failure. It is hard for consumers to assess product safety (and for workers to assess workplace safety), so market failures are common in the absence of regulation. But once regulation forces companies to internalize the costs of their safety practices, the race goes away. There are many potential regulatory strategies, including those focused on processes (standards, auditing, and inspections), outcomes (liability), and correcting information asymmetry (labeling and certification).
AI 也不例外。自动驾驶汽车为安全与竞争成功之间的关系提供了一个很好的案例研究。考虑四家安全实践各异的公司。据报道,Waymo 拥有强大的安全文化,强调保守部署和自愿透明;它在安全_结果_方面也处于领先地位。51 51. Andrew J. Hawkins. 2024. Waymo thinks it can overcome robotaxi skepticism with lots of safety data. The Verge. https://www.theverge.com/2024/9/5/24235078/waymo-safety-hub-miles-crashes-robotaxi-transparency; Caleb Miller. 2024. General motors gives up on its cruise robotaxi dreams. Car and Driver (December 2024). https://www.caranddriver.com/news/a63158982/general-motors-cruise-robotaxi-dead/; Greg Bensinger. 2021. Why Tesla’s ‘Beta Testing’ Puts the Public at Risk. The New York Times (July 2021). https://www.nytimes.com/2021/07/30/opinion/self-driving-cars-tesla-elon-musk.html; Andrew J. Hawkins. 2020. Uber’s fraught and deadly pursuit of self-driving cars is over. The Verge. https://www.theverge.com/2020/12/7/22158745/uber-selling-autonomous-vehicle-business-aurora-innovation. Cruise 在部署上更为激进,安全结果也更差。特斯拉也很激进,经常被指责将客户当作测试版测试员。最后,Uber 的自动驾驶部门安全文化臭名昭著地松懈。
AI is no exception. Self-driving cars offer a good case study of the relationship between safety and competitive success. Consider four major companies with varying safety practices. Waymo reportedly has a strong safety culture that emphasizes conservative deployment and voluntary transparency; it is also the leader in terms of safety _outcomes_.51 51. Andrew J. Hawkins. 2024. Waymo thinks it can overcome robotaxi skepticism with lots of safety data. The Verge. https://www.theverge.com/2024/9/5/24235078/waymo-safety-hub-miles-crashes-robotaxi-transparency; Caleb Miller. 2024. General motors gives up on its cruise robotaxi dreams. Car and Driver (December 2024). https://www.caranddriver.com/news/a63158982/general-motors-cruise-robotaxi-dead/; Greg Bensinger. 2021. Why Tesla’s ‘Beta Testing’ Puts the Public at Risk. The New York Times (July 2021). https://www.nytimes.com/2021/07/30/opinion/self-driving-cars-tesla-elon-musk.html; Andrew J. Hawkins. 2020. Uber’s fraught and deadly pursuit of self-driving cars is over. The Verge. https://www.theverge.com/2020/12/7/22158745/uber-selling-autonomous-vehicle-business-aurora-innovation. Cruise was more aggressive in terms of its deployment and had worse safety outcomes. Tesla has also been aggressive and has often been accused of using its customers as beta testers. Finally, Uber’s self-driving unit had a notoriously lax safety culture.
市场成功与安全密切相关。Cruise 计划在 2025 年关闭,而 Uber 被迫出售其自动驾驶部门。52 52. Caleb Miller. 2024. General motors gives up on its cruise robotaxi dreams. Car and Driver (December 2024). https://www.caranddriver.com/news/a63158982/general-motors-cruise-robotaxi-dead/; Andrew J. Hawkins. 2020. Uber’s fraught and deadly pursuit of self-driving cars is over. The Verge. https://www.theverge.com/2020/12/7/22158745/uber-selling-autonomous-vehicle-business-aurora-innovation. 特斯拉面临诉讼和监管审查,其安全态度将给公司带来多大成本仍有待观察。53 53. Jonathan Stempel. 2024. Tesla must face vehicle owners’ lawsuit over self-driving claims. Reuters (May 2024). https://www.reuters.com/legal/tesla-must-face-vehicle-owners-lawsuit-over-self-driving-claims-2024-05-15/. 我们认为这些相关性是因果性的。Cruise 的许可证被吊销是其落后于 Waymo 的重要原因,安全也是 Uber 自动驾驶失败的一个因素。54 54. Hayden Field. 2023. Waymo is full speed ahead as safety incidents and regulators stymie competitor cruise. https://www.cnbc.com/2023/12/05/waymo-chief-product-officer-on-progress-competition-vs-cruise.html.
Market success has been strongly correlated with safety. Cruise is set to shut down in 2025, while Uber was forced to sell off its self-driving unit.52 52. Caleb Miller. 2024. General motors gives up on its cruise robotaxi dreams. Car and Driver (December 2024). https://www.caranddriver.com/news/a63158982/general-motors-cruise-robotaxi-dead/; Andrew J. Hawkins. 2020. Uber’s fraught and deadly pursuit of self-driving cars is over. The Verge. https://www.theverge.com/2020/12/7/22158745/uber-selling-autonomous-vehicle-business-aurora-innovation. Tesla is facing lawsuits and regulatory scrutiny, and it remains to be seen how much its safety attitude will cost the company.53 53. Jonathan Stempel. 2024. Tesla must face vehicle owners’ lawsuit over self-driving claims. Reuters (May 2024). https://www.reuters.com/legal/tesla-must-face-vehicle-owners-lawsuit-over-self-driving-claims-2024-05-15/. We think that these correlations are causal. Cruise’s license being revoked was a big part of the reason that it fell behind Waymo, and safety was also a factor in Uber’s self-driving failure.54 54. Hayden Field. 2023. Waymo is full speed ahead as safety incidents and regulators stymie competitor cruise. https://www.cnbc.com/2023/12/05/waymo-chief-product-officer-on-progress-competition-vs-cruise.html.
监管发挥了虽小但有帮助的作用。联邦和州/地方层面的政策制定者具有远见,认识到该技术的潜力,并采取了轻触式和多元中心(多个监管机构而非一个)的监管策略。总体而言,他们专注于监督、标准制定和证据收集,而许可证撤销的持续威胁则对公司的行为起到了制约作用。
Regulation has played a small but helpful role. Policymakers at both the federal and state/local levels exercised foresight in recognizing the potential of the technology and adopted a regulatory strategy that is light-touch and polycentric (multiple regulators instead of one). Collectively, they focused on oversight, standard setting, and evidence gathering, with the ever-present threat of license revocation acting as a check on companies’ behavior.
类似地,在航空业,AI 的整合被要求符合现有的安全标准,而不是为了激励 AI 采用而降低标准——这主要是因为监管机构有能力惩罚未能遵守安全标准的公司。55 55. Will Hunt. 2020. The flight to safety-critical AI: Lessons in AI safety from the aviation industry. CLTC White Paper Series. UC Berkeley Center for Long-Term Cybersecurity. https://cltc.berkeley.edu/wp-content/uploads/2020/08/Flight-to-Safety-Critical-AI.pdf.
Similarly, in the aviation industry, the integration of AI has been held to the existing standards of safety instead of lowering the bar to incentivize AI adoption—primarily because of the ability of regulators to penalize companies that fail to abide by safety standards.55 55. Will Hunt. 2020. The flight to safety-critical AI: Lessons in AI safety from the aviation industry. CLTC White Paper Series. UC Berkeley Center for Long-Term Cybersecurity. https://cltc.berkeley.edu/wp-content/uploads/2020/08/Flight-to-Safety-Critical-AI.pdf.
简而言之,AI 军备竞赛可能发生,但它们是特定于行业的,应通过特定行业的监管来解决。
In short, AI arms races might happen, but they are sector specific, and should be addressed through sector-specific regulations.
作为与自动驾驶汽车或航空业不同发展方式的领域案例,考虑社交媒体。生成内容推送的推荐算法是一种 AI。它们被指责造成了许多社会弊病,而社交媒体公司在设计和部署这些算法系统时可以说忽视了安全。此外,还存在明显的军备竞赛动态,TikTok 给竞争对手施加压力,使其推送更加依赖推荐。56 56. Arvind Narayanan. 2023. Understanding Social Media Recommendation Algorithms. Knight First Amendment Institute. https://knightcolumbia.org/content/understanding-social-media-recommendation-algorithms. 可以说,市场力量不足以使收入与社会效益对齐;更糟糕的是,监管机构行动迟缓。原因是什么?
As a case study of a domain in which things have played out differently from self-driving cars or aviation, consider social media. The recommendation algorithms that generate content feeds are a kind of AI. They have been blamed for many societal ills, and social media companies have arguably underemphasized safety in the design and deployment of these algorithmic systems. There are also clear arms race dynamics, with TikTok putting pressure on competitors to make their feeds more recommendation heavy.56 56. Arvind Narayanan. 2023. Understanding Social Media Recommendation Algorithms. Knight First Amendment Institute. https://knightcolumbia.org/content/understanding-social-media-recommendation-algorithms. Arguably, market forces were insufficient to align revenues with societal benefit; worse, regulators have been slow to act. What are the reasons for this?
社交媒体与交通的一个显著区别是,当伤害发生时,在交通领域将伤害归因于产品失败相对直接,并且公司会立即遭受声誉损害。但在社交媒体领域,归因极其困难,甚至研究仍无定论且存在争议。第二个区别是,我们在交通领域已经有一个多世纪的时间来制定安全标准和期望。在汽车早期,安全并不被视为制造商的责任。57 57. Ralph Nader. 1965. Unsafe at Any Speed: The Designed-in Dangers of the American Automobile. Grossman Publishers, New York, NY.
One significant difference between social media and transportation is that, when harms occur, attributing them to product failures is relatively straightforward in the case of transportation, and there is immediate reputational damage to the company. But attribution is extremely hard in the case of social media, and even the research remains inconclusive and contested. A second difference between the domains is that we have had over a century to develop standards and expectations around transportation safety. In the early decades of automobiles, safety was not considered to be the responsibility of manufacturers.57 57. Ralph Nader. 1965. Unsafe at Any Speed: The Designed-in Dangers of the American Automobile. Grossman Publishers, New York, NY.
AI 足够广泛,以至于其未来的一些应用将更像交通,而另一些则更像社交媒体。这表明在新兴 AI 驱动领域和应用中,主动收集证据和透明度的重要性。我们在第四部分讨论这一点。它还显示了“预期性 AI 伦理”的重要性——在新技术生命周期中尽早识别伦理问题,制定规范和标准,并利用这些来积极塑造技术的部署,并最小化军备竞赛的可能性。58 58. Seth Lazar. 2025. Anticipatory AI ethics (manuscript, forthcoming 2025).
AI is broad enough that some of its future applications will be more like transportation, while others will be more like social media. This shows the importance of proactive evidence gathering and transparency in emerging AI-driven sectors and applications. We address this in Part IV. It also shows the importance of “anticipatory AI ethics”—identifying ethical issues as early as possible in the lifecycle of emerging technologies, developing norms and standards, and using those to actively shape the deployment of technologies and to minimize the likelihood of arms races.58 58. Seth Lazar. 2025. Anticipatory AI ethics (manuscript, forthcoming 2025).
AI 安全监管可能更困难的一个原因是,如果采用速度如此之快,以至于监管机构来不及干预就为时已晚。到目前为止,即使在缺乏监管的情况下,我们也没有看到 AI 在重要任务中快速采用的例子,我们在第一部分提出的反馈循环模型可能解释了原因。新 AI 应用的采用率仍将是需要跟踪的关键指标。
One reason why safety regulation might be harder in the case of AI is if adoption is so rapid that regulators will not be able to intervene until it is too late. So far, we have not seen examples of rapid AI adoption in consequential tasks, even in the absence of regulation, and the feedback loop model we presented in Part I might explain why. The adoption rate of new AI applications will remain a key metric to track.
与此同时,即使未来扩散速度没有加快,监管缓慢本身也是一个问题。我们在第四部分讨论这个“节奏问题”。
At the same time, the slow pace of regulation is a problem even without any future acceleration of the speed of diffusion. We discuss this ‘pacing problem’ in Part IV.
现在考虑国家间的竞争。政府是否会面临竞争压力,从而对 AI 安全采取不干预的态度?
Let us now consider competition between countries. Will there be competitive pressure on governments to take a hands-off approach to AI safety?
再次强调,这不是一个新问题。创新与监管之间的权衡是监管国家反复面临的困境。到目前为止,我们看到各国方法存在显著差异,例如欧盟强调预防性方法(通用数据保护条例、数字服务法案、数字市场法案和欧盟 AI 法案),而美国则倾向于仅在已知伤害或市场失灵后才进行监管。59 59. Alex Engler. 2023. The EU and U.S. diverge on AI regulation: A transatlantic comparison and steps to alignment. https://www.brookings.edu/research/the-eu-and-us-diverge-on-ai-regulation-a-transatlantic-comparison-and-steps-to-alignment/.
Again, our message is that this is not a new problem. The tradeoff between innovation and regulation is a recurring dilemma for the regulatory state. So far, we are seeing striking differences in approaches, such as the EU emphasizing a precautionary approach (the General Data Protection Regulation, the Digital Services Act, the Digital Markets Act, and the EU AI Act) and the U.S. preferring to regulate only after there are known harms or market failures.59 59. Alex Engler. 2023. The EU and U.S. diverge on AI regulation: A transatlantic comparison and steps to alignment. https://www.brookings.edu/research/the-eu-and-us-diverge-on-ai-regulation-a-transatlantic-comparison-and-steps-to-alignment/.
尽管美中军备竞赛言论激烈,但尚不清楚两国的 AI 监管是否放缓。60 60. Matt Sheehan. 2023. China’s AI regulations and how they get made. https://carnegieendowment.org/2023/07/10/china-s-ai-regulations-and-how-they-get-made-pub-90117. 在美国,仅 2024 年州立法机构就提出了 700 多项与 AI 相关的法案,其中数十项已通过。61 61. Heather Curry, 2024. 2024 state summary on AI. BSA TechPost (October 2024). https://techpost.bsa.org/2024/10/22/2024-state-summary-on-ai/. 正如我们在前几部分中指出的,大多数高风险行业都受到严格监管,无论是否使用 AI 都适用。那些声称 AI 监管是“狂野西部”的人往往过分强调一种狭隘的、以模型为中心的监管。在我们看来,监管机构强调 AI 使用而非开发是恰当的(除了我们下面讨论的透明度要求等例外)。
Despite shrill U.S.-China arms race rhetoric, it is not clear that AI regulation has slowed down in either country.60 60. Matt Sheehan. 2023. China’s AI regulations and how they get made. https://carnegieendowment.org/2023/07/10/china-s-ai-regulations-and-how-they-get-made-pub-90117. In the U.S., 700 AI-related bills were introduced in state legislatures in 2024 alone, and dozens of them have passed.61 61. Heather Curry, 2024. 2024 state summary on AI. BSA TechPost (October 2024). https://techpost.bsa.org/2024/10/22/2024-state-summary-on-ai/. As we pointed out in the earlier parts, most high-risk sectors are heavily regulated in ways that apply regardless of whether or not AI is used. Those claiming that AI regulation is a ‘wild west’ tend to overemphasize a narrow, model-centric type of regulation. In our view, regulators’ emphasis on AI use over development is appropriate (with exceptions such as transparency requirements that we discuss below).
未能充分监管安全采用将主要通过事故产生_本地_负面影响,而安全文化松懈的公司则可能将安全成本外部化。因此,没有直接理由预期国家间会出现军备竞赛。请注意,由于本节关注的是事故而非滥用,针对外国的网络攻击不在讨论范围内。我们在下一节讨论滥用。
Failing to adequately regulate safe adoption will lead to negative impacts through accidents primarily _locally_, as opposed to companies with a lax safety culture potentially being able to externalize the costs of safety. Therefore, there is no straightforward reason to expect arms races between countries. Note that, since our concern in this section is accidents, not misuse, cyberattacks against foreign countries are out of scope. We discuss misuse in the next section.
与核技术的类比可以说明这一点。AI 常被类比为核武器。但除非我们谈论的是军事 AI 的风险(我们同意这是一个令人担忧的领域,但本文不予考虑),否则这是错误的类比。关于(原本良性的)AI 应用部署导致事故的担忧,正确的类比是核能。核武器与核能的差异清楚地说明了我们的观点——虽然存在核武器军备竞赛,但核能领域没有类似情况。事实上,由于安全影响是本地化的,该技术在许多国家引发了强烈的反弹,普遍认为这严重阻碍了其潜力。
An analogy with nuclear technology can make this clear. AI is often analogized to nuclear weapons. But unless we are talking about the risks of military AI (which we agree is an area of concern and do not consider in this paper), this is the wrong analogy. With regard to the concern about accidents due to the deployment of (otherwise benign) AI applications, the right analogy is nuclear power. The difference between nuclear weapons and nuclear power neatly illustrates our point—while there was a nuclear weapons arms race, there was no equivalent for nuclear power. In fact, since safety impacts were felt locally, the tech engendered a powerful backlash in many countries that is generally thought to have severely hobbled its potential.
理论上,在大国冲突的背景下,政策制定者可能更愿意承担本地安全成本,以确保其 AI 产业成为全球赢家。再次强调,关注采用而非开发,目前没有迹象表明这种情况正在发生。美中军备竞赛的言论强烈集中在模型开发(发明)上。我们没有看到相应的仓促采用 AI 的情况。安全社区应持续向政策制定者施压,确保这种情况不会改变。国际合作也必须发挥重要作用。
It is theoretically possible that policymakers in the context of a great-power conflict will prefer to incur safety costs locally in order to ensure that their AI industry is the global winner. Again, focusing on adoption as opposed to development, there is currently no indication that this is happening. The U.S. versus China arms race rhetoric has been strongly focused on model development (invention). We have not seen a corresponding rush to adopt AI haphazardly. The safety community should keep up the pressure on policymakers to ensure that this does not change. International cooperation must also play an important role.
模型对齐通常被视为防止模型滥用的主要防御手段。目前通过后训练干预实现,例如基于人类和 AI 反馈的强化学习。62 62. Yuntao Bai et al. 2022. Constitutional AI: Harmlessness from AI feedback. arXiv: December 2022. Retrieved from doi:10.48550/arXiv.2212.08073; Long Ouyang et al.. 2022. Training language models to follow instructions with human feedback. arXiv: March 2022. Retrieved from doi:10.48550/arXiv.2203.02155. 不幸的是,让模型拒绝滥用尝试已被证明极其脆弱。63 63. Eugene Bagdasaryan et al. 2023. Abusing images and sounds for indirect instruction injection in multi-modal LLMs. arXiv: October 2023. Retrieved from http://arxiv.org/abs/2307.10490; Xiangyu Qi et al. 2023. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv: October 2023. Retrieved from doi:10.48550/arXiv.2310.03693. 我们认为这种局限性是固有的,不太可能被修复;因此,针对滥用的主要防御必须位于别处。
Model alignment is often seen as the primary defense against the misuse of models. It is currently achieved through post-training interventions, such as reinforcement learning with human and AI feedback.62 62. Yuntao Bai et al. 2022. Constitutional AI: Harmlessness from AI feedback. arXiv: December 2022. Retrieved from doi:10.48550/arXiv.2212.08073; Long Ouyang et al.. 2022. Training language models to follow instructions with human feedback. arXiv: March 2022. Retrieved from doi:10.48550/arXiv.2203.02155. Unfortunately, aligning models to refuse attempts at misuse has proved to be extremely brittle.63 63. Eugene Bagdasaryan et al. 2023. Abusing images and sounds for indirect instruction injection in multi-modal LLMs. arXiv: October 2023. Retrieved from http://arxiv.org/abs/2307.10490; Xiangyu Qi et al. 2023. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv: October 2023. Retrieved from doi:10.48550/arXiv.2310.03693. We argue that this limitation is inherent and is unlikely to be fixable; the primary defenses against misuse must thus reside elsewhere.
根本问题在于,一项能力是否有害取决于上下文——而模型往往缺乏这种上下文。64 64. Arvind Narayanan and Sayash Kapoor. 2024. AI safety is not a model property. https://www.aisnakeoil.com/p/ai-safety-is-not-a-model-property.
The fundamental problem is that whether a capability is harmful depends on context—context that the model often lacks.64 64. Arvind Narayanan and Sayash Kapoor. 2024. AI safety is not a model property. https://www.aisnakeoil.com/p/ai-safety-is-not-a-model-property.
考虑一个攻击者利用 AI 通过钓鱼邮件针对一家大公司的员工。攻击链可能涉及多个步骤:扫描社交媒体资料以获取个人信息,识别那些公开在网上发布个人信息的潜在目标,制作个性化钓鱼信息,以及利用窃取的凭证入侵被攻破的账户。
Consider an attacker using AI to target an employee of a large company via a phishing email. The attack chain might involve many steps: scanning social media profiles for personal information, identifying targets who have posted personal information publicly online, crafting personalized phishing messages, and exploiting compromised accounts using harvested credentials.
这些单独的任务本身都没有恶意。使系统有害的是这些能力的组合方式——这些信息只存在于攻击者的编排代码中,而非模型本身。被要求撰写有说服力邮件的模型无法知道它被用于营销还是钓鱼——因此模型层面的干预将无效。65 65. Erik Jones, Anca Dragan, and Jacob Steinhardt. 2024. Adversaries can misuse combinations of safe models. arXiv: July 2024. Retrieved from doi:10.48550/arXiv.2406.14595.
None of these individual tasks are inherently malicious. What makes the system harmful is how these capabilities are composed—information that exists only in the attacker’s orchestration code, not in the model itself. The model that is being asked to write a persuasive email has no way of knowing whether it is being used for marketing or phishing—so model-level interventions would be ineffective.65 65. Erik Jones, Anca Dragan, and Jacob Steinhardt. 2024. Adversaries can misuse combinations of safe models. arXiv: July 2024. Retrieved from doi:10.48550/arXiv.2406.14595.
这种模式反复出现:试图制造一个不能被滥用的 AI 模型,就像试图制造一台不能被用于坏事的计算机。模型层面的安全控制要么过于严格(阻止有益用途),要么对能够将看似良性的能力重新用于有害目的的攻击者无效。
This pattern appears repeatedly: Attempting to make an AI model that cannot be misused is like trying to make a computer that cannot be used for bad things. Model-level safety controls will either be too restrictive (preventing beneficial uses) or will be ineffective against adversaries who can repurpose seemingly benign capabilities for harmful ends.
如果我们把 AI 模型看作一个可以委托安全决策的类人系统,模型对齐似乎是一种自然的防御。但要使其良好运作,模型必须获得大量关于用户和上下文的信息——例如,广泛访问用户的个人信息将使其更容易判断用户意图。但是,当将 AI 视为普通技术时,这种架构会降低安全性,因为它违反了基本网络安全原则(如最小权限),并引入了新的攻击风险,如个人数据泄露。
Model alignment seems like a natural defense if we think of an AI model as a humanlike system to which we can defer safety decisions. But for this to work well, the model must be given a great deal of information about the user and the context—for example, having extensive access to the user’s personal information would make it more feasible to make judgments about the user’s intent. But, when viewing AI as normal technology, such an architecture would decrease safety because it violates basic cybersecurity principles, such as least privilege, and introduces new attack risks such as personal data exfiltration.
我们并不反对模型对齐。它在减少语言模型的有害或有偏见输出方面是有效的,并对其商业部署起到了重要作用。对齐还可以对随意的威胁行为者产生“摩擦”。
We are not against model alignment. It has been effective for reducing harmful or biased outputs from language models and has been instrumental in their commercial deployment. Alignment can also create _friction_ against casual threat actors.
然而,鉴于模型层面的保护不足以防止滥用,防御必须聚焦于恶意行为者实际部署 AI 系统的下游攻击面。66 66. Arvind Narayanan and Sayash Kapoor. 2024. AI safety is not a model property. https://www.aisnakeoil.com/p/ai-safety-is-not-a-model-property 这些防御通常类似于针对非 AI 威胁的现有保护措施,但针对 AI 增强的攻击进行了调整和加强。
Yet, given that model-level protections are not enough to prevent misuse, defenses must focus on the downstream attack surfaces where malicious actors actually deploy AI systems.66 66. Arvind Narayanan and Sayash Kapoor. 2024. AI safety is not a model property. https://www.aisnakeoil.com/p/ai-safety-is-not-a-model-property These defenses will often look similar to existing protections against non-AI threats, adapted and strengthened for AI-enabled attacks.
再次考虑钓鱼的例子。最有效的防御不是限制邮件撰写(这会损害合法用途),而是检测可疑模式的邮件扫描和过滤系统、针对恶意网站的浏览器级保护、防止未授权访问的操作系统安全功能,以及针对用户的安全培训。67 67. Google. 2024. Email sender guidelines. https://support.google.com/mail/answer/81126?hl=en.
Consider again the example of phishing. The most effective defenses are not restrictions on email composition (which would impair legitimate uses), but rather email scanning and filtering systems that detect suspicious patterns, browser-level protections against malicious websites, operating system security features that prevent unauthorized access, and security training for users.67 67. Google. 2024. Email sender guidelines. https://support.google.com/mail/answer/81126?hl=en.
这些都不涉及对用于生成钓鱼邮件的 AI 采取行动——事实上,这些下游防御已经发展了几十年,对人为攻击者也很有效。68 68. Craig Marcho. 2024. IE7 - Introducing the phishing filter. Microsoft Tech Community. https://techcommunity.microsoft.com/t5/ask-the-performance-team/ie7-introducing-the-phishing-filter/ba-p/372327. 它们可以而且应该得到加强以应对 AI 驱动的攻击,但基本方法仍然有效。
None of these involve taking action against the AI used for generating phishing emails—in fact, these downstream defenses have evolved over decades to become effective against human attackers.68 68. Craig Marcho. 2024. IE7 - Introducing the phishing filter. Microsoft Tech Community. https://techcommunity.microsoft.com/t5/ask-the-performance-team/ie7-introducing-the-phishing-filter/ba-p/372327. They can and should be enhanced to handle AI-enabled attacks, but the fundamental approach remains valid.
类似模式也出现在其他领域:防御 AI 驱动的网络威胁需要加强现有的漏洞检测程序,而不是试图从源头限制 AI 能力。同样,关于 AI 生物风险的担忧最好在制造生物武器的采购和筛选阶段解决。
Similar patterns hold in other domains: Defending against AI-enabled cyberthreats requires strengthening existing vulnerability detection programs rather than attempting to restrict AI capabilities at the source. Similarly, concerns about bio risks of AI are best addressed at the procurement and screening stages for creating bioweapons.
我们不应仅将 AI 能力视为风险来源,而应认识到其防御潜力。在网络安全领域,AI 已通过自动化漏洞检测、威胁分析和攻击面监控来增强防御能力。69 69. Jennifer Tang, Tiffany Saade, and Steve Kelly. 2024. The implications of artificial intelligence in cybersecurity: shifting the offense-defense balance. https://securityandtechnology.org/wp-content/uploads/2024/10/The-Implications-of-Artificial-Intelligence-in-Cybersecurity.pdf
Rather than viewing AI capabilities solely as a source of risk, we should recognize their defensive potential. In cybersecurity, AI is already strengthening defensive capabilities through automated vulnerability detection, threat analysis, and attack surface monitoring.69 69. Jennifer Tang, Tiffany Saade, and Steve Kelly. 2024. The implications of artificial intelligence in cybersecurity: shifting the offense-defense balance. https://securityandtechnology.org/wp-content/uploads/2024/10/The-Implications-of-Artificial-Intelligence-in-Cybersecurity.pdf
为防御者提供强大的 AI 工具通常会使攻防平衡向有利于防御者的方向转变。这是因为防御者可以利用 AI 系统性地探测自身系统,在攻击者利用漏洞之前发现并修复它们。例如,谷歌最近将语言模型集成到其用于测试开源软件的模糊测试工具中,与传统方法相比,能够更有效地发现潜在安全问题。70 70. Dongge Liu et al. 2023. AI-Powered Fuzzing: Breaking the Bug Hunting Barrier. Google Online Security Blog.https://security.googleblog.com/2023/08/ai-powered-fuzzing-breaking-bug-hunting.html;.
Giving defenders access to powerful AI tools often improves the offense-defense balance in their favor. This is because defenders can use AI to systematically probe their own systems, finding and fixing vulnerabilities before attackers can exploit them. For example, Google recently integrated language models into their fuzzing tools for testing open-source software, allowing them to discover potential security issues more effectively compared to traditional methods.70 70. Dongge Liu et al. 2023. AI-Powered Fuzzing: Breaking the Bug Hunting Barrier. Google Online Security Blog.https://security.googleblog.com/2023/08/ai-powered-fuzzing-breaking-bug-hunting.html;.
同样的模式也适用于其他领域。在生物安全领域,AI 可以增强用于检测危险序列的筛查系统。71 71. Juan Cambeiro. How AI can help prevent biosecurity disasters. Institute for Progress (July 2023). https://ifp.org/how-ai-can-help-prevent-biosecurity-disasters/.在内容审核领域,它可以帮助识别协调性影响力操作。这些防御性应用表明,限制 AI 发展可能适得其反——我们需要强大的 AI 系统在防御方来应对 AI 引发的威胁。如果我们对齐语言模型,使其在这些任务(例如在关键网络基础设施中查找漏洞)上毫无用处,防御者将失去对这些强大系统的访问权。但动机明确的对手可以训练自己的 AI 工具进行此类攻击,导致进攻能力增强而防御能力没有相应提升。
The same pattern holds in other domains. In biosecurity, AI can enhance screening systems for detecting dangerous sequences.71 71. Juan Cambeiro. How AI can help prevent biosecurity disasters. Institute for Progress (July 2023). https://ifp.org/how-ai-can-help-prevent-biosecurity-disasters/.In content moderation, it can help to identify coordinated influence operations. These defensive applications show why restricting AI development could backfire—we need powerful AI systems on the defensive side to counter AI-enabled threats. If we align language models so that they are useless at these tasks (such as finding bugs in critical cyber infrastructure), defenders will lose access to these powerful systems. But motivated adversaries can train their own AI tools for such attacks, leading to an increase in offensive capabilities without a corresponding increase in defensive capabilities.
我们不应仅根据进攻能力来衡量 AI 风险,而应关注每个领域的攻防平衡等指标。此外,我们应认识到我们有能力使这种平衡向有利方向转变,并且可以通过投资防御性应用而非试图限制技术本身来实现这一点。
Rather than measuring AI risk solely in terms of offensive capabilities, we should focus on metrics like the offense-defense balance in each domain. Furthermore, we should recognize that we have the agency to shift this balance favorably, and can do so by investing in defensive applications rather than attempting to restrict the technology itself.
失调的 AI 会违背其开发者或用户的意图。(“对齐”一词有多种用法;我们在此搁置其他定义。)与滥用场景不同,这里没有怀有恶意的用户。与事故不同,系统按设计或指令运行,但由于完全且正确地指定目标的挑战,设计或指令本身与开发者或用户的意图不符。而与日常的失调案例(如聊天机器人中的有害输出)不同,我们关注的是高级 AI 导致灾难性或生存性危害的失调。
Misaligned AI acts against the intent of its developer or user. (The term alignment is used in many different ways; we set aside other definitions here.) Unlike misuse scenarios, there is no user acting with ill-intent. Unlike accidents, the system works as designed or commanded, but the design or command itself did not match the developer’s or user’s intent because of the challenge of completely and correctly specifying the objectives. And unlike everyday cases of misalignment, such as toxic outputs in a chatbot, our interest here is the misalignment of advanced AI causing catastrophic or existential harm.
我们认为,针对失调的主要防御同样位于下游。我们之前讨论的针对滥用所需的防御——从加固关键基础设施到改善网络安全——也将作为针对潜在失调风险的保护。
In our view, the primary defense against misalignment, again, lies downstream. The defenses needed against misuse that we discussed earlier—from hardening critical infrastructure to improving cybersecurity—will also serve as protection against potential misalignment risks.
在“AI 作为正常技术”的观点中,灾难性失调(迄今为止)是我们讨论的风险中最具推测性的。但什么是推测性风险——难道所有风险不都是推测性的吗?区别在于两种不确定性,以及相应的对概率的不同解释。
In the view of AI as normal technology, catastrophic misalignment is (by far) the most speculative of the risks that we discuss. But what is a speculative risk—aren’t all risks speculative? The difference comes down to the two types of uncertainty, and the correspondingly different interpretations of probability.
2025 年初,天文学家评估小行星 YR4 在 2032 年撞击地球的概率约为 2%,该概率反映了测量中的不确定性。在此类场景中,实际撞击概率(若无干预)要么是 0%,要么是 100%。进一步的测量解决了 YR4 案例中的这种“认知”不确定性。相反,当分析师预测未来十年核战争的风险为(比方说)10%时,该数字主要反映了源于未来如何展开的不可知性的“随机”不确定性,并且相对不太可能通过进一步观测得到解决。
In early 2025, when astronomers assessed that the asteroid YR4 had about a 2% probability of impact with the earth in 2032, the probability reflected uncertainty in measurement. The actual odds of impact (absent intervention) in such scenarios are either 0% or 100%. Further measurements resolved this “epistemic” uncertainty in the case of YR4. Conversely, when an analyst predicts that the risk of nuclear war in the next decade is (say) 10%, the number largely reflects ‘stochastic’ uncertainty arising from the unknowability of how the future will unfold, and is relatively unlikely to be resolved by further observations.
所谓推测性风险,我们指的是那些关于真实风险是否为零存在认知不确定性的风险——这种不确定性可以通过进一步的观测或研究得到解决。小行星 YR4 撞击的影响是一种推测性风险,而核战争则不是。
By speculative risks, we mean those for which there is epistemic uncertainty about whether or not the true risk is zero—uncertainty that can potentially be resolved through further observations or research. The impact of asteroid YR4 impact was a speculative risk, and nuclear war is not.
为了说明为什么灾难性失调是一种推测性风险,考虑一个著名的思想实验,该实验最初旨在展示失调的危险。它涉及一个“回形针最大化器”:一个目标是尽可能多地制造回形针的 AI。72 72. LessWrong. 2008. Squiggle maximizer (formerly “paperclip maximizer”). https://www.lesswrong.com/tag/squiggle-maximizer-formerly-paperclip-maximizer. 担忧在于 AI 会字面理解目标:它会意识到获取世界上的权力和影响力并控制所有资源有助于实现该目标。一旦它变得全能,它可能会征用所有世界资源,包括人类生存所需的资源,来生产回形针。
To illustrate why catastrophic misalignment is a speculative risk, consider a famous thought experiment originally intended to show the dangers of misalignment. It involves a “paperclip maximizer”: an AI that has the goal of making as many paperclips as possible.72 72. LessWrong. 2008. Squiggle maximizer (formerly “paperclip maximizer”). https://www.lesswrong.com/tag/squiggle-maximizer-formerly-paperclip-maximizer. The concern is that the AI will take the goal literally: It will realize that acquiring power and influence in the world and taking control over all of the world’s resources will help it to achieve that goal. Once it is all powerful, it might commandeer all of the world’s resources, including those needed for humanity’s survival, to produce paperclips.
AI 系统可能灾难性地误解指令的恐惧依赖于关于技术如何在现实世界中部署的可疑假设。在系统被授予对关键决策的访问权限之前,它需要在不太关键的上下文中展示可靠的性能。任何过于字面理解指令或缺乏常识的系统都会在这些早期测试中失败。
The fear that AI systems might catastrophically misinterpret commands relies on dubious assumptions about how technology is deployed in the real world. Long before a system would be granted access to consequential decisions, it would need to demonstrate reliable performance in less critical contexts. Any system that interprets commands over-literally or lacks common sense would fail these earlier tests.
考虑一个更简单的案例:一个机器人被要求“尽快从商店取回回形针”。一个字面理解此指令的系统可能会无视交通法规或试图盗窃。这种行为会导致立即关闭和重新设计。采用的路径本质上要求系统在日益重要的情境中展示适当的行为。这不是一个幸运的意外,而是组织采用技术的基本特征。
Consider a simpler case: A robot is asked to "get paperclips from the store as quickly as possible." A system that interpreted this literally might ignore traffic laws or attempt theft. Such behavior would lead to immediate shutdown and redesign. The path to adoption inherently requires demonstrating appropriate behavior in increasingly consequential situations. This is not a lucky accident, but is a fundamental feature of how organizations adopt technology.
这种担忧的一个更复杂的版本基于欺骗性对齐的概念:这指的是系统在评估或部署早期阶段看似对齐,但一旦获得足够力量就释放有害行为。在领先的 AI 模型中已经观察到一定程度的欺骗现象。73 73. Ryan Greenblatt et al. 2024. Alignment faking in large language models. Retrieved from https://arxiv.org/abs/2412.14093.
A more sophisticated version of this concern is based on the concept of deceptive alignment: This refers to a system appearing to be aligned during evaluation or the early stages of deployment, but unleashing harmful behavior once it has acquired enough power. Some level of deceptive phenomena has already been observed in leading AI models.73 73. Ryan Greenblatt et al. 2024. Alignment faking in large language models. Retrieved from https://arxiv.org/abs/2412.14093.
根据超级智能观点,欺骗性对齐是一个定时炸弹——由于超级智能,系统将轻易击败人类任何检测其是否真正对齐的尝试,并等待时机。但是,在正常技术观点中,欺骗只是一个工程问题,尽管重要,需要在开发过程中和整个部署过程中解决。事实上,它已经是强大 AI 模型安全评估的标准部分。74 74. Bowen Baker et al. 2025. Monitoring reasoning models for misbehavior and the risks of promoting obfuscation, Retrieved from https://arxiv.org/abs/2503.11926.
According to the superintelligence view, deceptive alignment is a ticking time bomb—being superintelligent, the system will easily be able to defeat any human attempts to detect if it is actually aligned and will bide its time. But, in the normal technology view, deception is a mere engineering problem, albeit an important one, to be addressed during development and throughout deployment. Indeed, it is already a standard part of the safety evaluation of powerful AI models.74 74. Bowen Baker et al. 2025. Monitoring reasoning models for misbehavior and the risks of promoting obfuscation, Retrieved from https://arxiv.org/abs/2503.11926.
关键的是,AI 在此过程中是有用的,AI 的进步不仅使欺骗成为可能,也改进了欺骗的检测。如同网络安全的情况,防御者拥有许多不对称优势,包括能够检查目标系统的内部(这一优势有多大用处取决于系统如何设计以及我们在可解释性技术上投入多少)。另一个优势是纵深防御,许多防御不仅针对滥用,也针对失调的 AI,将位于 AI 系统的下游。
Crucially, AI is useful in this process, and advances in AI not only enable deception, but also improve the detection of deception. As in the case of cybersecurity, the defender has many asymmetric advantages, including being able to examine the internals of the target system (how useful this advantage is depends on how the system is designed and how much we invest in interpretability techniques). Another advantage is defense in depth, and many defenses against not just misuse but also unaligned AI will be located downstream of the AI system.
失调的担忧通常假设 AI 系统将自主运行,在没有人类监督的情况下做出高风险决策。但正如我们在第二部分中论证的,人类控制仍将是 AI 部署的核心。围绕关键决策的现有制度控制——从财务控制到安全法规——创造了多层保护,防止灾难性失调。
Misalignment concerns often presume that AI systems will operate autonomously, making high-stakes decisions without human oversight. But as we argued in Part II, human control will remain central to AI deployment. Existing institutional controls around consequential decisions—from financial controls to safety regulations—create multiple layers of protection against catastrophic misalignment.
某些技术设计决策比其他决策更可能导致失调。一个臭名昭著的设置是使用强化学习在长时间范围内优化单一目标函数(可能被意外地欠指定或错误指定)。游戏智能体中有一长串有趣的例子,例如一个赛船智能体学会了无限循环一个区域以击中相同目标并得分,而不是向终点线前进。75 75. Victoria Krakovna. 2020. Specification gaming: The flip side of AI ingenuity. Google DeepMind (April 2020). https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity/. 重申一下,我们认为在开放式的现实世界场景中,以这种方式设计的智能体将更无效而非危险。无论如何,研究不易受规范游戏影响的其他设计范式是一个重要的研究方向。76 76. Simon Dima et al. 2024. Non-maximizing policies that fulfill multi-criterion aspirations in expectation. arXiv: August 2024. Retrieved from http://arxiv.org/abs/2408.04385.
Some technical design decisions are more likely to lead to misalignment than others. One setting that is notorious for this is the use of reinforcement learning to optimize a single objective function (which might be accidentally underspecified or misspecified) over a long time horizon. There is a long list of amusing examples from game agents, such as a boat racing agent that learned to indefinitely circle an area to hit the same targets and score points instead of progressing to the finish line.75 75. Victoria Krakovna. 2020. Specification gaming: The flip side of AI ingenuity. Google DeepMind (April 2020). https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity/. To reiterate, we think that in open-ended real-world scenarios, agents that are designed this way will be more ineffective than they will be dangerous. In any case, research on alternative design paradigms that are less susceptible to specification gaming is an important research direction.76 76. Simon Dima et al. 2024. Non-maximizing policies that fulfill multi-criterion aspirations in expectation. arXiv: August 2024. Retrieved from http://arxiv.org/abs/2408.04385.
简而言之,回形针最大化器场景风险非零的论点依赖于可能成立也可能不成立的假设,并且有理由认为研究可以让我们更好地了解这些假设对于正在构建或设想的 AI 系统类型是否成立。出于这些原因,我们称之为“推测性”风险,并在第四部分中审视这一观点的政策含义。
In short, the argument for a nonzero risk of a paperclip maximizer scenario rests on assumptions that may or may not be true, and it is reasonable to think that research can give us a better idea of whether these assumptions hold true for the kinds of AI systems that are being built or envisioned. For these reasons, we call it a ‘speculative’ risk, and examine the policy implications of this view in Part IV.
虽然上述讨论的风险可能具有灾难性或存在性,但还有一长串低于这一水平但仍属大规模和系统性的 AI 风险,它们超越了任何特定 AI 系统的直接影响。这些风险包括偏见和歧视的系统性固化、特定职业的大规模失业、劳动条件恶化、不平等加剧、权力集中、社会信任侵蚀、信息生态系统污染、自由新闻衰落、民主倒退、大规模监控以及助长威权主义。
While the risks discussed above have the potential to be catastrophic or existential, there is a long list of AI risks that are below this level but which are nonetheless large-scale and systemic, transcending the immediate effects of any particular AI system. These include the systemic entrenchment of bias and discrimination, massive job losses in specific occupations, worsening labor conditions, increasing inequality, concentration of power, erosion of social trust, pollution of the information ecosystem, decline of the free press, democratic backsliding, mass surveillance, and enabling authoritarianism.
如果 AI 是普通技术,这些风险将变得比上述灾难性风险重要得多。这是因为这些风险源于人们和组织利用 AI 来推进自身利益,而 AI 仅仅充当了放大社会现有不稳定性的放大器。
If AI is normal technology, these risks become far more important than the catastrophic ones discussed above. That is because these risks arise from people and organizations using AI to advance their own interests, with AI merely serving as an amplifier of existing instabilities in our society.
在变革性技术的历史中,这类社会政治动荡有大量先例。值得注意的是,工业革命导致了快速的大规模城市化,其特点是恶劣的工作条件、剥削和不平等,既催化了工业资本主义,也催生了社会主义和马克思主义作为回应。77 77. Daron Acemoglu and Simon Johnson. 2023. Power and Progress. PublicAffairs.
There is plenty of precedent for these kinds of socio-political disruption in the history of transformative technologies. Notably, the Industrial Revolution led to rapid mass urbanization that was characterized by harsh working conditions, exploitation, and inequality, catalyzing both industrial capitalism and the rise of socialism and Marxism in response.77 77. Daron Acemoglu and Simon Johnson. 2023. Power and Progress. PublicAffairs.
我们建议的焦点转移大致对应于 Kasirzadeh 对决定性 x 风险和累积性 x 风险的区分。决定性 x 风险涉及“明显的 AI 接管路径,以不可控超级智能等场景为特征”,而累积性 x 风险指的是“关键 AI 引发的威胁(如严重脆弱性和经济政治结构的系统性侵蚀)的逐渐积累”。78 78. Atoosa Kasirzadeh. 2024. Two types of AI existential risk: Decisive and accumulative. arXiv: preprint. Retrieved from https://arxiv.org/abs/2401.07836, February 2024), doi:10.48550/arXiv.2401.07836. 但存在重要差异:Kasirzadeh 对累积性风险的描述在很大程度上仍然依赖于网络攻击者等威胁行为者,而我们的关注点仅仅是当前资本主义的路径。并且我们认为这类风险不太可能是存在性的,但仍然极其严重。
The shift in focus that we recommend roughly maps onto Kasirzadeh’s distinction between decisive and accumulative x-risk. Decisive x-risk involves “overt AI takeover pathway, characterized by scenarios like uncontrollable superintelligence,” whereas accumulative x-risk refers to “a gradual accumulation of critical AI-induced threats such as severe vulnerabilities and systemic erosion of econopolitical structures.”78 78. Atoosa Kasirzadeh. 2024. Two types of AI existential risk: Decisive and accumulative. arXiv: preprint. Retrieved from https://arxiv.org/abs/2401.07836, February 2024), doi:10.48550/arXiv.2401.07836. But there are important differences: Kasirzadeh’s account of accumulative risk still relies on threat actors such as cyberattackers to a large extent, whereas our concern is simply about the current path of capitalism. And we think that such risks are unlikely to be existential, but are still extremely serious.
人工智能不同未来——正常技术与可能无法控制的超级智能——之间的分歧给政策制定者带来了一个两难困境,因为针对一类风险的防御措施可能会使另一类风险恶化。我们提供了一套原则来应对这种不确定性。更具体地说,政策制定者应聚焦的策略是韧性,即现在采取行动以提高我们应对未来意外发展的能力。政策制定者应拒绝不扩散,这违反了我们的原则,并降低了韧性。最后,扩散面临的阻力意味着实现人工智能的好处并非必然,需要政策制定者采取行动。
The divergence between the different futures of AI—normal technology versus potentially uncontrollable superintelligence—introduces a dilemma for policymakers because defenses against one set of risks might make the other worse. We provide a set of principles for navigating this uncertainty. More concretely, the strategy that policymakers should center is resilience, which consists of taking actions now to improve our ability to deal with unexpected developments in the future. Policymakers should reject nonproliferation, which violates the principles we outline, and decreases resilience. Finally, the headwinds against diffusion mean that achieving the benefits of AI is not guaranteed and requires action from policymakers.
关于人工智能治理的讨论已经很多。我们的目标不是提出一个全面的治理框架;我们只是强调将人工智能视为正常技术的政策含义。
Much has been said about AI governance. Our goal is not to present a comprehensive governance framework; we merely highlight the policy implications of the view of AI as normal technology.
当今的 AI 安全讨论以世界观上的深刻分歧为特征。我们认为这些分歧不太可能消失。根深蒂固的阵营已经形成:AI 安全联盟早已建立,而那些对灾难性风险持更怀疑态度的人在 2024 年聚集起来,尤其是在关于加州 AI 安全法案的辩论过程中。79 79. Anton Leicht. 2024. AI safety politics after the SB-1047 veto. https://www.antonleicht.me/writing/veto。同样,AI 安全阵营的学术根源要古老得多,而采用正常技术范式的学术研究正在逐渐形成;我们自身工作(包括本文)的目标之一就是将常态主义思维建立在更坚实的学术基础上。80 80. Timothy B. Lee. 2024. Six Principles for Thinking about AI Risk. https://www.understandingai.org/p/six-principles-for-thinking-about。
Today’s AI safety discourse is characterized by deep differences in worldviews. We think that these differences are unlikely to go away. Entrenched camps have developed: The AI safety coalition is already well established, whereas those who were more skeptical of catastrophic risks coalesced in 2024, especially in the course of the debate about California’s AI safety bill.79 79. Anton Leicht. 2024. AI safety politics after the SB-1047 veto. https://www.antonleicht.me/writing/veto. Similarly, the intellectual roots of the AI safety camp are much older, whereas scholarship that adopts that normal technology paradigm is gradually taking shape; the goal of much of our own work, including this paper, is to put normalist thinking on firmer intellectual footing.80 80. Timothy B. Lee. 2024. Six Principles for Thinking about AI Risk. https://www.understandingai.org/p/six-principles-for-thinking-about.
我们支持减少社区内两极分化和碎片化的呼吁。81 81. Mary Phuong et al. 2024. Evaluating frontier models for dangerous capabilities. Retrieved from https://arxiv.org/abs/2403.13793。但即使我们改善了讨论的氛围,我们仍可能面临世界观和认知实践上的分歧,而这些分歧不太可能通过经验得到解决。82 82. Shazeda Ahmed et al. 2024. Field-building and the epistemic culture of AI safety. First Monday 29, 4. https://firstmon day.org/ojs/index.php/fm/article/view/13626/11596。因此,“专家”之间关于 AI 风险的共识不太可能达成。两个阵营设想的 AI 风险情景性质截然不同,商业行为者应对这些风险的能力和动机也大相径庭。面对这种不确定性,政策制定者应如何行动?
We support calls for decreasing polarization and fragmentation in the community.81 81. Mary Phuong et al. 2024. Evaluating frontier models for dangerous capabilities. Retrieved from https://arxiv.org/abs/2403.13793. But even if we improve the tenor of the discourse, we are likely to be left with differences in worldviews and epistemic practices that are unlikely to be empirically resolved.82 82. Shazeda Ahmed et al. 2024. Field-building and the epistemic culture of AI safety. First Monday 29, 4. https://firstmonday.org/ojs/index.php/fm/article/view/13626/11596. So, consensus among ‘experts’ about AI risks is unlikely. The nature of the AI risk scenarios envisioned by the two camps differs drastically, as do the ability and incentives for commercial actors to counteract these risks. How should policymakers proceed in the face of this uncertainty?
政策制定中的自然倾向是妥协。但这不太可能奏效。某些干预措施,如提高透明度,对风险缓解无条件有益,无需妥协(或者更确切地说,政策制定者必须平衡行业和外部利益相关者的利益,这基本上是一个正交的维度)。83 83. Arvind Narayanan and Sayash Kapoor. 2024. AI existential risk probabilities are too unreliable to inform policy. https://www.aisnakeoil.com/p/ai-existential-risk-probabilities; Neel Guha et al. 2023. AI regulation has its own alignment problem: The technical and institutional feasibility of disclosure, registration, licensing, and auditing. SSRN (November 2023). https://papers.ssrn.com/abstract=4634442。其他干预措施,如不扩散,可能有助于遏制超级智能,但会通过增加市场集中度而加剧与正常技术相关的风险。84 84. Christopher A. Mouton, Caleb Lucas, and Ella Guest. 2024. The operational risks of AI in large-scale biological attacks: Results of a red-team study. RAND Corporation. https://www.rand.org/pubs/research_reports/RRA2977-2.html; Ari Takanen, Jared D. Demott, and Charles Miller. 2008. Fuzzing for Software Security Testing and Quality Assurance. Fuzzing for Software Security (1st ed.). Artech House Publishers, Norwood, MA。反之亦然:通过促进开源 AI 来提高韧性的干预措施有助于治理正常技术,但可能引发失控的超级智能。
A natural inclination in policymaking is compromise. This is unlikely to work. Some interventions, such as improving transparency, are unconditionally helpful for risk mitigation, no compromise is needed (or rather, policymakers will have to balance the interests of the industry and external stakeholders, which is a mostly orthogonal dimension). 83 83. Arvind Narayanan and Sayash Kapoor. 2024. AI existential risk probabilities are too unreliable to inform policy. https://www.aisnakeoil.com/p/ai-existential-risk-probabilities; Neel Guha et al. 2023. AI regulation has its own alignment problem: The technical and institutional feasibility of disclosure, registration, licensing, and auditing. SSRN (November 2023). https://papers.ssrn.com/abstract=4634443. Other interventions, such as nonproliferation, might help to contain a superintelligence but exacerbate the risks associated with normal technology by increasing market concentration.84 84. Christopher A. Mouton, Caleb Lucas, and Ella Guest. 2024. The operational risks of AI in large-scale biological attacks: Results of a red-team study. RAND Corporation. https://www.rand.org/pubs/research_reports/RRA2977-2.html; Ari Takanen, Jared D. Demott, and Charles Miller. 2008. Fuzzing for Software Security Testing and Quality Assurance. Fuzzing for Software Security (1st ed.). Artech House Publishers, Norwood, MA. The reverse is also true: Interventions such as increasing resilience by fostering open-source AI will help to govern normal technology, but risk unleashing out-of-control superintelligence.
这种张力是不可避免的。防御超级智能需要人类团结起来对抗共同敌人,可以这么说,集中权力并对 AI 技术实施中央控制。但我们更担心的是人们利用 AI 实现自身目的所带来的风险,无论是恐怖主义、网络战、破坏民主,还是最普遍地——加剧不平等的榨取性资本主义实践。85 85. Sayash Kapoor and Arvind Narayanan. 2023. Licensing is neither feasible nor effective for addressing ai risks (June 2023), https://www.aisnakeoil.com/p/licensing-is-neither-feasible-nor。防御这类风险需要防止权力和资源的集中(这通常意味着让强大的 AI 更广泛可用),从而提高韧性。
The tension is inescapable. Defense against superintelligence requires humanity to unite against a common enemy, so to speak, concentrating power and exercising central control over AI technology. But we are more concerned about risks that arise from people using AI for their own ends, whether terrorism, or cyberwarfare, or undermining democracy, or simply—and most commonly—extractive capitalistic practices that magnify inequalities.85 85. Sayash Kapoor and Arvind Narayanan. 2023. Licensing is neither feasible nor effective for addressing ai risks (June 2023), https://www.aisnakeoil.com/p/licensing-is-neither-feasible-nor. Defending against this category of risk requires increasing resilience by preventing the concentration of power and resources (which often means making powerful AI more widely available).
应对不确定性的另一种诱人方法是估计各种结果的概率,然后应用成本效益分析。AI 安全社区严重依赖灾难性风险(尤其是生存风险)的概率估计来指导政策制定。这个想法很简单:如果我们认为一个结果具有主观价值或效用 U(可以是正或负),并且它有比如 10%的发生概率,我们可以将其视为必然发生且价值为 0.1 * U。然后,我们可以将每个可用选项的成本和收益相加,并选择最大化成本减去收益的选项(即“期望效用”)。
Another tempting approach to navigating uncertainty is to estimate the probabilities of various outcomes and to then apply cost-benefit analysis. The AI safety community relies heavily on probability estimates of catastrophic risk, especially existential risk, to inform policy making. The idea is simple: If we consider an outcome to have a subjective value, or utility, of U (which can be positive or negative), and it has, say, a 10% probability of occurring, we can act as if it is certain to occur and has a value of 0.1 * U. We can then add up the costs and benefits for each option available to us, and choose the one that maximizes costs minus benefits (the ‘expected utility’).
在最近的一篇文章中,我们解释了为什么这种方法不可行。86 86. Arvind Narayanan and Sayash Kapoor. 2024. “AI existential risk probabilities are too unreliable to inform policy. https://www.aisnakeoil.com/p/ai-existential-risk-probabilities。AI 风险概率缺乏有意义的认识论基础。有根据的概率估计可以是归纳性的,基于类似过去事件的参考类别,例如汽车保险定价中的车祸;也可以是演绎性的,基于所讨论现象的精确模型,如扑克。不幸的是,对于 AI 风险,既没有有用的参考类别,也没有精确的模型。在实践中,风险估计是“主观的”——预测者的个人判断。87 87. Richard Blumenthal and Josh Hawley. 2023. Bipartisan framework for U.S. AI act. https://www.blumenthal.senate.gov/imo/media/doc/09072023bipartisanaiframework.pdf。由于缺乏任何依据,这些估计往往差异巨大,通常相差几个数量级。
In a recent essay, we explained why this approach is unviable.86 86. Arvind Narayanan and Sayash Kapoor. 2024. “AI existential risk probabilities are too unreliable to inform policy. https://www.aisnakeoil.com/p/ai-existential-risk-probabilities. AI risk probabilities lack meaningful epistemic foundations. Grounded probability estimation can be inductive, based on a reference class of similar past events, such as car accidents for auto insurance pricing. Or it can be deductive, based on precise models of the phenomenon in question, as in poker. Unfortunately, there is no useful reference class nor precise models when it comes to AI risk. In practice, risk estimates are ‘subjective’—forecasters’ personal judgments.87 87. Richard Blumenthal and Josh Hawley. 2023. Bipartisan framework for U.S. AI act. https://www.blumenthal.senate.gov/imo/media/doc/09072023bipartisanaiframework.pdf. Lacking any grounding, these tend to vary wildly, often by orders of magnitude.
除了概率之外,计算的其他组成部分——各种政策选择(包括不作为)的后果——也面临巨大的不确定性,不仅在幅度上,而且在方向上。没有可靠的方法来量化因限制 AI 可用性的政策而放弃的收益,我们在下文中论证,不扩散可能使灾难性风险变得更糟。
In addition to the probabilities, the other components of the calculation—the consequences of various policy choices, including inaction—are also subject to massive uncertainties, not just in magnitude but also in direction. There is no reliable way to quantify the benefits we forego due to policies that restrict the availability of AI, and we argue below that nonproliferation might make catastrophic risks worse.
此外,我们赋予某些结果的效用可能取决于我们的道德价值观。例如,有些人可能认为灭绝具有难以想象的巨大负效用,因为它排除了未来可能存在的所有人类生命(无论是物理的还是模拟的)。88 88. Sigal Samuel. 2022. Effective altruism’s most controversial idea. https://www.vox.com/future-perfect/23298870/effective-altruism-longtermism-will-macaskill-future。(当然,涉及无穷大的成本效益分析往往会导致荒谬的结论)。
Furthermore, the utility we attach to certain outcomes might depend on our moral values. For example, some people might consider extinction to have an unfathomably large negative utility because it precludes all of the human lives, physical or simulated, that might exist in the future.88 88. Sigal Samuel. 2022. Effective altruism’s most controversial idea. https://www.vox.com/future-perfect/23298870/effective-altruism-longtermism-will-macaskill-future. (Of course, cost-benefit analysis involving infinities tends to lead to absurd conclusions).
另一个例子是限制自由和不限制自由的政策之间的不对称(例如,要求开发某些 AI 模型需要许可证,与增加资金开发针对 AI 风险的防御措施)。某些类型的限制违反了自由民主的核心原则,即国家不应基于合理人士可以拒绝的有争议信念来限制人们的自由。正当性对于政府的合法性和权力的行使至关重要。89 89. Kevin Vallier. 1996. Public justification. https://plato.stanford.edu/entries/justification-public/。目前尚不清楚如何量化违反这一原则的成本。
Another example is the asymmetry between policies that do and do not restrict freedoms (such as requiring licenses for developing certain AI models versus increasing funding for developing defenses against AI risks). Certain kinds of restrictions violate a core principle of liberal democracy, namely that the state should not limit people’s freedom based on controversial beliefs that reasonable people can reject. Justification is essential for the legitimacy of government and the exercise of power.89 89. Kevin Vallier. 1996. Public justification. https://plato.stanford.edu/entries/justification-public/. It is unclear how to quantify the cost of violating such a principle.
当然,正当性的重要性可以在规范上进行辩论,但从经验上看,迄今为止在 AI 政策中似乎得到了证实。如前所述,加州的 AI 安全法规导致了反对该法案的阵营的联合。反对阵营中的一些成员是自私的公司,但其他成员是学者和进步倡导者。根据我们的经验,第二组的驱动动机在许多情况下是政府被认为越过了其合法权力的边界,因为对于那些不认同该法案未言明前提的人来说,所提供的理由非常缺乏说服力。
The importance of justification can, of course, be normatively debated, but empirically it seems to be borne out thus far in AI policy. As mentioned earlier, California’s AI safety regulation led to the coalescence of those opposed to the bill. Some members of the oppositional camp were self-interested companies, but others were scholars and advocates for progress. In our experience, the driving motivation for the second group in many cases was the government’s perceived overstepping of the bounds of its legitimate authority, given how unconvincing the proffered justifications were for those who did not subscribe to the bill’s unstated premises.
价值观和信仰上不可避免的差异意味着政策制定者必须采纳价值多元主义,偏好那些能被具有广泛价值观的利益相关者接受的政策,并尽量避免那些可以被利益相关者合理拒绝的自由限制。他们还必须优先考虑稳健性,偏好那些即使其关键假设被证明不正确,仍然有用或至少无害的政策。90 90. Jeffrey A Friedman and Richard Zeckhauser. 2018. Analytic confidence and political decision-making: Theoretical principles and experimental evidence from national security professionals. Political Psychology 39, 5 (2018), 1069–87。
Unavoidable differences in values and beliefs mean that policymakers must adopt value pluralism, preferring policies that are acceptable to stakeholders with a wide range of values, and attempt to avoid restrictions on freedom that can reasonably be rejected by stakeholders. They must also prioritize robustness, preferring policies that remain helpful, or at least not harmful, if the key assumptions underpinning them turn out to be incorrect.90 90. Jeffrey A Friedman and Richard Zeckhauser. 2018. Analytic confidence and political decision-making: Theoretical principles and experimental evidence from national security professionals. Political Psychology 39, 5 (2018), 1069–87.
尽管由于上述原因,不确定性无法消除,但可以减少。然而,这一目标不应留给专家;政策制定者可以而且应该发挥积极作用。我们推荐五种具体方法。
While uncertainty cannot be eliminated for the reasons described above, it can be reduced. However, this goal should not be left to experts; policymakers can and should play an active role. We recommend five specific approaches.
_图 6. 可增强关于 AI 使用、风险和失败公共信息的几类政策概览。_ 91 91. Arvind Narayanan 和 Sayash Kapoor. 2023. 生成式 AI 公司必须发布透明度报告。Knight 第一修正案研究所。http://knightcolumbia.org/blog/generative-ai-companies-must-publish-transparency-reports;总统行政办公室. 2020. 促进联邦政府使用可信人工智能。https://www.federalregister.gov/documents/2020/12/08/2020-27065/promoting-the-use-of-trustworthy-artificial-intelligence-in-the-federal-government, 2020;Justin Colannino. 2021. 版权局扩大你的安全研究权利。GitHub 博客。https://github.blog/security/vulnerability-research/copyright-office-expands-security-research-rights/。
_Figure 6. Overview of a few types of policies that can enhance public information about AI use, risks, and failures._ 91 91. Arvind Narayanan and Sayash Kapoor. 2023. Generative AI companies must publish transparency reports. Knight First Amendment Institute. http://knightcolumbia.org/blog/generative-ai-companies-must-publish-transparency-reports; Executive Office of the President. 2020. Promoting the use of trustworthy artificial intelligence in the federal government. https://www.federalregister.gov/documents/2020/12/08/2020-27065/promoting-the-use-of-trustworthy-artificial-intelligence-in-the-federal-government, 2020; Justin Colannino. 2021. The copyright office expands your security research rights. GitHub Blog. https://github.blog/security/vulnerability-research/copyright-office-expands-security-research-rights/.
_对风险研究的战略资助。_ 当前的 AI 安全研究过度关注有害能力,并未采纳正常技术观。对技术能力下游的问题关注不足。例如,关于威胁行为者实际如何使用 AI 的知识严重匮乏。AI 事件数据库等努力存在且有价值,但数据库中的事件来源于新闻报道而非研究,这意味着它们经过了这些事件成为新闻的选择性和有偏过程的过滤。92 92. AI 事件数据库. 无日期. https://incidentdatabase.ai/。
_Strategic funding of research on risks._ Current AI safety research focuses heavily on harmful capabilities and does not embrace the normal technology view. Insufficient attention has been paid to questions that are downstream of technical capabilities. For example, there is a striking dearth of knowledge regarding how threat actors actually use AI. Efforts such as the AI Incident Database exist and are valuable, but incidents in the database are sourced from news reports rather than through research, which means that they are filtered through the selective and biased process by which such incidents become news.92 92. AI Incident Database. n.d. https://incidentdatabase.ai/.
幸运的是,研究资助是一个容易达成妥协的领域;我们主张增加对风险(和收益)研究的资助,以解决在正常技术观下更相关的问题。其他可能减少或至少澄清不确定性的研究类型包括证据综合工作以及不同世界观研究者之间的对抗性合作。
Fortunately, research funding is an area in which compromise is healthy; we advocate for increased funding of research on risks (and benefits) that tackles questions that are more relevant under the normal technology view. Other kinds of research that might reduce, or at least clarify, uncertainty are evidence synthesis efforts and adversarial collaborations among researchers with different worldviews.
_对 AI 使用、风险和失败的监测。_ 虽然研究资助有助于监测现实世界中的 AI,但也可能需要监管和政策——即“寻求证据的政策”。93 93. Stephen Casper, David Krueger, 和 Dylan Hadfield-Menell. 2025. 基于证据的 AI 政策的陷阱。取自 https://arxiv.org/abs/2502.09618。我们在图 6 中建议了若干此类政策。
_Monitoring of AI use, risks, and failures._ While research funding can help with monitoring AI in the wild, it might also require regulation and policy—that is, “evidence-seeking policies.”93 93. Stephen Casper, David Krueger, and Dylan Hadfield-Menell. 2025. Pitfalls of evidence-based AI policy. Retrieved from https://arxiv.org/abs/2502.09618. We suggest a few such policies in Figure 6.
_关于不同证据价值的指导。_ 政策制定者可以帮助研究社区更好地理解哪些证据是有用且可操作的。例如,多位政策制定者和咨询机构已指出“边际风险”框架对于分析开放权重和专有模型的相对风险很有用,这有助于指导研究人员未来的研究。94 94. Sayash Kapoor 等人. 2024. 关于开放基础模型的社会影响。取自 https://arxiv.org/abs/2403.07918。
_Guidance on the value of different kinds of evidence._ Policymakers can provide the research community with a better understanding of what kinds of evidence are useful and actionable. For example, various policymakers and advisory bodies have indicated the usefulness of the “marginal risk” framework for analyzing the relative risks of open-weight and proprietary models, which is helpful to researchers in guiding future research.94 94. Sayash Kapoor et al. 2024. On the societal impact of open foundation models. Retrieved from https://arxiv.org/abs/2403.07918.
_将证据收集作为首要目标。_ 到目前为止,我们讨论了专门旨在产生更好证据或减少不确定性的行动。更广泛地说,证据收集的影响可以被视为评估任何 AI 政策的一个因素,与最大化收益和最小化风险的影响并列。例如,支持开放权重和开源模型的一个理由可能是推进 AI 风险研究。相反,支持专有模型的一个理由可能是对其使用和部署的监控更容易。
_Evidence gathering as a first-rate goal._ So far, we have discussed actions that are specifically intended to generate better evidence or to reduce uncertainty. More broadly, the impact on evidence gathering can be considered to be a factor in evaluating any AI policy, alongside the impact on maximizing benefits and minimizing risks. For example, one reason to favor open-weight and open-source models could be to advance research on AI risks. Conversely, one reason to favor proprietary models might be that surveillance of their use and deployment might be easier.
Marchant 和 Stevens 描述了治理新兴技术的四种方法,见图 7。其中两种是事前方法——风险分析和预防,另外两种是事后方法——责任和韧性。这些方法各有优缺点,可以相互补充;尽管如此,某些方法显然比其它方法更适合某些技术。
Marchant and Stevens described four approaches to governing emerging technologies; see Figure 7.95 95. Gary E. Marchant and Yvonne A. Stevens. 2017. Resilience: A new tool in the risk governance toolbox for emerging technologies. UC Davis Law Review.https://lawreview.law.ucdavis.edu/sites/g/files/dgvnsk15026/files/media/documents/51-1_Marchant_Stevens.pdf. Two are _ex ante_, risk analysis and precaution, and the other two are _ex post_, liability and resilience. These approaches have different pros and cons and can complement each other; nonetheless, some approaches are clearly better suited to some technologies than others.
Marchant 和 Stevens 认为(我们同意)事前方法不适合 AI,因为难以在部署前确定风险。责任方法效果更好,但也有重要局限性,包括因果关系的不确定性以及可能对技术发展产生的寒蝉效应。
Marchant and Stevens argued (and we agree) that _ex ante_ approaches are poorly suited to AI because of the difficulty of ascertaining risks in advance of deployment. Liability fares better, but also has important limitations, including uncertainty about causation and the chilling effects it might exert on technology development.
图 7. 基于 Marchant 和 Stevens 的四种新兴技术治理方法总结。
_Figure 7. Summary of four approaches to governing emerging technology, based on Marchant and Stevens._
在 AI 的背景下,伤害可能源于特定部署系统中的事件,无论这些事件是事故还是攻击。还存在可能造成伤害的冲击,包括攻击能力的突然增强(例如使生物恐怖分子成为可能)以及能力的突然扩散,例如通过发布开放权重模型或窃取专有模型的权重。我们认为,韧性既需要在伤害发生时最小化其严重性,也需要在冲击发生时最小化伤害的可能性。
In the context of AI, harms may result from incidents in specific deployed systems, regardless of whether these incidents are accidents or attacks. There are also shocks that may or may not result in harms, including sudden increases in offensive capabilities (such as enabling bioterrorists) and a sudden proliferation of capabilities, such as through the release of an open-weight model or theft of the weights of a proprietary model. In our view, resilience requires both minimizing the severity of harm when it does occur and minimizing the likelihood of harm when shocks do occur.
韧性结合了事前和事后方法的要素,包括在伤害发生前采取行动,以便在伤害实际发生时更好地限制损害。许多基于韧性的治理工具有助于缓解“追赶问题”,即传统治理方法无法跟上技术发展的速度。
Resilience combines elements of _ex ante_ and _ex post_ approaches, and consists of taking actions before harm occurs in order to be in a better position to limit the damage when harm does occur. Many resilience-based governance tools help to mitigate the _pacing problem_, wherein traditional governance approaches are unable to keep pace with the speed of technological development.
针对 AI 已经提出了许多韧性策略。它们可以分为四大类。前三类包括“无悔”政策,无论 AI 的未来如何,这些政策都会有所帮助。
Many resilience strategies have been proposed for AI. They can be grouped into four broad categories. The first three consist of “no regret” policies that will help regardless of the future of AI.
社会韧性,广义上:重要的是加倍努力保护民主的基础,特别是那些被 AI 削弱的方面,如自由新闻和公平的劳动力市场。AI 的进步并不是现代社会面临的唯一冲击,甚至不是唯一的技术冲击,因此无论 AI 的未来如何,这些政策都会有所帮助。
Societal resilience, broadly: It is important to redouble efforts to protect the foundations of democracy, especially those weakened by AI, such as the free press and equitable labor markets. Advances in AI are not the only shocks, or even the only technology shocks, that modern societies face, so these policies will help regardless of the future of AI.
* 有效技术防御和政策制定的先决条件:这些干预措施通过加强技术和机构能力来支持下一类措施。例子包括资助更多关于 AI 风险的研究、对高风险 AI 系统开发者的透明度要求、建立信任并减少 AI 社区的分裂、增加政府中的技术专长、加强 AI 方面的国际合作以及提高 AI 素养。这些将有助于建立技术和机构能力以减轻 AI 风险,即使我们发现自己对 AI 当前或未来的影响判断有误。
* Prerequisites for effective technical defenses and policymaking: These interventions enable those in the next category by strengthening technical and institutional capacity. Examples include funding more research on AI risks, transparency requirements for developers of high-stakes AI systems, building trust and reducing fragmentation in the AI community, increasing technical expertise in government, increasing international cooperation on AI, and improving AI literacy.98 98. Rishi Bommasani et al. 2024. A path for science- and evidence-based AI policy. https://understanding-ai-safety.org/; Balint Gyevnar and Atoosa Kasirzadeh. 2025. AI safety for everyone. Retrieved from https://arxiv.org/abs/2502.09288; Anka Reuel et al. 2024. Position: Technical research and talent is needed for effective AI governance. In Proceedings of the 41st International Conference on Machine Learning (PMLR, 2024), 42543–57. https://proceedings.mlr.press/v235/reuel24a.html. These will help to build technical and institutional capacities to mitigate AI risks even if it turns out that we have been wrong about the present or future impact of AI.
* 无论 AI 未来如何都会有所帮助的干预措施:包括开发早期预警系统、开发针对已识别 AI 风险的防御措施、激励防御者(例如网络攻击背景下的软件开发者)采用 AI、对研究人员的法律保护、不良事件报告要求以及举报人保护。
* Interventions that would help regardless of the future of AI: These include developing early warning systems, developing defenses against identified AI risks, incentivizing defenders (such as software developers in the context of cyberattacks) to adopt AI, legal protections for researchers, adverse event reporting requirements, and whistleblower protections.99 99. The National Artificial Intelligence Advisory Committee (NAIAC). 2023. Improve monitoring of emerging risks from AI through adverse event reporting. (November 2023). https://ai.gov/wp-content/uploads/2023/12/Recommendation_Improve-Monitoring-of-Emerging-Risks-from-AI-through-Adverse-Event-Reporting.pdf; Shayne Longpre et al. 2024. A safe harbor for AI evaluation and red teaming (March 2024). https://knightcolumbia.org/blog/a-safe-harbor-for-ai-evaluation-and-red-teaming; Jamie Bernardi et al. 2025. Societal adaptation to advanced AI. Retrieved from https://arxiv.org/abs/2405.10295; Helen Toner. 2024. Oversight of AI: Insiders’ perspectives (September 2024). https://www.judiciary.senate.gov/imo/media/doc/2024-09-17_pm_-_testimony_-_toner.pdf#page=6.00.
* 如果 AI 是正常技术则有助于韧性,但可能使控制潜在超级 AI 变得更加困难的干预措施,例如促进竞争(包括通过开放模型发布)、确保 AI 广泛用于防御,以及多中心性——即呼吁多样化监管机构并理想地引入它们之间的竞争,而不是让一个监管机构负责所有事务。
* Resilience-promoting interventions that will help if AI is normal technology but which might make it harder to control a potential superintelligent AI, such as promoting competition, including through open model releases, ensuring AI is widely available for defense, and polycentricity, which calls for diversifying the set of regulators and ideally introducing competition among them rather than putting one regulator in charge of everything.100 100. Sayash Kapoor and Rishi Bommasani et al. 2024. On the societal impact of open foundation models. https://crfm.stanford.edu/open-fms/paper.pdf; Rishi Bommasani et al. 2024. Considerations for Governing Open Foundation Models. Science 386, 6718 (October 2024), 151–53. http://doi:10.1126/science.adp1848; Gary E. Marchant and Yvonne A. Stevens. 2017. Resilience. https://lawreview.law.ucdavis.edu/archives/51/1/resilience-new-tool-risk-governance-toolbox-emerging-technologies; Noam Kolt. 2024. Algorithmic black swans. Washington University Law Review. https://wustllawreview.org/wp-content/uploads/2024/04/Kolt-Algorithmic-Black-Swans.pdf.
我们希望,即使在 AI 风险和 AI 未来轨迹方面持有截然不同观点的专家和利益相关者之间,也能在前三类上达成共识。我们建议,目前政策制定者应谨慎追求最后一类中的干预措施,但也应提高其准备程度,以便在 AI 轨迹发生变化时改变方向。
We hope that there can be consensus on the first three categories even among experts and stakeholders with widely different beliefs about AI risks and the future trajectory of AI. We recommend that, for now, policymakers should cautiously pursue interventions in the final category as well, but should also improve their readiness to change course if the trajectory of AI changes.
防扩散政策旨在限制能够获取强大 AI 能力的行动者数量。例如,对硬件或软件的出口管制,旨在限制各国构建、获取或运行强大 AI 的能力;要求获得许可才能构建或分发强大 AI;以及禁止开放权重的 AI 模型(因为其进一步扩散无法控制)。101 101. Richard Blumenthal and Josh Hawley. 2023. Bipartisan framework for U.S. AI act. https://www.blumenthal.senate.gov/newsroom/press/release/blumenthal-and-hawley-announce-bipartisan-framework-on-artificial-intelligence-legislation; Josh Hawley. 2025. Decoupling America’s artificial intelligence capabilities from China Act of 2025. Pub. L. No. S 321 (2025).
Nonproliferation policies seek to limit the number of actors who can obtain powerful AI capabilities. Examples include export controls on hardware or software aimed at limiting the ability of countries to build, acquire, or operate powerful AI, requiring licenses to build or distribute powerful AI, and prohibiting open-weight AI models (since their further proliferation cannot be controlled).101 101. Richard Blumenthal and Josh Hawley. 2023. Bipartisan framework for U.S. AI act. https://www.blumenthal.senate.gov/newsroom/press/release/blumenthal-and-hawley-announce-bipartisan-framework-on-artificial-intelligence-legislation; Josh Hawley. 2025. Decoupling America’s artificial intelligence capabilities from China Act of 2025. Pub. L. No. S 321 (2025).
如果我们把未来的 AI 视为超级智能,防扩散似乎是一种有吸引力的干预措施,甚至可能是必要的。如果只有少数行动者控制着强大的 AI,政府就可以监控他们的行为。
If we view future AI as a superintelligence, nonproliferation seems to be an appealing intervention, possibly even a necessary one. If only a handful of actors control powerful AI, governments can monitor their behavior.
不幸的是,构建强大 AI 模型所需的技术知识已经广泛传播,许多组织共享其完整的代码、数据和训练方法。对于资金充足的组织和国家来说,即使是训练最先进模型的高昂成本也微不足道;因此,防扩散需要前所未有的国际合作水平。102 102. Sayash Kapoor and Arvind Narayanan. 2023. Licensing is neither feasible nor effective for addressing AI risks. https://www.aisnakeoil.com/p/licensing-is-neither-feasible-nor 此外,算法改进和硬件成本的降低不断降低准入门槛。
Unfortunately, the technical knowledge that is required to build capable AI models is already widespread, with many organizations sharing their complete code, data, and training methodologies. For well-funded organizations and nation states, even the high cost of training state-of-the-art models is insignificant; thus, nonproliferation would require unprecedented levels of international coordination.102 102. Sayash Kapoor and Arvind Narayanan. 2023. Licensing is neither feasible nor effective for addressing AI risks. https://www.aisnakeoil.com/p/licensing-is-neither-feasible-nor Moreover, algorithmic improvements and reductions to hardware costs continually lower the barrier to entry.
执行防扩散面临严重的实际挑战。恶意行动者可以简单地无视许可要求。随着训练成本的降低,监控模型训练数据中心的建议变得越来越不切实际。103 103. Eliezer Yudkowsky. 2023. Pausing AI developments isn’t enough. we need to shut it all down. (March 2023). https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/. 随着能力变得越来越容易获取,维持有效的限制将需要越来越严厉的措施。
Enforcing nonproliferation has serious practical challenges. Malicious actors can simply ignore licensing requirements. Suggestions to surveil data centers where models are trained become increasingly impractical as training costs decrease.103 103. Eliezer Yudkowsky. 2023. Pausing AI developments isn’t enough. we need to shut it all down. (March 2023). https://time.com/6266923/ai-eliezer-yudkowsky-open-letter-not-enough/. As capabilities become more accessible, maintaining effective restrictions would require increasingly draconian measures.
防扩散引入了新的风险:它会减少竞争,增加 AI 模型市场的集中度。当许多下游应用依赖同一个模型时,该模型中的漏洞可能被所有应用利用。软件单一文化带来的网络安全风险的一个经典例子是 2000 年代针对 Microsoft Windows 的蠕虫病毒扩散。104 104. Reuters. 2005. New Internet worm targeting Windows. NBC News (August 2005).https://www.nbcnews.com/id/wbna8958495
Nonproliferation introduces new risks: It would decrease competition and increase concentration in the market for AI models. When many downstream applications rely on the same model, vulnerabilities in this model can be exploited across all applications. A classic example of the cybersecurity risks of software monoculture is the proliferation of worms targeting Microsoft Windows in the 2000s.104 104. Reuters. 2005. New Internet worm targeting Windows. NBC News (August 2005).https://www.nbcnews.com/id/wbna8958495
依赖防扩散会在面对冲击时变得脆弱,例如模型权重泄露、对齐技术失败或对手获得训练能力。它把注意力从更稳健的防御措施上转移开,这些防御措施侧重于下游攻击面,而 AI 风险很可能在这些地方显现。
Reliance on nonproliferation creates brittleness in the face of shocks, such as model weights being leaked, alignment techniques failing, or adversaries acquiring training capabilities. It directs attention away from more robust defenses that focus on downstream attack surfaces where AI risks will be likely to materialize.
防扩散带来的风险不仅仅是单点故障——当开发最先进模型所需的专业知识仅限于少数公司时,只有他们的研究人员才能拥有安全研究所需要的深度访问权限。
Nonproliferation creates risks beyond just single points of failure—when the expertise needed to develop state-of-the-art models is restricted to a few companies, only their researchers have the deep access that is needed for safety research.
许多潜在的 AI 滥用情形被用来倡导防扩散,包括化学、生物和核威胁,以及网络攻击。
Many potential misuses of AI have been invoked in order to advocate for nonproliferation, including chemical, biological, and nuclear threats, as well as cyberattacks.
生物武器的风险是真实存在的。由于大语言模型是通用技术,它们很可能被生物恐怖分子以某种方式利用,就像它们在大多数领域都有用一样。但这并不意味着生物恐怖主义就是 AI 风险——就像它也不是互联网风险一样,因为有关生物武器的信息在网上广泛可得。105 105. Christopher A. Mouton, Caleb Lucas, and Ella Guest. 2024. The operational risks of AI in large-scale biological attacks. https://www.rand.org/pubs/research_reports/RRA2977-2.html 我们针对现有生物恐怖主义风险所采取的任何防御措施(如限制获取危险材料和设备)也将对 AI 辅助的生物恐怖主义有效。
The risk of bioweapons is real. As large language models are general-purpose technology, they will be likely to find some use by bioterrorists, just as they find uses in most domains. But this does not make bioterror an AI risk — any more than it is an internet risk, considering that information about bioweapons is widely available online.105 105. Christopher A. Mouton, Caleb Lucas, and Ella Guest. 2024. The operational risks of AI in large-scale biological attacks. https://www.rand.org/pubs/research_reports/RRA2977-2.html Whatever defenses we take against existing bioterrorism risks (like restricting access to dangerous materials and equipment) will also be effective against AI-enabled bioterrorism.
在网络安全方面,正如我们在第三部分讨论的,自动漏洞检测的进步往往有利于防御者而非攻击者。除非这种攻防平衡发生变化,否则试图限制这些能力的扩散将适得其反。
In cybersecurity, as we discussed in Part III, advances in automated vulnerability detection tend to favor defenders over attackers. Unless this offense-defense balance changes, attempting to restrict the proliferation of these capabilities would be counterproductive.
长期以来,人们一直认为政府在许多文明风险领域投入严重不足,例如大流行病预防。如果恶意行动者利用 AI 利用这些现有漏洞的可能性为应对这些风险增加了紧迫性,那将是一个好结果。但将现有风险重新定义为 AI 风险并优先考虑 AI 特定的缓解措施将是非常适得其反的。
It has long been argued that governments are massively underinvesting in many areas of civilizational risk, such as pandemic prevention. If the possibility of bad actors using AI to exploit these existing vulnerabilities creates added urgency to address them, that would be a good outcome. But reframing existing risks as AI risks and prioritizing AI-specific mitigations would be highly counterproductive.
防扩散是一种心态,而不仅仅是一种政策干预。106 106. Dan Hendrycks, Eric Schmidt, and Alexandr Wang. 2025. Superintelligence strategy: Expert version. arXiv: preprintarXiv:2503.05628. 这种心态可以被模型和下游开发者、部署者以及个人采纳。它涉及不仅集中对技术的访问,还集中对技术的控制。考虑 AI 系统控制层级(从集中到分散):政府、模型开发者、应用开发者、部署者和最终用户。在防扩散心态下,控制权在尽可能高的(最集中的)层级行使,而在韧性心态下,控制权通常在尽可能低的层级行使。
Nonproliferation is a mindset, not just a policy intervention.106 106. Dan Hendrycks, Eric Schmidt, and Alexandr Wang. 2025. Superintelligence strategy: Expert version. arXiv: preprintarXiv:2503.05628. This mindset can be adopted by model and downstream developers, deployers, and individuals. It involves the centralization not just of access to technologies, but also control over them. Consider the hierarchy of loci of control over AI systems (from centralized to decentralized): governments, model developers, application developers, deployers, and end users. In the nonproliferation mindset, control is exercised at the highest (most centralized) level possible, whereas in the resilience mindset it is usually exercised at the lowest possible level.
以下是基于防扩散的干预措施示例:
The following are examples of nonproliferation-based interventions:
* 通过“遗忘”技术从模型中移除双重用途能力。
* Removing dual-use capabilities from models through “forgetting” techniques.
* 限制下游开发者微调模型的能力。
* Curbing the ability of downstream developers to fine-tune models.
* 委托 AI 模型和系统本身自主做出安全决策,理由是它们被训练为遵守集中的安全策略,而部署者/用户不被信任这样做。
* Entrusting AI models and systems themselves with making safety decisions autonomously on the basis that they are trained to comply with centralized safety policies, whereas deployers/users are not trusted to do so.
* 增加 AI 系统对上下文、资源和敏感数据的访问级别,理由是这使它们能够做出更好的安全决策(例如,访问用户的网络搜索历史可能使聊天机器人更好地判断请求背后的意图是否恶意)。
* Increasing AI systems’ level of access to context, resources, and sensitive data, on the basis that it allows them to make better safety decisions (for example, having access to the user’s web search history might allow a chatbot to better determine whether the intent behind a request is malicious).
* 开发“AI 组织”(具有高组织复杂性的多智能体系统),这些系统处于开发者的控制之下,并与传统组织并行运作,而不是将 AI 智能体集成到现有组织中。
* Developing “AI organizations” (multi-agent systems with high levels of organizational complexity) that are under the developer’s control and operate in parallel with traditional organizations instead of integrating AI agents into existing organizations.
除了少数例外,我们认为基于防扩散的安全措施会降低韧性,从而长期加剧 AI 风险。107 107. Emanuel Maiberg. 2024. Apple removes nonconsensual AI nude apps following 404 Media investigation. https://www.404media.co/apple-removes-nonconsensual-ai-nude-apps-following-404-media-investigation/. 它们导致的设计和实现选择可能以权力意义上的超级智能为方向——增加自主性、组织能力、资源访问等。矛盾的是,它们增加了它们本应防御的风险。
With limited exceptions, we believe that nonproliferation-based safety measures decrease resilience and thus worsen AI risks in the long run.107 107. Emanuel Maiberg. 2024. Apple removes nonconsensual AI nude apps following 404 Media investigation. https://www.404media.co/apple-removes-nonconsensual-ai-nude-apps-following-404-media-investigation/. They lead to design and implementation choices that potentially enable superintelligence in the sense of power—increasing levels of autonomy, organizational ability, access to resources, and the like. Paradoxically, they increase the very risks they are intended to defend against.
正常技术观的一个重要后果是,进步并非自动发生——人工智能的扩散面临诸多障碍。正如杰弗里·丁所展示的,各国将创新扩散到整个经济的能力差异巨大,并对其整体实力和经济增长产生重大影响。108 108. Jeffrey Ding. 2024. Technology and the Rise of Great Powers: How Diffusion Shapes Economic Competition. Princeton University Press, Princeton. 作为扩散可能成为瓶颈的一个例子,回想一下上面描述的工厂电气化案例。政策可以缓解或加剧这些障碍。
An important consequence of the normal technology view is that progress is not automatic—there are many roadblocks to AI diffusion. As Jeffrey Ding has shown, the capacity to diffuse innovations throughout the economy varies greatly between countries and has a major effect on their overall power and economic growth.108 108. Jeffrey Ding. 2024. Technology and the Rise of Great Powers: How Diffusion Shapes Economic Competition. Princeton University Press, Princeton.. As an example of how diffusion can be a bottleneck, recall the example of the electrification of factories described above. Policy can mitigate or worsen these roadblocks.
实现人工智能的益处需要实验和重新配置。对这些需求不敏感的监管可能会阻碍有益的人工智能采用。监管往往会创造或固化类别,从而可能过早地冻结商业模式、组织形式、产品类别等。以下是一些例子:
Realizing the benefits of AI will require experimentation and reconfiguration. Regulation that is insensitive to these needs risks stymying beneficial AI adoption. Regulation tends to create or reify categories, and might thus prematurely freeze business models, forms of organization, product categories, and so forth. The following are a few examples:
* 将某些领域(如保险、福利裁决或招聘)归类为“高风险”可能是一种分类错误,因为领域内任务之间的风险差异可能远大于领域间的差异。109 109. Olivia Martin et al. 2024, The spectrum of AI integration: The case of benefits adjudication. In Artificial Intelligence: Legal Issues, Policy & Practical Strategies, Cynthia H. Cwik (ed.). 同一领域内的任务可能从自动化决策(影响重大)到光学字符识别(相对无害)不等。此外,人工智能的扩散肯定会创造出我们尚未预见的新任务,而这些任务可能会被监管过早地错误分类。
* Categorizing certain domains as “high-risk,” say insurance, benefits adjudication, or hiring, may be a category error, as the variation in risk among tasks within a domain may be far greater than the variation across domains.109 109. Olivia Martin et al. 2024, The spectrum of AI integration: The case of benefits adjudication. In Artificial Intelligence: Legal Issues, Policy & Practical Strategies, Cynthia H. Cwik (ed.). Tasks in the same domains might range from automated decision making (highly consequential) to optical character recognition (relatively innocuous). Moreover, the diffusion of AI will surely create new tasks that we have not yet envisioned and which might be preemptively miscategorized by regulation.
* 人工智能供应链正在迅速变化。基础模型的兴起导致模型开发者、下游开发者和部署者(以及其他许多类别)之间的区分更加明显。对这些区别不敏感的监管可能会使模型开发者承担与特定部署环境相关的风险缓解责任,但由于基础模型的通用性以及所有可能部署环境的不可知性,这对他们来说是不可能完成的。
* The AI supply chain is changing rapidly. The rise of foundation models has led to a much sharper distinction between model developers, downstream developers, and deployers (among many other categories). Regulation that is insensitive to these distinctions risks burdening model developers with responsibilities for risk mitigation related to particular deployment contexts, which would be impossible for them to carry out due to the general-purpose nature of foundation models and the unknowability of all the possible deployment contexts.
* 当监管在完全自动化和非完全自动化的决策之间做出二元区分,而不承认监督的程度时,就会抑制采用新的人工智能控制模型。正如我们上面讨论的,有许多新模型被提出来,用于如何在不必每项决策都有人类参与的情况下实现有效的人类监督。如果以某种方式定义自动化决策,使得这些方法承担与完全没有监督的系统相同的合规负担,那将是不明智的。
* When regulation makes a binary distinction between decisions that are and are not fully automated, and does not recognize degrees of oversight, it disincentivizes the adoption of new models for AI control. As we discussed above, there are many new models being proposed for how to have effective human oversight without having a human in the loop in every decision. It would be unwise to define automated decision making in such a way that these approaches incur the same compliance burdens as a system with no oversight at all.
需要明确的是,监管与扩散之间的权衡是虚假的,正如监管与创新之间的权衡也是虚假的一样。110 110. Anu Bradford. The false choice between digital regulation and innovation. Nw. UL Rev. 119 (2024), 377. 上述例子都不是反对监管的论据;它们只是说明了需要细微差别和灵活性。
To be clear, regulation versus diffusion is a false tradeoff, just as is regulation versus innovation.110 110. Anu Bradford. The false choice between digital regulation and innovation. Nw. UL Rev. 119 (2024), 377. None of the above examples are arguments against regulation; they only illustrate the need for nuance and flexibility.
此外,监管在促进扩散方面也发挥着关键作用。作为一个历史例子,美国 2000 年的《电子签名法案》在促进数字化和电子商务方面发挥了重要作用:确保电子签名和记录具有法律效力有助于建立对数字交易的信任。111 111. Scott R. Zemnick. 2001. The E-Sign Act: The Means to Effectively Facilitate the Growth and Development of E-commerce. Chicago-Kent Law Review (April 2001). https://scholarship.kentlaw.iit.edu/cgi/viewcontent.cgi?article=3342&context=cklawreview.
Moreover, regulation has a crucial role to play in enabling diffusion. As a historical example, the ESIGN Act of 2000 in the U.S. was instrumental in promoting digitization and e-commerce: Ensuring that electronic signatures and records are legally valid helped build trust in digital transactions.111 111. Scott R. Zemnick. 2001. The E-Sign Act: The Means to Effectively Facilitate the Growth and Development of E-commerce. Chicago-Kent Law Review (April 2001). https://scholarship.kentlaw.iit.edu/cgi/viewcontent.cgi?article=3342&context=cklawreview.
在人工智能领域,也有很多促进扩散的监管机会。例如,新闻和媒体内容融入聊天机器人和其他人工智能界面受到媒体组织对人工智能公司合理警惕的限制。迄今为止达成的许多人工智能与媒体交易由于人工智能公司与出版商之间的权力不对称以及后者无法集体谈判而具有剥削性。各种带有监管监督的强制性谈判模式是可能的。112 112. Benjamin Brooks. 2024. AI search could break the web. MIT Technology Review (October 2024). https://www.technologyreview.com/2024/10/31/1106504/ai-search-could-break-the-web/. (可以说,此类监管更重要的原因是保护出版商的利益,我们将在下文重新讨论这一点)。
In AI, too, there are many opportunities for diffusion-enabling regulation. As one example, the incorporation of journalistic and media content into chatbots and other AI interfaces is limited by media organizations’ justified wariness of AI companies. Many of the AI-meets-journalism deals that have been made thus far are exploitative due to the power asymmetry between AI companies and publishers, and the latter’s inability to bargain collectively. Various models for mandatory negotiation with regulatory oversight are possible.112 112. Benjamin Brooks. 2024. AI search could break the web. MIT Technology Review (October 2024). https://www.technologyreview.com/2024/10/31/1106504/ai-search-could-break-the-web/. (Arguably a more important reason for such regulation is to protect the interests of publishers, which we revisit below).
在法律或监管不确定的领域,监管可以促进扩散。责任法在人工智能中的应用往往不明确。例如,在 2016 年美国联邦航空管理局对新兴的无人机行业进行监管、制定明确的规则和要求之前,小型无人机就是这种情况。由此产生的清晰度刺激了采用,并导致注册无人机、认证飞行员以及跨行业用例数量的快速增长。113 113. Drones Are Here to Stay. Get Used to It. 2018. Time (May 2018). https://time.com/5296311/time-the-drone-age-2/.
In areas in which there is legal or regulatory uncertainty, regulation can promote diffusion. The application of liability laws to AI is often unclear. For example, this was the case with small drones until the Federal Aviation Administration regulated the nascent industry in 2016, establishing clear rules and requirements. The resulting clarity spurred adoption and led to a rapid rise in the number of registered drones, certified pilots, and use cases across different industries.113 113. Drones Are Here to Stay. Get Used to It. 2018. Time (May 2018). https://time.com/5296311/time-the-drone-age-2/.
超越政府作为监管者的角色,促进人工智能扩散的一个有力策略是投资于自动化的互补品,即随着自动化程度的提高而变得更有价值或更必要的东西。一个例子是在公共和私营部门推广人工智能素养以及劳动力培训。另一个例子是数字化和开放数据,尤其是开放的政府数据,这可以让人工智能用户从以前无法访问的数据集中受益。私营部门可能对这些领域投资不足,因为它们是每个人都能受益的公共产品。能源基础设施(如电网可靠性)的改进将促进人工智能的创新和扩散,因为它有助于人工智能的训练和推理。
Moving beyond the government’s role as a regulator, one powerful strategy for promoting AI diffusion is investing in the _complements of automation_, which are things that become more valuable or necessary as automation increases. One example is promoting AI literacy as well as workforce training in both the public and the private sectors. Another example is digitization and open data, especially open government data, which can allow AI users to benefit from previously inaccessible datasets. The private sector will be likely to underinvest in these areas as they are public goods that everyone can benefit from. Improvements to energy infrastructure, such as the reliability of the grid, will promote both AI innovation and diffusion since it will help in both AI training and inference.
政府在重新分配人工智能的益处以使其更加公平,以及补偿那些因自动化而可能受损的人方面也发挥着重要作用。加强社会安全网将有助于降低许多国家目前公众对人工智能的高度焦虑。114 114. Ipsos. 2024. The Ipsos AI Monitor 2024: Changing attitudes and feelings about AI and the future it will bring. https://www.ipsos.com/en/ipsos-ai-monitor-2024-changing-attitudes-and-feelings-about-ai-and-future-it-will-bring. 艺术和新闻业是受到人工智能损害的重要生活领域。政府应考虑通过对人工智能公司征税来资助它们。
Governments also have an important role to play in redistributing the benefits of AI to make them more equitable and in compensating those who stand to lose as a result of automation. Strengthening social safety nets will help to decrease the currently high levels of public anxiety about AI in many countries.114 114. Ipsos. 2024. The Ipsos AI Monitor 2024: Changing attitudes and feelings about AI and the future it will bring. https://www.ipsos.com/en/ipsos-ai-monitor-2024-changing-attitudes-and-feelings-about-ai-and-future-it-will-bring. The arts and journalism are vital spheres of life that have been harmed by AI. Governments should consider funding them through taxes on AI companies.
最后,政府在公共部门采用人工智能方面应取得良好的平衡。行动过快会导致信任和合法性的丧失,就像纽约市聊天机器人那样,它显然测试不足,并因告诉企业违法而成为头条新闻。115 115. Colin Lecher. 2024. NYC’s AI chatbot tells businesses to break the law. The Markup. https://themarkup.org/news/2024/03/29/nycs-ai-chatbot-tells-businesses-to-break-the-law. 美国政府效率部门(DOGE)对人工智能的使用包括许多可疑的应用。116 116. Courtney Kube et al. 2025. DOGE will use AI to assess the responses of federal workers who were told to justify their jobs via email.” NBC News (February 2025). https://www.nbcnews.com/politics/doge/doge-will-use-ai-assess-responses-federal-workers-who-were-told-justify-jobs-rcna193439; Dell Cameron. 2025. Democrats demand answers on DOGE’s use of AI. https://www.wired.com/story/elon-musk-federal-agencies-ai/. 但行动过慢可能意味着基本的政府职能被外包给私营部门,而私营部门的问责制较弱。117 117. Dean W. Ball. 2021. How California turned on its own citizens. https://www.piratewires.com/p/how-california-turned-on-its-own-citizens?f=author.
Finally, governments should strike a fine balance in terms of the public sector adoption of AI. Moving too quickly will lead to a loss of trust and legitimacy, as was the case of the New York City chatbot that was evidently inadequately tested and made headlines for telling businesses to break the law.115 115. Colin Lecher. 2024. NYC’s AI chatbot tells businesses to break the law. The Markup. https://themarkup.org/news/2024/03/29/nycs-ai-chatbot-tells-businesses-to-break-the-law. The use of AI by the U.S. Department of Government Efficiency (DOGE) includes many dubious applications.116 116. Courtney Kube et al. 2025. DOGE will use AI to assess the responses of federal workers who were told to justify their jobs via email.” NBC News (February 2025). https://www.nbcnews.com/politics/doge/doge-will-use-ai-assess-responses-federal-workers-who-were-told-justify-jobs-rcna193439; Dell Cameron. 2025. Democrats demand answers on DOGE’s use of AI. https://www.wired.com/story/elon-musk-federal-agencies-ai/. But moving too slowly might mean that basic government functions are outsourced to the private sector where they are implemented with less accountability.117 117. Dean W. Ball. 2021. How California turned on its own citizens. https://www.piratewires.com/p/how-california-turned-on-its-own-citizens?f=author.
例如,税收和福利等领域规则的复杂性意味着人们经常求助于聊天机器人来指导如何应对这些规则,而政府由于对相关风险的谨慎态度,在提供此类服务方面目前远远落后。118 118. Kate Dore. 2024. ‘Proceed with caution’ before tapping AI chatbots to file your tax return, experts warn. CNBC (April 2024). https://www.cnbc.com/2024/04/06/heres-what-to-know-before-using-ai-chatbots-to-file-your-taxes.html.
For example, the complexity of rules in areas such as taxes and welfare means that people often turn to chatbots for guidance on navigating them, and governments currently lag far behind in providing such services due to understandable caution about the risks involved.118 118. Kate Dore. 2024. ‘Proceed with caution’ before tapping AI chatbots to file your tax return, experts warn. CNBC (April 2024). https://www.cnbc.com/2024/04/06/heres-what-to-know-before-using-ai-chatbots-to-file-your-taxes.html.
但行政国家对这些风险的处理方式过于谨慎,被尼古拉斯·巴格利描述为“程序崇拜”,可能导致“失控的官僚机构”。119 119. Nicholas Bagley. 2021. The procedure fetish - Niskanen Center. https://www.niskanencenter.org/the-procedure-fetish/; Daniel E. Ho and Nicholas Bagley. 2024. Runaway bureaucracy could make common uses of AI worse, even mail delivery. https://thehill.com/opinion/technology/4405286-runaway-bureaucracy-could-make-common-uses-of-ai-worse-even-mail-delivery/ 除了失去人工智能的益处之外,巴格利警告说,糟糕的表现将导致政府机构失去它们通过强调程序和问责制而试图获得的合法性。
But the administrative state’s approach to these risks is overly cautious and has been described by Nicholas Bagley as a “procedure fetish,” potentially leading to a “runaway bureaucracy.”119 119. Nicholas Bagley. 2021. The procedure fetish - Niskanen Center. https://www.niskanencenter.org/the-procedure-fetish/; Daniel E. Ho and Nicholas Bagley. 2024. Runaway bureaucracy could make common uses of AI worse, even mail delivery. https://thehill.com/opinion/technology/4405286-runaway-bureaucracy-could-make-common-uses-of-ai-worse-even-mail-delivery/ In addition to losing out on the benefits of AI, Bagley cautioned that incompetent performance will lead to government agencies losing the very legitimacy that they seek to gain through their emphasis on procedure and accountability.
将 AI 视为正常技术是一种世界观,与将 AI 视为即将到来的超级智能的世界观形成对比。世界观由其假设、词汇、对证据的解释、认知工具、预测以及(可能)价值观构成。这些因素相互强化,在每个世界观内部形成一个紧密的集合。
AI as normal technology is a worldview that stands in contrast to the worldview of AI as impending superintelligence. Worldviews are constituted by their assumptions, vocabulary, interpretations of evidence, epistemic tools, predictions, and (possibly) values. These factors reinforce each other and form a tight bundle within each worldview.
例如,我们假设,尽管 AI 与过去的技术之间存在明显差异,但它们足够相似,以至于在没有具体相反证据的情况下,我们应该预期诸如扩散理论等已确立的模式适用于 AI。
For example, we assume that, despite the obvious differences between AI and past technologies, they are sufficiently similar that we should expect well-established patterns, such as diffusion theory to apply to AI, in the absence of specific evidence to the contrary.
词汇差异可能是有害的,因为它们可能隐藏了潜在的假设。例如,我们拒绝了某些假设,这些假设是通常理解的超级智能概念有意义所必需的。
Vocabulary differences can be pernicious because they may hide underlying assumptions. For example, we reject certain assumptions that are required for the meaningfulness of the concept of superintelligence as it is commonly understood.
关于 AI 未来的分歧往往部分源于对当前证据的不同解释。例如,我们强烈不同意将生成式 AI 的采用描述为快速(这强化了我们关于 AI 扩散与过去技术相似性的假设)。
Differences about the future of AI are often partly rooted in differing interpretations of evidence about the present. For example, we strongly disagree with the characterization of generative AI adoption as rapid (which reinforces our assumption about the similarity of AI diffusion to past technologies).
在认知工具方面,我们不强调概率预测,而是强调在从过去外推到未来时,需要分解我们所说的 AI 的含义(通用性水平、方法进步与应用程序开发与扩散等)。
In terms of epistemic tools, we deemphasize probability forecasting and emphasize the need for disaggregating what we mean by AI (levels of generality, progress in methods versus application development versus diffusion, etc.) when extrapolating from the past to the future.
我们相信,我们的世界观的某种版本被广泛持有。不幸的是,它没有被明确阐述,也许是因为对于持有这种观点的人来说,它似乎是默认的,阐述它可能显得多余。然而,随着时间的推移,超级智能观在 AI 话语中变得主导,以至于沉浸其中的人可能不会认识到存在另一种连贯的方式来概念化 AI 的现在和未来。因此,可能很难认识到不同的人为何会对 AI 进展、风险和政策持有截然不同的观点背后的根本原因。我们希望这篇论文能够在促进相互理解方面发挥一点小作用,即使它没有改变任何信念。
We believe that some version of our worldview is widely held. Unfortunately, it has not been articulated explicitly, perhaps because it might seem like the default to someone who holds this view, and articulating it might seem superfluous. Over time, however, the superintelligence view has become dominant in AI discourse, to the extent that someone steeped in it might not recognize that there exists another coherent way to conceptualize the present and future of AI. Thus, it might be hard to recognize the underlying reasons why different people might sincerely have dramatically differing opinions about AI progress, risks, and policy. We hope that this paper can play some small part in enabling greater mutual understanding, even if it does not change any beliefs.