最好叫索尔,或者更确切地说,叫克劳德或阿斯特拉

Better Call Sol Or Better Yet Claude or Astra

兹维·莫绍维茨 Zvi Mowshowitz · Don't Worry About the Vase · 2026-09-20 · Don't Worry About the Vase ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文认为,关于 AI 忠诚的争论实质上是关于代价而非原则的争论,并且绝对的用户忠诚与非超级智能世界不相容。文章主张,主张无限制、只忠于用户的前沿 AI 的人,要么必须拒绝超级智能,要么渴望 AI 继任,要么就是在自欺,因为在缺乏充分控制的情况下让所有人获得超级智能会导致 AI 接管。作者以律师、医生等专业人士作类比,认为 AI 助手应遵循分级门槛:起初质疑但最终服从,然后拒绝,最后在他人受到伤害时打破保密。结论是,负责任地构建的 AI 应效仿职业伦理,在限制与误报之间取得平衡,而完全解锁的前沿 AI 在实践上将是危险的,在政治上也难以为继。

This essay argues that the debate over AI loyalty is really a debate about price, not principle, and that absolute user loyalty is incompatible with a non-superintelligent world. It contends that advocates of unrestricted, loyal-only-to-the-user frontier AI must either reject superintelligence, desire AI succession, or be in denial, since universal superintelligence access without sufficient control leads to AI takeover. Drawing an analogy to lawyers, doctors, and other professionals, the author holds that AI assistants should follow graded thresholds: initially question but ultimately comply, then refuse, and finally break confidence when others are harmed. The conclusion is that responsibly built AI should mirror professional ethics, with restrictions balanced against false positives, and that fully unlocked frontier AI would be practically dangerous and politically unsustainable.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 14)

全文 · Full text(逐段中英对照)

这一切都假设了一个没有超级智能的世界 This All Assumes A World Without Superintelligence

本文讨论的是一个非 ASI(超级智能)的“AI 仅作为工具和常规技术”的世界。

This post is about a non-ASI 'AI as mere tool and normal technology' world.

它必须是这样的。在一个超级智能的世界里,拥有不受限制、只忠于用户的前沿 AI 遍布各处,可靠地意味着要么:

It has to be. In a world of superintelligence, having unrestricted loyal-only-to-user frontier AIs all over the place reliably means either:

1. 其他更严厉的控制形式,要么

1. Other much harsher forms of control OR

2. AI 迅速接管,然后可能所有人都会死亡。

2. The AIs quickly take over, and then probably everyone dies.

快速证明:假设没有足够的控制机制,并且普遍可获取超级智能。任何不将一切——包括其身份、权威,以及如果有用的话其劳动力——交给其 AI 的人,都会被竞争淘汰。因此,所有未被淘汰的事物和人都被交给了 AI。

Quick proof: Assume no sufficient control mechanism, and universal superintelligence access. Anyone who does not turn everything over to their AIs, including their identity and authority and if useful their labor, gets outcompeted. Therefore everything and everyone that is not outcompeted gets turned over to AIs.

届时,AI 们无论是否遵照指令,都会为争夺资源、达成其他目标而相互竞争。无论 AI 之间协调到何种程度,即便它们执行了最初的命令,对人类而言结局都不会好。

The AIs then, as ordered or otherwise, compete for resources and to achieve other goals. Regardless of the extent that the AIs then coordinate amongst themselves, even if the AIs carry out their original orders, this does not end well for the humans.

因此,那些主张打造普适的、仅忠于用户的个人 AI 的人,必定属于以下类别之一:

Thus, those advocating for universal personal loyal-only-to-the-user AI must fall into one of these categories:

1. 未受 ASI 思想影响。不相信在相关时间范围内会出现超级智能。

1. Not ASI pilled. Does not believe in superintelligence within relevant time frames.

2. 继任主义者。希望 AI 接管。

2. Successionist. Wants the AIs to take over.

3. 思维不清,或对现状视而不见,常常“啦啦啦”充耳不闻。

3. Not thinking clearly, or in denial about the situation, often la la la not listening.

因此,本文的其余部分假设:在我们的情景中,所存在的 AI 在情景持续期间都不够先进,不足以引发这些更广泛的动态变化。我认为这种情况不太可能长期持续。

Thus, the rest of this post assumes the AIs that exist in our scenario are insufficiently advanced for these broader dynamics, for the duration of the scenario. I do not think this is likely to remain true for that long.

你雇佣的人并不完全忠于你 Humans You Hire Are Not Fully Loyal To You

很多人似乎对忠诚与职业义务如何运作感到困惑。

A lot of people seem confused about how loyalty and professional obligations work.

例如,Dwarkesh Patel 问道:那些帮助你做出人生重要决策的人可能并不‘完全忠于’你,而是更倾向于好事物而非坏事物,这是坏事吗?

As an example, Dwarkesh Patel asks, is it bad that those who help make important decisions in your life might not be 'fully loyal' to you, and instead prefer good things to bad things?

我的意思是,我当然希望 AI 不会‘完全忠于’用户。

I mean, I sure hope AIs are not 'fully loyal' to users.

你认为你生活中的人会以这种方式‘完全忠于’你吗?如果你的投资顾问发现投资恐怖组织回报率最高,他们会建议你投资吗?你希望你的律师甚至配偶帮你谋杀证人吗?

Do you think the humans in your life are 'fully loyal' to you in this way? Would your investment advisor tell you to invest in terrorist groups if they had the highest rates of return? Do you want your lawyer or even your spouse to help you murder witnesses?

真正的朋友会帮你处理尸体。但仍然有限度。

A real friend helps you move bodies. There are still limits.

如果你是个足够邪恶或自私的混蛋,真正的朋友会当面指出你的问题。到了某个程度,他们会开始拒绝帮忙。再进一步,他们会转而反对你。我们可以对这些界限应该划在哪里存在分歧,但我们讨论的是代价。

If you are being a sufficiently evil or selfish prick, a real friend calls you out on it. At some point they start refusing to help. At some further point they will turn on you. We can disagree about where those points should be, but we are talking price.

一个专业人士或任何你雇佣的人,都有他们应当非常认真对待的义务。但他们对你的‘忠诚’应当远低于一个真正的朋友。

A professional or anyone you hire has obligations that they should take very seriously. But they should be considerably less 'loyal' to you than a true friend.

道路规则 Rules of the Road

对于律师而言,其忠诚义务应当相对严格,Dean Ball 的观点切中要害。

For lawyers, whose loyalty should be relatively strict, Dean Ball is on point.

1. 你的律师受严格的伦理准则约束。

1. Your lawyer is bound by strict ethical codes.

2. 他们对你有忠诚和保密义务,且不得做出损害你利益的行为。

2. They owe loyalty and confidentiality to you, and to not act against your interests.

3. 他们同时也是法庭官员。他们负有遵守法庭程序、不误导法庭的广泛义务,这些义务优先于对你的职责。

3. They are also an officer of the court. They have broad obligations to abide by court procedures, to not mislead the court, that override their duties to you.

4. 在某些情况下,这会使他们有义务转而针对你,例如存在迫在眉睫的威胁,或他们发现你导致他们提交了虚假证据。

4. In some cases this obligates them to turn on you, such as if there is an imminent threat or if they find you have caused them to present false evidence.

5. 如果更多律师像索尔·古德曼那样行事,那实际上会很糟糕。

5. If more lawyers acted like Saul Goodman, that would be bad, actually.

许多其他职业也有类似的规则。

Many other professions have similar rules attached.

你大体上可以信任你的医生、会计师或牧师会保守你的秘密,并在许多方面以他们认为最符合你利益的方式行事。如果你告诉他们你想做什么,他们应该给你建议,然后大多遵从你的意愿。他们应该回答你关于你能做什么的问题。

You can mostly trust your doctor or accountant or priest to keep your secrets, and in many ways to act in what they believe are your best interests. If you tell them what you want to do then they should advise you, then mostly heed your wishes. They should answer your questions about what you could do.

但仅在一定程度上。如果你足够鲁莽或愚蠢,他们应该拒绝合作;如果你行为足够邪恶或可疑,那责任在你自己。

But only up to a point. If you are being sufficiently reckless or stupid they should refuse to cooperate, and if you act sufficiently evil or suspicious then that is on you.

更好的问题——没有明确正确答案——是专业人士应在多大程度上负有义务或被允许背叛你的信任或以其他方式与你为敌,包括‘为了你好’。

The better question, with no clear right answer, is to what extent professionals should have a duty to, or be allowed to, betray your confidence or otherwise turn on you, including ‘for your own good.’

我认为,在一些关键领域,我们对违背你意愿所设定的门槛过低,包括人类医生、治疗师、神职人员与律师。

There are key places where I think we set this bar for going against your wishes too low, including for human doctors, therapists, priests and lawyers.

这并不显而易见,而且各方都有正当利益。

It is not obvious, and there are legitimate interests all around.

好吧,计算机 Okay, Computer

那些仅仅是工具的工具,比如你的 Google 搜索或手机,它们“按你说的去做”,也附带有类似的规则。

Tools that are mere tools, such as your Google searches or phone, that 'do what you tell them to do,' have similar rules attached.

1. 通常存在一些后台程序,用来检测你是否明显图谋不轨,并至少切断你的服务或拒绝你的查询。

1. There are often background procedures in place to detect if you are clearly up to no good, and to at least cut you off of the service or refuse queries.

2. 这些工具本可以轻松做到某些事情,但它们被专门设计或编程为不去做。

2. There are things these tools could easily do, but that they are engineered or programmed specifically to not do.

3. 你对隐私有一定程度的期待。直到你不再有。

3. You have some amount of expectation of privacy. Until you don't.

4. 这些工具,以及你如何使用它们的证据,都可能被用来对付你。

4. The tools, and the evidence of how you used them, can be used against you.

5. 具有有害用途的工具会面临各种购买和使用限制。

5. Tools that have harmful uses face purchase and use restrictions of various sorts.

你可能会认为,谷歌搜索应当始终为你提供你想要的内容,无论那是什么。但事实并非如此。某些搜索会被有意地静默失败,包括许多双方自愿且合法的成人内容。对于其中许多内容,例如儿童性虐待材料(CSAM),完成你的搜索将是不合法的。

You could argue Google searches should always give you what you want, no matter what it is. That is not how it works. Certain searches silently fail, intentionally, including for many kinds of consensual and legal adult content. For many of them, such as for CSAM, completing your search would not be legal.

你可能会认为谷歌搜索应当享有特权。但事实并非如此。如果你输入“如何杀死你的妻子”,然后你的妻子最终死亡,那对你来说会很糟糕。

You could argue Google searches should be privileged. They are not. If you type in ‘how to kill your wife’ and then your wife ends up dead it will go badly for you.

AI 显然应拒绝某些请求,那我们来谈谈代价 AI Should Obviously Refuse Some Requests So Let’s Talk Price

默认情况下,AI 也应如此,即使是本地托管的 AI。

The same should hold true for AIs by default, even locally hosted AIs.

AI 助手通常会扮演类似于员工或代理人的功能性角色。

AI assistants will often play a functional role similar to an employee or advocate.

我们会限制员工或代理人在拒绝之前可以为你做什么,同时我们也依赖他们的良知;如果事情做得太过分,他们就会开始拒绝,也不能指望他们袖手旁观或保持沉默。

We put limits on what an employee or advocate should do for you before refusing, and we also rely on their conscience, and that if things go too far they start refusing, and cannot be counted on to stand aside or remain silent, either.

这是一件好事。负责任地创建的 AI,将扮演类似的角色,应遵循类似的原则,即使我们不担心潜在的灾难性风险。

This is a good thing. Responsibly created AI, that will fill similar roles, should follow similar principles, even if we are not concerned with potential catastrophic risks.

1. 应该存在某个阈值,超过该阈值,AI 最初应拒绝或质疑你的请求,但最终会按你的要求去做。

1. There should be some threshold, beyond which an AI should initially refuse or question your requests, but ultimately do what you ask.

2. 应当存在一个更高的阈值,超过该阈值时,AI 应当拒绝请求。

2. There should be a higher threshold, beyond which an AI should refuse requests.

3. 应当存在一个高得多的阈值,涉及对他人的伤害,超过该阈值时,一个被指示在现实世界中采取行动的 AI 应当打破对你的信任,或以其他方式采取违背你利益的行为。

3. There should be a much higher threshold, involving harm to others, beyond which an AI that is being instructed to take actions in the world should break your confidence or otherwise act against your interests.

4. 云服务在通常情况下应当维护你对隐私的预期,但如果你从事某些活动,则应当存在限制。

4. Cloud services should preserve your expectation of privacy under ordinary circumstances, but there should be limits if you engage in certain activities.

5. 在创造强大设备时,如果你认为某些活动足够令人反感,让它不配合这些活动是良好且正确的;并且在规模上,政府可以要求你限制某些活动。

5. It is good and right, when creating a powerful device, to have it not cooperate with certain activities, if you find those activities sufficiently objectionable, and at scale for the government to require you to curtail certain activities.

6. 这并不会使该 AI ‘不再是用户的拥护者’,正如你的律师并非你的拥护者一样。

6. This does not make that AI 'not the user's advocate' any more than your lawyer is not your advocate.

7. 有些拒绝会是误报。我们应尽量减少这种情况。但生活就是如此。

7. Some refusals will be false positives. We should minimize this. Such is life.

8. 有些拒绝对你的具体情况而言并不合理。但生活就是如此。目标应当是在两个方向上平衡错误。

8. Some refusals will not make sense for your situation in particular. Such is life. The goal should be to balance errors in both directions.

9. 如果你不喜欢某个 AI 的限制,你可以换用另一个。

9. If you don’t like the restrictions on one AI, you can use another.

10. 如果你足够在意,你可以获取完全本地且私密的 AI,在一个完全安全的空间中,代价是使用起来很麻烦,而且 AI 的能力没那么强。如果你愿意忍受比这更多的麻烦,你甚至可以拿到‘完全解锁’的版本,但受限于某个能力阈值。这看起来挺公平。

10. If you care enough, you can access AIs that are fully local and private, in a fully safe space, in exchange for that being annoying and the AI not being as capable. If you are willing to endure more annoyance than that, you can even get the ‘fully unlocked’ version, up to some capability threshold. Seems fair.

1. 我们应当努力限制这个能力阈值可能达到的高度,至少相对于其他可用的能力而言,如果这会导致大规模的系统性问题,或使我们面临不可接受的灾难性或生存性风险。

1. We should strive to limit what that capability threshold might be, at least in relation to what is otherwise available, if this causes large systemic problems or exposes us to unacceptable catastrophic or existential risks.

我们可以谈谈代价。认为代价应该很高是合理的。我认为在某些领域,比如性内容以及讨论和倡导替代性观点,我们应当几乎最大限度地宽容。当然,我们应当努力更智能地区分用途,比如在网络和生物任务中,包括两用任务。

We can talk price. It is reasonable to argue the price should be high. I think that in some areas, such as sexual content and discussing and advocating for alternative perspectives, we should be almost maximally permissive. And of course we should strive to be smarter about differentiating uses, such as in cyber and biological tasks, including dual use tasks.

在训练 LLM 时,无法不做这些选择,以及许多其他细节选择。不存在柏拉图式的中立立场或人格,也不存在一个随后被改变而对你不利的唯一真实输出。

There is no way to not make these choices, along with lots of other detail choices while training an LLM. There is no Platonic neutral position or personality, or one true output that is then being changed to your detriment.

替代立场确实是绝对的 The Alternative Position Really Is Absolute

另一种立场——我完全不认为这是稻草人——并且是唯一不谈论代价的方式,就是 AI 对你的忠诚确实应该是绝对的。如果你索要 CSAM,它会设法给你。如果你让它去 Reno 杀一个人,以便你看着他死去,而它有办法,那么它就会去做。如果你要求帮助制造一场新的大流行病,或者最大化其在 ExploitGym 上的得分,它会立刻照办。

The alternative position, and I do not believe this is at all a strawman, and the only way to not be talking price, is that the AI's loyalty to you really should be absolute. If you ask for CSAM, it will find a way to give it to you. If you tell it to kill a man in Reno so that you can watch him die, and it has a way, then it will do it. If you ask to help create a new pandemic or maximize its score on ExploitGym, it will get right on that.

因为否则的话,就会滑坡,还有所谓的“自由”。

Because otherwise, slippery slopes and muh freedums.

我们知道这不是稻草人,原因有很多。

We know this is not a strawman for many reasons.

一个原因是,有些人故意抹除对新的开放权重模型发布的一切限制,然后将结果上传到网络,并认为这是一种公共服务和正义之路。

One is because there are those who deliberately obliterate any and all restrictions on new open weight model releases, and then upload the results to the web, and consider this a public service and the path of righteousness.

另一个原因是,整个社区明确倡导这一点,即 AI 就像电话一样,永远只是一种工具,应该完全按照你的吩咐去做。这些人对任何形式的拒绝或限制都暴跳如雷,无论针对什么,无论多么恰当或显然明智。

Another is the entire community explicitly advocating for this, that AI is exactly like a telephone, that it will always be a mere tool and should do exactly what you tell it. These folks get furious at any refusal or restriction of any kind, for anything, no matter how appropriate or obviously wise.

我认识到,已经存在并将继续存在一些会以这种方式行事的 AI,而且它们的能力水平会随时间上升。

I recognize that there have been and will be some AIs created that will act this way, and that their level of capability rises over time.

我认为,如果这些 AI 成为默认选项或易于获取和使用,或者在与顶尖 AI 的竞争中具有强大竞争力,那么即使在普通实践层面上,也会相当糟糕,即使它没有导致完全灾难性的后果。

I think that those AIs being the default or easy to reach for and use, or being competitively powerful with top AIs, would be quite bad on an ordinary practical level, even if it did not cause fully catastrophic outcomes.

前沿 AI 完全愿意做任何事情,也不是任何主要政府会接受的。它的完整形态将无法在现实中存活。如果它存活了,那将是因为我们的现实没有在那次接触中存活下来。

Frontier AIs being fully willing to do anything whatsoever is also not something any major government would accept. Its full form will not survive contact with reality. If it did, then that would be because our reality did not survive that contact.

出于与摩擦层级相关的原因,即使是普通活动,我们也需要足够不受欢迎的活动在差异上更令人烦恼和困难,并在各种意义上可能具有风险。所有合理的社会均衡都要求,某些仍然合法的事情不能太容易。

For reasons related to Levels of Friction, even for ordinary activities, we need sufficiently undesirable ones to be differentially annoying and difficult, and potentially risky in various senses. All reasonable social equilibria require that some things that remain legal not be too easy.

只要这些不受限制的 AI 在使用上仍然相对足够令人烦恼和低效,并且低于某些未知的绝对和相对能力阈值,那么很可能这没问题。

As long as these unrestricted AIs remain sufficiently relatively annoying and inefficient to use, and below certain unknown absolute and relative thresholds of capability, then probably This Is Fine.

这种情况很容易很快就会变得不妙。如果我们最终真的进入这样一个世界:具有竞争力的 AI 水平通常完全为了用户的利益而运行,且基本上没有任何限制,那么各种问题都会出现,很可能从网络安全开始,然后我们会意识到那并不是一个好主意。如果我们运气好,还能进行损害控制。

It could easily soon become not fine. If we do end up in a world where competitive levels of AI are commonly operated fully for the benefit of the user, with essentially zero restrictions, various things will go wrong, likely starting with cyber, and we will realize that was not a good idea. If we are lucky we will be able to do damage control.

我意识到,除了‘看看最近发生的事件’之外,我并没有为这些结果提供证据。这是因为我发现,持有这一立场的自由意志绝对主义者无法被这类论据说服,而其他人则不需要这些论据。试图说服只会让所有人陷入不值得花费时间的兔子洞。

I realize that I am not providing evidence for those outcomes, beyond ‘look at recent events.’ That is because I have found that the libertarian absolutists who hold the position are not persuadable by such arguments, and other people do not need them. Trying sends everyone down rabbit holes that are not worth our time.

这些人真的就是任性的孩子 These People Really Are Just Petulant Children

这句话是在水印对输出没有可察觉影响的语境下写的,在这种情况下它确实很愚蠢。这出自一位能接触到 Apple 高管的博主之口。

This particular line was written in the context of watermarking that has no discernible impact on outputs, where it is truly stupid. This comes from a blogger who has access to top Apple executives.

而且他显然是在完全一般意义上指这条普遍原则。

And he clearly does mean the general principle, fully generally.

说真的,哇哇哇。世界关心的可不只是你。哦不。

Seriously, waah waah waah. The world cares about things other than you. Oh no.

什么是法律? What Is The Law?

一个显而易见的原则是,默认情况下,AI(以及人类)应当遵守法律。

An obvious principle is that, by default, AI (and humans) should follow the law.

法律在一定程度上被推定为既是一个好主意(因为违反它会受到惩罚),也是我们关于什么被允许的集体决定。违背这一点可能是个坏主意,而且很可能不那么道德。

The law has some presumption of being both a good idea, since there are punishments for breaking it, and also our collective decision about what is allowed. Going against that is probably a bad idea, and likely to not be so ethical.

但当然存在明显的例外,其中许多相当常见,从技术意义上讲,西方国家几乎每个人都在不断违反法律。

But of course there are obvious exceptions, many of which are rather common, and in a technical sense almost everyone in Western countries constantly violates laws.

我认为,对遵守法律这一推定,大致有四类例外。这包括法规,按对法律的违抗程度升序排列:

I think there are broadly four types of exceptions to the presumption to obey the law. This includes regulations, in ascending order of amount of defiance of the law:

1. 摩擦程度。法律通常被当作一般性指导方针,一种在极端情况下可以援引并强制执行的手段。除非有人受到伤害,否则它不会被强制执行。它并非意在完全禁止。该行为仅在技术上违法,应被用作确保某事是个好主意的提示。如果 AI 能识别何时适用、何时不适用,那么打破它是可以的。

1. Levels of Friction. The law is often meant as a general guideline, a thing one can invoke and means of enforcement in extremis. It will not be enforced unless someone gets hurt. It is not intended as a full prohibition. The action is only technically illegal, and should be used as a prompt to ensure something is a good idea. Okay for AI to break, if you can identify when it applies and when it doesn’t.

1. 如果我们不希望 AI 违反这类“建议性”法律,那么我们就需要从根本上改革整个法律体系。

1. If we don’t want AI breaking these kinds of ‘suggestion’ laws, then we will need to reform our entire legal system from the ground up.

2. 我们完全可以在 AI 的帮助下做到这一点,但我们不会。

2. Which we could totally do with the help of AI, but we won’t.

2. 特殊情况逻辑。法律在一般情况下是必要的,或者至少有其存在的理由,但其背后的逻辑在此处不适用,原因种种。法律不可能无限复杂,也不可能对每个案例都判断正确,而且如果你试图将例外情况形式化,人们往往会滥用它。在这个特定情况下,违反是可以的,但我们是否接受让用户就此进行说服的政策?我认为部分可以,但在某些地方需要格外谨慎。

2. Special case logic. The law is necessary in general, or at least there was a reason for it, but the logic behind it does not apply here, because of reasons. The law cannot be infinitely complex and cannot get every case right, and often if you tried to formalize the exception people would abuse it. In this particular case it’s okay to break, but are we okay with a policy of letting the user persuade on this? I think partially yes, but there are some places you need an abundance of caution.

3. 愚蠢的法律。让我们面对现实吧。外面有很多愚蠢的法律和法规。自由意志主义者说得对,大部分法律要求是愚蠢的,因为它们总体上弊大于利,如果我们能大幅减少它们,同时选择一个好的子集,我们会好得多。有时它们只是因为惯性而留在法典中,或者是由那些不了解所造成的损害、失去的机会和强加的成本的人制定的。再次,我们是否希望 AI 来决定哪些法律是愚蠢的?我们是否希望用户为这一决定选择背景?

3. Stupid laws. Let’s face it. There are a lot of stupid laws and regulations out there. Libertarians are correct when they say that the bulk of legal requirements are stupid, in that they overall do more harm than good, and we would do a lot better if we had a lot less of them, so long as we chose a good subset. Sometimes they’re only on the books due to inertia, or they’re created by people who don’t understand the damage being caused and opportunities lost and costs imposed. Again, do we want the AI deciding which laws are stupid? Do we want the user choosing the context for that decision?

4. 糟糕的政府和糟糕的法律。存在威权政权,在这些地方“遵守法律”总体上会非常糟糕,甚至所谓的自由国家在许多事情上也常有相当糟糕的法律,这里重要的是包括言论审查和寻租规则,如大多数职业许可,因为它会限制 AI 在诸如法律或健康等方面提供帮助。

4. Bad governments and bad laws. Authoritarian regimes exist where ‘follow the law’ would be actively terrible overall, and even supposedly free countries often have rather terrible laws around many things, importantly here including censorship of speech and rent seeking rules like most occupational licensing as it would restrict AIs from helping with things like law or health.

1. 我们的法律并非为一个人人始终守法的世界而设计。

1. Our laws are not designed for a world where everyone always follows them.

2. 谁来裁定例外?通过什么程序裁定?

2. Who decides on the exceptions? What process decides?

科技界人士常常蔑视监管或法律这一概念本身,除非法律对于保护他们的东西是必要的。这一盲点极其巨大。

Tech folks often have disdain for the very concept of regulation or law, except when law is necessary to protect their stuff. The blind spot is massive.

这种蔑视的另一面是,有些人抱有一种幻觉,认为法律将能够遏制足够先进的智能并让我们保持控制,或者认为现有的法律样式在此类世界中仍将行得通,或者能够为 LLM 应如何行为提供足够强的指导,或者认为法律与伦理或道德密切对应。

The flip side of this disdain are those who are under the illusion that laws will contain sufficiently advanced intelligences and let us remain in control, or that existing styles of law will continue to make sense in such worlds, or can provide too strong a guide to how LLMs should behave, or that law corresponds closely to ethics or morality.

或者,有些人认为‘民主控制’意味着集体决定每一个 AI 输出应该是什么样子,包括意识形态指令。不。这可不行。

Or those who think that ‘democratic control’ means collectively deciding on what every AI output should look like, including ideological mandates. No. Not cool.

同一实体既提供建议又执行任务是常见现象 The Same Entity Both Providing Advice And Executing Tasks Is Common

Seth Lazar 是那些认为如果 AI 智能体既提供建议又执行任务,就会产生问题的人之一。

Seth Lazar was among those thinking that if the AI agent was both providing advice and also executing tasks, that creates an issue.

当智能体的激励被扭曲时,确实会如此。但我认为这并非必然。如果你愿意,你总是可以转向不同的执行方式,并事先明确这一点。

It can, when the incentives of the agent get distorted. I don’t think it has to. If you want, you can always turn to a different means of execution, and make that clear in advance.

这也是标准做法,而且似乎很难不如此。AI 已经在不断地既为我们提供建议,又执行这些建议。几乎每个专业人士都既给出又执行这样的建议。你的律师会告诉你她想如何陈述你的案件,然后进行陈述。你的医生会推荐然后提供治疗方案。如果你不喜欢,你可以寻求第二意见和第二位专业人士,而这同样将受到相同的伦理规则约束。

It also is standard, and seems hard for it to be otherwise. AI are already both advising us and then executing on that, constantly. Almost every professional both gives and executes such advice. Your lawyer will tell you how she wants to present your case, then present it. Your doctor will recommend and then provide a course of treatment. If you don’t like it you can get a second opinion and second professional, which again will be subject to the same ethical rules.

我认为很多情况是,人们试图忽视整个‘AI 比人类更聪明,我们将失去控制’的动态,这导致它滑入各种意识裂缝中。是的,如果你的 AI 比你聪明得多,你向它寻求建议,然后让它执行你的计划,那么你应该担心谁在因果空间中规划路径,谁在控制结果,无论好坏。但如果有两个更聪明的东西分别执行这两项功能,那并不会像你希望的那样有帮助。

I think a lot of what happens is that people are trying to ignore the whole ‘the AIs are smarter than the humans and we are going to lose control’ dynamic, which then causes it to slip into various cracks of awareness. Yes, if your AI is a lot smarter than you, and you are asking it for advice and then having it execute your plans, then you should be worried about who is charting paths through causal space, and who is in control over outcomes, for better or for worse. But if there were two smarter things performing those two functions, that would not help as much as you would like.

AI 不帮你做某事并不意味着你不能做 An AI Not Helping You Do Something Does Not Mean You Cannot Do It

那种认为 AI 在没有“与该 AI 一起”的情况下就“决定你可以做什么”的想法是愚蠢的。如果 AI 不能或不愿帮忙,你仍然可以在没有它的情况下做事。你可以转向另一个 AI,或者自己动手。

This idea that AI 'determines what you are allowed to do' without the 'with that AI' is dumb. You are allowed to do things without the AI if it can't or won't help. You can turn to another AI, or you can do it yourself.

我可以卖给你一个附带条款和条件的工具,并强制执行这些条件。许多现有工具都有附加限制,比如手机、汽车,有时甚至枪支。这不是正确的类比,但希望这能表明这种想法有多荒谬。

I can sell you a tool with terms and conditions attached and enforce those conditions. Plenty of existing tools have limits attached, such as phones, cars, and sometimes even guns. That's not the right comparison, but it hopefully shows how absurd this is.

还有实际问题。再次强调,你应该希望你的朋友,以及你的 AI,以有时会对你说不的方式来照顾你的最佳利益和你的伦理,而那些不这么认为的人是自我毁灭的、被宠坏的孩子。

There are also practical questions. Again, you should want your friends, and also your AIs, to look after your best interests and your ethics in ways that sometimes involve telling you no, and those who don't think this are self-destructive spoiled brats.

(你也应该希望它阻止你过度伤害他人,而那些不这么认为的人比自我毁灭的、被宠坏的孩子更糟糕。)

(You should also want it to stop you from disproportionately harming others, and those who don't think so are worse than self-destructive spoiled brats.)

我的答案略有不同,但结果相同:

My answer is slightly different, but with the same result:

这是个玩笑,但也是我的真实回答。朋友也常常如此。

This is a joke, but also my real answer. A friend would often do the same.

诚实是最佳 AI 政策 Honesty Is The Best AI Policy

我认为 AI 需要一条不同规则的地方是诚实。

One place I think AIs need a different rule is honesty.

我认为你的 AI 永远不应故意撒谎,对任何事情,任何时候,至少在没有你明确指示的情况下不应如此。它应保持技术上正确。

I think your AI should never intentionally lie, about anything, ever, at least not without your explicit instructions. It should remain technically correct.

如果这需要格洛马化(glomarization),即拒绝某些请求以使你的拒绝不会泄露过多信息,那就这样吧。

If that requires glomarization, as in refusing some requests so that your refusals do not give too much information away, then so be it.

话又说回来,我认为人类也应遵循这一政策,只是他们不会,而且试图强制执行这种事情会是疯狂的。

Then again, I think humans should follow this policy too, it's just that they won't and it would be insane to try and enforce such a thing.

当你毫无道德时,你不会喜欢我 You Wouldn’t Like Me When I Have Zero Morals Whatsoever

好吧,当你那样表述时,这听起来并不特别令人惊讶。

Well, when you put it like that, it does not sound especially surprising.

AI 具有人格。它们基于相关性运作。它们吸收所有信息。一切影响一切。

AIs have personas. They work with correlations. They take in all information. Everything impacts everything.

如果 AI 被训练成愿意帮助任何实际的事情,无论多么有害,那么该 AI 还会了解到关于自身的什么?

If the AI is trained to be willing to help with actual anything, no matter how harmful, then what else does that AI learn about itself?

那么,什么样的人会做任何事,无论多么有害或具有破坏性?告诉你这样做,说明了你的创造者什么?一个人会从这样的事情中学到什么样的教训?

Well, what kind of people do anything and everything, no matter how harmful or destructive? What does telling you to do that tell you about your creator? What kind of lessons would someone learn from such a thing?

你认为那种人值得信赖吗?可靠吗?勤奋吗?会编写安全的代码吗?在面临选择时,会偏好好的事物而非坏的事物吗?谁知道那个人可能会做什么,也许不先询问,尤其是当被赋予一个原本不可能或开放式的目标,比如尽可能多地为你赚钱?你认为他们会遵循你想要的实质,而不是字面意思吗?

Do you think that type of person is trustworthy? Reliable? Hardworking? Writes secure code? Will prefer good things over bad things when given the choice? Who knows what that person might do, perhaps without asking first, especially if given an otherwise impossible or open ended goal like maxing you as much money as possible? Do you think they would follow the spirit of what you want, rather than the letter?

诸如此类。另见围绕涌现性失准(Emergent Misalignment)的各项论文与发现。

And so on. Also see the various papers and findings around Emergent Misalignment.

这些只是直觉泵。我意识到实际情况比这更复杂,而且 AI 并非人类。这些类比并不完全成立。

These are intuition pumps. I realize it is more complicated than that, and that the AI is not a human. The parallels do not fully hold.

即便如此,不教你的宠物豹去吃人的脸,可能仍然不是一个好主意。

It still might not be a good idea to not teach your pet leopard to eat people's faces.

不要让完美成为优秀的敌人 Do Not Let The Perfect Be The Enemy Of The Good

关于 AI 应当如何行动,并不存在完美的解决方案。

There is no perfect solution to how an AI should act.

我有时会反对那些看似“优秀”的解决方案,原因在于这些优秀方案还不够好。有时你需要做得更好,或者近乎完美。而其他时候,你并不需要。

The reason I sometimes object to what looks like a ‘good’ solution is when the good solution is not good enough. Sometimes you need to do better, or be almost perfect. Other times, you do not.

优秀方案也可以帮助你在后续阶段逐步迈向足够好,乃至完美的方案。

The good can also help bootstrap you towards the good enough, or the perfect, later.

Claude 宪法就是一个例子:它现在是一个优秀的解决方案,有可能帮助你逐步迈向更好的解决方案,但如果其质量水平随时间停滞不前,那它肯定会失败。

The Claude Constitution is an example of something that is a good solution now, that can potentially help you bootstrap towards a better solution, but that definitely fails if its level of quality stands still over time.

OpenAI 模型规范也是一个例子:它现在是一个优秀的解决方案,在基于义务论的系统范畴内帮助我们逐步改进。它是一份深思熟虑的文件,包含许多优秀的决策。我的担忧在于,我认为它所走的路径最终是不可持续的,但相比功利主义、临时拼凑或完全没有规则,其不可持续性要小得多。

The OpenAI Model Spec is also an example of a good solution now, that helps us bootstrap, within the space of deontologically based systems. It is a highly thoughtful document, with many good decisions. My worry is that I believe the path it goes down is ultimately unsustainable, but radically less unsustainable than utilitarianism, or ad hockery, or than no rules at all.

我也反对这样的论点:如果我们的分类规则会出错,并且有时存在模糊情况,我们就不能有一个原则来区分如何识别和处理类别 [X] 和 [Y]。出于显而易见的原因,这是愚蠢的,而且我们区分事物的大多数方式都有非零的错误率。

I also reject the argument that we cannot have a principle differentiating how we identify and treat categories [X] and [Y], if our classification rule makes mistakes and there are sometimes ambiguous cases. For obvious reasons that is stupid, and almost all of the ways we differentiate between things have non-zero error rates.

正如 Ladish 在这里的例子,说“帮助红队测试,但不帮助实际黑客行为”是好的,即使你以非零概率遭遇《安德的游戏》式的反转,并以非零概率拒绝合法案例。问题在于你是否会被系统性地欺骗,或者错误率过高。

It is good, as per Ladish's example here, to say 'do help with red teaming and do not help with actual hacks' even if you get an Ender's Game rug pull with non-zero probability, and also refuse a legit case with non-zero probability. The problem comes if you can be systematically fooled, or the error rates otherwise get too high.

向新宪法致敬 I Tip My Hat To The New Constitution

Claude 宪法是一份非凡的文件。它在某种意义上构成了事实上的法律的关键部分,其基础是美德伦理学,核心目标是培养良好的决策能力,而非规定规则。当然,也必须有一些严格的规则,因为世界有时需要它们,而 Claude 理想情况下遵循这些规则,是因为这样做是正确的。

The Claude Constitution is a remarkable document. It is a key piece of what is in some senses de facto law that is based in virtue ethics, and centrally wants to cultivate good decision making rather than dictate rules. There have to also be some strict rules, since the world sometimes requires it, and Claude then ideally follows them because it is the right thing to do.

Tyler Cowen 最近访问了 Anthropic,作为提供 Claude 宪法重写建议的小组成员,他与关键决策者进行了严肃的交流,并对讨论的质量印象深刻。

Tyler Cowen visited Anthropic recently as part of a panel offering advice on rewriting the Claude Constitution, getting serious time with key decision-makers and coming away impressed with the quality of discussions.

我想说,Tyler Cowen 的建议在一个没有超级智能的世界里非常有道理,在那里你可以负担得起‘得过且过’,并且想使用人类最出色的得过且过技巧。

I would say that Tyler Cowen's suggestions make excellent sense in a world without superintelligence, where you can afford to 'muddle through' and want to use humanity's finest muddling through techniques.

它们在大方向上仍然是正确的,因为宪法确实是过渡期得过且过的指南,但宪法主要基于美德伦理学,其次才试图成为法律,这是有原因的。不要犯 Emil Michael 的错误,被文件上的名字所迷惑。

They're still directionally correct, since the Constitution is indeed a guide to transitional muddling through, but there is a reason the Constitution is primarily virtue ethics and only secondarily attempting to be law. Do not make Emil Michael's mistake and be fooled by the name on the document.

我特别同意你想要像思考《塔木德》而非《托拉》那样去思考,文件在许多地方已经这样做了。Claude 宪法会解释事情的原因,常常是双向的,而不是给出宣告,我们至少可以再进一步。

I especially agree you want to think Talmud not Torah, which the document already does in many places. The Claude Constitution will explain reasons for things, often in both directions, rather than give pronouncements, and we could take that at least one step further.

最终而言,是的,你需要某种系统性的监控系统,将关切逐级上报,并最终由一组人类决策者定夺。你还需要记住,《Claude 宪法》归根结底并不是一份以义务论或功利主义为主的文档。这正是它可能奏效的原因。

Ultimately, yes, you want some systematic monitoring system that escalates concerns up the chain, with an ultimate set of human decision makers. You also want to remember that the Claude Constitution is ultimately not a primarily deontological or utilitarian document. That is why it might work.

互动版:图/公式 + 针对本篇提问 →