嵌入式评估者的探索

The Quest for Embedded Evaluators

兹维·莫绍维茨 Zvi Mowshowitz · Don't Worry About the Vase · 2026-09-27 · Don't Worry About the Vase ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文考察了为前沿 AI 公司建立可信嵌入式评估者这一难题,指出尽管独立性、专业性、透明度和免受报复在普遍意义上都是可取的,但在当前条件下无法同时实现。必须有人为评估者提供资金并聘用他们,而合格的人才不可避免地会通过顶级实验室积累经验,从而产生利益冲突,即便是成熟的审计机制也难以解决。作者梳理了杰弗里·辛顿、斯图尔特·罗素和阿文德·纳拉亚南的提议,以及加布里埃尔·韦尔的强制性责任保险构想,结论是保险只能起到边际作用,无法直接解决最重要的问题。Anthropic 与埃森哲的合作以及计划纳入 METR 的做法,被评估为非营利专业知识与正式企业审计的结合,其中 METR 虽非完美,但被视为最佳单一选择。结论是,OpenAI 提出的利用现有国家 AI 安全研究所的方案聊胜于无,但仍显软弱无力,而嵌入式评估的现有选项都不尽如人意。

This article examines the challenge of establishing credible embedded evaluators for frontier AI companies, arguing that while independence, expertise, transparency, and protection from retaliation are universally desirable, they cannot all be achieved simultaneously under current conditions. Someone must fund and hire the evaluators, and qualified people inevitably gain experience through the top labs, creating conflicts of interest that even mature auditing schemes struggle to resolve. The author reviews proposals from Geoffrey Hinton, Stuart Russell, and Arvind Narayanan, as well as Gabriel Weil's mandatory liability insurance idea, concluding that insurance would be only marginally helpful and would not directly solve the most important problems. Anthropic's partnership with Accenture and planned inclusion of METR is assessed as a combination of nonprofit expertise and formal corporate auditing, with METR favored as the best single choice despite imperfect alternatives. The conclusion is that OpenAI's proposal to leverage existing national AI safety institutes is better than nothing but remains milquetoast, and that the available options for embedded evaluation are not great.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 5)

全文 · Full text(逐段中英对照)

目录 Table of Contents

1. 听着,我所要求的只是你找到一个资质极高、经验丰富、值得信赖、没有利益冲突的嵌入式评估方,它将完全免费地工作,不接受政府援助,也不接受并非由实验室选择的 EA 来源的资金。

1. Look, All I'm Asking For Is That You Find A Highly-Qualified, Experienced, Trustworthy, Non-Conflicted Source of Embedded Evaluators That Will Work Entirely For Free, Without Government Assistance or Money from EA Sources Not Chosen By the Lab.

2. Anthropic 与埃森哲合作开展嵌入式评估,还计划纳入 METR。

2. Anthropic Partners with Accenture for Embedded Evaluation, also Plans to Include METR.

4. OpenAI 建议做我们能做的最少的事。

4. OpenAI Suggests Doing The Least We Can Do.

看,我所要求的只是:你找到一个完全合格、经验丰富、值得信赖、无利益冲突的嵌入式评估者来源,他们完全免费工作,不依赖政府援助,也不接受实验室未选择的 EA 来源的资金 Look, All I’m Asking For Is That You Find A Highly-Qualified, Experienced, Trustworthy, Non-Conflicted Source of Embedded Evaluators That Will Work Entirely For Free, Without Government Assistance or Money from EA Sources Not Chosen By the Lab

你为什么告诉我这很难?

Why are you telling me that this is hard?

显然,节标题中的所有要求单独来看和合在一起都是可取的。它们都是锦上添花的东西。问题在于,在当前条件下,你显然无法同时接近满足所有这些要求。

Obviously all of the requests in the section title are individually and collectively desirable. They are all nice-to-haves. The problem is, you obviously cannot get that close to having all of them at once under present conditions.

1. 总得有人在某个地方雇用这些人并支付账单。

1. Someone, somewhere, has to hire the people and pay the bill.

2. 合格、经验丰富且值得信赖的人需要以某种方式获得这种经验和信任,而这将涉及顶级实验室。

2. The qualified, experienced and trustworthy people need to have gained that experience and trust in some way, which is going to involve the top labs.

我并不是说这些抱怨没有道理,但即使是成熟且高度信任的审计方案,也大多缺乏好的解决方案。旋转门和资金来源都是难题。你仍然需要进行审计。

I am not saying the complaints are invalid, but even well-established, highly trusted auditing schemes mostly lack good solutions. The revolving door and the source of funding are hard problems. You still need to audit.

由 Geoffrey Hinton、Stuart Russell 和 Arvind Narayanan 领导的一个小组在一封公开信中列出了他们对可信嵌入式评估者的最低标准:

A group led by Geoffrey Hinton, Stuart Russell and Arvind Narayanan lays out their minimum standards for credible embedded evaluators in a public letter:

1. 前沿 AI 公司应依赖具有实质独立性的评估者。

1. Frontier AI companies should rely on evaluators that are meaningfully independent.

2. 报酬不能以评估发现为条件,不应存在所有权、治理或其他商业关系,也不应有编辑控制权。任何利益冲突都应予以披露。

2. Payments cannot be contingent on findings, there should be no ownership or governing or other commercial relation, and no editorial control. Any conflicts of interest should be disclosed.

3. 前沿 AI 公司应纳入不同的观点和专业领域。

3. Frontier AI companies should incorporate differing viewpoints and areas of expertise.

4. 嵌入式评估者应保持透明。

4. Embedded evaluators should be transparent.

4. 嵌入式评估员应受到保护,免遭其嵌入公司报复。

4. Embedded evaluators should be shielded from retaliation from the companies they embed with.

5. 前沿 AI 公司应授予嵌入式评估员与其自身高度特权员工同等的访问权限,但保护数据的情况除外。

5. Frontier AI companies should grant embedded evaluators access equivalent to that of their own highly privileged employees, with exceptions for protecting data.

Gabriel Weil 在 7 月指出,你不希望让 AI 开发者雇佣自己的裁判。但如果政府不愿支付他们费用,而且任何方面都对整件事相当敌视,那么还有谁会来挑选和雇佣他们呢?

Gabriel Weil pointed out in July that you do not want to let the AI developer hire their own referees. But who else is going to pick and hire them, if the government wants nothing to do with paying them and if anything is looking rather hostile to the whole thing?

我们生活在愚蠢的时间线里。Weil 担心政府将难以监督那些执行审计的 IVO。他没有考虑到白宫可能会基于原则而对审计 actively 发怒。Weil 的提议是强制性责任保险,这虽然是个好主意,但能完成全部工作的版本将无法购买,而实际售出的版本又无法覆盖我们最担心的那类事件。

We live in the stupid timeline. Weil worried that government will struggle to oversee the IVOs that would do the audits. He had not considered that the White House might actively get angry over the audits, on principle. Weil's proposal was mandatory liability insurance, which would be a good idea although the version of it that did the full job would be impossible to buy, and the version that got sold would not cover the kinds of events we most worry about.

如果 Anthropic 价值 2.2 万亿美元,而我们担心他们可能‘无清偿能力可执行’,那么究竟谁才不是无清偿能力可执行的呢?

If Anthropic is worth $2.2 trillion and we are worried they might be 'judgment-proof' then who exactly is not judgment proof?

我完全支持对超过某个能力阈值的人工智能模型要求购买普通(或常规)保险,包括在内部部署或权重发布之前,而不仅仅是封闭的外部部署。这会有边际帮助。但它不会直接解决最重要的问题。

I would be totally up for requiring mundane (or prosaic) insurance for AI models above some capability threshold, including prior to internal deployment or weights release, not only closed external deployment. It would be marginally helpful. What it would not do is directly solve the most important problems.

一种可能性是,你要求购买比如说 1000 亿或 1 万亿美元的保险,不是因为这样能修正激励,而是因为这样我们就可以要求他们公布保险费用的确切金额,并要求保险公司披露他们用于定价的相关风险因素。

One possibility is that you require buying let's say $100 billion or $1 trillion in insurance, not because that fixes the incentives, but because then we can require that they publish exactly what that insurance cost them, and for the insurance company to disclose the associated risk factors they used to price it.

Anthropic 与埃森哲合作开展嵌入式评估,并计划纳入 METR Anthropic Partners with Accenture for Embedded Evaluation, also Plans to Include METR

他们宣布与埃森哲建立合作伙伴关系。

They announced a partnership with Accenture.

根据背景讨论和其他信息,加上这里的字面解读,我的理解是 Anthropic 计划结合使用这两种方法。

From background discussions and other information, plus the plain reading here, my understanding is that Anthropic plans to use a combination of both approaches.

1. 自带资金和专业知识的非营利组织,能够提供更多专家内部视角,并关注安全问题。

1. Nonprofits that come with their own funding and expertise, and can provide more of an expert inside view combined with a focus on safety concerns.

2. 大型、正式、清晰的企业方法,采取更多外部视角,就像地方法官那样。

2. The big, formal, legible corporate approach that takes more of an outside view the way you would if you were a magistrate.

使用埃森哲的明显问题包括:埃森哲和 Anthropic 之前有过合作,而且 Anthropic 将为这次行动提供资金。两者都不理想,但再次强调,替代方案是什么?这应该是 Faculty 分支的工作,如果它能保持自身文化并保持独立,那是一个合理的选择。它们在多个方面清晰可信,并且确实有业绩记录。

The obvious problems for using Accenture include that Accenture and Anthropic have previously partnered, and that Anthropic will be funding this operation. Neither is ideal, but again what is the alternative? This should be the job of the Faculty subdivision, which if it keeps its culture and stays independent is a reasonable option. They are in a number of ways legible and credible, and they do bring a track record.

Luke Muehlhauser 持乐观态度,并指出参与其中的一些优秀人士。Oliver Habryka 和 Dave Kasten 则持怀疑态度。

Luke Muehlhauser is optimistic and points to some good people involved. Oliver Habryka and Dave Kasten are skeptical.

如果必须选择一个评估机构进行嵌入式合作,我会选择 METR。

If I had to pick one evaluator to embed, I would choose METR.

如果必须选择两个评估机构进行嵌入式合作,我会选择 METR,然后我会选择一个更知名且更稳妥的机构。可选项并不理想。

If I had to pick two evaluators to embed, I would choose METR, and then I would choose someone more legible and boring. The options are not great.

理论上,对于知名且稳妥的席位,我的首选是英国 AISI,但显然有理由认为这大概不会得到白宫的支持。理论上,如果白宫支持并愿意资助,CAISI 会是不错的选择。然后你会考虑四大审计公司,但它们都没有相关专业知识,而且其中三家是像埃森哲这样的 Claude 商店。也许是 MITRE 或 RAND,但它们的资金状况各有问题。没有很好的选择。

In theory my first pick for the legible slot would be UK AISI, but there are obvious reasons why that presumably wouldn’t play with the White House. CAISI would be great in theory if the White House was on board and wanted to fund that. Then you look at the Big Four auditing firms, but none have the expertise and three of them are Claude shops like Accenture. Maybe MITRE or RAND, but the funding situations there have their own issues. There weren’t great options.

你或许可以通过今天创办自己的审计公司来帮助解决这个问题。

You could perhaps help solve this problem by starting your own auditing firm today.

解读 METR Reading the METR

Drake Thomas 将关于嵌入式第三方评估者的各种观点从最差到最好进行了排序。这一分类法及其排序在我看来是正确的。

Drake Thomas ranks the takes on embedded third-party evaluators from worst to best. This taxonomy and its order seem correct to me.

从数量上看,这类观点中的绝大多数都远低于零点,而且没有任何观点真正论证了“不,我们不应该设置嵌入式评估者”。

The vast majority of such takes by volume are far below the zero point, and no takes make an actual case for 'no, we should not have embedded evaluators.'

OpenAI 建议采取最低限度的行动 OpenAI Suggests Doing The Least We Can Do

OpenAI 就此主题发表了一篇新文章《为 AI 下一阶段构建标准》。他们强调,其目标仍然是实现自动化的 AI 研究员,然后让这个研究员来完成我们的对齐作业,但前提是这能安全地完成。

OpenAI has offered a new essay on the subject, Building standards for the next phase of AI. They emphasize that their goal is still an automated AI researcher, which then would be asked to do our alignment homework, but only if it can be done safely.

我们已经到了这样一个节点:OpenAI 表示他们将“寻求在改进过程中保持人类参与”的方式,因为否则这种情况就不会发生。

We are at the point where OpenAI says they will 'seek ways to keep humans in the loop' of improvement, because otherwise that would not happen.

他们接着警告说,碎片化的国际标准或缺乏集体行动可能导致糟糕的结果。是的。

They next warn that fragmented international standards, or a lack of collective action, could lead to poor outcomes. Yes.

这么多人觉得有必要像念咒语一样不断重复“避免权力集中”,这是你被允许拥有的担忧。如果我们无法克服这一点,我看不出我们如何能到达一个好的境地。

So many people feel the need to keep repeating 'avoiding the concentration of power' like it is a mantra, the concern that you are allowed to have. If we cannot get over that, I don't see how we get to a good place.

好了,铺垫够了。提案是什么?

Okay, enough preliminaries. What is the proposal?

他们建议利用许多国家现有的 AISI 网络。

They suggest leveraging the network of existing AISIs in many countries.

好吧,诚然,这总比什么都没有强,但这是一个非常温和、缺乏力度的提议。

Okay, sure, that's all better than nothing, but this is a very milquetoast proposal.

互动版:图/公式 + 针对本篇提问 →