AI 文本水印是免费且有益的

AI Text Watermarking Is Free And Good

兹维·莫绍维茨 Zvi Mowshowitz · Don't Worry About the Vase · 2026-08-21 · Don't Worry About the Vase ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文探讨了围绕 Anthropic 实施 AI 文本水印的争议,认为反对声浪与实际影响不成比例。作者解释说,水印是无成本的,因为 AI 输出本质上是随机的,可以在不降低质量的情况下进行编码。文章指出了引发公众愤怒的几个因素,包括对随机性的误解、对任何改动的怀疑,以及对 Anthropic 的普遍不信任。文章还回应了合理的担忧,例如恶意行为者可能移除水印,以及人工编辑文本可能产生误报。作者总结道,虽然有些担忧是合理的,但大多数反对源于偏执和缺乏技术理解,水印是提高透明度的有益工具。

This article explores the controversy surrounding Anthropic's implementation of AI text watermarking, arguing that the backlash is largely disproportionate to the actual impact. The author explains that watermarking is costless because AI outputs are inherently random, allowing for encoding without quality degradation. The piece identifies several factors driving public anger, including misunderstanding of randomness, suspicion of any alteration, and a general distrust of Anthropic. It also addresses legitimate concerns, such as the potential for watermark removal by malicious actors and false positives on human-edited text. The author concludes that while some concerns are valid, most objections stem from paranoia and a lack of technical understanding, and that watermarking is a beneficial tool for transparency.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 13)

全文 · Full text(逐段中英对照)

目录 Table of Contents

3. 人们不理解 LLM 的输出已经是随机的。

3. People Don’t Understand LLM Outputs Are Already Random.

4. 人们不相信这种方法是无成本的。

4. People Don’t Trust The Method To Be Costless.

5. 人们原则上对任何改动都持怀疑态度。

5. People Are Suspicious Of Any Alteration On Principle.

6. 也许部分原因在于“水印”这个词。

6. Maybe It’s Partly The Word Watermark.

7. 很多人不想被抓到。

7. A Lot Of People Don’t Want To Get Caught.

8. 有些时候你宁愿不被认出。

8. There Are Some Times You Prefer Not To Be Recognized.

9. 有些担忧是有充分理由的。

9. There Are Some Good Reasons To Be Concerned.

11. 中间部分的写作与错误率。

11. The Writing In The Middle and Error Rates.

12. 百万用于防御,但不为进贡出一分钱。

12. Millions For Defense But Not One Cent For Tribute.

这没什么 This Is Fine

不,并非如此。很多人的反应是变得非常愤怒。

No. Not so much. A lot of people responded by getting Big Mad.

实际效果无非是:会有一个 API 告诉你某段文字是否出自 Claude。仅此而已。然而。

The entire practical effect is: There will be an API that will tell you if a given piece of writing comes from Claude. That’s it. And yet.

这篇文章的其余部分将探讨为什么人们对此非常愤怒,这在很大程度上是一个关于人们如何为几乎不存在的事情而激动的实例。

The rest of this post is about exploring why people are Big Mad about this, in large part as a worked example of how people get worked up over approximately nothing.

我的结论是,多种不同的因素正在汇聚在一起。

My conclusion is that a bunch of different factors are coming together.

Anthropic 错乱综合征 Anthropic Derangement Syndrome

这是主要原因。我们不要假装不是这样。

This is the main reason. Let’s not pretend otherwise.

你不会看到人们因为这件事对 Google 大动肝火。你也不会看到他们对 OpenAI 大动肝火,尽管这项技术是在 OpenAI 发明的,而且他们已承诺未来会继续这样做。诸如此类。

You don’t see people getting Big Mad at Google over this. You don’t see them getting Big Mad at OpenAI, even though that’s where this was invented and they have committed to doing this going forward. And so on.

不幸的是,Anthropic 是第一个宣布将实施水印以遵守欧盟《实践准则》的公司。

It is unfortunate that Anthropic was the first to announce they were implementing watermarking to comply with the EU Code of Practice.

因为这现在与 Anthropic 联系在一起,各种负面情绪试图附着其上。某些类型的人会寻找理由来感到不满。

Because this is now associated with Anthropic, all sorts of bad vibes try to attach themselves. Certain types of people look for reasons to be upset.

Anthropic 怎么做都不对。如果他们一开始就高调宣传,人们会把水印与 Anthropic 联系起来。因为他们一开始不够高调——因为这件事实际上不是什么大事——人们反而对此感到不满,然后仍然把它与 Anthropic 联系起来,尽管 Google 已经在两年前实施了水印并发布了公开检测器。

Anthropic can’t win. If they were initially louder about it, they would tie watermarking to Anthropic. Because they started out insufficiently loud, due to this not actually being a big deal, people get mad about that instead, and then still tie it to them, despite Google having implemented it and shipped a public detector over two years ago.

人工智能并不被普遍视为人,但功劳应归于应得之处。

AIs are not widely considered people, but credit should go where credit is due.

关于水印的其他一些抱怨是否合理、可以理解,或是源于真正的误解?当然有。但很大一部分是,人们认为 Anthropic 的氛围不好,因此寻找理由感到不满,并且原则上拒绝相信我上面给出的解释,而假设一定有某种险恶的事情在发生。

Are some of the other complaints about watermarking legitimate, understandable, or born of genuine misunderstandings? Sure. But a lot is that people think Anthropic vibes are bad, and thus look for reasons to be upset, and on principle refuse to believe the explanation I put up top, and assume something sinister must be going on.

有些人的偏执和妄想指向的是抽象的“开放性”概念,而非特别针对 Anthropic,但在上下文中这等同于同一件事。这是一个例子,说明人们迷恋于“这不是开放的,因此必定是险恶的”,尽管在这种情况下,实际上这是数学性的,任何人都可以验证。当然,通过询问封闭模型 Grok,因为埃隆·马斯克有着模糊的开放氛围,尽管他所有的竞争模型都是封闭的。

Some people’s paranoia and derangement is directed at abstract notions of ‘openness’ rather than Anthropic in particular, which amounts to the same thing in context. This is an example of fetishizing that this is not ‘open’ therefore must be sinister, even though in this case actually it is mathematical and anyone could verify it. By asking closed model Grok, of course, because Elon Musk has vaguely open vibes despite keeping all its competitive models closed.

人们并不理解大语言模型的输出本就是随机的 People Don’t Understand LLM Outputs Are Already Random

如果 AI 的输出是确定性的,那么在不使输出至少略微变差的情况下嵌入水印是不可能的。

If AI outputs were deterministic, it would be impossible to encode a watermark without making them at least marginally worse.

许多人直觉上认为 AI 的输出不是随机的。他们认为他们得到的是唯一正确的输出,尽管你可以重新生成输出,而且它确实会有所不同。因此,如果你没有得到“真正的”或原始的输出,那意味着你的输出一定变得更糟了。

Many people intuitively think that AI outputs are not random. They believe that what they get is the One True Output, even though you can regenerate the output and it will reliably be somewhat different. Thus, if you’re not getting the ‘real’ or original output, that means your output must have gotten worse.

当人们听到水印时,他们认为这会使输出变差,因为要留下标记,你必须做出不同的选择。你必须改变一些东西。

When people hear watermark, they think it will make the outputs worse, because to leave a mark you have to make different choices. You have to change something.

同样,他们会认为,如果你在优化额外的东西,那将会花费更多。

Similarly, they would assume that if you are optimizing for an additional thing, it is going to cost more.

但在这里,这种想法被证明是错误的,因为你有足够的随机性可以利用,你可以在保持相同有效分布的同时,仍然嵌入水印。

It turns out to be wrong here, because you have enough randomness to play with that you can get the same effective distribution, and still encode the watermark.

而且,并非所有人都了解其原理。几乎没有人仔细阅读过 Scott Aaronson 的文章,因为几乎没有人阅读。真正属于常识的东西少之又少。如果你在讨论中期望人们了解基本的技术事实,你将会一直感到深深的困惑。

And no, everyone does not know how this works. Almost no one reads Scott Aaronson in detail, because almost no one reads. Very few things are actually common knowledge. If you follow discourse expecting people to know basic technical facts, you will keep being deeply confused.

人们不相信该方法是无成本的 People Don’t Trust The Method To Be Costless

我认为,这里的怀疑在很大程度上被“Anthropic 错乱综合征”以及对 Anthropic 的普遍不信任所放大,即一种“他们有所图谋”的感觉。

I believe the skepticism here is greatly enhanced by Anthropic Derangement Syndrome, and by general distrust of Anthropic, a sense that 'they're up to something.'

其中有多少怀疑是源于对该方法的怀疑?

How much of the skepticism is due to skepticism of the method?

一项快速调查表明,这是主要的担忧。

A quick survey suggests that this is the majority of the concern.

[](https://substackcdn.com/image/fetch/$s_!9zmr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbacc0c5c-47cc-46ba-9ae0-84483ae8c3ac_1153x498.png)

[](https://substackcdn.com/image/fetch/$s_!9zmr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbacc0c5c-47cc-46ba-9ae0-84483ae8c3ac_1153x498.png)

这个帖子就是一个例子,有人难以接受输出实际上没有受到影响,因为分布没有改变。

This thread is an example of someone finding it difficult to accept that there is effectively no impact on outputs, because the distribution does not change.

我确实理解这种感受。这是一个魔术师的把戏,一个数学证明,它确实有效。'你可以标记它而不产生任何影响'这一说法非常反直觉,人们身体的每一根纤维都想说不,直到某个瞬间他们恍然大悟,意识到数学就是数学,实际上它确实有效。

I do sympathize. This is a magician's trick, a math proof, that works. There is something highly counterintuitive about 'you can mark it while having no impact' and every fiber in people's bodies wants to say no, until something clicks and they realize that the math is math and actually yes it works.

如果你像这样公开'认输',我会对你刮目相看,即使我从未见过原来的'输'。这里有一个明显的博弈论问题:如果每个人都预测到大家会那样反应,但事实并非如此,所以这一举动是安全且明智的。

I will think actively better of you if you 'take the L' like this in public, even if I never saw the original L. There is an obvious game theoretic issue with that if everyone predicted everyone would react that way, but they don't, so this play is safe and wise.

这是一个例子,有人非常非常努力地试图说这种方法在技术上'有权衡',以暗示它并非无成本,因为有一种强烈的驱动力不希望它是无成本的,以便对此感到愤怒,而实际上在实践中它显然是无成本的,事实上谷歌进行了大量实验来证明它是无成本的。

This is an example of someone trying really, really hard to say that this method technically 'has tradeoffs' to imply it is not costless, because there is a strong drive to not want it to be costless in order to be mad about it, when obviously in practice it is costless, indeed Google ran extensive experiments to prove it is costless.

这是一个例子,尽管数学非常清晰,但有人原则上直接说'不,我不相信'。而这一回应则是一个例子,基于'嗯,我仔细思考我的选择'的逻辑,拒绝理解选择无论如何都是随机的,从而幻觉出声称观察到的质量损失。

This is an example of flat out 'nope, I don't believe it' on principle, despite the math being very clear. And this response is an example of hallucinating a loss in quality, that is claimed to be observed, based on logic of 'well I think hard about my choices' and refusing to understand the choice is random either way.

这是一个(更为恶劣的)例子,有人假装不理解随机的含义,将'有两种可能的续写'解读为声称两种续写意味着相同的东西。

This is a (much more egregious) example of someone pretending not to understand what random means, reading 'there are two possible continuations' as a claim that the two continuations mean the same thing.

人们原则上对任何改动都持怀疑态度 People Are Suspicious Of Any Alteration On Principle

这种态度很常见,比如:

There is a lot of this kind of attitude, things like:

1. 任何不是为了帮助我而做的改动,本质上都可疑,可能会伤害我。

1. Any alteration done not to help me is inherently suspicious and might hurt me.

2. 如果我是客户,你无权乱动我的东西。

2. If I am the customer, you have no right to mess with my stuff.

当然,无论模型是开源的还是闭源的,它都是成千上万甚至数百万个决策的产物,其中许多决策并非基于客户的需求。所有这类产品都在不断被“改动”。通常,改动并不会“满足你的需求”,无论是你的个人需求,还是用户或客户的普遍需求。此外,还有许多其他问题需要解决,包括法律要求和防止伤害。

Except of course, whether the model be open or closed, it is the product of thousands or millions of decisions, many of which are not about what you the customer wanted. All such products are 'altered' constantly. Often the change will not 'serve your needs,' either yours in particular or those of users or customers in general. There are many other problems to solve as well, including legal requirements and harm prevention.

当这些改动不可见时,人们并不在意。当某一改动变得突出时,人们就会愤怒。

When this is invisible, people do not care. When one in particular becomes salient, people get angry.

与 DRM 的类比也在这里毒化了讨论环境。DRM 在所有实现中都很糟糕,因为它让你的产品变得更差。最好的情况是消耗资源,它还可能主动阻止你使用产品,有时甚至会搞乱你的机器。我们所有玩家都深知这种痛苦。有时你不得不忍受一些,因为版权不会自动执行,即使你个人仍然会尊重它。

The parallel to DRM is also poisoning the well here. DRM sucks, in all its implementations, because it makes your product worse. At best it eats resources, and it can actively prevent you from using the product, and sometimes it can mess up your machine. All us gamers know the pain well. You sometimes have to do some of it, because copyright does not enforce itself even if you in particular would still honor it.

这不是那种情况。但它与 DRM 略有相似之处,而这可能就足够了。

This is not like that. But it slightly vibes with it, and that can be enough.

也许部分原因在于“水印”这个词 Maybe It’s Partly The Word Watermark

我的意思是,我猜?我推测“秘密”会因同样的原因而糟糕十倍。

I mean, I guess? I presume 'secret' would be ten times worse for the same reasons.

这里基本上没有任何一个名称不会以同样的方式产生错误的观感,并给人留下你做了额外工作去改变某些东西、从而使事情变得更糟的印象。我认为“水印”是一个相对较好的名称,并且不希望陷入委婉语的循环之中。

There is basically no handle here that won't have the wrong vibes in the same way, and give the impression that you did extra work to change something, and therefore made things worse. I think watermark is a relatively good name and would prefer not to go on a euphemism treadmill.

很多人不想被抓到 A Lot Of People Don’t Want To Get Caught

当然,你不能直接这么说。嗯,有些人可以,但更多的人选择不这样做。这往往甚至不是完全有意识的。

Of course, you can’t come out and put it like that. Well, some people can, but a lot more of them choose not to. Often this will not even be fully conscious.

不要过度思考这种情况。很多人重视将 AI 写作冒充为自己的能力,或者他们希望人们无法证明这一点。

Don’t overthink the situation. A lot of people value the ability to pass off AI writing as their own, or they want people to be unable to prove it.

而且不,你不能切换到 ChatGPT,它们也会有标记。

And no, you won’t be able to switch to ChatGPT, they will have marks too.

还有一堆‘不是你会,但你可以’的说法。

There is also a bunch of ‘not that you would, but you could.’

有些时候你更希望不被识别 There Are Some Times You Prefer Not To Be Recognized

Anthropic 或 API 能否识别出生成文本的用户?

Can Anthropic or the API figure out which user generated the text?

不能。Anthropic 在 FAQ 中确认,此方法无法做到这一点。

No. Anthropic confirms in the FAQ that this method cannot do that.

有一些值得担忧的正当理由 There Are Some Good Reasons To Be Concerned

任何影响社会动态的变化,即使反对意见大多是错误的,且事物本身显然积极,仍然会存在一些弊端和担忧。

With any change that impacts social dynamics, even when the objections are primarily wrong, and the thing seems clearly positive, there are still going to be some downsides and concerns.

以下是我认为至少有些合理的一些担忧。

Here are the ones that seem at least somewhat legitimate to me.

作弊 作弊 作弊 作弊 作弊 Cheat Cheat Cheat Cheat Cheat

这是我最大的实际担忧,即最坏的人会移除水印。

This is my top real concern, which is that the worst people will remove the watermark.

移除水印并非易事,但既然你有一个答案键来训练和检查每种情况,无疑可以构建 AI 工具,以某种方式打乱小的选择,从而降低或消除水印。你可以验证它是否有效。

Removing the watermark is non-trivial, but given you have an answer key to train with and check in each case, one can doubtless build AI tools that scramble small choices in ways that degrade or erase the watermark. You can verify that it worked.

他们也可以使用本地或其他不带水印的模型,或者使用不太可能被检查的模型。

They could also use a local or other model that does not carry a watermark, or not one that anyone is likely to check.

这可能使你处于更不利的位置,因为他们可以利用假阴性作为辩护。当你有一个足够好的测试,默认信任它,但在关键时刻可能被伪造时,那可能非常糟糕。

This potentially puts you in a worse position, since they can then use the false negative as a defense. When you have a test that is good enough that you trust it by default, but that is possible to fake when it counts, that can be pretty bad.

我的回应是,这将会很烦人,而且这样做明显是内疚的表现,比最初使用 AI 要糟糕得多,而且其他检测方法仍然有效。Pangram 仍会将其标记为 AI,至少如果按照我设想的方式去做,它不会混淆原始内容来自哪个 AI(Pangram 算法可以区分不同的 AI,但不会告诉用户)。它仍然会被人类读者视为 AI。

My response is that this is going to be annoying, and also doing it is clear consciousness of guilt and far worse than the initial AI use, and that other methods of detection will still work. Pangram will still mark it as AI, at least if you do it the way I’m imagining, and won’t be confused about which AI the original came from (the Pangram algorithm can differentiate different AIs, but it doesn’t tell the user). It would also still read to a human as AI.

因此,如果例如一个大学生使用了它,而大学在没有水印证据的情况下无法采取行动,这个问题可能会滚雪球式地扩大,但现在学生需要付出更多努力,并且有明显的犯罪意图,实际上我预计这大多不会发生。

Thus this would be an issue if, for example, a college student used it, and the university couldn't act without the proof from the watermark, and that issue could snowball, but now the student needs to put in more effort, and has clear mens rea, and in practice I expect this to be mostly not something that is done.

现实世界的测试是另一个不必担心的理由。每个人都可以使用 Pangram。所以理论上,任何人都可以反复检查 Pangram 的结果,直到它被判定为人类。但我们已经观察到人们在现实中的反应,几乎没有人这样做。人们只是……被抓住了。

The real world test is another reason not to worry. Everyone has access to Pangram. So in theory, anyone could iteratively check the Pangram result until it comes back as human. But we have observed how people react in the wild, and approximately no one does. People just… get caught.

总的来说,如果你提高了提交 AI 工作的努力程度,这大多能起到作用。如果你通过用自己的话重写整个内容来去除水印,当然,这没问题,并且算作“任务圆满完成”。

In general, if you raise the effort level of submitting AI work, that mostly does do the job. If you remove the watermark by rewriting the whole thing in your own words, of course, that is fine and counts as Mission Fucking Accomplished.

中间态写作与错误率 The Writing In The Middle and Error Rates

当写作在一定程度上使用 AI,但大部分仍由人类完成时,水印会怎样?如果你使用了一些 AI 的惯用表达,即使作品大部分是你自己的,人们会将其归类为 AI 生成吗?会不会出现零容忍政策和不幸的案例?这一切都是概率性的,那么当答案出错时会发生什么?

When writing uses AI to some extent, but is still largely written by a human, what happens with the watermark? If you use some of the AI turns of phrase, will people classify your work as AI, even if it is largely your own? Will there be zero tolerance policies and unfortunate cases? This is all probabilistic, so what happens when the answer comes back wrong?

答案是水印衡量的是 AI 处理的程度。因此,翻译和文件转换可能会触发水印,这很不幸,但人们可以意识到这一点。校对也可能触发水印,如果你让 Claude 自动进行相关编辑的话;但如果你用它来发现错误并自己修正,就不会触发。

The answer is that the watermark measures AI processing. So translations and file conversions might trigger the watermark, which is unfortunate, but one can be aware of that. So could proofreading, if you let Claude automatically make the related edits, but if you use it to find errors and correct them yourself, it won’t.

因此,我确信强水印阳性在 AI 写作或文字处理方面始终是真正的阳性。我的帖子上的水印特征量不会是零,因为我引用了其他人,他们有时会使用 Claude,或者直接引用 Claude。直接引用有明确标记,但水印无法区分。在明显有益的情况下使用 AI 系统的任何文字,都会带来一定程度的偏执。

Thus I am confident that strong watermark positives will consistently be true positives in terms of AI writing or processing of the words. The amount of watermark signature on my posts will not be zero, since I am quoting others who sometimes will have used Claude, or sometimes directly quoting Claude. The direct quotes are clearly marked, but the watermark won’t know the difference. There will be some degree of paranoia when using any words from an AI system, in places where such use is clearly good.

所以,是的,这种情况会发生一些,但我预计人们会迅速习惯这一点,并且应该有一个高度假设:给人们更多信息是好的,如何反应取决于他们自己。如果你很偏执,并且身处那些对哪怕一丝 AI 使用都大为光火的人群中,而且那些人可能真的会使用 API,你可以先通过 API 检查自己的输出。或者你可以尊重那些人的偏好,即使在某些你我都认为有益的方式下也不使用 AI。

So yeah, a little of that will happen, but I expect people to rapidly get used to this, and there should be a high presumption that giving people more info is good and it is up to them how to react to it. If you are paranoid and among people who are Big Mad about even a sliver of AI use, and also that might actually use the API, you can check your own output first via the API. Or you can honor the preferences of those folks, and not use AI even in some of the ways that you and I would agree are good.

像往常一样,在实际错误和误报方面,人们对自动化系统的错误和潜在错误的容忍度远低于对人类的容忍度。错误率将非常非常低。如果一个人试图判断你是否使用了 AI,他们的错误率会很高,远高于水印。而 AI 使用是一个要求“证明”过度的领域,尤其是在学术界,常常让人们“逃脱”大家都知道他们做过的事情。这不是刑法,如果没有人会进监狱,你就不需要同样超高的置信度。

As usual, in terms of actual mistakes and false positives, people have vastly lower tolerance for errors and potential errors by automated systems, than they do for humans. The error rate is going to be very, very low. If a human is trying to decide if you used AI, they’re going to have a substantial error rate, far higher than the watermark. And AI use is one place where demands to ‘prove’ things go too far, especially in academia, often letting people often ‘get away with’ things that everybody knows they did. This is not criminal law, if no one is going to jail you should not need the same super high level of confidence.

百万军费,不纳贡赋 Millions For Defense But Not One Cent For Tribute

最后一个担忧并非关于水印本身,而是关于这一强制要求及其在欧盟《行为准则》中的来源。

The last concern is not about the watermark itself, but about the mandate and its origin in the EU Code of Practice.

Anthropic 因欧盟法律在全球范围内实施了水印,因为在全球范围内实施比仅在欧盟实施要容易得多。这意味着欧盟正在违背我们的意愿,侵蚀并污染我们宝贵的人工智能输出令牌。我们必须反击。谁知道他们接下来还会针对什么?

Anthropic implemented watermarking worldwide due to an EU law, since it is a lot easier to do it everywhere than only for the EU. So that means that the EU is sapping and impurifying our precious AI output tokens, against our will. We must fight back. Who knows what else they might target next?

从客观层面回应,不,他们并没有那样做。他们确实在技术上改变了输出,就像蝴蝶扇动翅膀改变天气一样,但并非以任何系统性或方向性的方式。如上所述,输出并未降级。这没问题。

The object level response is that no, they are not doing that. They are technically changing the outputs, the same way that a butterfly flaps its wings and changes the weather, but not in any systematic or directional way. As discussed above, the outputs are not degraded. This Is Fine.

真正的担忧在于原则,以及接下来可能发生的事情。是的,目前这没问题,强迫苹果使用 USB-C 目前也没问题,但他们对我们科技公司征收的罚款基本上是现代海盗行为,谁知道接下来会发生什么。

The real concern is the principle, and what might come next. Yes, this is fine for now, and forcing Apple to use USB-C was fine for now, but the fines they impose on our tech companies are basically modern piracy and who knows what comes next.

对此我要说,他们是一个巨大的市场,是的,他们确实有一定发言权,而且一直都有,限制因素在于美国是否会反击,或者极端情况下我们进行地理围栏,带着我们的球回家,正如一些人工智能服务在欧盟和其他地方因类似问题已经发生的那样。

To which I say, they are a huge market, and yes they get some say, and have always gotten some say, and the limiting factor is America pushes back or in extremis we geofence, take our ball and go home, as indeed has already happened with some AI services, in the EU and elsewhere, over similar issues.

理论上,接下来可能发生的另一件事是,对水印的要求没有规模下限。因此,理论上他们可以针对任何没有水印的 AI,包括小型开放模型。如果他们对此大做文章,那将是不利的。我预计不会发生这种情况,但他们做过更蠢的事。

Another thing that might come next, in theory, is that there is no size minimum on requiring watermarks. So in theory they could come after any AI without them, including tiny open models. If they made a sufficiently large fuss about this it would be bad. I do not expect this, but they’ve done stupider.

这就是博弈。我担心欧盟或其他方面会以这种方式强加审查制度,迫使我们的科技公司配合。这可能会扩展到对 AI 输出的意识形态要求。如果他们或其他方面真的试图这样做,我相信我们的前沿实验室不会在全球范围内应用此类变更,而且如果他们尝试,确实会面临强烈的反弹。相反,我预测实验室至少会威胁进行地理围栏,或对欧盟的查询使用激进的分类器或类似技术。

This is the dance. I am worried about the EU or others imposing a censorship regime in this way, forcing our tech companies to play along. That could plausibly extend to ideological requirements on AI outputs. If they or others did try that, I believe our frontier labs would not apply such changes globally, and indeed would face severe backlash if they tried. Instead, I predict the labs would at least threaten to geofence, or use aggressive classifiers or similar tech on EU queries.

是的,我们确实需要警惕布鲁塞尔的越权行为。但这件事?这没问题。

Yes, we do have to keep an eye out for Brussels overreaching. But this? This Is Fine.

互动版:图/公式 + 针对本篇提问 →