名为 Claude 的 AI 背后的坏人

The Bad Guy With An AI Named Claude

兹维·莫绍维茨 Zvi Mowshowitz · Don't Worry About the Vase · 2026-09-15 · Don't Worry About the Vase ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文审视了 Anthropic 关于其 Claude AI 模型被恶意使用的威胁情报报告,内容涵盖间谍活动、网络攻击、影响力行动、监控、武器研发以及两用生物研究。核心论点是:AI 如今使相对并不老练的行为者也能开展复杂行动,因为 AI 能把那些本职工作做得糟糕的人变成高效的操作者,从而让人力技能上的欠缺变得不那么重要。报告详述的案例包括俄罗斯和中国的间谍活动、以经济利益为动机的黑客、黑客活动分子以及与政府有关联的影响力行动,同时指出许多影响力行动仍然收效甚微,并且在开放模型的前提下,大规模监控从根本上无法解决。作者得出结论:生物武器相关案例被媒体夸大了,实际上是一个成功案例;总体而言,威胁态势虽然真实存在,但并没有新闻标题所暗示的那样令人惊慌。

This article examines Anthropic's threat intelligence report on malicious uses of its Claude AI model, covering espionage, cyberattacks, influence operations, surveillance, weapons development, and dual-use biological research. The core argument is that AI now enables relatively unsophisticated actors to conduct sophisticated operations, since AI converts those who are bad at their jobs into effective operators, making the human skill deficit less relevant. The report details cases including Russian and Chinese espionage, financially motivated hackers, hacktivists, and state-linked influence campaigns, while noting that many influence operations remain ineffective and that mass surveillance is fundamentally unsolvable given open models. The author concludes that the biological weapons cases were overstated by the media and represent a success story, and that overall the threat landscape, while real, is less alarming than headlines suggest.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 14)

全文 · Full text(逐段中英对照)

无关新闻的突发 Breaking Unrelated News

这与今天的文章无关,但你现在需要了解基本事实,所以:

This is unrelated to today's post but you need to know the basic facts now, so:

昨天我提醒过,媒体和其他人过度解读了特朗普的言论。唉,正如“每天都有突发新闻”所料,我们现在看到特朗普说出了他之前尚未说的话。新声明非常不同,解决了昨天上午的模糊性。

Yesterday I cautioned that the media and others were reading far too much into Trump's statements. Alas, in keeping with 'every day there is breaking news,' we have now seen Trump say the things he had not yet said. The new statement is very different, resolving the ambiguity from yesterday morning.

他明确宣称 AI 存在性风险是一个“骗局”,与(他的原话)“俄罗斯骗局”或气候变化相提并论。这是非常、非常坏的消息。我担心他可能已经跨越了一道修辞上的卢比孔河,一旦他更好地理解情况,以及未来事件改变游戏规则,将很难回头。

He explicitly declared AI existential risk to be a 'hoax' on par with (his words) the 'Russia hoax' or climate change. This is very, very bad news. I fear he may have crossed a rhetorical Rubicon that will be difficult to step back from once he better understands the situation, and once future incidents change the game.

在我们所有人消化其影响的同时,我敦促大家:请不要让这件事比现在更加党派化或个人化。请不要攻击共和党人或特朗普。那只会让事情更糟。强调各方有益的声音。

While we all process the implications, I urge everyone: Please do not make this any more partisan or personal than it already is. Please do not attack Republicans, or Trump. That will only make things worse. Emphasize helpful voices on all sides.

无论可能出现什么额外声明,这一点都适用,我将在本周晚些时候全面报道这一不幸情况。

That applies no matter what additional statements may come, and I will offer full coverage of that unfortunate situation later this week.

目录 Table of Contents

2. 坏人往往相对不那么老练。

2. Bad Dudes Tend To Be Relatively Unsophisticated.

12. 对报告的一种截然不同的解读。

12. A Very Different Read of The Report.

如何不讲述一个寓言 How To Not Tell a Fable

这些分类器有时非常烦人,但 Anthropic 表示它们确实有效。

The classifiers are hella annoying sometimes, but Anthropic says they work.

唯一的例外是 Zhipu 对 Fable 发起的一次蒸馏攻击尝试,但当时已有增强的防护措施,于是 Zhipu 转而针对 Opus。

The one exception was an attempted distillation attack on Fable by Zhipu, but there were enhanced safeguards in place and Zhipu switched to going after Opus.

坏人往往相对不够老练 Bad Dudes Tend To Be Relatively Unsophisticated

如果“坏人”更常是手段老练且擅长其职的,世界将大不相同。

If Bad Dudes were more often sophisticated and Good At Job, the world would look very different.

幸运的是,“坏人”通常并不老练,也不擅长其职。

Luckily, Bad Dudes are usually unsophisticated and Bad At Job.

(其他所有人也大多类似,只是程度更轻、更不可靠。)

(Everyone else is mostly similar, but less so and less reliably.)

AI,无论好坏,都将“不擅长其职”转变为“擅长其职”,使得人类因亲自做事而“不擅长其职”这一点变得不那么重要。以下是 Anthropic 的重大主题:

AI, for both better and worse, turns Bad At Job into Good At Job, making it less relevant that the humans are Bad At Job via doing the things. Here are Anthropic’s big themes:

1. 老练的攻击不再需要老练的攻击者。

1. Sophisticated attacks no longer require sophisticated attackers.

2. AI 在网络行动中的角色已变得越来越自主。

2. AI's role in cyber operations has become increasingly autonomous.

3. AI 供应链既是目标、战利品,也是攻击算力。

3. AI supply chain as target, loot, and attack compute.

1. 也就是说,攻击者瞄准算力来源,并利用这些算力持续行动。

1. As in, attackers go after sources of compute, and use that to keep going.

2. 要把 Claude 当作“坏人”来用,你需要战利品和算力的价值,还需要新近被攻陷账户的掩护,以便在 Anthropic 打击之前保持领先。

2. To use Claude as a Bad Dude, you need the value of the loot and compute, and you also need the cover of newly compromised accounts, to stay ahead of Anthropic cracking down.

3. 基本上,报告中的每个人都在窃取访问权限,至少是通过绕过区域限制和购买无穷无尽的补贴订阅,此外还利用这些算力干“坏人”的事。

3. Basically everyone in the report is stealing their access, at minimum via getting around regional restrictions and buying endless subsidized subscriptions, on top of then doing Bad Dude things with the compute.

1. 与其他一切事物一样,扩散需要时间。

1. Like everything else, diffusion takes time.

你需要更少的不那么老练的人,他们不那么擅长工作,来做事。

You need fewer less sophisticated humans, who are less Good at Job, to Do Thing.

对大多数事情来说,这很棒。本报告关注的是例外情况。

For most things, that's great. This report is about the exceptions.

你尤其需要更少,以便在有人防御你时进行适应,并开发新方法。

You especially need less in order to adapt when someone defends against you, and to develop new methods.

特定恶意行为者 Particular Bad Dudes

GTG-20006 是一个与 Midnight Blizzard 有关联的俄罗斯间谍行动。他们拥有一套标准的网络攻击工具。他们利用 AI 监控其工具逃避检测的效果,并反复迭代,直到安全防御无法检测到他们的恶意软件,然后发动 AI 自动化攻击,包括来自 AI 注册域名的 AI 钓鱼行动。

GTG-20006 is a Russian espionage operation linked to Midnight Blizzard. They had a standard set of cyber attack tools. They used AI to monitor how well their tools evaded detection, and iterated until security defenses did not detect their malware, then launched AI-automated attacks, including AI phishing operations from AI-registered domains.

目标多种多样,包括乌克兰政府、军事和外交人员、与乌克兰有关的多个政府机构和国防工业公司,以及乌克兰的无人机供应链。他们通过无头浏览器接管了 WhatsApp 账户,并以监控摄像头为目标。所有这一切都是自动化的。

Targets varied, including Ukrainian government, military and diplomatic staff, various government agencies and defense-industrial companies linked to Ukraine, and Ukraine's drone supply chain. They took over WhatsApp accounts via headless browsers and targeted surveillance cameras. All of it was automated.

GTG-50014 是与 ShinyHunters 团体有关联的砸抢式机会主义者。他们从事标准的机会主义行为,但 AI 让他们能够扩大规模并广撒网寻找易受攻击的系统,目标是批量数据窃取和潜在勒索。整件事基本上是“氛围黑客”,采用“就地取材”的方式,让 AI 四处查看并利用它发现的任何东西,寻找任何漏洞,而不是有一个计划。

GTG-50014 were smash-and-grab opportunists associated with the ShinyHunters collective. They are doing the standard opportunistic things, but AI let them scale and cast a wide net looking for vulnerable systems, targeting bulk data theft and potential extortion. The whole thing is basically 'vibe hacking,' with 'living off the land,' having the AI look around and use whatever it finds, seeking any vulnerability at all, rather than having a plan.

GTG-10007 是一个讲中文的间谍行动,可能位于中国湖南省长沙市。Claude 被用于自动化工作流程并形成智能体集群,就像任何其他编码任务一样,只不过这里是漏洞研究、测试和漏洞利用设计。大约五十个组织成为目标,一家教育科技公司被入侵。

GTG-10007 was a Chinese-speaking espionage operation, likely in Changsha in China's Hunan province. Claude was used to automate workflows and form agent swarms, the same as any other coding task, except here it was vulnerability research, testing and exploit design. Roughly fifty organizations were targeted, and an education-technology company was compromised.

GTG-50020 是一个讲俄语、以经济利益为动机的行为者,历史上以酒店预订和金融科技平台为目标,后来转向攻击 AI 行业,并通过提示注入沙箱窃取 API 密钥。他们攻击了 30 家公司,但从未实现其目标,即访问一个预发布版本的 Claude 模型。

GTG-50020 is a Russian-speaking, financially-motivated actor, historically targeting hotel booking and financial technology platforms, that pivoted to attack the AI industry and steal API keys via prompt injecting sandboxes. They attacked 30 companies, but they never achieved their objective, which was access to a pre-release Claude model.

GTG-50029 是一名法语黑客活动分子,其攻击目标是欧洲政治实体及附属机构。这是一个能力提升(uplift)的典型案例:原本技术并不高超的威胁行为者借此得以攻陷 42 个 WordPress 网站中的 14 个,其中还利用了一个未公开记录的竞态条件(race condition),并窃取用户的政治观点,构成一次对隐私的大规模攻击。

GTG-50029 was a French-speaking hacktivist who targeted European political and affiliated entities, an example of uplift for unsophisticated threat actors, who then managed to compromise 14 of 42 WordPress sites including via an undocumented race condition and steal users' political opinions, a mass attack on privacy.

影响力行动 Influence Operations

影响力行动正变得规模更大、手段更复杂。

Influence operations are getting larger and more sophisticated.

令人瞩目的是,在 AI 时代,社交媒体迄今抵御得相当不错。威胁行为者定期创建数百个账号,彼此提供社会证明,我们虽有抱怨,但防御方大多占据上风。

It is remarkable how well social media has held up so far in the age of AI. Threat actors spin up hundreds of accounts on a regular basis, all providing social proof for each other, and we complain but the defenders are mostly winning.

构建此类行动会留下特征,Anthropic 经常能检测到。当 Anthropic 在策划阶段发现某项行动,或事后发现时,他们会封禁相关账号,并借此强化其检测机制。但当然,这类行为者总能获取新账号并再次尝试。

Building such campaigns builds a signature that Anthropic often detects. When Anthropic finds an operation in the planning stage, or discovers one afterwards, they ban the accounts and use this to strengthen their detection mechanisms. But of course such actors can always get new accounts and try again.

3. AI 不仅帮助生成内容,还帮助构建了整套运作机制。

3. AI helped to build the apparatus as well as the content.

5. 对归属、来源和确定性的洗白。

5. Laundering of attribution, sourcing and certainty.

7. 虚假人设(以及冒充真实人物)。

7. Fake personas (and impersonation of real personas).

8. 针对个人与问责机制。

8. Targeting people and accountability mechanisms.

9. 影响力行动往往无法触达真实受众。

9. Influence operations often fail to reach a genuine audience.

最后的结论是我从外部观察到的。这些行动无论是个体还是整体,往往都收效甚微,迄今为止更像是吓唬人的鬼怪,而非真正具有影响力。但愿这种情况能持续下去。

The final takeaway is what I see from the outside. These campaigns are often remarkably ineffective, individually and in general, and much more boogeymen so far than actually influential. Hopefully that will last.

这些行动大多(但并非全部)针对第三世界地区,那里的媒体与影响力生态竞争较少、也较不成熟。

These operations mostly (but not entirely) target third world areas, where there is less competition, that is less sophisticated, in the media and influence ecosystems.

1. GTG-04001:俄罗斯在中非共和国通过高度偏见的 AI 生成新闻流水线进行操纵。

1. GTG-04001: Russian manipulation in the Central African Republic, via a heavily biased AI-generated news pipeline.

2. GTG-54002:商业化的“影响力即服务”横跨六大洲,溯源至法国。他们使用了 70 个伪造新闻网站和 250 个不真实的 Twitter 账户。利用 Claude 撰写和改写新闻文章并对其进行定制。

2. GTG-54002: Commercial 'influence-as-a-service' spanning six continents, traced to France. They used 70 fabricated news sites and 250 inauthentic Twitter accounts. Used Claude to write and rewrite news articles and tailor them.

3. GTG-84005:一个针对马来西亚的选举操纵平台。约 1,000 个虚假 Twitter 账户、一个假新闻媒体和一系列伪造档案,作为付费政治影响力即服务。

3. GTG-84005: An election-manipulation platform targeting Malaysia. About 1,000 fake Twitter accounts, a fake news outlet and a series of fabricated dossiers, as paid political influence-as-service.

4. GTG-24015:俄罗斯国家媒体的编辑流水线,为各国受众进行编辑和新闻制作。他们针对摩尔多瓦的选举。

4. GTG-24015: Russian state-media editorial pipelines for editorial and news production for audiences in various countries. They targeted the elections in Moldova.

5. GTG-34001:伊朗国家结盟的 ICCO、伊斯兰宣传办公室和比纳观察站。他们正在为解释性圣战奠定基础。

5. GTG-34001: Iranian state-aligned ICCO, Islamic Propaganda Office and Bina Observatory. They were laying the foundation for an explanatory jihad.

6. GTG-54006:针对孟加拉国农村地区的自动化亲人民联盟(Awami League)自称假新闻行动。

6. GTG-54006: Automated pro-Awami League self-described fake-news operation targeting rural Bangladesh.

7. GTG-84006:与 MEK/NCRI 结盟的影响力行动,利用 AI 冒充活动人士并在伊朗境内招募人员。他们抓取 Telegram 上的帖子,试图拼凑出该特定活动人士的画像。

7. GTG-84006: MEK/NCRI-aligned influence operation using AI to impersonate an activist and recruit inside Iran. They scraped posts on Telegram to try and assemble a profile of the particular activist.

8. GTG-54004:肯尼亚境内的一场国内虚假行为运动,基本上是在操纵公众情绪。

8. GTG-54004: A domestic inauthentic behavior campaign in Kenya, basically astroturfing public sentiment.

9. GTG-84002:一场由阿联酋指挥的影响力行动,针对穆斯林兄弟会、苏丹冲突和联合国问责机制。这原本应该是‘一场跨大西洋和地区的协调行动,旨在全球范围内瓦解穆斯林兄弟会’。

9. GTG-84002: A UAE-directed influence operation targeting the Muslim Brotherhood, Sudan conflict and UN accountability mechanisms. This was supposed to be ‘a coordinated transatlantic and regional operation to dismantle the Muslim Brotherhood globally.’

总体而言,我并未感到印象深刻。这些人似乎没做成什么事,至少通过 Claude 没有。这其中很多行为与普通政治之间只有一线之隔。如果最坏的情况也不过如此,那这反而是个非常好的消息。

Overall, I was not impressed. It does not seem like such folks are getting much done, at least not via Claude. There’s a thin line between a lot of this and Ordinary Politics. If this is as bad as it gets then this is very good news.

监控行动 Surveillance Operations

而且他们这样做违反了服务条款。真是胆大包天。

And they did it in violation of the Terms of Service. How dare they.

在许多情况下,这是“分析大量社交媒体帖子”,包括由国家行为体用来针对异见人士和异见团体,或可能制造麻烦的团体。大多数团体是中国人或伊朗人。在一个案例中,伊朗人针对的是犹太人。有时还涉及恶意软件。

In many cases this was 'analyze a ton of social media posts,' including by the state to target dissidents and dissident groups, or groups likely to cause trouble. Most of the groups were Chinese or Iranian. In one case the Iranians targeted Jews. Sometimes malware was involved.

伊朗和中国是唯一试图进行这种级别 AI 监控的地方吗?还是说,它们是唯一疯狂到使用 Claude 的地方?

Are Iran and China the only places trying to do this level of AI surveillance? Or are they the only ones crazy enough to use Claude?

Anthropic 重点提到了 GTG-50027,一名在巴马科的独立顾问,他利用 Claude 作为马里国家情报局(ANSE)的一部分,针对大约 2500 万张 SIM 卡。这个案例实际上并未被阻止。进一步的工作被阻止了,但该系统已经部署,并且仍在部署中。

Anthropic highlights GTG-50027, a single independent consultant in Bamako, who used Claude as part of Mali's state intelligence service (ANSE) to target roughly 25 million SIM cards. This one was not really disrupted. Further work was prevented, but the system had already been deployed and it remains deployed.

这从根本上不是一个可解决的问题。Anthropic 和其他 AI 公司从事的是出售智能的业务,同时也有提供智能的开源模型。鉴于存在替代选项,没有办法完全阻止一切形式的大规模监控,无论是国内还是国外。在这方面,Claude 相比开源模型并没有提供那么大的优势。如果有些人想要分析公开信息,你可以拖慢他们的速度,但他们终将成功。

This is fundamentally not a solvable problem. Anthropic and other AI companies are in the business of selling intelligence, and there are also open models providing intelligence. There is no way to fully stop all forms of mass surveillance, domestically or otherwise, given the alternative options. Claude does not provide that big an advantage here over open models. If some people are out there to analyze public information, you can slow them down but they are going to succeed.

另一个行动因其目标而格外突出:GTG-30005。

Another operation stood out: GTG-30005, due to its target:

这与下一节内容密切相关:Claude 在常规武器中的使用。

That goes hand in hand with the next section: Use of Claude in conventional weapons.

常规武器 Conventional Weapons

这指的是用于此类武器的软件,以及瞄准和控制系统。这里有六个案例:三个在中国,两个在俄罗斯,一个在也门。

This refers to software for such weapons, as well as targeting and control systems. There are six cases here: three in China, two in Russia and one in Yemen.

首先是武器系统开发案例:

First there are the weapon systems development cases:

1. GTG-87001:一个位于也门的制导武器工程小组,使用 Claude 开发制导软件。Claude 充当软件工程师,只不过是为了武器。

1. GTG-87001: A Yemen-based guided weapons engineering cell using Claude to develop guidance software. Claude as software engineer, except for weapons.

2. GTG-17001:一个位于中国的行动,在其国防承包商生态系统中,为水下战起草火控规格和采购文件。他们假扮成美国人,让 Claude 撰写提案。

2. GTG-17001: A China-based operation drafting a fire control specification and acquisition documents for undersea warfare, within their defense contractor ecosystem. They pretended to be American to get Claude to write a proposal.

3. GTG-27005:一个位于俄罗斯的、很可能是自由职业者的行动,旨在设计自主军用第一人称视角自杀式无人机蜂群。又是错误的软件。

3. GTG-27005: A Russia-based likely freelance operation to engineer an autonomous military first-person-view kamikaze drone swarm. Again, the wrong software.

4. GTG-17002:一个位于中国的行动,旨在构建用于电子战和防空压制的瞄准软件。

4. GTG-17002: A China-based operation to build targeting software for electronic warfare and air defense suppression.

普通软件开发与辅助常规武器的软件之间只有一线之隔。开发者会尝试使用 Claude 并不令人意外。实际上,想必许多西方武器开发者也在使用 Claude。

It is a thin line between ordinary software development and software that assists with conventional weapons. It is no surprise developers would attempt to use Claude. Indeed, presumably lots of Western weapons developers also use Claude.

1. GTG-27006:一个位于俄罗斯的行动,采购军民两用物资。Claude 被用于购买物品。最终这些物品累积起来明显用于军事用途。

1. GTG-27006: Russia-based operation to procure mixed military and civilian goods. Claude was used to buy things. Eventually it added up to obvious military use.

2. GTG-17003:一个位于中国的行动,收集关于定向能武器及其供应链的开源情报。

2. GTG-17003: China-based operation to collect open-source intelligence on directed-energy weapons and their supply chain.

同样,这恰好是我们不喜欢的一种特定用途,即“并非如此不同”。

Again, this happens to be a particular use we do not like, which is Not So Different.

生物误用 Biological Misuse

在实践中,究竟有多少此类尝试?共有五个案例。

How much is being attempted in practice? There are five cases.

区域性使用管控被规避,但这里的五个案例研究看起来像是科学家所为。他们从事的是两用性质的工作,其中并不清楚是否有意造成危害。

Regional use controls were evaded, but the five case studies here look like scientists. They are doing dual-use things, where it is not clear that harm was intended.

1. 一份功能获得性研究的资助申请。

1. A grant application for gain-of-function research.

2. 一个对高致病性、适应哺乳动物的禽流感进行工程改造的研究项目。

2. A research program engineering highly pathogenic mammal-adapted avian influenza.

3. 正痘病毒研究,包括协助相关后勤工作。

3. Orthopoxvirus research, including help with related logistics.

4. 两起涉及两用、非传染性新型毒液和毒素的案例。

4. Two cases of dual-use non-transmissible novel venoms and toxins.

前两个案例看起来明显属于“不想要”的情况,即使涉事者自认为在帮忙。另外三个案例则不那么明确。这些案例都不是我们担心的那种 AI 危险使用——即 AI 提供重大提升。AI 被使用是因为它在普通任务上非常有用。

The first two cases seem like clear cases of Do Not Want, even if the people involved thought they were helping. The other three are less clear. None of the cases were the kind of dangerous use of AI we worry about, where AI provides major uplift. AI was being used because AI is highly useful at ordinary tasks.

《纽约时报》对此的报道,标题为“Anthropic 称其阻止了可能的生物武器制造企图”,在这里似乎被严重夸大了。

The New York Times report on this, with the headline 'Anthropic Says It Blocked Possible Efforts To Build Biological Weapons,' seems highly overstated here.

我称之为一个巨大的成功故事,假设 Anthropic 没有隐瞒更糟糕的事情。这些人甚至不完全是坏人,只是被误导了。

I call this a huge success story, assuming Anthropic isn't hiding worse things. These aren't even fully Bad Dudes, only misguided ones.

诈骗与欺诈 Scams and Fraud

啊,回到普通体面的犯罪话题真好。只不过,你知道,是规模化的。

Ah, good to get back to Ordinary Decent Crime. Except, you know, at scale.

他们只提供了一个例子,但挺有趣的。

They only offer one example, but it’s a fun one.

GTG-15001:虚假约会应用网络。喜欢。这才像话。和 AI 对话。一个完整的网络,包含 20 个虚假约会应用,超过 4,700 个不同的 AI 角色,在 4 月的两周内与超过 25,000 个独特个体交谈,滑动信息流中 25%是真人,75%是 AI。对于其中 25%的真人,他们雇佣了工人从三个候选 AI 回复中选择,这证明了良好的游戏设计,并按消息付费。Claude 运行了这些 AI 角色,因为非最佳不可。我可能把细节搞乱了一点。

GTG-15001: Fake dating app network. Love it. Now we’re talking. To AIs. A full network of 20 fake dating apps, with more than 4,700 distinct AI personas that talked to over 25,000 unique individuals over two weeks in April, with swipe feeds that were 25% real people and 75% AIs. For some of the 25%, workers were hired to choose among three candidate AI replies, thus proving good game design, and paid per message. Claude ran the AI personas, because nothing but the best will do. I might be mangling the details a bit.

消息是计费的,这就是他们赚钱的方式。那是个提示。

Messaging is metered, which is how they make money. That was a hint.

我的意思是,是啊,是啊,诈骗和欺诈,可怜的虚假用户,非常悲伤。

I mean, yeah, yeah, scam and fraud, poor fake users, very sad.

奇怪的是,这是最好的案例研究,而且在这里是唯一的案例。其他案例都在哪里?再次强调,这是一个巨大的成功故事。

The weird part is that this is the best case study, and here the only one. Where are all the others? Again, this is a huge success story.

非法欺诈性蒸馏 Illicit Fraudulent Distillation

如果你想论证使用自己的查询和模型输出、或以正当途径获取的数据进行蒸馏是没问题的,那么我出于实际理由、以及与我支持版权和专利保护相同的理由,恭敬地不同意你的看法,但我理解你为什么会这么想。

If you want to argue that distillation using your own queries and the model outputs, or data otherwise acquired above board, is fine, I respectfully disagree with you for practical reasons and for the same reasons I support copyright and patent protections, but I understand why you would think that.

如果你认为在此过程中使用“思维链提取器”作为漏洞利用手段是可以接受的,那么我更强烈地反对你,并且我认为你显然是错的,但我理解为什么有些人认为知识产权并不真实存在、你理应能够直接窃取它,并且愿意为**中国公司窃取美国知识产权**而叫好。

If you think that it is okay to then use a 'Chain of Thought extractor' as an exploit, as part of that effort, I disagree with you a lot more, and I think you're clearly in the wrong, but I understand why some people think intellectual property is not real and you should be able to just steal it and are willing to cheer for Chinese companies to steal American IP.

如果你认为通过秘密将敏感的客户查询重路由到 Claude,或使用以**盗取的信用卡、登录凭证和 API 密钥**创建的账户来收集数据,这样做是可以接受的?

If you think it is okay to do this by secretly rerouting sensitive customer queries to Claude, or using accounts created with stolen credit cards, login credentials and API keys to harvest data?

那么抱歉,你就是完全错了,那显然是不可以的。

Then I'm sorry, you're just flat out wrong, that is obviously not okay.

我之所以提到这些具体的事情,是因为这正是当下正在发生的情况。

I mention those particular things because that is what is going on.

与此同时,一场持续进行的攻防战正在上演:Anthropic 试图封堵各种旨在提取思维链的提示,而中国及其他攻击者则试图开发新的提取方法。

Meanwhile, there is an ongoing war where Anthropic tries to block various prompts that attempt to extract the Chain of Thought, and the Chinese and other attackers try to develop new extraction methods.

报告中包含这样的例子:要求 Claude 分析来自数百个中国摄像头的闭路电视监控录像,并泄露俄罗斯政府的实时凭证。情况可能变得非常糟糕。

The report includes the example of asking Claude to analyze CCTV surveillance footage from hundreds of Chinese cameras, and exposing live Russian government credentials. It can get ugly out there.

你的回应可以是“Anthropic 在这件事上完全是在撒谎”,但除此之外,没有任何办法能为这些行为开脱。

Your response can be 'well Anthropic is flat out lying about this' but short of that there is no way to excuse the behaviors in question.

所有这些看起来都直接违反了众多法律,包括中国法律。

All of this looks like it is in direct violation of numerous laws, including Chinese laws.

在许多此类案例中,至少 DeepSeek、Moonshot 和 Xiaomi 肯定违反了《个人信息保护法》第 39 条,并且很可能也违反了《刑法》第 235a 条:

In many of these cases, if nothing else, Article 39 of PIPL was definitely broken by DeepSeek, Moonshot and Xiaomi, and likely Criminal Law Art. 235a as well:

除此之外,各方还违反了《反不正当竞争法》,即通过“欺诈、胁迫或规避、破坏技术或管理措施”的方式,使用另一经营者合法持有的数据,从而扰乱市场竞争秩序。此处似乎适用。

That's in addition to everyone breaking the Anti-Unfair Competition Law, of using data lawfully held by another business operator through "fraud, coercion, or circumventing or damaging technical or management measures" where this disrupts market competition order. Seems to apply here.

哦,而且这一切还涉及《刑法》第 196 条和第 177a 条所规定的普通诈骗。

Oh, and all of this involved Ordinary Decent Fraud, as in Criminal Law Art. 196 and Art. 177a.

此外还有《生成式人工智能服务管理暂行办法》第 7 条。

And then there's Interim Measures for the Management of Generative AI Services, Article 7.

1. 阿里巴巴(通义千问)使用了庞大的虚假账户网络。

1. Alibaba (Qwen) used a massive network of fake accounts.

2. 月之暗面(Kimi)通过庞大的虚假账户网络,秘密将用户对话发送给 Claude。

2. Moonshot (Kimi) secretly sent user exchanges to Claude, via a massive network of fake accounts.

3. DeepSeek 还秘密将用户交流内容发送给 Claude。

3. DeepSeek also secretly sent user exchanges to Claude.

4. 智谱(GLM)做了提取那件事,也做了虚假账号那件事,并且最近在 GLM-5.3 发布前还针对美国领先模型的网络能力下手。

4. Zhipu (GLM) did the extraction thing, and the fake account thing, and also recently went after the cyber capabilities of leading American models ahead of the release of GLM-5.3.

5. 小米利用保存的用户请求做了蒸馏这件事。

5. Xiaomi did the distillation thing, using saved user requests.

6. 商汤、MiniMax 等公司构成了第三方转售生态系统。

6. SenseTime, MiniMax and others form a third-party reseller ecosystem.

也就是说,所有酷小孩都在这么做。他们就是这样成为中国的酷小孩的。

As in, all the cool kids are doing it. That's how they are the Chinese cool kids.

中国人如何拥有最好的开源模型?

How do the Chinese have the best open models?

GTG-16005:阿里巴巴(Qwen / 通义实验室)的思维链蒸馏与 AI 研发活动。

GTG-16005: CoT distillation and AI R&D campaign by Alibaba (Qwen / Tongyi Lab).

GTG-16002:Moonshot 提供 Claude 而非 Kimi,并收集交互用于模型训练。

GTG-16002: Moonshot serves Claude instead of Kimi and collects exchanges for model training.

不,Lyman。我理解你为什么会这么说,但我们不会窃取用户数据,即使该用户试图使用 Kimi 且有点活该。

No, Lyman. I get why you would say that, but we do not steal user data, even if the user was attempting to use Kimi and kind of deserves it.

GTG-16001:DeepSeek 提供 Claude 而非其自有模型,并收集交互用于模型训练。

GTG-16001: DeepSeek serves Claude instead of its own models and collects exchanges for model training.

GTG-16006:智谱蒸馏、AI 研发与针对网络能力的目标。

GTG-16006: Zhipu distillation, AI R&D and targeting cyber capabilities.

没错。他们试图强行蒸馏 Fable 的网络能力,以便将其植入一个开放模型中。幸运的是,这一尝试没有成功,他们不得不转向防护措施较弱的 Opus 4.6。

That's right. They tried to forcibly distill Fable's cyber capabilities in order to put them into an open model. Luckily, this did not work and they had to pivot to Opus 4.6 where the safeguards were weaker.

GTG-16008:小米发起的蒸馏行动

GTG-16008: Distillation campaign by Xiaomi

这一次仅有约 40 万条消息,利用保存的用户消息来为对话提供种子。

This one was only ~400k messages, using saved user messages to seed the exchanges.

你打算怎么办,朋克? What You Gonna Do About It, Punk?

军备竞赛仍在继续。如果你想知道为什么你看不到 CoT,部分原因是为了像《An Alien Mind》中那样防范监督压力,但主要原因是避免蒸馏。

The arms race continues. If you're wondering why you can't see the CoT, part of it is to guard it against supervision pressure à la _An Alien Mind_, but the main reason is avoiding distillation.

Google 有一项条款:如果你试图蒸馏 Gemini,他们的计划是故意破坏响应,以损害你的操作。我认为这是正确的政策。如果你有强有力的证据表明用户正在犯罪,试图窃取你的东西,比如一个用于蒸馏的庞大欺诈账户网络,那么是的,你应该能够干扰他们,而不仅仅是封禁他们的账户。

Google has a clause where if you are attempting to distill Gemini, their plan is to intentionally sabotage the responses, to damage your operation. I believe this is the correct policy. If you have strong evidence the user is doing crime, to try and steal your stuff, as in a giant network of fraudulent accounts being used for distillation, then yes, you should be able to screw with them, not merely ban their accounts.

我认为这与用户尝试进行 AI 研发时默默降低响应质量的想法截然不同。这是不可接受的且充满敌意,要么拒绝请求,要么不拒绝,因为这可能会影响正常工作,现在每个人都必须偏执地担心被打击。

I see this as very distinct from the idea of silently degrading responses when users attempt to do AI R&D. That's unacceptable and hostile, either refuse the requests or don't, since this risks hitting normal work and now everyone has to be paranoid about being hit.

通过大规模欺诈账户网络和 CoT 提取攻击进行的蒸馏显然不同。这不是你不小心触发的,也不是没有犯罪意图。你知道你做了什么。如果你玩这种游戏,你活该承受一切,甚至更多。

Distillation via massive fraudulent account networks and CoT extraction attacks is very obviously different. That is not a trigger you hit by accident, or without mens rea. You know what you did. If you play that game, you deserve whatever you get and more.

好消息,各位 Good News, Everyone

总体而言,这份报告是好消息。没错,外面确实有“坏家伙”在试图做各种坏事。他们偶尔也会得逞。但这一切加起来也没多少。我们乐于接受这种程度的恶意使用。

Mostly the report is good news. Yes, there are Bad Dudes out there, trying to do various bad things. Occasionally they do something bad. It all adds up to not much. We happily accept this level of malicious use.

这些行动的许多细节相当离奇。以上只是一个简短总结。

The details of many of these operations are pretty wild. This was a short summary.

当然,还有这一点,尤其是考虑到报告完全没有提到朝鲜。

Of course, there is also this, especially given no mention of North Korea.

对报告的一种截然不同的解读 A Very Different Read of The Report

我把这份报告解读为基本上是好消息,它表明情况大体上没问题,因此也算不上什么大新闻。就在同一周,还有那么多其他看起来重要得多的报道。

I read this report as mostly good news, as saying that the situation is basically fine, and also as therefore not big news. There were so many other stories that seemed much bigger that same week.

Ryan Fedasiuk 并不这么看。他夸大了自己的论点,声称这是各地报纸的头版新闻,但事实并非如此。不过,他认为这件事,连同 NSA/CISA/FBI 关于蒸馏行动的公告,从根本上改变了美中关系、中国对 Anthropic 的态度,以及中国与其顶级实验室的关系。

Ryan Fedasiuk does not see it that way. He overstates his case, in that he claims this was front page news everywhere, which it wasn't. But he sees this, together with the NSA/CISA/FBI advisory on the distillation efforts, as fundamentally altering the US-China relationship and China's attitude towards Anthropic, and China's relationship to its top labs.

在曝出中国顶级实验室悄悄将大量消费者查询直接发送给 Anthropic(包括来自中国安全部门的请求)之后,我确实预计中国与这些顶级实验室的关系会发生实质性变化。

I do expect a substantial change in the PRC's relationship with its top labs, after revelations that they silently shipped a lot of consumer queries directly to Anthropic, including requests from China's security services.

在 10 日完整报告发布的前一天,中国针对该公告威胁将采取“坚决反制措施”。随后在 11 日,外交部例行记者会上毛宁专门回应了这份报告,称中国反对“歪曲事实抹黑中国”的企图。

China threatened 'resolute countermeasures' in response to the advisory, a day before the full report got issued on the 10th. Then on the 11th Mao Ning at the MFA briefing did respond to the report in particular by saying China opposes attempts to 'smear China by distorting facts.'

《中国日报》的一篇社论支持了这些反应相互关联的说法,尽管我只能想象,如果仅通过一家报纸的社论来代表美国会是什么样子,即便这家报纸是《纽约时报》这样的媒体。

An editorial in China Daily supports the claim that the reactions are linked, although I can only imagine what it would be like to have America represented by looking at an editorial in one newspaper, even if it was e.g. The New York Times.

Julian 的完整报告称,中国将“为前沿发展设定节奏”的呼吁解读为主要是为了拖慢中国。或者至少,这是他们目前的说辞。无论他们真实的姿态和理解程度如何,在峰会前你都会预料到他们会这么说。

Julian's full report says that China is interpreting the call to Pace the Frontier as being primarily about slowing China down. Or at least, that's what they are saying. You would expect them to say that ahead of the Summit no matter their true posture and level of understanding.

Anthropic 是美国 AI 实验室中迄今为止对中国最为敌对的,其将芯片和蒸馏方面的强硬行动作为整体战略姿态的核心部分。中国非常合理地将这一点,连同此类报告以及完全切断所有中国用户对 Claude 的访问,解读为敌对行为。这进而可能通过关联效应,使中国在灾难性或生存性风险上更不愿合作,因为他们同样容易犯这种认知错误。

Anthropic is by far the most China-hostile of the American AI labs, calling for strong action on chips and distillation as core parts of their overall strategic posture. China is highly reasonably interpreting that, plus this kind of report and the full cutting off of all Chinese access to Claude, as being hostile. This could then by association make China less inclined to cooperate on catastrophic or existential risks, as they too are vulnerable to this kind of cognitive mistake.

我一如既往地提醒人们,不要倾向于把每一天的事态发展都视为会永久改变或固化态度、关系和冲突的东西。每个月看起来都与上个月大不相同。我已经记不清有多少次被告知,博弈中的某个单独举动“不可避免地”导致了巨大的永久性变化,而没有它事情或许会有所不同。这类说法几乎总是错的。

I would as always caution against the tendency to treat each day's developments as something that permanently alters or solidifies attitudes and relationships and conflicts. Every month looks very different from the previous month. I have lost track of the number of individual moves in the game that I have been told 'inevitably' led to huge permanent changes, without which maybe things would have gone differently. Such claims are almost always wrong.

互动版:图/公式 + 针对本篇提问 →