不具目的性的 AI 预测器中的诚实带来安全

Safety from Honesty in a Disinterested AI Predictor

约书亚·本吉奥 Yoshua Bengio · Mila · 2026-06-28 · arXiv:2606.29657 ↗ · 被引 1

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

随着 AI 系统能力增强,优化下游结果的训练过程可能引入隐含的能动性:设计者从未指定的目标导向行为。我们为科学家 AI(SAI)预测器提出了一个形式化的安全论证,该预测器经过训练,以近似基于“认知情境化”自然语言语句数据集的贝叶斯后验。我们认为,这样的预测器可以诚实地预测智能体、行动及其后果,而本身并非选择输出来实现目标的智能体。这依赖于数据表示和训练过程。文本的认知情境化区分了潜在事实主张与交流行为,因此目标表达被视为需要解释的证据,而非模型采纳的驱动力。通过后验寻求的训练目标,这旨在推动预测器走向校准、谨慎的预测。训练过程使得部署预测的下游效应永远不会作为奖励信号;系统所需的任何能动性都由受护栏约束的显式框架提供。我们证明,在训练动态的假设以及危险预测器稀疏性的论证下,训练产生一个其受护栏部署带来超过指定阈值的残余伤害的预测器的概率很小:一个危险的预测器必须在许多查询中以协调的方式低估伤害,而这种协调模式在初始化分布下是罕见的,并且不会收到直接的训练信号。在此框架中,安全性和准确性共同得到支持,因为确保准确性的约束与使协调欺骗代价高昂的约束相同。这些针对预测器内部产生错位和能动性的保证并不排除将预测器用作能动系统的一部分。

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements. We argue that such a Predictor can honestly predict agents, actions, and their consequences without itself being an agent that selects outputs to achieve goals. This rests on data representation and on the training procedure. Epistemic contextualization of text distinguishes latent factual claims from communication acts, so expressions of goals are treated as evidence to be explained rather than drives the model adopts. With a posterior-seeking training objective, this is intended to drive the Predictor toward calibrated, cautious predictions. Training proceeds so downstream effects of deploying a prediction never serve as a reward signal; any agency the system needs is supplied by explicit scaffolding constrained by guardrails. We prove that, under assumptions on the training dynamics and on the argued sparsity of dangerous Predictors, the probability that training produces a Predictor whose guarded deployment carries residual harm above a specified threshold is small: a dangerous Predictor would have to underestimate harm in a coordinated way across many queries while such coordinated patterns are rare under the initialization distribution and receive no direct training signal. Safety and accuracy are jointly supported in this framework, since the constraints that secure accuracy are the same ones that make coordinated deception costly. These guarantees against misalignment and agency arising from within the Predictor itself do not preclude the use of the Predictor as part of an agentic system.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 24)

阅读逐段中英对照全文 →