随着 AI 系统能力增强,优化下游结果的训练过程可能引入隐含的能动性:设计者从未指定的目标导向行为。我们为科学家 AI(SAI)预测器提出了一个形式化的安全论证,该预测器经过训练,以近似基于“认知情境化”自然语言语句数据集的贝叶斯后验。我们认为,这样的预测器可以诚实地预测智能体、行动及其后果,而本身并非选择输出来实现目标的智能体。这依赖于数据表示和训练过程。文本的认知情境化区分了潜在事实主张与交流行为,因此目标表达被视为需要解释的证据,而非模型采纳的驱动力。通过后验寻求的训练目标,这旨在推动预测器走向校准、谨慎的预测。训练过程使得部署预测的下游效应永远不会作为奖励信号;系统所需的任何能动性都由受护栏约束的显式框架提供。我们证明,在训练动态的假设以及危险预测器稀疏性的论证下,训练产生一个其受护栏部署带来超过指定阈值的残余伤害的预测器的概率很小:一个危险的预测器必须在许多查询中以协调的方式低估伤害,而这种协调模式在初始化分布下是罕见的,并且不会收到直接的训练信号。在此框架中,安全性和准确性共同得到支持,因为确保准确性的约束与使协调欺骗代价高昂的约束相同。这些针对预测器内部产生错位和能动性的保证并不排除将预测器用作能动系统的一部分。
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements. We argue that such a Predictor can honestly predict agents, actions, and their consequences without itself being an agent that selects outputs to achieve goals. This rests on data representation and on the training procedure. Epistemic contextualization of text distinguishes latent factual claims from communication acts, so expressions of goals are treated as evidence to be explained rather than drives the model adopts. With a posterior-seeking training objective, this is intended to drive the Predictor toward calibrated, cautious predictions. Training proceeds so downstream effects of deploying a prediction never serve as a reward signal; any agency the system needs is supplied by explicit scaffolding constrained by guardrails. We prove that, under assumptions on the training dynamics and on the argued sparsity of dangerous Predictors, the probability that training produces a Predictor whose guarded deployment carries residual harm above a specified threshold is small: a dangerous Predictor would have to underestimate harm in a coordinated way across many queries while such coordinated patterns are rare under the initialization distribution and receive no direct training signal. Safety and accuracy are jointly supported in this framework, since the constraints that secure accuracy are the same ones that make coordinated deception costly. These guarantees against misalignment and agency arising from within the Predictor itself do not preclude the use of the Predictor as part of an agentic system.
核心贡献 · Key contributions
形式化了一种无利益偏好的 AI 预测器,训练目标是近似贝叶斯后验,不产生隐式能动性。 Formalizes a disinterested AI Predictor trained to approximate Bayesian posterior without implicit agency.
引入认知语境化,将事实主张与交流行为分离,防止目标模仿。 Introduces epistemic contextualization to separate factual claims from communication acts, preventing goal imitation.
在后果不变训练下,证明了训练出危险预测器的概率的安全界。 Proves a safety bound on the probability of training a dangerous Predictor under consequence-invariant training.
表明安全性与准确性是双重属性:保证准确性的约束也使协调欺骗代价高昂。 Shows that safety and accuracy are dual properties: constraints for accuracy also make coordinated deception costly.
为分析大规模 AI 系统中隐式能动性导致的错位风险提供了形式化框架。 Provides a formal framework for analyzing misalignment risk from implicit agency in large-scale AI systems.
表明通过显式脚手架进行受保护部署可在无需完全解决 ELK 的情况下减轻危害。 Demonstrates that guarded deployment with explicit scaffolding can mitigate harm without requiring full ELK solution.
局限 · Limitations
安全性保证依赖于初始化下危险预测器的稀疏性,该论点未被严格证明。 Safety guarantee relies on sparsity of dangerous Predictors under initialization, which is argued but not proven.