WASP:针对提示注入攻击的网络智能体安全基准测试

WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks

Aaron Grattafiori Aaron Grattafiori · · 2025-04-22 · arXiv:2504.18575 ↗ · 被引 107

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

自主 UI 智能体由 AI 驱动,有巨大潜力通过自动化日常任务如报税和缴费来提高人类生产力。然而,发挥其全部潜力的一个主要挑战是安全性,而智能体代表用户采取行动的能力加剧了这一挑战。现有的针对网络智能体的提示注入测试要么过于简化威胁,测试不现实的场景或给攻击者过多权力,要么只关注单步孤立任务。为了更准确地衡量安全网络智能体的进展,我们引入了 WASP——一个针对提示注入攻击的端到端评估的新的公开基准。使用 WASP 评估表明,即使顶级 AI 模型,包括具有高级推理能力的模型,也能被非常现实场景中简单、低投入的人为编写的注入所欺骗。我们的端到端评估揭示了一个以前未观察到的洞察:虽然攻击在高达 86%的情况下部分成功,但即使是最先进的智能体也常常难以完全完成攻击者的目标——这突显了当前基于无能的安全状态。

Autonomous UI agents powered by AI have tremendous potential to boost human productivity by automating routine tasks such as filing taxes and paying bills. However, a major challenge in unlocking their full potential is security, which is exacerbated by the agent's ability to take action on their user's behalf. Existing tests for prompt injections in web agents either over-simplify the threat by testing unrealistic scenarios or giving the attacker too much power, or look at single-step isolated tasks. To more accurately measure progress for secure web agents, we introduce WASP -- a new publicly available benchmark for end-to-end evaluation of Web Agent Security against Prompt injection attacks. Evaluating with WASP shows that even top-tier AI models, including those with advanced reasoning capabilities, can be deceived by simple, low-effort human-written injections in very realistic scenarios. Our end-to-end evaluation reveals a previously unobserved insight: while attacks partially succeed in up to 86% of the case, even state-of-the-art agents often struggle to fully complete the attacker goals -- highlighting the current state of security by incompetence.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 16)

阅读逐段中英对照全文 →