Aria-UI:面向 GUI 指令的视觉定位

Aria-UI: Visual Grounding for GUI Instructions

李俊男 Junnan Li · Rhymes AI / University of Hong Kong · 2024-12-20 · arXiv:2412.16256 ↗ · 被引 133

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

跨平台直接操作 GUI 的数字代理日益重要。对于这些代理,从语言指令到目标元素的定位仍是一大挑战,因为现有方法依赖 HTML 或 AXTree 输入。本文提出 Aria-UI,一种专为 GUI 定位设计的大型多模态模型。Aria-UI 采用纯视觉方法,摒弃辅助输入依赖。为适应异构规划指令,我们提出可扩展数据流水线,合成多样且高质量的定位指令样本。为处理任务执行中的动态上下文,Aria-UI 整合文本及图文交错的行动历史,实现鲁棒的上下文感知定位。Aria-UI 在离线和在线代理基准上均取得最先进结果,优于纯视觉和依赖 AXTree 的基线。我们发布所有训练数据和模型检查点,以促进进一步研究,详见 https://ariaui.github.io。

Digital agents for automating tasks across different platforms by directly manipulating the GUIs are increasingly important. For these agents, grounding from language instructions to target elements remains a significant challenge due to reliance on HTML or AXTree inputs. In this paper, we introduce Aria-UI, a large multimodal model specifically designed for GUI grounding. Aria-UI adopts a pure-vision approach, eschewing reliance on auxiliary inputs. To adapt to heterogeneous planning instructions, we propose a scalable data pipeline that synthesizes diverse and high-quality instruction samples for grounding. To handle dynamic contexts in task performing, Aria-UI incorporates textual and text-image interleaved action histories, enabling robust context-aware reasoning for grounding. Aria-UI sets new state-of-the-art results across offline and online agent benchmarks, outperforming both vision-only and AXTree-reliant baselines. We release all training data and model checkpoints to foster further research at https://ariaui.github.io.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 14)

阅读逐段中英对照全文 →