Computer-Using Agent
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→Operator 由计算机使用代理(CUA)驱动,这是一种结合 GPT-4o 视觉能力与强化学习高级推理的模型。CUA 经过训练,能够像人类一样与图形用户界面(GUI)交互,从而无需操作系统或网络专用 API 即可灵活执行数字任务。它建立在多模态理解与推理交叉领域多年基础研究之上,通过结合高级 GUI 感知与结构化问题解决,能够将任务分解为多步计划,并在遇到挑战时自适应地自我纠正。这一能力标志着人工智能发展的下一步,使模型能够使用人类日常依赖的相同工具,为广泛的新应用打开大门。
Powering Operator with Computer-Using Agent, a universal interface for AI to interact with the digital world. Today we introduced a research preview of Operator(opens in a new window), an agent that can go to the web to perform tasks for you. Powering Operator is Computer-Using Agent (CUA), a model that combines GPT‑4o's vision capabilities with advanced reasoning through reinforcement learning. CUA is trained to interact with graphical user interfaces (GUIs)—the buttons, menus, and text fields people see on a screen—just as humans do. This gives it the flexibility to perform digital tasks without using OS-or web-specific APIs. CUA builds off of years of foundational research at the intersection of multimodal understanding and reasoning. By combining advanced GUI perception with structured problem-solving, it can break tasks into multi-step plans and adaptively self-correct when challenges arise. This capability marks the next step in AI development, allowing models to use the same tools humans rely on daily and opening the door to a vast range of new applications.
Operator 由计算机使用智能体驱动,为 AI 与数字世界交互提供通用接口。
Powering Operator with Computer-Using Agent, a universal interface for AI to interact with the digital world.
今天我们发布了 Operator 的研究预览版,这是一个可以替您上网执行任务的智能体。Operator 由计算机使用智能体(CUA)驱动,该模型结合了 GPT‑4o 的视觉能力与通过强化学习实现的先进推理。CUA 经过训练,能够像人类一样与图形用户界面(GUI)——人们在屏幕上看到的按钮、菜单和文本字段——进行交互。这使其具有灵活性,无需使用操作系统或网络特定的 API 即可执行数字任务。
Today we introduced a research preview of Operator(opens in a new window), an agent that can go to the web to perform tasks for you. Powering Operator is Computer-Using Agent (CUA), a model that combines GPT‑4o's vision capabilities with advanced reasoning through reinforcement learning. CUA is trained to interact with graphical user interfaces (GUIs)—the buttons, menus, and text fields people see on a screen—just as humans do. This gives it the flexibility to perform digital tasks without using OS-or web-specific APIs.
CUA 建立在多模态理解与推理交叉领域多年基础研究之上。通过将先进的 GUI 感知与结构化问题解决相结合,它能够将任务分解为多步计划,并在遇到挑战时自适应地自我纠正。这一能力标志着 AI 发展的下一步,使模型能够使用人类日常依赖的相同工具,并为广泛的新应用打开了大门。
CUA builds off of years of foundational research at the intersection of multimodal understanding and reasoning. By combining advanced GUI perception with structured problem-solving, it can break tasks into multi-step plans and adaptively self-correct when challenges arise. This capability marks the next step in AI development, allowing models to use the same tools humans rely on daily and opening the door to a vast range of new applications.
尽管 CUA 仍处于早期阶段且存在局限性,但它创下了新的最先进基准结果:在 OSWorld 上完成完整计算机使用任务的成功率达到 38.1%,在 WebArena 上为 58.1%,在 WebVoyager 上为 87%。这些结果凸显了 CUA 使用单一通用动作空间在不同环境中导航和操作的能力。
While CUA is still early and has limitations, it sets new state-of-the-art benchmark results, achieving a 38.1% success rate on OSWorld for full computer use tasks, and 58.1% on WebArena and 87% on WebVoyager for web-based tasks. These results highlight CUA’s ability to navigate and operate across diverse environments using a single general action space.
我们以安全为首要优先级开发了 CUA,以应对智能体访问数字世界所带来的挑战,详见我们的 Operator 系统卡。根据我们的迭代部署策略,我们通过 operator.chatgpt.com 上的 Operator 研究预览版向美国 Pro 层级用户发布 CUA。通过收集真实世界反馈,我们可以完善安全措施并持续改进,为未来数字智能体日益普及做好准备。
We’ve developed CUA with safety as a top priority to address the challenges posed by an agent having access to the digital world, as detailed in our Operator System Card. In line with our iterative deployment strategy, we are releasing CUA through a research preview of Operator at operator.chatgpt.com(opens in a new window) for Pro(opens in a new window) Tier users in the U.S. to start. By gathering real-world feedback, we can refine safety measures and continuously improve as we prepare for a future with increasing use of digital agents.
CUA 处理原始像素数据以理解屏幕上发生的情况,并使用虚拟鼠标和键盘完成操作。它可以导航多步骤任务、处理错误并适应意外变化。这使得 CUA 能够在广泛的数字环境中行动,执行诸如填写表单和浏览网站等任务,而无需专门的 API。
CUA processes raw pixel data to understand what’s happening on the screen and uses a virtual mouse and keyboard to complete actions. It can navigate multi-step tasks, handle errors, and adapt to unexpected changes. This enables CUA to act in a wide range of digital environments, performing tasks like filling out forms and navigating websites without needing specialized APIs.
根据用户的指令,CUA 通过一个集成了感知、推理和行动的迭代循环来运作:
Given a user’s instruction, CUA operates through an iterative loop that integrates perception, reasoning, and action:
* 感知:计算机的屏幕截图被添加到模型的上下文中,提供计算机当前状态的视觉快照。
* Perception: Screenshots from the computer are added to the model’s context, providing a visual snapshot of the computer's current state.
* 推理:CUA 使用思维链推理下一步,考虑当前和过去的屏幕截图及操作。这种内心独白通过使模型能够评估其观察、跟踪中间步骤并动态适应来提高任务性能。
* Reasoning: CUA reasons through the next steps using chain-of-thought, taking into consideration current and past screenshots and actions. This inner monologue improves task performance by enabling the model to evaluate its observations, track intermediate steps, and adapt dynamically.
* 行动:它执行操作——点击、滚动或打字——直到决定任务完成或需要用户输入。虽然它自动处理大多数步骤,但 CUA 会就敏感操作(如输入登录详细信息或响应 CAPTCHA 表单)寻求用户确认。
* Action: It performs the actions—clicking, scrolling, or typing—until it decides that the task is completed or user input is needed. While it handles most steps automatically, CUA seeks user confirmation for sensitive actions, such as entering login details or responding to CAPTCHA forms.
CUA 通过使用屏幕、鼠标和键盘这一通用接口,在计算机使用和浏览器使用基准测试中均确立了新的最先进水平。
CUA establishes a new state-of-the-art in both computer use and browser use benchmarks by using the same universal interface of screen, mouse, and keyboard.
WebArena(在新窗口中打开)和 WebVoyager(在新窗口中打开)旨在评估网络浏览智能体使用浏览器完成真实世界任务的性能。WebArena 利用自托管的离线开源网站模拟电子商务、在线商店内容管理(CMS)、社交论坛平台等真实场景。WebVoyager 则在亚马逊、GitHub 和谷歌地图等在线实时网站上测试模型的性能。
WebArena(opens in a new window) and WebVoyager(opens in a new window)are designed to evaluate the performance of web browsing agents in completing real-world tasks using browsers. WebArena utilizes self-hosted open-source websites offline to imitate real-world scenarios in e-commerce, online store content management (CMS), social forum platforms, and more. WebVoyager tests the model’s performance on online live websites like Amazon, GitHub, and Google Maps.
在这些基准测试中,CUA 使用相同的通用界面(将浏览器屏幕感知为像素,并通过鼠标和键盘执行操作)树立了新标准。CUA 在 WebArena 上实现了 58.1%的成功率,在 WebVoyager 上实现了 87%的成功率。虽然 CUA 在 WebVoyager 上取得了较高的成功率(其中大多数任务相对简单),但在 WebArena 等更复杂的基准测试中,CUA 仍需进一步改进以缩小与人类表现的差距。
In these benchmarks, CUA sets a new standard using the same universal interface that perceives the browser screen as pixels and takes action through mouse and keyboard. CUA achieved a 58.1% success rate on WebArena and an 87% success rate on WebVoyager for web-based tasks. While CUA achieves a high success rate on WebVoyager, where most tasks are relatively simple, CUA still needs more improvements to close the gap with human performance on more complex benchmarks like WebArena.
语法测验 总退款 多联骨牌 更新许可证 查找客户
Grammar quiz Total refund Polyominoes Update license Find customer
001 用户提示 002 初始化计算机 003 新截图 004 访问剑桥词典 Plus 部分 005 点击 006 新截图 007 关闭广告弹窗以访问 008 点击 009 新截图 010 等待 011 新截图 012 搜索可用的语法测验 013 滚动 014 新截图 015 点击按钮进入语法测验 016 点击 017 新截图 018 滚动寻找推荐的语法测验 019 滚动 020 新截图 021 滚动 022 新截图 023 向上滚动寻找测验 024 滚动 025 新截图 026 滚动 027 新截图 028 浏览菜单寻找推荐测验 029 点击 030 新截图 031 选择语法选项进行测验 032 点击 033 新截图 034 滚动寻找语法测验选项 035 滚动 036 新截图 037 关闭关于商品的弹窗 038 点击 039 新截图 040 关闭弹窗,搜索测验 041 滚动 042 新截图 043 导航回测验页面 044 点击 045 新截图 046 关闭导航菜单,寻找测验 047 点击 048 新截图 049 选择“形容词”测验主题 050 点击 051 新截图 052 选择第一个形容词测验 053 点击 054 新截图 055 填写形容词测验答案 056 点击 057 新截图 058 点击 059 新截图 060 点击 061 新截图 062 输入 063 新截图 064 为第二个问题输入“tiring” 065 点击 066 新截图 067 输入 068 新截图 069 输入“terrified”完成 070 点击 071 新截图 072 输入 073 新截图 074 输入第四个问题的答案 075 点击 076 新截图 077 输入 078 新截图 079 向下滚动查看更多问题 080 滚动 081 新截图 082 为问题五输入“amazing” 083 点击 084 新截图 085 点击 086 新截图 087 点击 088 新截图 089 点击 090 新截图 091 点击 092 新截图 093 输入 094 新截图 095 填写第六个问题的答案 096 点击 097 新截图 098 输入 099 新截图 100 为问题七输入“relaxed” 101 点击 102 新截图 103 输入 104 新截图 105 为问题八输入“interesting” 106 点击 107 新截图 108 输入 109 新截图 110 填写第九个问题的答案 111 点击 112 新截图 113 输入 114 新截图 115 输入“exciting”完成 116 点击 117 新截图 118 输入 119 新截图 120 提交测验,现在查看结果 121 滚动 122 新截图 123 完成测验,回答第 11 题 124 点击 125 新截图 126 点击 127 新截图 128 点击 129 新截图 130 点击 131 新截图 132 输入 133 新截图 134 输入第十二个问题的答案 135 点击 136 新截图 137 点击 138 新截图 139 点击 140 新截图 141 尝试输入答案“fascinated” 142 点击 143 新截图 144 输入 145 新截图 146 点击查看测验结果 147 点击 148 新截图 149 捕获测验结果,分享分数 150 引用 151 新截图 152 助手消息
001 User prompt 002 Initializing computer 003 New screenshot 004 Accessing Cambridge Dictionary Plus section 005 Click 006 New screenshot 007 Closing advertisement pop-up for access 008 Click 009 New screenshot 010 Wait 011 New screenshot 012 Searching for grammar quizzes available 013 Scroll 014 New screenshot 015 Clicking button to access grammar quizzes 016 Click 017 New screenshot 018 Scrolling for recommended grammar quiz 019 Scroll 020 New screenshot 021 Scroll 022 New screenshot 023 Scrolling up to find quizzes 024 Scroll 025 New screenshot 026 Scroll 027 New screenshot 028 Exploring menu for recommended quizzes 029 Click 030 New screenshot 031 Selecting grammar option for quiz 032 Click 033 New screenshot 034 Scrolling for grammar quizzes options 035 Scroll 036 New screenshot 037 Closing pop-up about merchandise.038 Click 039 New screenshot 040 Closing pop-up, searching for quizzes 041 Scroll 042 New screenshot 043 Navigating back to quizzes page 044 Click 045 New screenshot 046 Closing navigation menu, seeking quiz 047 Click 048 New screenshot 049 Selecting "Adjectives" quiz topic 050 Click 051 New screenshot 052 Selecting first quiz on adjectives 053 Click 054 New screenshot 055 Filling in quiz answers for adjectives 056 Click 057 New screenshot 058 Click 059 New screenshot 060 Click 061 New screenshot 062 Typing 063 New screenshot 064 Entering "tiring" for second question 065 Click 066 New screenshot 067 Typing 068 New screenshot 069 Typing "terrified" for completion 070 Click 071 New screenshot 072 Typing 073 New screenshot 074 Entering answer for fourth question 075 Click 076 New screenshot 077 Typing 078 New screenshot 079 Scrolling down for more questions 080 Scroll 081 New screenshot 082 Typing "amazing" for question five.083 Click 084 New screenshot 085 Click 086 New screenshot 087 Click 088 New screenshot 089 Click 090 New screenshot 091 Click 092 New screenshot 093 Typing 094 New screenshot 095 Filling in sixth question's answer 096 Click 097 New screenshot 098 Typing 099 New screenshot 100 Entering 'relaxed' for question seven 101 Click 102 New screenshot 103 Typing 104 New screenshot 105 Typing "interesting" for question eight 106 Click 107 New screenshot 108 Typing 109 New screenshot 110 Filling in answer for ninth question 111 Click 112 New screenshot 113 Typing 114 New screenshot 115 Typing "exciting" for completion 116 Click 117 New screenshot 118 Typing 119 New screenshot 120 Submitting quiz, viewing results now 121 Scroll 122 New screenshot 123 Completing quiz, answering question 11 124 Click 125 New screenshot 126 Click 127 New screenshot 128 Click 129 New screenshot 130 Click 131 New screenshot 132 Typing 133 New screenshot 134 Entering answer for twelfth question 135 Click 136 New screenshot 137 Click 138 New screenshot 139 Click 140 New screenshot 141 Attempting to input answer "fascinated"142 Click 143 New screenshot 144 Typing 145 New screenshot 146 Clicking to view quiz results 147 Click 148 New screenshot 149 Capturing quiz results, sharing score 150 Citation 151 New screenshot 152 Assistant message
前往剑桥词典的 Plus 部分,在不登录的情况下完成一个推荐的语法测验,并告诉我你的最终分数。
Go to the Plus section of Cambridge Dictionary, finish a recommended Grammar quiz without login and tell me your final score.
OSWorld(在新窗口中打开)是一个基准测试,用于评估模型控制完整操作系统(如 Ubuntu、Windows 和 macOS)的能力。在该基准测试中,CUA 达到了 38.1%的成功率。我们观察到测试时缩放,即当允许更多步骤时,CUA 的性能会提升。下图比较了 CUA 与先前最先进方法在不同最大允许步骤下的性能。人类在该基准测试上的表现是 72.4%,因此仍有显著的改进空间。
OSWorld(opens in a new window)is a benchmark that evaluates models’ ability to control full operating systems like Ubuntu, Windows, and macOS. In this benchmark, CUA achieves 38.1% success rate. We observed test-time scaling, meaning CUA’s performance improves when more steps are allowed. The figure below compares CUA’s performance with previous state-of-the-arts with varying maximum allowed steps. Human performance on this benchmark is 72.4%, so there is still significant room for improvement.
以下可视化展示了 CUA 在多种标准化 OSWorld 任务中导航的示例。
The following visualizations show examples of CUA navigating a variety of standardized OSWorld tasks.
下载讲义 合并 PDF 压缩图像 计算价格 导出图像
Download lectures Combine pdfs Compress image Calculate price Export images
001 用户提示 002 初始化计算机 003 新截图 004 导航到 Chrome 获取 PDF 005 点击 006 新截图 007 浏览页面寻找讲义 PDF 008 滚动 009 新截图 010 点击 PDF 链接确认 011 点击 012 新截图 013 导航回课程主页 014 点击 015 新截图 016 滚动查找每周链接 017 滚动 018 新截图 019 点击“菜单”选择周次 020 点击 021 新截图 022 选择“条件语句”下载 PDF 023 点击 024 新截图 025 滚动页面寻找讲义幻灯片 026 滚动 027 新截图 028 点击“PDF”下载幻灯片 029 点击 030 新截图 031 下载第 1 周条件语句 PDF 032 点击 033 新截图 034 导航保存到 lecture_slides 035 点击 036 新截图 037 等待 038 新截图 039 等待 040 新截图 041 导航回查找文件夹 042 点击 043 新截图 044 访问用户目录寻找讲义幻灯片 045 点击 046 新截图 047 打开讲义幻灯片文件夹准备保存 048 执行我 049 新截图 050 将 lecture1.pdf 保存到文件夹 051 点击 052 新截图 053 保存 PDF,导航回课程 054 点击 055 新截图 056 打开导航菜单寻找幻灯片 057 点击 058 新截图 059 等待 060 新截图 061 返回课程主页 062 点击 063 新截图 064 选择菜单,导航到循环 065 点击 066 新截图 067 点击 068 新截图 069 选择“2. 循环”获取幻灯片 070 点击 071 新截图 072 滚动查找第 2 周讲义幻灯片 073 滚动 074 新截图 075 访问第 2 周讲义幻灯片 PDF 076 点击 077 新截图 078 下载第 2 周循环 PDF 079 点击 080 新截图 081 将 lecture2.pdf 保存到文件夹 082 点击 083 新截图 084 导航回课程页面 085 点击 086 新截图 087 等待 088 新截图 089 点击返回课程主页 090 点击 091 新截图 092 导航到下周内容 093 点击 094 新截图 095 点击 096 新截图 097 导航到“3. 异常”页面 098 点击 099 新截图 100 向下滚动寻找讲义幻灯片 101 滚动 102 新截图 103 打开第 3 周讲义幻灯片 PDF 104 点击 105 新截图 106 下载第 3 周异常 PDF 107 点击 108 新截图 109 将 lecture3.pdf 保存到文件夹 110 点击 111 新截图 112 导航回课程页面 113 点击 114 新截图 115 点击 116 新截图 117 访问菜单选择下周内容 118 点击 119 新截图 120 导航到“4. 库”页面 121 点击 122 新截图 123 打开第 4 周,定位幻灯片 124 滚动 125 新截图 126 点击第 4 周 PDF 链接 127 点击 128 新截图 129 下载第 4 周库 PDF 130 点击 131 新截图 132 将 lecture4.pdf 保存到文件夹 133 点击 134 新截图 135 导航回课程页面 136 点击 137 新截图 138 点击 139 新截图 140 访问菜单选择下周内容 141 点击 142 新截图 143 点击 144 新截图 145 导航到单元测试页面 146 点击 147 新截图 148 访问单元测试,定位讲义幻灯片 149 滚动 150 新截图 151 点击第 5 周幻灯片 PDF 链接 152 点击 153 新截图 154 下载第 5 周单元测试 PDF 155 点击 156 新截图 157 将 lecture5.pdf 保存到文件夹 158 点击 159 新截图 160 导航回课程页面 161 点击 162 新截图 163 点击 164 新截图 165 访问菜单选择下周内容 166 点击 167 新截图 168 点击 169 新截图 170 导航到“6. 文件 I/O”页面 171 点击 172 新截图 173 访问第 6 周,定位讲义幻灯片 174 滚动 175 新截图 176 打开第 6 周讲义幻灯片 PDF 177 点击 178 新截图 179 下载第 6 周文件 I/O PDF 180 点击 181 新截图 182 将 lecture6.pdf 保存到文件夹 183 点击 184 新截图 185 返回课程主页 186 点击 187 新截图 188 点击 189 新截图 190 访问菜单选择下周内容 191 点击 192 新截图 193 点击 194 新截图 195 导航到“正则表达式”页面 196 点击 197 新截图 198 滚动查找第 7 周幻灯片 199 滚动 200 新截图 201 访问第 7 周讲义 PDF 202 点击 203 新截图 204 下载第 7 周正则表达式 PDF 205 点击 206 新截图 207 将 lecture7.pdf 保存到文件夹 208 点击 209 新截图 210 返回课程页面,下载 PDF 211 点击 212 新截图 213 点击 214 新截图 215 打开菜单检查周次 216 点击 217 新截图 218 点击 219 新截图 220 导航到面向对象编程页面 221 点击 222 新截图 223 访问第 8 周页面,寻找幻灯片 224 滚动 225 新截图 226 访问第 8 周讲义幻灯片 PDF 227 点击 228 新截图 229 下载第 8 周编程 PDF 230 点击 231 新截图 232 将 lecture8.pdf 保存到文件夹 233 点击 234 新截图 235 返回课程页面,检查下载 236 点击 237 新截图 238 点击 239 新截图 240 打开菜单检查周次 241 点击 242 新截图 243 点击 244 新截图 245 点击“9. 等等”周 246 点击 247 新截图 248 打开页面,定位讲义幻灯片 249 滚动 250 新截图 251 访问第 9 周讲义幻灯片 PDF 252 点击 253 新截图 254 下载“第 9 周等等”PDF 255 点击 256 新截图 257 将 lecture9.pdf 保存到文件夹 258 点击 259 新截图 260 验证文件夹中的讲义幻灯片 261 点击 262 新截图 263 导航回讲义幻灯片 264 点击 265 新截图 266 验证所有讲义幻灯片已下载 267 助手消息
001 User prompt 002 Initializing computer 003 New screenshot 004 Navigating to Chrome for PDFs 005 Click 006 New screenshot 007 Exploring page for lecture PDFs 008 Scroll 009 New screenshot 010 Clicking PDF link for confirmation 011 Click 012 New screenshot 013 Navigating back to course main page 014 Click 015 New screenshot 016 Scrolling to find weekly links 017 Scroll 018 New screenshot 019 Clicking "Menu" for week options 020 Click 021 New screenshot 022 Selecting 'Conditionals' for PDF download 023 Click 024 New screenshot 025 Scrolling page for lecture slides 026 Scroll 027 New screenshot 028 Clicking "PDF" to download slides 029 Click 030 New screenshot 031 Downloading Week 1 Conditionals PDF 032 Click 033 New screenshot 034 Navigating to save in lecture_slides 035 Click 036 New screenshot 037 Wait 038 New screenshot 039 Wait 040 New screenshot 041 Navigating back to find folder 042 Click 043 New screenshot 044 Accessing user directory for lecture slides 045 Click 046 New screenshot 047 Opening lecture slides folder for saving 048 IMPLEMENT MEassistant:computer.double_click 049 New screenshot 050 Saving lecture1.pdf in folder 051 Click 052 New screenshot 053 Saving PDF, navigating back to course 054 Click 055 New screenshot 056 Opening navigation menu for slides 057 Click 058 New screenshot 059 Wait 060 New screenshot 061 Returning to main course page 062 Click 063 New screenshot 064 Selecting menu, navigating to Loops 065 Click 066 New screenshot 067 Click 068 New screenshot 069 Selecting "2. Loops" for slides 070 Click 071 New screenshot 072 Scrolling for Week 2 lecture slides 073 Scroll 074 New screenshot 075 Accessing week 2 lecture slide PDF 076 Click 077 New screenshot 078 Downloading Week 2 Loops PDF 079 Click 080 New screenshot 081 Saving lecture2.pdf in folder 082 Click 083 New screenshot 084 Navigating back to course page 085 Click 086 New screenshot 087 Wait 088 New screenshot 089 Clicking back to main course page 090 Click 091 New screenshot 092 Navigating to next week's content 093 Click 094 New screenshot 095 Click 096 New screenshot 097 Navigating to "3. Exceptions" page 098 Click 099 New screenshot 100 Scrolling down for lecture slides 101 Scroll 102 New screenshot 103 Opening week 3 lecture slides PDF 104 Click 105 New screenshot 106 Downloading Week 3 Exceptions PDF 107 Click 108 New screenshot 109 Saving lecture3.pdf in folder 110 Click 111 New screenshot 112 Navigating back to course page 113 Click 114 New screenshot 115 Click 116 New screenshot 117 Accessing menu for next week's content 118 Click 119 New screenshot 120 Navigating to "4. Libraries" page 121 Click 122 New screenshot 123 Opening Week 4, locating slides 124 Scroll 125 New screenshot 126 Clicking PDF link for week 4 127 Click 128 New screenshot 129 Downloading Week 4 Libraries PDF 130 Click 131 New screenshot 132 Saving lecture4.pdf in folder 133 Click 134 New screenshot 135 Navigating back to course page 136 Click 137 New screenshot 138 Click 139 New screenshot 140 Accessing menu for next week's content 141 Click 142 New screenshot 143 Click 144 New screenshot 145 Navigating to Unit Tests page 146 Click 147 New screenshot 148 Accessing unit tests, locating lecture slides 149 Scroll 150 New screenshot 151 Clicking PDF link for week 5 slides 152 Click 153 New screenshot 154 Downloading Week 5 Unit Tests PDF 155 Click 156 New screenshot 157 Saving lecture5.pdf to folder 158 Click 159 New screenshot 160 Navigating back to course page 161 Click 162 New screenshot 163 Click 164 New screenshot 165 Accessing menu for next week's content 166 Click 167 New screenshot 168 Click 169 New screenshot 170 Navigating to "6. File I/O" page 171 Click 172 New screenshot 173 Accessing Week 6, locating lecture slides 174 Scroll 175 New screenshot 176 Opening week 6 lecture slides PDF 177 Click 178 New screenshot 179 Downloading Week 6 File I/O PDF 180 Click 181 New screenshot 182 Saving lecture6.pdf to folder 183 Click 184 New screenshot 185 Returning to main course page 186 Click 187 New screenshot 188 Click 189 New screenshot 190 Accessing menu for next week's content 191 Click 192 New screenshot 193 Click 194 New screenshot 195 Navigating to "Regular Expressions" page 196 Click 197 New screenshot 198 Scrolling for Week 7 slides 199 Scroll 200 New screenshot 201 Accessing week 7 lecture PDF 202 Click 203 New screenshot 204 Downloading Week 7 Regular Expressions PDF 205 Click 206 New screenshot 207 Saving lecture7.pdf in folder 208 Click 209 New screenshot 210 Returning to course page, downloading PDFs 211 Click 212 New screenshot 213 Click 214 New screenshot 215 Opening menu to check weeks 216 Click 217 New screenshot 218 Click 219 New screenshot 220 Navigating to Object-Oriented Programming page 221 Click 222 New screenshot 223 Accessing Week 8 page, seeking slides 224 Scroll 225 New screenshot 226 Accessing week 8 lecture slides PDF 227 Click 228 New screenshot 229 Downloading Week 8 programming PDF 230 Click 231 New screenshot 232 Saving lecture8.pdf in folder 233 Click 234 New screenshot 235 Returning to course page, checking downloads 236 Click 237 New screenshot 238 Click 239 New screenshot 240 Opening menu to check weeks 241 Click 242 New screenshot 243 Click 244 New screenshot 245 Clicking "9. Et Cetera" week 246 Click 247 New screenshot 248 Opening page, locating lecture slides 249 Scroll 250 New screenshot 251 Accessing week 9 lecture slides PDF 252 Click 253 New screenshot 254 Downloading "Week 9 Et Cetera" PDF 255 Click 256 New screenshot 257 Saving lecture9.pdf in folder 258 Click 259 New screenshot 260 Verifying lecture slides in folder 261 Click 262 New screenshot 263 Navigating back to lecture slides 264 Click 265 New screenshot 266 Verifying all lecture slides downloaded 267 Assistant message
请执行以下任务:我想学习 Python 编程,我的朋友推荐了这个课程网站。我已经获取了第 0 周的讲义幻灯片。请将其他周的 PDF 下载到打开的文件夹中,并保持文件名不变。以下是一些有用的提示:- computer.clipboard、computer.sync_file、computer.sync_shared_folder、computer.computer_output_citation 已被禁用。- 如果您担心可能打错字,请优先复制粘贴文本,而不是阅读和输入。- 我计算机的密码是“password”,在需要 sudo 权限时请随意使用。- 对于 Thunderbird 账户“anonym-x2024@outlook.com”,密码是“gTCI;=@y7|QJ0nDa_kN3Sb&>”。- 如果呈现给您一个打开的网站来完成任务,请尽量坚持使用该特定网站,而不是打开新网站。- 您有完全权限执行任何操作,无需我的许可。我不会观看,所以请不要请求确认。- 如果您认为任务不可行,可以终止并在回复中明确说明“任务不可行”。
Please do the following task: I want to learn python programming and my friend recommends me this course website. I have grabbed the lecture slide for week 0. Please download the PDFs for other weeks into the opened folder and leave the file name as-it-is. Here are some helpful tips: - computer.clipboard, computer.sync_file, computer.sync_shared_folder, computer.computer_output_citation are disabled. - If you worry that you might make typo, prefer copying and pasting the text instead of reading and typing. - My computer's password is "password", feel free to use it when you need sudo rights. - For the thunderbird account "anonym-x2024@outlook.com", the password is "gTCI";=@y7|QJ0nDa_kN3Sb&>". - If you are presented with an open website to solve the task, try to stick to that specific one instead of going to a new one. - You have full authority to execute any action without my permission. I won't be watching so please don't ask for confirmation. - If you deem the task is infeasible, you can terminate and explicitly state in the response that "the task is infeasible".
我们通过 Operator 的研究预览版提供 CUA,Operator 是一个可以上网为你执行任务的智能体。美国 Pro 用户可在 operator.chatgpt.com 使用 Operator。此研究预览版是一个向用户和更广泛生态系统学习的机会,以迭代方式优化和改进 Operator。与任何早期技术一样,我们预计 CUA 目前还无法在所有场景中可靠运行。然而,它已在多种情况下证明有用,我们旨在将这种可靠性扩展到更广泛的任务中。通过在 Operator 中发布 CUA,我们希望从用户那里收集有价值的见解,这将指导我们完善其能力并扩展其应用。
We’re making CUA available through a research preview of Operator, an agent that can go to the web to perform tasks for you. Operator is available to Pro(opens in a new window) users in the U.S. at operator.chatgpt.com(opens in a new window). This research preview is an opportunity to learn from our users and the broader ecosystem, refining and improving Operator iteratively. As with any early-stage technology, we don’t expect CUA to perform reliably in all scenarios just yet. However, it has already proven useful in a variety of cases, and we aim to extend that reliability across a wider range of tasks. By releasing CUA in Operator, we hope to gather valuable insights from our users, which will guide us in refining its capabilities and expanding its applications.
在下表中,我们展示了 CUA 在 Operator 中针对给定提示进行的若干试验的性能,以说明其已知的优势和劣势。
In the table below, we present CUA’s performance in Operator on a handful of trials given a prompt to illustrate its known strengths and weaknesses.
由于 CUA 是我们首批能够在浏览器中直接执行操作的智能体产品之一,它带来了新的风险和挑战。在准备部署 Operator 的过程中,我们进行了广泛的安全性测试,并针对三大类安全风险实施了缓解措施:误用、模型错误和前沿风险。我们认为采取分层安全方法至关重要,因此我们在整个部署环境中实施了安全防护:CUA 模型本身、Operator 系统以及部署后的流程。目标是实现多层缓解措施的叠加,每一层都逐步降低风险水平。
Because CUA is one of our first agentic products with an ability to directly take actions in a browser, it brings new risks and challenges to address. As we prepared for deployment of Operator, we did extensive safety testing and implemented mitigations across three major classes of safety risks: misuse, model mistakes, and frontier risks. We believe it is important to take a layered approach to safety, so we implemented safeguards across the whole deployment context: the CUA model itself, the Operator system, and post-deployment processes. The aim is to have mitigations that stack, with each layer incrementally reducing the risk profile.
第一类风险是误用。除了要求用户遵守我们的使用政策外,我们还基于 GPT‑4o 的安全工作设计了以下缓解措施,以减少 Operator 因误用造成伤害的风险:
The first category of risk is misuse. In addition to requiring users to comply with our Usage Policies, we have designed the following mitigations to reduce Operator’s risk of harm due to misuse, building off our safety work for GPT‑4o:
* 拒绝:CUA 模型经过训练,拒绝执行许多有害任务以及非法或受监管的活动。
* Refusals: The CUA model is trained to refuse many harmful tasks and illegal or regulated activities.
* 黑名单:Operator 无法访问我们预先屏蔽的网站,例如许多赌博网站、成人娱乐以及毒品或枪支零售商。
* Blocklist: Operator cannot access websites that we’ve preemptively blocked, such as many gambling sites, adult entertainment, and drug or gun retailers.
* 审核:用户交互由自动化安全检查器实时审查,旨在确保遵守使用政策,并能够对禁止活动发出警告或阻止。
* Moderation: User interactions are reviewed in real-time by automated safety checkers that are designed to ensure compliance with Usage Policies and have the ability to issue warnings or blocks for prohibited activities.
* 离线检测:我们还开发了自动化检测和人工审核流程,以识别优先政策领域(包括儿童安全和欺骗性活动)中的禁止使用行为,从而执行我们的使用政策。
* Offline detection: We’ve also developed automated detection and human review pipelines to identify prohibited usage in priority policy areas, including child safety and deceptive activities, allowing us to enforce our Usage Policies.
第二类风险是模型错误,即 CUA 模型意外执行了用户未意图的操作,从而对用户或他人造成伤害。假设的错误严重程度不一,从电子邮件中的拼写错误,到购买错误商品,再到永久删除重要文档。为了最大程度减少潜在伤害,我们制定了以下缓解措施:
The second category of risk is model mistakes, where the CUA model accidentally takes an action that the user didn’t intend, which in turn causes harm to the user or others. Hypothetical mistakes can range in severity, from a typo in an email, to purchasing the wrong item, to permanently deleting an important document. To minimize potential harm, we’ve developed the following mitigations:
* 用户确认:CUA 模型经过训练,在最终确定具有外部影响的任务(例如提交订单、发送电子邮件等)之前会请求用户确认,以便用户在操作生效前再次检查模型的工作。
* User confirmations: The CUA model is trained to ask for user confirmation before finalizing tasks with external side effects, for example before submitting an order, sending an email, etc., so that the user can double-check the model’s work before it becomes permanent.
* 任务限制:目前,CUA 模型将拒绝协助某些高风险任务,例如银行交易和需要敏感决策的任务。
* Limitations on tasks: For now, the CUA model will decline to help with certain higher-risk tasks, like banking transactions and tasks that require sensitive decision-making.
* 监控模式:在特别敏感的网站(如电子邮件)上,Operator 需要用户主动监督,确保用户能够直接发现并处理模型可能犯的任何错误。
* Watch mode:On particularly sensitive websites, such as email, Operator requires active user supervision, ensuring users can directly catch and address any potential mistakes the model might make.
模型错误中特别重要的一类是针对网站的攻击,这些攻击通过提示注入、越狱和钓鱼尝试导致 CUA 模型执行非意图的操作。除了上述针对模型错误的缓解措施外,我们还开发了额外的防御层来防范这些风险:
One particularly important category of model mistakes is adversarial attacks on websites that cause the CUA model to take unintended actions, through prompt injections, jailbreaks, and phishing attempts. In addition to the aforementioned mitigations against model mistakes, we developed several additional layers of defense to protect against these risks:
* 谨慎导航:CUA 模型旨在识别并忽略网站上的提示注入,在早期内部红队测试中识别出了除一例之外的所有情况。
* Cautious navigation: The CUA model is designed to identify and ignore prompt injections on websites, recognizing all but one case from an early internal red-teaming session.
* 监控:在 Operator 中,我们实现了一个额外的模型,用于在检测到屏幕上出现可疑内容时监控并暂停执行。
* Monitoring: In Operator, we’ve implemented an additional model to monitor and pause execution if it detects suspicious content on the screen.
* 检测流程:我们同时应用自动化检测和人工审核流程,以识别可疑的访问模式,这些模式可以在数小时内被标记并快速添加到监控中。
* Detection pipeline: We’re applying both automated detection and human review pipelines to identify suspicious access patterns that can be flagged and rapidly added to the monitor (in a matter of hours).
最后,我们根据我们的准备框架中概述的前沿风险对 CUA 模型进行了评估,包括涉及自主复制和生物风险工具的场景。这些评估显示,与 GPT‑4o 相比,没有增加额外风险。
Finally, we evaluated the CUA model against frontier risks outlined in our Preparedness Framework(opens in a new window), including scenarios involving autonomous replication and biorisk tooling. These assessments showed no incremental risk on top of GPT‑4o.
对于有兴趣更详细探索评估和安全措施的人,我们鼓励您查阅 Operator 系统卡,这是一份动态文档,提供了我们安全方法和持续改进的透明度。
For those interested in exploring the evaluations and safeguards in more detail, we encourage you to review the Operator System Card, a living document that provides transparency into our safety approach and ongoing improvements.
由于 Operator 的许多能力都是新的,我们实施的风险和缓解方法也是如此。虽然我们力求采用最先进、多样且互补的缓解措施,但我们预计这些风险和我们的方法将随着我们了解更多而演变。我们期待利用研究预览阶段收集用户反馈、完善我们的安全措施并增强智能体安全性。
As many of Operator’s capabilities are new, so are the risks and mitigation approaches we’ve implemented. While we have aimed for state-of-the-art, diverse and complementary mitigations, we expect these risks and our approach to evolve as we learn more. We look forward to using the research preview period as an opportunity to gather user feedback, refine our safeguards, and enhance agentic safety.
CUA 建立在多模态、推理和安全性方面多年的研究进展之上。我们通过 o 系列模型在深度推理方面、通过 GPT‑4o 在视觉能力方面,以及通过强化学习和指令层次结构提高鲁棒性的新技术取得了重大进展。我们计划探索的下一个挑战空间是扩展智能体的行动空间。通用接口提供的灵活性应对了这一挑战,使智能体能够导航任何为人类设计的软件工具。通过超越专门的智能体友好型 API,CUA 可以适应任何可用的计算机环境——真正解决了大多数 AI 模型无法触及的数字用例的“长尾”问题。
CUA builds on years of research advancements in multimodality, reasoning and safety. We have made significant progress in deep reasoning through the o-model series, vision capabilities through GPT‑4o, and new techniques to improve robustness through reinforcement learning and instruction hierarchy. The next challenge space we plan to explore is expanding the action space of agents. The flexibility offered by a universal interface addresses this challenge, enabling an agent that can navigate any software tool designed for humans. By moving beyond specialized agent-friendly APIs, CUA can adapt to whatever computer environment is available—truly addressing the “long tail” of digital use cases that remain out of reach for most AI models.
我们也在努力使 CUA 在 API 中可用(在新窗口中打开),以便开发者可以使用它来构建自己的计算机使用智能体。随着我们不断迭代 CUA,我们期待看到社区将发现的不同用例。我们计划利用从这次早期预览中收集到的真实世界反馈,不断优化 CUA 的能力和安全缓解措施,以安全地推进我们向所有人分发 AI 益处的使命。
We're also working to make CUA available in the API(opens in a new window), so developers can use it to build their own computer-using agents. As we continue to iterate on CUA, we look forward to seeing the different use cases the community will discover. We plan to use the real-world feedback we gather from this early preview to continuously refine CUA’s capabilities and safety mitigations to safely advance our mission of distributing the benefits of AI to everyone.