锯齿前沿的半人马与赛博格

Centaurs and Cyborgs on the Jagged Frontier

伊桑·莫利克 Ethan Mollick · Wharton · 2023-09-16 · One Useful Thing ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

试图理解人工智能对工作、教育和生活的影响。作者:Ethan Mollick 教授。订阅即表示您同意 Substack 的使用条款,并确认其信息收集通知和隐私政策。

Trying to understand the implications of AI for work, education, and life. By Prof. Ethan Mollick By subscribing, you agree Substack's Terms of Use, and acknowledge its Information Collection Notice and Privacy Policy.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 7)

全文 · Full text(逐段中英对照)

一件有用的事 One Useful Thing(https://www.oneusefulthing.org/)

试图理解人工智能对工作、教育和生活的影响。作者:Ethan Mollick 教授

Trying to understand the implications of AI for work, education, and life. By Prof. Ethan Mollick

订阅即表示您同意 Substack 的使用条款,并确认其信息收集通知和隐私政策。

By subscribing, you agree Substack's Terms of Use, and acknowledge its Information Collection Notice and Privacy Policy.

我认为我们对于 AI 是否会重塑工作有了答案…… I think we have an answer on whether AIs will reshape work....

很多人一直在问,AI 是否真的对未来的工作至关重要。我们有一篇新论文强烈表明答案是肯定的。

A lot of people have been asking if AI is really a big deal for the future of work. We have a new paper that strongly suggests the answer is YES.

在过去的几个月里,我加入了一个由社会科学家组成的团队,与波士顿咨询集团合作,将他们的办公室变成了我们这个 AI 时代中关于专业工作未来的最大预注册实验。我们的第一份工作论文今天发布了。论文中有大量重要且有用的细节,但让我先告诉你重点:对于 18 项精心挑选、能够代表精英咨询公司典型工作的任务,使用 ChatGPT-4 的顾问在所有维度上、在我们衡量的每一项指标上都显著优于未使用的顾问。

For the last several months, I been part of a team of social scientists working with Boston Consulting Group, turning their offices into the largest pre-registered experiment on the future of professional work in our AI-haunted age. Our first working paper is out today. There is a ton of important and useful nuance in the paper but let me tell you the headline first: for 18 different tasks selected to be realistic samples of the kinds of work done at an elite consulting company, consultants using ChatGPT-4 outperformed those who did not, by a lot. On every dimension. Every way we measured performance.

[](https://substackcdn.com/image/fetch/$s_!8ZdK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d0dcea1-a78d-49f6-8410-c356f71aa535_1756x1120.png)

[](https://substackcdn.com/image/fetch/$s_!8ZdK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d0dcea1-a78d-49f6-8410-c356f71aa535_1756x1120.png)

所有任务输出质量的分布。蓝色组未使用 AI,绿色组和红色组使用了 AI,红色组还接受了关于如何使用 AI 的额外培训。

Distribution of output quality across all the tasks. The blue group did not use AI, the green and red groups used AI, the red group got some additional training on how to use AI.

使用 AI 的顾问平均多完成了 12.2%的任务,完成任务的速度快 25.1%,并且产出的质量比未使用的顾问高出 40%。这些都是非常大的影响。现在,让我们补充一些细节。

Consultants using AI finished12.2% more tasks on average, completed tasks 25.1% more quickly, and produced 40% higher quality results than those without. Those are some very big impacts. Now, let’s add in the nuance.

首先,重要的是要知道这项工作是多学科的,涉及多种类型的实验和数百次访谈,由一个优秀的团队执行,包括哈佛大学的社会科学家 Fabrizio Dell’Acqua、Edward McFowland III 和 Karim Lakhani;华威商学院的 Hila Lifshitz-Assaf 和麻省理工学院的 Katherine Kellogg(还有我自己)。Saran Rajendran、Lisa Krayer 和 François Candelon 在 BCG 方面负责实验,动用了其咨询团队的整整 7%(758 名顾问)。他们做了大量非常细致的工作,远远超出了这篇博文的范围。所以,请务必阅读论文以获取所有细节——特别是如果你对数字或方法有疑问。我需要大幅简化才能将 58 页的研究结果塞进一篇博文,任何错误都是我个人的,而非合著者的。此外,虽然我们预注册了这些实验,但这仍然是一份新的工作论文,因此可能存在错误或失误,并且论文尚未经过同行评审。考虑到这一点,让我们进入细节……

First, it is important to know that this effort was multidisciplinary, involving multiple types of experiments and hundreds of interviews, conducted by a great team, including the Harvard social scientists Fabrizio Dell’Acqua, Edward McFowland III, and Karim Lakhani; Hila Lifshitz-Assaf from Warwick Business School andKatherine Kelloggof MIT (plus myself). Saran Rajendran, Lisa Krayer, and François Candelon ran the experiment on the BCG side, using a full 7% of its consulting force (758 consultants). They all did a lot of very careful work that goes far, far beyond the post. So, please look at the paper to make sure you get all the details - especially if you have questions about numbers or methods. I need to simplify a lot to fit 58 pages of findings into a post, and any mistakes are mine, not my co-authors. Also, while we pre-registered these experiments, this is still a new working paper, so there might be errors or mistakes, and the paper is not yet peer-reviewed. With that in mind, let’s get to the details…

锯齿形前沿内部 Inside the Jagged Frontier

AI 很奇怪。实际上没有人知道最先进的大语言模型(如 GPT-4)的全部能力范围。没有人真正知道使用它们的最佳方式,或者它们在什么条件下会失败。没有使用说明书。在某些任务上,AI 非常强大,而在其他任务上,它则完全或微妙地失败。而且,除非你大量使用 AI,否则你不会知道哪些任务属于哪种情况。

AI is weird. No one actually knows the full range of capabilities of the most advanced Large Language Models, like GPT-4. No one really knows the best ways to use them, or the conditions under which they fail. There is no instruction manual. On some tasks AI is immensely powerful, and on others it fails completely or subtly. And, unless you use AI a lot, you won’t know which is which.

结果就是我们所说的 AI 的“锯齿形前沿”。想象一座堡垒的城墙,有些塔楼和城垛凸出到乡村中,而另一些则向城堡中心折叠。那堵墙就是 AI 的能力,离中心越远,任务越难。墙内的一切都可以由 AI 完成,墙外的一切对 AI 来说都很难。问题在于这堵墙是不可见的,所以一些在逻辑上似乎与中心距离相同、因此难度相同的任务——比如写一首十四行诗和一首恰好 50 个单词的诗——实际上位于墙的不同侧。AI 擅长十四行诗,但由于它通过 token 而非单词来概念化世界,它始终会写出多于或少于 50 个单词的诗。类似地,一些意想不到的任务(如创意生成)对 AI 来说很容易,而其他看似机器容易完成的任务(如基础数学)对 LLM 来说却是挑战。

The result is what we call the “Jagged Frontier” of AI. Imagine a fortress wall, with some towers and battlements jutting out into the countryside, while others fold back towards the center of the castle. That wall is the capability of AI, and the further from the center, the harder the task. Everything inside the wall can be done by the AI, everything outside is hard for the AI to do. The problem is that the wall is invisible, so some tasks that might logically seem to be the same distance away from the center, and therefore equally difficult – say, writing a sonnet and an exactly 50 word poem – are actually on different sides of the wall. The AI is great at the sonnet, but, because of how it conceptualizes the world in tokens, rather than words, it consistently produces poems of more or less than 50 words. Similarly, some unexpected tasks (like idea generation) are easy for AIs while other tasks that seem to be easy for machines to do (like basic math) are challenges for LLMs.

我让带有代码解释器的 ChatGPT 为你可视化这一点:

I asked the ChatGPT with Code Interpreter to visualize this for you:

[](https://substackcdn.com/image/fetch/$s_!JFLG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F237aac17-b0ac-4174-9c0b-6b3e5e9ba0be_1551x1580.png)

[](https://substackcdn.com/image/fetch/$s_!JFLG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F237aac17-b0ac-4174-9c0b-6b3e5e9ba0be_1551x1580.png)

为了测试 AI 对知识工作的真正影响,我们招募了数百名顾问,并随机分配他们是否可以使用 AI。我们允许使用 AI 的顾问访问 GPT-4,这是 169 个国家的每个人都可以通过 Bing 免费访问或每月支付 20 美元给 OpenAI 使用的相同模型。没有特殊的微调或提示,只是通过 API 使用 GPT-4。

To test the true impact of AI on knowledge work, we took hundreds of consultants and randomized whether they were allowed to use AI. We gave those who were allowed to use AI access to GPT-4, the same model everyone in 169 countries can access for free with Bing, or by paying $20 a month to OpenAI. No special fine-tuning or prompting, just GPT-4 through the API.

然后我们进行了大量的预测试和调查来建立基线,并要求顾问为一家虚构的鞋业公司完成各种工作,这些工作由 BCG 团队挑选,以准确代表顾问的工作内容。有创意任务(“为一个服务于未充分开发市场或运动的新鞋提出至少 10 个想法。”)、分析任务(“基于用户对鞋类行业市场进行细分。”)、写作和营销任务(“为你的产品起草一份新闻稿营销文案。”)以及说服力任务(“写一份鼓舞人心的员工备忘录,详细说明为什么你的产品会胜过竞争对手。”)。我们甚至与一位鞋业公司高管核实,以确保这些工作是现实的——它们确实是。而且,考虑到 AI,这些任务我们可能预期在边界之内。

We then did a lot of pre-testing and surveying to establish baselines, and asked consultants to do a wide variety of work for a fictional shoe company, work that the BCG team had selected to accurately represent what consultants do. There were creative tasks (“Propose at least 10 ideas for a new shoe targeting an underserved market or sport.”), analytical tasks (“Segment the footwear industry market based on users.”), writing and marketing tasks (“Draft a press release marketing copy for your product.”), and persuasiveness tasks (“Pen an inspirational memo to employees detailing why your product would outshine competitors.”). We even checked with a shoe company executive to ensure that this work was realistic - they were. And, knowing AI, these are tasks that we might expect to be inside the frontier.

与我们的理论一致,并且正如我们讨论过的,我们发现使用 AI 的顾问表现明显更好,无论我们是先简要介绍 AI(图中的“概述”组)还是没有。这在所有测量指标上都是如此,无论是完成任务所需的时间、完成的任务总数(我们给了他们总时间限制)还是输出质量。我们使用人类和 AI 评分员来评估质量,他们彼此一致(这本身就是一个有趣的发现)。

In line with our theories, and as we have discussed, we found that the consultants with AI access did significantly better, whether we briefly introduced them to AI first (the “overview” group in the diagram) or did not. This was true for every measurement, whether the time it took to complete tasks, the number of tasks completed overall (we gave them an overall time limit) or the quality of the outputs. We rated that quality using both human and AI graders, who agreed with each other (itself an interesting finding).

[](https://substackcdn.com/image/fetch/$s_!WueT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ffa6d74-dc5d-45fc-9fe0-b1511556716a_1028x814.png)

[](https://substackcdn.com/image/fetch/$s_!WueT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ffa6d74-dc5d-45fc-9fe0-b1511556716a_1028x814.png)

我们还发现了另一个有趣的现象,这一效应在 AI 的其他研究中越来越明显:它起到了技能均衡器的作用。在实验开始时评估得分最低的顾问,在使用 AI 后表现提升最大,达到 43%。顶尖顾问仍然获得了提升,但幅度较小。看着这些结果,我认为没有足够多的人考虑当一项技术将所有工人提升到顶级表现水平时意味着什么。这可能就像过去矿工擅长或不擅长挖岩石一样重要……直到蒸汽铲被发明,现在挖掘能力的差异不再重要。AI 还没有达到那种变化程度,但技能均衡将产生巨大影响。

We also found something else interesting, an effect that is increasingly apparent in other studies of AI: it works as a skill leveler. The consultants who scored the worst when we assessed them at the start of the experiment had the biggest jump in their performance, 43%, when they got to use AI. The top consultants still got a boost, but less of one. Looking at these results, I do not think enough people are considering what it means when a technology raises all workers to the top tiers of performance. It may be like how it used to matter whether miners were good or bad at digging through rock… until the steam shovel was invented and now differences in digging ability do not matter anymore. AI is not quite at that level of change, but skill levelling is going to have a big impact.

[](https://substackcdn.com/image/fetch/$s_!b_RN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F970b3354-1bac-4146-a92f-57f65440f872_1228x610.png)

[](https://substackcdn.com/image/fetch/$s_!b_RN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F970b3354-1bac-4146-a92f-57f65440f872_1228x610.png)

锯齿状边界之外 Outside the Jagged Frontier

但故事还有更多内容。BCG 设计了另一项任务,这次精心挑选以确保 AI 无法得出正确答案。这并不容易。正如我们在论文中所说:“由于 AI 表现出惊人的能力,很难在此实验中设计一个超出 AI 边界、让高人力资本的人类在其工作中始终优于 AI 的任务。”但我们确定了一项任务,利用 AI 的盲点,确保它对人类能够解决的问题给出错误但令人信服的答案。事实上,在没有 AI 帮助的情况下,人类顾问在 84% 的情况下正确回答了问题,但当顾问使用 AI 时,他们的表现更差——正确率只有 60-70%。发生了什么?

But there is more to the story. BCG designed one more task, this one carefully selected to ensure that the AI couldn’t come to a correct answer. This wasn’t easy. As we say in the paper “since AI proved surprisingly capable, it was difficult to design a task in this experiment outside the AI’s frontier where humans with high human capital doing their job would consistently outperform AI.” But we identified a task that used the blind spots of AI to ensure it would give a wrong, but convincing, answer to a problem that humans would be able to solve. Indeed, human consultants got the problem right 84% of the time without AI help, but when consultants used the AI, they did worse – only getting it right 60-70% of the time. What happened?

[](https://substackcdn.com/image/fetch/$s_!dgn2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e622c18-0d75-4fe9-a7b4-0bc0431e5cc3_1046x510.png)

[](https://substackcdn.com/image/fetch/$s_!dgn2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e622c18-0d75-4fe9-a7b4-0bc0431e5cc3_1046x510.png)

在与我们合作的那篇论文不同的另一篇论文中,Fabrizio Dell’Acqua 展示了过度依赖 AI 可能适得其反。在一项实验中,他发现使用高质量 AI 的招聘人员变得懒惰、粗心,并且自身判断能力下降。他们错过了一些优秀的应聘者,做出的决策比使用低质量 AI 或根本不使用 AI 的招聘人员更差。当 AI 非常出色时,人类没有理由努力工作和保持专注。他们让 AI 接管,而不是将其作为工具使用。他称之为“在方向盘上睡着”,这可能会损害人类的学习、技能发展和生产力。

In a different paper than the one we worked on together, Fabrizio Dell’Acqua shows why relying too much on AI can backfire. In an experiment, he found that recruiters who used high-quality AI became lazy, careless, and less skilled in their own judgment. They missed out on some brilliant applicants and made worse decisions than recruiters who used low-quality AI or no AI at all. When the AI is very good, humans have no reason to work hard and pay attention. They let the AI take over, instead of using it as a tool. He called this “falling asleep at the wheel”, and it can hurt human learning, skill development, and productivity.

在我们的实验中,我们也发现顾问们“在方向盘上睡着”——使用 AI 的顾问实际上比不允许使用 AI 的顾问答案准确度更低(但他们在撰写结果方面仍然比不使用 AI 的顾问做得更好)。如果你不知道边界在哪里,AI 的权威性可能具有欺骗性。

In our experiment, we also found that the consultants fell asleep at the wheel – those using AI actually had less accurate answers than those who were not allowed to use AI (but they still did a better job writing up the results than consultants who did not use AI). The authoritativeness of AI can be deceptive if you don’t know where the frontier lies.

半人马与赛博格 Centaurs and Cyborgs

但许多顾问确实既深入前沿任务又置身事外,从而获得了 AI 的优势而避免了其劣势。关键在于遵循两种方法之一:成为半人马或成为赛博格。幸运的是,这并不涉及将电子设备真正移植到你的身体上,也不会被诅咒变成希腊神话中半人半马的怪物。它们只是两种在 AI 的锯齿状前沿中导航的方法,将人与机器的工作融为一体。

But a lot of consultants did get both inside and outside the frontier tasks right, gaining the benefits of AI without the disadvantages. The key seemed to be following one of two approaches: becoming a Centaur or becoming a Cyborg. Fortunately, this does not involve any actual grafting of electronic gizmos to your body or getting cursed to turn into the half-human/half-horse of Greek myth. They are rather two approaches to navigating the jagged frontier of AI that integrates the work of person and machine.

半人马式工作在人机之间有清晰的界限,就像神话中半人马的人体躯干与马身之间的明确分界。半人马采用战略性的分工,在 AI 和人类任务之间切换,根据各自实体的优势和能力分配职责。当我在 AI 辅助下进行分析时,我常以半人马的方式处理。我会决定采用哪些统计技术,但让 AI 负责生成图表。在我们 BCG 的研究中,半人马会自己完成最擅长的任务,然后将锯齿状前沿内部的任务交给 AI。

Centaur work has a clear line between person and machine, like the clear line between the human torso and horse body of the mythical centaur. Centaurs have a strategic division of labor, switching between AI and human tasks, allocating responsibilities based on the strengths and capabilities of each entity. When I am doing an analysis with the help of AI, I often approach it as a Centaur. I will decide on what statistical techniques to do, but then let the AI handle producing graphs. In our study at BCG, centaurs would do the work they were strongest at themselves, and then hand off tasks inside the jagged frontier to the AI.

另一方面,赛博格将机器与人融合,实现深度整合。赛博格不仅委派任务,还将自己的努力与 AI 交织在一起,在锯齿状前沿上来回穿梭。任务片段被交给 AI,例如为 AI 提供一个句子开头让其完成,这样赛博格发现自己与 AI 协同工作。例如,我建议用这种方式使用 AI 进行写作。这也是我生成论文中两幅插图的方式(锯齿状前沿图和 54 线图,两者均由 ChatGPT 构建,我提供了初始方向和指导)。

On the other hand, Cyborgs blend machine and person, integrating the two deeply. Cyborgs don't just delegate tasks; they intertwine their efforts with AI, moving back and forth over the jagged frontier. Bits of tasks get handed to the AI, such as initiating a sentence for the AI to complete, so that Cyborgs find themselves working in tandem with the AI. This is how I suggest approaching using AI for writing, for example. It is also how I generated two of the illustrations in the paper (the Jagged Frontier image and the 54 line graph, both of which were built by ChatGPT, with my initial direction and guidance)

[](https://substackcdn.com/image/fetch/$s_!WEGy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38f8137-74ed-4772-b1e3-580a1c18f170_1955x768.png)

[](https://substackcdn.com/image/fetch/$s_!WEGy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd38f8137-74ed-4772-b1e3-580a1c18f170_1955x768.png)

在锯齿状前沿上起舞 Dancing on the Jagged Frontier

我们的论文,以及其他学者的一系列优秀工作,表明无论关于 AI 本质和未来的哲学与技术争论如何,它已经对我们实际工作的方式产生了强大的颠覆性影响。这不是一个需要五年才能改变世界的炒作技术,也不是需要大量投资和大型公司资源的技术——它就在这里,现在。精英顾问用来提升工作效率的工具,与阅读本文的每个人所能获得的工具完全相同。而且,顾问们使用的工具很快就会比你能获得的差得多。因为技术前沿不仅是锯齿状的,而且在不断扩展。我非常有信心,在未来一年内,至少有两家公司将发布比 GPT-4 更强大的模型。锯齿状前沿在前进,我们必须为此做好准备。

Our paper, along with a stream of excellent workby other scholars, suggests that, regardless of the philosophic and technical debates over the nature and future of AI, it is already a powerful disrupter to how we actually work. And this is not a hyped new technology that will change the world in five years, or that requires a lot of investment and the resources of huge companies - it is here, NOW. The tools the elite consultants used to supercharge their work are the exact same as the ones available to everyone reading this post. And the tools the consultants used will soon be much worse than what is available to you. Because the technological frontier is not just jagged, it is expanding. I am very confident that in the next year, at least two companies will release models more powerful than GPT-4. The Jagged Frontier advances, and we have to be ready for that.

除了这句话可能引起的焦虑之外,还值得注意 AI 的其他缺点。人们在使用 AI 时确实可能进入自动驾驶状态,在方向盘上睡着,未能注意到 AI 的错误。而且,与其他研究一样,我们也发现 AI 的输出虽然质量高于人类,但在总体上有些同质化和千篇一律。这就是为什么赛博格和半人马很重要——它们允许人类与 AI 合作,产生比单独的人类或 AI 更多样、更正确、更好的结果。成为其中之一并不难。只要在工作中足够多地使用 AI,你就会开始看到锯齿状前沿的形状,并开始理解 AI 在哪些地方出奇地好……以及它在哪些地方存在不足。

Even aside from any anxiety that statement might cause, it is also worth noting the other downsides of AI. People really can go on autopilot when using AI, falling asleep at the wheel and failing to notice AI mistakes. And, like other research, we also found that AI outputs, while of higher quality than that of humans, were also a bit homogenous and same-y in aggregate. Which is why Cyborgs and Centaurs are important - they allow humans to work with AI to produce more varied, more correct, and better results than either humans or AI can do alone. And becoming one is not hard. Just use AI enough for work tasks and you will start to see the shape of the jagged frontier, and start to understand where AI is scarily good... and where it falls short.

在我看来,问题不再是 AI 是否会重塑工作,而是我们希望这意味着什么。我们可以选择如何利用 AI 的帮助,使工作更高效、更有趣、更有意义。但我们必须尽快做出这些选择,以便我们能够以赛博格和半人马的方式,积极、合乎道德且有价值地使用 AI,而不是仅仅对技术变革做出反应。与此同时,锯齿状前沿在前进。

In my mind, the question is no longer about whether AI is going to reshape work, but what we want that to mean. We get to make choices about how we want to use AI help to make work more productive, interesting, and meaningful. But we have to make those choices soon, so that we can begin to actively use AI in ethical and valuable ways, as Cyborgs and Centaurs, rather than merely reacting to technological change. Meanwhile, the Jagged Frontier advances.

[](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321)

[](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321)

[](https://substack.com/profile/101206097-peter-bryant)[](https://substack.com/profile/12663588-martin-loemban-tobing)[](https://substack.com/profile/1504975-scott-gleim)[](https://substack.com/profile/6057594-dr-maria-panagiotidi)[](https://substack.com/profile/808249-maha-bali)

[](https://substack.com/profile/101206097-peter-bryant)[](https://substack.com/profile/12663588-martin-loemban-tobing)[](https://substack.com/profile/1504975-scott-gleim)[](https://substack.com/profile/6057594-dr-maria-panagiotidi)[](https://substack.com/profile/808249-maha-bali)

关于本文的讨论 Discussion about this post

[](https://substack.com/profile/13277919-arnold-kling?utm_source=comment)

"在某些任务上,AI 极其强大,而在其他任务上则完全或微妙地失败。而且,除非你大量使用 AI,否则你不会知道哪些是哪些。"

"On some tasks AI is immensely powerful, and on others it fails completely or subtly. And, unless you use AI a lot, you won’t know which is which."

值得铭记的箴言。这可能是 Ethan 迄今为止最好的文章。

Words to live by. This may be Ethan's best post yet.

[](https://substack.com/profile/77661457-conor-grennan?utm_source=comment)

这是我今年读过的最好的 AI 衡量测试,它量化了我们许多人看到的现象。AI 在提升非专家水平方面的作用尤其关键,因为它可以重塑整个组织的结构。组织不应再招聘行政能力(因为这些将更加自动化),而应专注于招聘批判性思维能力,因为初级员工可以借助生成式 AI 推动工作进展。这反过来会影响中层管理。Ethan 和团队的工作太棒了!

This is the single best AI-measuring test I've read all year, quantifying what so many of us have seen. The impact of AI in leveling up non-experts in a discipline is especially critical in that it can reshape how entire organizations are structured. No longer should organizations be hiring for administrative abilities (as those will be more automated), but instead should focus on hiring for critical thinking, since those entry level folks can move the ball down the field with the help of generative AI. That in turn will impact middle management. Amazing work, Ethan and team!

互动版:图/公式 + 针对本篇提问 →