梁文锋 4 小时投资人会议实录:持续说不,DeepSeek 不做下一个 BAT

DeepSeek's Consistent 'No' to Being the Next BAT: A 4-Hour Investor Meeting

梁文锋 Wenfeng Liang · DeepSeek · 2026-05-20 · 投资人会议实录 ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

录音:5 月 20 日 · 整理:2026 年 7 月 16 日_音频时长约 3 小时 44 分 钟_

Recording: May 20 · Organized: July 16, 2026_Audio duration about 3 hours 44 minutes_

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 3)

全文 · Full text(逐段中英对照)

录音信息 Recording Info

录音:5 月 20 日 · 整理:2026 年 7 月 16 日_音频时长约 3 小时 44 分 钟_

Recording: May 20 · Organized: July 16, 2026_Audio duration about 3 hours 44 minutes_

全文转写(上) Transcript (Part 1)

以及我们公司的其他同事,我们一开始来做这个公司,初衷是没有想到说我最后要赚多少钱,要到资本市场上去,要上市,要怎么样的,所以我们是没有这个初衷的。

As for the other colleagues in our company, when we first started this company, our original intention was not about how much money we would eventually make, or going to the capital market, or going public, or anything like that. So we didn't have that intention from the start.

最开始的几十个人完全没有这么想过。如果他这么想,他就不会来。所以总体讲,我们是怀着一个对这个世界非常大的善意来做这个事情,然后我们觉得这是对人类有用的,这是 一个金钱以 外的事情。

The first few dozen people never thought about it. If they had, they wouldn't have come. So overall, we started this with great goodwill toward the world, and we felt this was something useful for humanity, something beyond money.

当然,到了后面,这个事情发现利益非常大之后,又有其他的诱惑,这个是另外的事情。但是我们出发的初衷、我们的愿景,以及我们保持到现在的这个愿景,不是按照一个商业利益最大化的方式来做的 。我觉得这 点比较关键。

Of course, later on, when it became clear that the benefits were huge, other temptations emerged—that's a different story. But our original intention, our vision, and the vision we maintain to this day were not built around maximizing commercial interests. I think this point is crucial.

我大概二十年前的时候,在管理上我最崇拜的是杰克·韦尔奇,就是 GE 的前 CEO。现在回来看,他说的大部分东西可能都已经不对了,但是他最重要的一点说对了:一个公司最重要 的是它的愿景。

About twenty years ago, the person I most admired in management was Jack Welch, the former CEO of GE. Looking back now, much of what he said may no longer hold true, but one of his most important points was correct: the most important thing for a company is its vision.

管理一个大公司,靠的不是你的规章制度,靠的是愿景。愿景是什么呢?愿景不是挂在墙上的标语,愿景是你怎么做,不是怎么说,就是你怎么实际运行。反正杰克·韦尔奇说的原话我忘了,大概是这个意思。

Managing a large company doesn't rely on rules and regulations; it relies on vision. What is vision? Vision is not a slogan hanging on the wall; vision is how you act, not what you say—that is, how you actually operate. Anyway, I've forgotten Jack Welch's exact words, but that's the gist.

所以,我们是怎么管理这么多人、怎么组织起来的?其实我们没有组织,就是愿景驱动的,靠一个愿 景来组织。我们是没有组织的。

So how do we manage so many people and organize ourselves? Actually, we don't have an organization—it's vision-driven, organized by a shared vision. We have no organization.

这个有好处,也有坏处。未来我们会想办法扬长避短,但这个是我们的特色。我们并不是以一个“我要实现什么 KPI、没 有考核”的方式来做,只有愿景。

This has both advantages and disadvantages. In the future, we will find ways to leverage the strengths and overcome the weaknesses, but this is our characteristic. We don't operate in a way that says 'I need to achieve some KPI' or 'there's no assessment'; we only have a vision.

这个愿景甚至也不是成文的,并不是写出来的,没有写出来过任何东西。这个愿景是在我们做事情的方法、我们对待这个世界的态度里。可能我们公司每个人对这个愿景的理解也不一样,可能每一个人的愿景也是有差别的 ,但是在一 个大的方向上是一致的。

This vision isn't even written down—it hasn't been documented anywhere. This vision is embedded in the way we do things and our attitude toward the world. Maybe everyone in our company has a different understanding of this vision, and perhaps each person's vision differs, but the broad direction is aligned.

我觉得还是怀着对这个世界非常大的善意,然后想做一 点事情。我们是用这个来组织起来的。

I think it's about having great goodwill towards the world and wanting to do something. This is what unites us.

接下来我先讲,讲完之后大家再提问。我可能会围绕着这个愿景来讲后面的事情。这个愿景是真的,不是编出来的,我们是真的这么想,真的这么做的。否则 你没法解释我们的很多事情。

Let me speak first; you can ask questions afterwards. I'll probably talk about the vision and what follows. This vision is real, not fabricated—we truly think and act this way. Otherwise, you can't explain many of our actions.

这个愿景为什么我们这么坚持开源?因为这个愿景本身就要求开源。你没有这个愿景,你没法把人组织起来。

Why do we insist on open source? Because the vision itself requires open source. Without this vision, you can't bring people together.

比如说,智谱也开源,但是智谱的开源跟我们的开源不一样。智谱的开源有一种被迫的感觉,他们觉得这不是本意,但是对我们来讲,这就是我们的本意。

For example, Zhipu also opens sources, but their approach differs from ours. Zhipu's open source feels forced—it's not their intention. For us, it is exactly our intention.

然后在开源这个事情上面,我们一开始就想得非常清楚。首先,第一个就是愿景;第二个,我们认为要把 AI 这个事情在商业上做成,开源是有好处的。

Regarding open source, we had a clear understanding from the start. First, the vision; second, we believe that for AI to succeed commercially, open source is beneficial.

这听起来有点矛盾,有点违反直觉,因为在历史上,开源跟商业化都是有冲突的。但我觉得 AI 跟以前不一样。因为在历史上,一个软件公司,它的市场一年可能就几十亿美金,离开源了它就没有了,可能就只剩下几千万或几亿美金了。

This sounds somewhat contradictory and counterintuitive, since historically open source and commercialization have been in conflict. But I think AI is different. Historically, a software company's market might be worth a few billion dollars a year—lose the source and it shrinks to a few tens or hundreds of millions.

但是 AI 这个事情足够大,最终它可能得占掉人类社会 GDP 的百分之十,那么它其实是一个非常大的数字。一个人独占这个事情,你不可能独占这个事情,你一 定得跟别人分享,否则你肯定活不下来。

But AI is huge; it could eventually account for 10% of human society's GDP. That's an enormous number. You can't monopolize it—you must share, otherwise you won't survive.

这跟之前开源做一个软件可能是不一样的,因为那个软件的市场就没那么大。但是 AI 这个事情实在太大了。如果说我们想独占这个利益,那么一定是要被历史抛弃的。 我觉得最主要这是一个客观的规律,这是一个历史观。

This is different from previous open-source software, where the market wasn't that big. But AI is too vast. If we try to monopolize the benefits, history will abandon us. I think it's mainly an objective rule, a historical perspective.

并不是说我不开源,我就能够独占这个市场,这在理论上就不符合客观事实。你一定会遇 到很多阻力,一定会有其他的方法阻止你实现这个目标。

It's not that by not open-sourcing, I can monopolize this market; this is theoretically inconsistent with objective facts. You will inevitably encounter many obstacles, and there will always be other ways to prevent you from achieving this goal.

在这种情况下,我觉得不一定要按照传统的商业思维。你需要有一套机制来确保你自己能够获得的利益是有限的,那么你才有可能做成。需要克制,我觉得是需要克制。

In this situation, I think we don't necessarily have to follow traditional business thinking. You need a mechanism to ensure that the benefits you can obtain are limited, then you may succeed. Restraint is necessary; I think restraint is necessary.

如果说我们想在我们手上把 AI 这个事情做成,首先第一个,我觉得是需要克制的。不能够想着人类 GDP 的百分之多少都归我了,或者说中国 GDP 百分之多少都归我了。你越这么想,越做不成。

If we want to make AI happen in our hands, first of all, I think restraint is needed. You cannot think that a certain percentage of human GDP will belong to me, or a certain percentage of China's GDP will belong to me. The more you think that way, the less likely you are to succeed.

所以我们从一开始就觉得需要克制。你越克制,你越有可能能够做成这层事。这是一个商业上的考量,当然这是一个宏观的考虑。

So from the very beginning, we felt that restraint was needed. The more restrained you are, the more likely you are to succeed in this matter. This is a business consideration, and of course, it's a macro-level consideration.

我觉得这是符合直觉的,至少是符合我的直觉,或者说至少我是真的这么想的。我们并没有非常多的其他优势,我们没有什么本事,我们并没有比别人有 钱,也没有说我们人员比其他公司更好,其实没有的。

I think this is intuitive, at least to me, or at least I truly think so. We don't have many other advantages; we don't have extraordinary abilities; we are not richer than others, nor are our personnel better than other companies. In fact, we don't.

你想,我们两年前成立这个公司的时候,我们又没有很多钱,又没有很多卡,又没有什么知名 度 ,又没有什么号召力,我们就是一群非常平凡的人。

You see, when we founded this company two years ago, we didn't have much money, many GPUs, much visibility, or much influence. We were just a group of very ordinary people.

我真的就是一群平凡的人。如果说喜欢的一个叙事是一群平凡的人做出了不平凡的事情,而不是一群天才做出了不平凡的事情,这个跟我们的克制是很有关系的,跟我们的克制、跟我们的愿景是一脉相承的。

We really are just a group of ordinary people. If there is a narrative that a group of ordinary people accomplished extraordinary things, rather than a group of geniuses, this is very much related to our restraint, and it is consistent with our restraint and our vision.

那么,开源跟商业化会不会有冲突?我觉得在 AI 这个事情上面,你不克制,你就不起。开源是属于克制的一部分,那么我们的克制不仅表现在开源上面,我们还表现在很多方面 上。但总体上来讲,我们是不用考虑开源这个事,不用考虑克制这个事情。

So, would open source and commercialization conflict? I think in AI, if you are not restrained, you won't succeed. Open source is part of restraint. Our restraint is not only reflected in open source but also in many other aspects. But overall, we don't need to consider open source as a separate issue, or consider restraint as a separate issue.

你越克制,可能就越容易做成,或者说至少到目前为止是印证的,到目前为止是能解释得通的。否则没有办法能够解释为什么我们能够做成:我们并没有什么武器,起点又非常低,资源又非常少,我们的人其实也就是随机的一群平凡的人。我自 己也就是一个大学毕业的学生,也不是最顶级的那个学校毕业的。

The more restrained you are, the more likely you are to succeed, or at least that has been the case so far and can be explained. Otherwise, there would be no way to explain why we have succeeded: we had no special weapons, a very low starting point, very few resources, and our people were just a random group of ordinary individuals. I myself am just a college graduate, not even from a top-tier university.

这个克制也是我们愿景里的一部分。AI 这个事情太大了,利益太大了。我们非常克制,只要能够做成,最后利益都会非常大。你随便分一点,利益就非常大,所以现在根本不用考虑拿这里面的哪 一部分利益、怎么拿,我觉得根本不用考虑这件事情,因为这个利益足够大了。

This restraint is also part of our vision. AI is enormous, and the potential profits are enormous. We are very restrained; as long as we can succeed, the eventual profits will be huge. Even a tiny slice will be enormous, so there's no need to think about which portion to take or how to take it. I think we don't need to consider that at all, because the profit is large enough.

你只要分一点点,就已经很足够了。所以我们之前说,我们只赚取一个合理的利润,只看你的意愿,而不是利润之大,这个是不一样的。这不是我们的 API 定价。我们 API 定价觉得一个合理的利润,大概是我们到市场上买一批设备回来,十个月收回成本,我觉得这是一个合理的利润。

You only need a tiny portion, and it's already enough. So we said earlier that we only aim for a reasonable profit, driven by our intention, not by the size of the profit — that's different. This is our API pricing logic. We believe a reasonable profit is roughly recovering the cost of purchasing a batch of equipment from the market within ten months. I think that is a reasonable profit.

在当前的情况下,考虑到你有风险、还有前期的投入等等,如果一个服务器我们在财务上按照三年或者五年的摊销,但在商业上,我们觉得大概十个月收回成本,我们觉得够了,OK,我们觉得够了。所以这个是我们现在 API 定价的逻辑。我们的 V3.2 Flash 跟其他的,都是十个月收回设备的成本。

Under current circumstances, considering risks and upfront investments, even though we might amortize a server over three or five years financially, from a business perspective we think recovering costs in about ten months is enough. OK, we think that's enough. So that's the logic behind our current API pricing. Our V3.2 Flash and other models all recover equipment costs within ten months.

这就是我们的标准。它其实不是利润最大化的,如果利润最大化,应该把价格设得更高。因为在这个价格区间,用户的需求是没有弹性的, 就是我价格再翻一半,或者说我价格再抬高一倍,token 的消耗量区别不大的。

That's our standard. It's not profit-maximizing. If we were maximizing profit, we would set a higher price. Because in this price range, user demand is inelastic; even if we double the price, the token consumption would not change much.

我跟你们讲个故事,就是我们那个 DDCP,是我们那个模型。一开始我们担心需求太多,所以一开始把价格定得比较高,团队里大家不是很高兴。后来我把价格又降下来了,降到四分之一,大家就很开心。

Let me tell you a story: our model, DDCP. At first, we were worried about too much demand, so we set a relatively high price. The team wasn't very happy. Later, I lowered the price to a quarter, and everyone was delighted.

我觉得这个才是我们真实的想法。就是我前面说的愿景,我们还是让这个东西对人是有用的,而不是我们赚最多的钱,而是我们在能够赚到合理利润的情况下,大家都用得起。我觉得这次我们公 司其他人的想法 ,就是我们那时候降价的时候,公司群里面很多人是欢呼的,大家都觉得很开心。

I think this is our true intention. As I said earlier in the vision, we want this to be useful to people, not to make the most money. Under the condition of earning a reasonable profit, everyone can afford it. I think that's how others in our company felt at that time—when we lowered the price, many people in the company group cheered, and everyone was very happy.

因为这是我们花了这么多精力,这么用心把这个模型做好的目的。目的就是能够非常便宜、效果非常好,能够让大家都能够充分地用。我们就觉得这很开心,这是我们的动力, 这是我们的愿景 ,这是我们公司能够凝聚到一起去做这个事情的共识吧,这是我们公司内部的共识。

Because this is the purpose of our hard work and dedication in building this model. The goal is to make it very cheap, very effective, and accessible for everyone to use fully. That makes us happy; it's our motivation, our vision, and the consensus that brings our company together to do this. This is the internal consensus within our company.

这点应该是比较特殊的,因为降价这样对于我们其他的竞争对手来讲,肯定不是个好事,他们一定不是欢呼的。 因为你的收入、你的 ARR,那个你降一半,ARR 就掉一半。对,这是一个我们不一样的地方。

This point should be rather unique, because lowering prices is certainly not a good thing for our other competitors; they definitely wouldn't cheer. Because your revenue, your ARR, if you cut it in half, the ARR drops by half. Yes, this is where we are different.

我们觉得这就够了。从我们公司内部来讲,我十个月收回成本,这个商业上我已经非常满意了。对于公司外部来讲,我 们也觉得这个价 格是大家比较开心、比较乐意看到的,大家是双赢的,公司跟社会、跟所有人都双赢的。

We think that's enough. From within the company, recovering costs in ten months, I'm very satisfied with this business model. Externally, we also think this price is something everyone is happy and willing to see; it's a win-win for all, the company, society, and everyone.

我觉得,OK,刚刚有人在屏幕上留言说,十个月回本这个利润太高了。确 实是还有降价的 空间,确实还有降价空间。在模型的优化上也还有空间,所以整体降价空间还是比较大的。

I think, OK, someone just left a message on the screen saying that a ten-month payback period means the profit margin is too high. Indeed, there is still room for price reduction, indeed there is. There is also room for model optimization, so overall the room for price reduction is still relatively large.

但是这个成本,十个月回本,我们自己能做到,其他家做不到。像 可能阿里或者腾 讯,他不是我们的优化,他的成本应该是要比这个高好几倍的。这里还是有很多优化的工作。

But with this cost, a ten-month payback, we can achieve it, but others cannot. For example, Alibaba or Tencent, they don't have our optimization; their costs should be several times higher. There is still a lot of optimization work to be done.

刚刚说我们为什么不继续降价,是因为它没有弹性。就是我再降价,需求不会更多了,或者我再降价,需求增 加得很少了。因为这个价格所有人都用得起了,大家都觉得这个价格是满意的,不会因为这个价格贵而不用了。

As for why we don't continue lowering prices, it's because there's no elasticity. Even if I lower the price further, demand won't increase much, or the increase will be very small. Because at this price, everyone can afford it, and everyone is satisfied with the price; they won't stop using it because it's too expensive.

所以降价首先公司不会有更多的收入,对社会来讲也没有更多的价值,因为这个价格大家都满意了。你价格再低,对这个社会的幸福感也没有增加很多 。对,OK,不过刚刚这个问题上,就是我们在定价这个事情上面,我们肯定不是以 公司收入最高或者利润最高,不是最出发点的。

So lowering prices first of all won't bring the company more revenue, nor will it bring more value to society, because everyone is already satisfied with this price. Even if you lower the price further, it won't significantly increase society's happiness. Yes, OK, but on this issue, when it comes to pricing, we certainly do not take maximizing company revenue or profit as our starting point.

这是我们克制的一部分,因为从短期 来讲,你价格高一点,你可能收入多一点;但从长期来讲,还真不好说。因为我觉得克制是一种策略。

This is part of our restraint, because in the short term, if you set a higher price, you might have more revenue; but in the long term, it's hard to say. Because I think restraint is a strategy.

对我来讲,克制是一种战略。就在于有时候你可以舍 弃一些,来换更多其他的东西。不开源这个事,其实也是一样的,也可以认为是我们的压力,也可以认为是我们的让利。

For me, restraint is a strategy. It's about sometimes giving up something in exchange for more of other things. Not open-sourcing is actually the same; it can be seen as our pressure or our concession.

首先,这个让利对于我们公司内部来讲,我们很高兴,大家都很开心,员工觉得很有成就感,我们会因此有凝聚力。以及这个让利对社会是有好处的,社会也很高兴,其 他同行或者说普 通人都会很高兴。所以这种克制,我理解是这种克制从长远上来讲,能够增加我们做成 AGI 的概率。

First, this concession makes our company internally very happy—everyone is delighted, employees feel a strong sense of achievement, and it fosters cohesion. Moreover, this concession benefits society, making the public happy as well, including our peers and ordinary people. So, I understand this kind of restraint as something that, in the long run, increases our probability of achieving AGI.

在考虑一件事的时候,我是毫不怀疑 AGI 会有非常大的商业价值。那么在 这个基础上,我 优先考虑的不是我怎么多加一点份额,我怎么多拿一些份额,我优先考虑的是我怎么增加我能够做成的概率。

When considering any matter, I have no doubt that AGI will have tremendous commercial value. On that basis, my priority is not how to increase my share or grab a bigger piece, but rather how to increase the probability of making it happen.

这个克制可能还体现在很多其他的方面。比如说,我们去年春节用户突 然很多,但是我 们并没有去追求我要留这些用户,或者说拿这些用户来变现,或者说我要去抢这些商业利益,在用户上面兑现。

This restraint might also manifest in many other aspects. For example, during last year's Spring Festival, we saw a sudden surge in users, but we did not pursue retaining those users, monetizing them, or seizing commercial benefits by cashing in on them.

我们没有去抢用户,没有去赚钱,但我们很努力想办法把用户服务好。我们并不会有这样 的想法,说我要做成下一个超级 App,然后我要去跟谁竞争,我要做成下一个字节、做成下一个腾讯,完全没有这样的想法。

We did not chase users or money, but we worked hard to serve our users well. There was never any thought of becoming the next super app, competing with others, or becoming the next ByteDance or Tencent—none of that at all.

我们是可以这么做,但是我们没有这么做。我的理 解是,这也是克制的一部分。你不要想什么都要赚了,你好像有了用户之后,好像就能够做成下一个字节了,然后你就把那个吃了。

We could have done so, but we did not. In my understanding, this is also part of restraint. You don't have to try to profit from everything, nor should you think that having users means you can become the next ByteDance and then just swallow it all up.

我觉得这个在商业上是行得通的,是有可能的。如果去年时候我们用了大笔钱,就跟字节去抢 用户,也是一种 打法。但是我们选择是一种非常克制的做法,就是我不跟你去争这个东西,因为后面还有西瓜,前面的可能都是芝麻。

I think this approach is commercially viable and possible. If we had spent a lot of money last year to compete with ByteDance for users, that would have been one strategy. But we chose a very restrained approach—not contending for these things, because there might be a bigger watermelon later, and what's in front might just be sesame seeds.

我不应该什么芝麻都抢了。当然,可能这个芝麻比较大,但我觉得后面的 AI 可能前面 都不算大。现在 来看,去年我们没有在 C 端去发力,可能是对的。因为能看到后面真的是有更大的西瓜,前面真的只是一些小芝麻。

I shouldn't grab every sesame seed. Of course, maybe some sesame seeds are big, but I believe the AI opportunity ahead far surpasses everything that came before. Looking back, it might have been right not to focus on the consumer side last year, because you can see there's truly a much bigger watermelon behind, and what's in front is just small sesame seeds.

如果去年我就有很多钱,然后把这个事情做得非常大的话,有什么好处?你并没有得到什么 东西。这些是我的真实想法,因为我觉得后面的 AGI 机会应该是非常大的,后面的 AGI 机会永远是非常大的。

If I had had a lot of money last year and scaled this up significantly, what would have been the benefit? Nothing. These are my genuine thoughts, because I think the AGI opportunity ahead is enormous—it will always be enormous.

我甚至不用考虑我到时候在里面是占一个位置,或者说在那里面我的商业模式是什么,我们根本就不用考虑。只要有那么大的商 业机会,你一定是有办法的。那么前面的芝麻,我们也会捡,但是我们就随手捡一点,并不会说我停下来,或者说我把它当成重要事情来做。

I don't even need to consider whether I'll occupy a position there or what my business model will be inside it—we don't need to think about that at all. As long as there is such a huge business opportunity, you will definitely find a way. As for the small gains in front of us, we'll pick them up casually, but we won't stop or treat them as important tasks.

所以去年的 C 端日活这些,我觉得可能就是个小事。但我们也捡了,我们也用一个比较低 的成本维持了用 户的使用,因为有可能以后是有用的。虽然现在不知道这用户有什么用,现在它是一个纯成本的支出,但是以后可能是有用的。

So last year's daily active users on the C-end—I think it might be a trivial matter. But we still picked it up, and we maintained user engagement at a relatively low cost, because it might be useful in the future. Although we don't know what use these users will have now, and it's currently a pure cost, it might be useful later.

既然是随手能够拿到的,我们就会顺手拿到。包括今年来看,很有可能我们在 API 或者 AI 方面的 ARR 收入, 也是有一个机会的。就是如果这个需求可以继续扩大,如果继续扩大,显卡可以买到更多的显卡,那么 ARR 做到几个亿美金是很有可能的。

Since it's something we can get easily, we'll grab it casually. Looking at this year, there's a good chance that our ARR revenue from API or AI could also have an opportunity. If this demand continues to expand, and if we can buy more GPUs, then achieving an ARR of several hundred million USD is very possible.

如果说 AI 能做到十亿美金的话,那么基本上我公司的现金流可能就能 够回正了,能够 cover 我的研发费用,能够 cover 我的所有费用。所以说这个也是有可能的,但是我们没有把它当一个优先来做。

If AI can reach one billion USD, then basically my company's cash flow could break even, covering our R&D costs and all expenses. So this is also possible, but we haven't prioritized it.

这个我们会做,但我觉得这个是个重要事情。它不是我们第一优先考虑的,或者不是我们今 天真的关心的。更 大 的机会应该还在后面,前面的机会包括去年的 C 端、今年的 B 端,我觉得这事要做,要把它做好,但这不是我们 的目标。

We will do this, but I think it's an important matter. It's not our top priority, or not what we are really concerned about today. The bigger opportunities should still lie ahead. The opportunities in front include last year's C-end and this year's B-end. I think we should do this and do it well, but it's not our goal.

或者说,我们公司大部分人不认为这是一个非常重要的事情,不认为这是一个跟 AGI 比较起来同等重要的问题。

Or rather, most people in our company don't consider this a very important issue, nor an issue as important as AGI.

开源那个可以多说一下,因为之前有很多问题问的都是开源。首先,我觉得我们是会开源的, 然后我们最强的 模型可能也是会开源的。因为我看不到闭源什么好处,看不到必然的好处。字节它的模型是闭源的,它有什么好处?我看不到有什么好处。

We can talk more about open source, because many previous questions were about open source. First, I think we will open source, and our strongest models may also be open sourced. Because I don't see any benefit in being closed source—no inherent benefit. ByteDance's models are closed source—what benefit does that bring? I don't see any advantage.

哪怕是模型开源,你把所有东西都告诉别人,这个门槛也非常高。别人要用起来,这个门槛也非常高。他要用起来,就很难;其次,他要用起来,还要成本做得很低,也很难很难,没有那么容易。

Even if the model is open source and you tell everything to others, the barrier is still very high. For others to use it, the barrier is very high. It's hard for them to get started; moreover, for them to use it and keep costs low, it's also very, very difficult—not that easy.

并不是我开源了,他就能够轻易地做到跟我一样的部署成本。这里边还是有很多工作要做的。虽然说这些工作原理都明白,但是不是每一家公司都愿意,或者都有这个意愿和能力来组织人力去达到这个目标的。

It's not that just because I open source it, others can easily achieve the same deployment cost as me. There is still a lot of work to be done. Although the principles are understood, not every company is willing or has the will and ability to organize manpower to achieve this goal.

这点我也很习惯。它可能就不擅长做这个事情,因为它阻力太大了。它要控制这个成本很难, 它有很多管理上的 以 及物理上 的制约。这也是 创业公司的优势,因为创业公司如果太小的话,你没有这个力量来做这个事情;如果你是大公司的话,你很难组织。

I'm also used to this. It may not be good at this because the resistance is too great. It's hard for it to control costs due to many management and physical constraints. This is also an advantage of startups. If a startup is too small, it lacks the power to do this; if it's a large company, it's hard to organize.

这个事情都有难点,所以这是属 于我们这个规模的公司的一个甜区,sweet point。然后如果我们更大了,可能我们没有其他问题;如果更小的话,朋友们力量又会不足。

Every option has its difficulties, so this is the sweet spot for a company of our size. If we were larger, we might face other problems; if smaller, our friends would lack sufficient strength.

所以,至于开源,我觉得我们应该定价。 目前我们应该认识到,不会逼人。因为定价的模型,我也不会收一个非常高的费用,我也是可能按照十个月回本来收这个费 。按十个月回本,就已经能 够把竞争对手……

So, regarding open source, I think we should set a price. At present, we should recognize that it won't force anyone. Because in our pricing model, I won't charge a very high fee. I might charge based on a ten-month payback period. With a ten-month payback, we can already make competitors...

我按十个月回本这个,就能够让独立 部署的第三方无利可图。 第三方他做不到,他做不到这个成本, 他肯定做不到 。

With a ten-month payback, we can make independent deployment unprofitable for third parties. They cannot achieve this cost; they definitely cannot.

就是开源,我觉得对我们的商业模式是没有任何影响的。前提是我们只赚六倍的利润,十个月收回成本,大概对应的是六倍的利润。我们只赚六倍利润的情况 下,开源是不会有什么影响的。

Open source, I believe, has no impact on our business model. The premise is that we only take six times the profit, recouping costs in ten months, which corresponds to about six times the profit. Under the condition of only six times profit, open source will not have any effect.

但如果说你要赚一百倍利润,那么开源确实会影响你赚一百倍利润,因为第三方会部署,他可能是二十倍成本,就比你低了。

But if you want to make a hundred times profit, then open source will indeed affect that, because third parties can deploy and their cost might be twenty times, which is lower than yours.

这个模式是不是长 期 可以持续的?我觉得是可以的。就在我们这个愿景下,我觉得开源是长期可以持续的,或者我们 是打算这么做的。 你 可以认为,就是有克制,也是有让你比较长利。

Is this model sustainable in the long run? I think it is. Under our vision, I believe open source can be sustainable in the long run, or we intend to do this. You can think of it as being restrained, yet also allowing for long-term benefits.

这个策略是能够在技术前面让我们 有更多的机会, 我们能够做成 AGI 的概率也是更大的。我们更从容。

This strategy allows us to have more opportunities ahead of the technology, and our probability of achieving AGI is also greater. We are more relaxed.

你想,我们根本都不用加班,因为就没那么难。但是对其他家来讲,可能就很难,因为你想的也太多了。

You see, we don't even need to work overtime because it's not that difficult. But for other companies, it might be very difficult because they have too many things to consider.

其实不难的,根本就不难。才开始是一个……外面可能看起 来,我们选了一个 很 难的模式,我们要做研究,我们去做最难的事情,好像是个 hard 的模式。但其实我们在外地方舍弃了很多 ,使得我们还是 非常有力的,我们还是做得非常轻松的。

It's actually not difficult, not difficult at all. At first, it may seem from the outside that we chose a very hard mode—doing research, taking on the hardest tasks, like a hard mode. But in reality, we have given up many things elsewhere, which makes us still very powerful and we can do it very easily.

所以对开源这件事情,实际上我的判断是,它是持续的。开源和商业付费之间没有冲突,前提是六倍利润的情况下没有冲突的。

So regarding open source, my judgment is that it is sustainable. There is no conflict between open source and commercial payment, provided that there is a six-fold profit margin, there is no conflict.

六倍利润看起来很高,但其实不高。因 为现在 AI 的 效 率这么高的情况下,现在一个合理利润可能就是这么多。未来它可能会降到,比如说四倍、三倍,我觉得这已经是……再到底也不可能再降 了。但就是它还是 会 有很大的利润。

A six-fold profit margin seems high, but it's actually not. Given the current efficiency of AI, a reasonable profit margin might be just this much. In the future, it may drop to, say, four-fold or three-fold, and I think that's already the bottom; it won't go lower. But it will still have significant profits.

就光是看卖 API 这个事情,但我并不觉得卖 API 这个事 情有那么大吸引 力。但是就这个事情,你可以说明,没有,我没有看到有什么冲突。

Just looking at selling APIs, I don't think selling APIs is that attractive. But regarding this matter, it can be said that no, I don't see any conflict.

然后我也不担心别人部署我们的模型,然后跟我们来竞争,一点都不担心。我们还希望他们能够部署起来。

And I'm not worried about others deploying our models and competing with us—not at all. In fact, we hope they can deploy them.

我们尽可能给开 源社区提供帮助, 协助大家能够把我们的模型部 署起来。我不担心他会跟我抢这个生意,因为这个市场足够大。我只担心他部署不起来,他有些细节没做对,效果变得比较差,或者说他的成本会比较高。

We will do our best to help the open-source community, assisting everyone in deploying our models. I'm not worried that they will take away our business, because the market is big enough. I'm only worried that they might fail to deploy, get some details wrong, resulting in poor performance or higher costs.

然后去年的时候,我问的,去年有 To B 这个生意的时候会问得比较多的是:我们那个 C 端,我开源, 那么 C 端是不 是 跟我这个 C 端又会冲突 了呢?因为我没有 流 量的优势,或者说腾讯自己流量很多,他部署我们的开源模型,他就把所有的 C 端用户都接过去了,就把我的 C 端用户都抢走了。

Then last year, I was asked—when there was the To B business, the question that came up more often was: our C-end, I open source it, so will the C-end conflict with this C-end? Because I don't have the advantage of traffic, or Tencent itself has a lot of traffic. They deploy our open-source model and take all the C-end users, stealing my C-end users away.

然后问题是:我们给的开源模型,跟我们自己部署的模型是不是一样?是一样的。我们不会说开源一个差点的模型,然后我们自己部署的时候 用一个更好的模型,是不会的,是一样的。

Then the question is: Is the open-source model we give the same as the model we deploy ourselves? It is the same. We won't open source a worse model and then deploy a better one ourselves—no, it's the same.

这也说明它其实没有冲突。我们去年一整年,在 C 端我基本就是开源了,然后 C 端的服务没有看到冲突,真的没有 看到冲突。所以,这是开源的这部分。

This also shows that there is actually no conflict. Over the past year, on the C-end, I basically open-sourced it, and the C-end service has seen no conflict—truly no conflict. So that's the open-source part.

然后下面还有就是,公司的长期 vision,我觉得我们目标应该是 AGI。每个人对 AI 定义不一定一样,但是不妨碍我们把 AGI 当做我们的目标。

Then next, regarding the company's long-term vision, I think our goal should be AGI. Everyone's definition of AI may not be the same, but that doesn't prevent us from taking AGI as our goal.

从技术路线上来看,其实 AGI 这条 路线图是比较清 晰的。跟目前的这一代 AI 技术,如果说你能够把一个问题描述得很清楚,给它完整的上下文和指令,它已经超过人类了。但这里有个定义,有个前提是:你给它完整的上下文,你给它完整的指令。

From a technical roadmap perspective, the path to AGI is actually quite clear. Compared to the current generation of AI technology, if you can describe a problem very clearly, give it complete context and instructions, it already surpasses humans. But there is a definition, a premise: you give it complete context, you give it complete instructions.

然后这 个定义是很难达到 的。比如说我们今天开会,其实我们前面有很长很强的上下文,可能大家有几十年的上下文,然后这是 AI 不具备的。那么现在 AI 能具备的是,在一个局限的上下文里面,它能做得 比人更好。

And this definition is hard to achieve. For example, in our meeting today, we actually have a long and strong context beforehand—maybe decades of context for everyone—and this is something AI does not have. What AI can do now is, within a limited context, perform better than humans.

但它还是替代不了人类。这里面还差一个地方,是持续学习。因为人也能够持续学习。你招一个员工,他可能花两个月时间来熟悉这公司的环境,熟悉 他的工作,他谈 两个月的,你这样他就可以上手了。

But it still cannot replace humans. There is still one missing piece: continuous learning. Because humans can also learn continuously. When you hire an employee, he may spend two months getting familiar with the company environment and his work. After those two months, he can get started.

他就能做很多事情。他能够听懂你说的话,比如说你说,叫那个小王来,他就知道小王是谁。但如果 AI 的话,因为这个上下文,他前面没有这两个月的学习。

He can do many things. He can understand what you say—for example, if you say 'call that Xiao Wang,' he knows who Xiao Wang is. But for AI, because of this context, it lacks the prior two months of learning.

你跟他说,叫小王过来, 你得告诉他小王是谁、什么职务、他在哪里、要怎么去找他、要找他的时候注意什么。你得给 AI 所有 的上下文。那么在这 种情 况下 ,AI 是能做的 , 但是你不可能给他所有的上下文,也不现实。

You tell it to call Xiao Wang over. You have to tell it who Xiao Wang is, what his position is, where he is, how to find him, and what to be aware of when looking for him. You have to give the AI all the context. In that case, the AI can do it, but it's unrealistic to give it all the context.

所以 AI 并不能够替代你的员工。但如果说 AI 具有持续学习的能力,它跟你的员工一样,到公司学习两个月,那么

So AI cannot replace your employees. But if AI had the ability to continuously learn, like your employees learning at the company for two months, then...

AI 的发展, 我们可以理解成它是一个阶梯。去年走的阶梯是 CoT,就是思维链。因为我们发现,通过思维链的方式,可以让智能达到一个更高的水平。通过它自己思考,可以让这个上限,可 以 让这个 AI 能做更多的事情。

The development of AI can be understood as a ladder. The step last year was CoT, that is, chain of thought. Because we found that through chain of thought, intelligence can reach a higher level. Through its own thinking, this upper limit can be raised, allowing AI to do more things.

那么我们又跨过了一个阶梯。今年的阶梯就是 Agent,因为我们发现,用 Agent 的方式,即便多有事情也可以做,它的 能力范围会更大,它的智能上限会更高。

Then we've stepped up another rung. This year's step is Agent, because we found that with the Agent approach, even many things can be done, its capability scope is larger, and its intelligence ceiling is higher.

为什么是阶梯呢?因为后面的每一步都是基于前面的基础上的。Agent 要用到 CoT,然后 CoT 也要用到前面的阶梯,前面的阶梯就是语言模型,所以它并没有一步是白走的。

Why a ladder? Because each subsequent step is based on the previous one. Agent uses CoT, and CoT also uses the previous step, which is the language model. So no step is wasted.

所以,AI 的发展, 智能的走向是有迹可循的。说今年走的这个阶梯是 Agent,但是 Agent 这个阶梯,它也是会走完的。就是它把所有可能解决的问题都解决完之后,但是它还是不能替代你的员 工,但是它已经 到它能力的上限了。

So, the development of AI and the direction of intelligence have a traceable pattern. It is said that this year's step is Agent, but the Agent step will also reach its end. That is, after it solves all possible problems, it still cannot replace your employees, but it has reached the upper limit of its capability.

就好像 CoT 一样,CoT 到它的上限之后,它已经超过最顶尖的人类了,在做奥数题、写程序上面已经超过最顶尖的人类了。但是它还是停在那个地方,它那个技术并没有能够到达 AGI。

Just like CoT, after CoT reached its upper limit, it has surpassed the best humans in solving Olympiad math problems and writing programs. But it still stays there; that technology has not reached AGI.

所以你看,AI 这个智能的走向,它是有迹可循的。然后在 Agent 之后,我们觉得应该要解决的问题是持续学习,就是怎么让模型可以持续地学习,而不是说你要给它一个很强的训练,它应该能够像人一样做一个比较长时间的持续学习。

So you see, the direction of AI intelligence is traceable. And after Agent, we think the problem that should be solved is continuous learning, that is, how to make the model learn continuously, rather than giving it a strong training; it should be able to engage in long-term continuous learning like a human.

这个问题跟完成任务等等是同一回事,它们相关,解决的是同一个问题。我们现在站在 Agent 这个地方,能看到的是下一个瓶颈,就是持续学习。下一个要解决的问题,就是解决怎么持续学习,这个是看得到的, 是相对来讲比较明确的 。

This problem is the same as completing tasks, etc. They are related and solve the same issue. From where we stand with the Agent, the next bottleneck we can see is continuous learning. The next problem to solve is how to achieve continuous learning, which is visible and relatively clear.

就是卡在前面的这个障碍,你必须要跨过去,而且一定是有办法能跨过去的,但需要时间。持续学习之后,可能我们就会来到一 个奇点。这个奇点就是, 当这个模型能够持续学习之后,它已经能够做人类能做的所有事情了。

It is the obstacle ahead that you must overcome, and there must be a way to overcome it, but it takes time. After continuous learning, we may arrive at a singularity. This singularity is that once the model can continuously learn, it will be able to do everything humans can do.

它就能够自己开发自己的版本,能够自己再研究,然后开发自 己的下一个版本,开发 更 前的人工智能模型。所以它就会到一个奇点,能够实现自己的迭代。

It will be able to develop its own version, conduct its own research, and then develop its next version, developing more advanced AI models. So it will reach a singularity, achieving self-iteration.

但这个奇点,它并不是一个奇点,它也是一个渐进的过程。这个过程可能 也 是一个比较长的渐变,它不是个突变。但是习惯性地,我们都认为它可能是个奇点。

But this singularity is not a singularity; it is also a gradual process. This process may be a relatively long gradual change, not a sudden mutation. However, habitually, we all consider it a singularity.

因为在很早之前,那些预言家认为这里会有个奇点,但其实它不是一个奇点,它是一个连续的过程。然 后这一步走完之后,我 觉 得才是具身智能。

Because long ago, those prophets thought there would be a singularity here, but in fact it is not a singularity; it is a continuous process. After this step, I think it is embodied intelligence.

这是我们的推测,这是我们觉得这个时间表应该是:先解决学习去学习,然 后再到那个智能的奇点 , 能自我迭代的奇点,然后才是具身智能。到具身智能之后,它就走进现实世界,可以给你做家务,可以给你养老。

This is our speculation, and we think the timeline should be: first solve learning to learn, then reach the singularity of intelligence, the singularity of self-iteration, and then embodied intelligence. After embodied intelligence, it will enter the real world, can do housework for you, and provide elderly care.

我们觉得这是一个比较理想的路线图,但每个人的观点不一样,没有对错之分。只是我们觉得, 这个路线图是最轻松的。

We think this is a relatively ideal roadmap, but everyone's views are different, and there is no right or wrong. We just think this roadmap is the easiest.

这个路线图是因为每一步你要做新的都很少。这个路线图我们可以不用加班。但是如果路线图是反过来的,比如说它要先实现具身智能,那么这里自己做得很累,这是一个很苦的活。我们不希望是这样的路线图,我们希望做得轻松 一点。

This roadmap is because each step requires very little new work. With this roadmap, we don't need to work overtime. But if the roadmap is reversed, for example, if it first achieves embodied intelligence, then it will be very tiring to do it yourself; it's a hard job. We don't want such a roadmap; we want to do it more easily.

如果说 我 们先解决持续学习,再解决那个自我迭代的奇点,再解决具身智能,这个路上就很轻松 。因为到后面之后,你 可 以用前面的技术来帮助开发后面的技术。奇点之后再做具身智能,那个就不用人来做,就不用我们来做了,那个模型自己就可以出来了。

If we first solve continual learning, then the self-iterative singularity, and then embodied intelligence, the path will be much smoother. Because later on, you can use the earlier technologies to help develop the later ones. After the singularity, when we work on embodied intelligence, it won't require human effort—we won't have to do it; the model can do it itself.

所以这个是回答我们的长期目标是什么。我就跟他说,这就 是我们的长期目标,就 是 我们说的 AGI。

So this answers our long-term goal. I told him, this is our long-term goal, which is what we call AGI.

大家回到现实。去年最重要的现实就是,大家都要做 Chatbot,要抢 C 端的流量。那么今年的现实 是,大家都要抢 To B 的收入,要参与到这里面去,因为如果不参与的话,根本就不在牌桌上,对不对?

Back to reality. Last year, the most important reality was that everyone had to build a Chatbot and compete for C-end traffic. This year, the reality is that everyone is rushing to capture To B revenue and get involved, because if you don't, you're not even at the table, right?

但我 们并不认为它是一个重要事情,或者说,在我们公司内部,我们真正关心的,其实是刚刚我说的这些 AGI 的路线图,以及下一步这个技术怎么突破。

But we don't think it's a big deal. Internally, what we really care about is the AGI roadmap I just mentioned and how to achieve the next technological breakthrough.

但是有一点很奇怪的是,最想得到的东西,反而你就得不到。那个并没有那么在意的事情,反而是还蛮容易能够得到。

However, it's strange that what you desire most is exactly what you cannot get. The things you don't care about that much actually come quite easily.

这里边有一个战略上的优势,就是我们心里面想着 AGI,我们做的是 AGI。那我们再做那个应用,再做那个 C 端、B 端的时候,根本就不用太多的心思去 做那个事情,其实花很少的精力就可以了。

There's a strategic advantage here: we keep AGI in mind, and we are working on AGI. So when we work on applications or C-end and B-end, we don't need to put much thought into it; it takes very little effort.

我觉得是说,你站在一个技术的高位上来做相对低一级别的技术,是有这样降维打击的。至少在去年的 C 端,我们看到确实是这样。我们没有在 C 端上花很多力气,甚至一度我们都不想维 护那些用户了,但是用户赶都赶不走。

I think if you stand at a high technological level and work on a lower-level technology, it has a dimensionality reduction effect. At least in the C-end last year, we saw that was true. We didn't invest much effort on the C-end; at one point we didn't even want to maintain those users, but they just wouldn't leave.

因为真的是赶不走 ,所以最终都还在。但 稍微……然后今年 B 端的这个收入,现在看起来这个增长还比较乐观。我觉得可能这个数字跟同行比的话,应该也是比较好的,我估计。但是我们并没有花很大的精力去做这事情,我们根本就没有

Because they really wouldn't leave, so we still have them. But slightly... And this year's B-end revenue growth looks optimistic. I estimate the numbers might be good compared to peers. But we didn't put much effort into it; we simply didn't.

做互联网智能的上线,走的时候,AGI 这一步是我必须走的是个台阶。通往 AGI 的路上,我要经过这个台阶。那么我把这些技术都通过 API 给大家服务,我没有干额外的事情。

When deploying internet intelligence, stepping into AGI is a necessary milestone for me. On the path to AGI, I must pass this step. So I provide all these technologies to everyone via API, without doing anything extra.

我们还是在做 AI,这是一个副产品。我只要搞几个人维护这个 API 就可以了,甚至客服都没有,也不用销售,什么都不需要,用户自己就会来。

We are still working on AI; this is a byproduct. I just need a few people to maintain this API, without even customer service or sales. Nothing else is needed; users will come on their own.

或者说,被认为是 C 端的用户,C 端和 B 端,都是我们做 AGI 路上的副产物,都是一个中间的产出,跟我做 AGI 没有冲突。我并不是为了做 C 端,或者为了做 B 端而去做它,而是我们做 AGI 是为了做 AGI,刚好能产出这个东西,我就把它拿来做商业化了。

In other words, what are considered C-end users, C-end and B-end, are all byproducts on our path to AGI, intermediate outputs that do not conflict with my pursuit of AGI. I am not doing it for C-end or B-end; rather, we pursue AGI for AGI's sake, and this happens to be a product we can commercialize.

这跟其他公司不一样。其他公司就是为了做这个,为了服务 C 端用户,或者服务 B 端用户而去做这个模型。但对我们来讲,初衷不是这样的,我们初衷还是去追求 AGI。

This is different from other companies. Other companies build models specifically to serve C-end users or B-end users. For us, that is not the original intention; our original intention is still to pursue AGI.

我觉得这个在一定程度上是一种降维打击。AGI 是更大的愿景,这个愿 景能够凝聚起更多优秀 的人,它有更强的凝聚力。所以我在组织上有优势,然后我利用这个优势去……它就是一种降维打击。

I think this is, to some extent, a form of dimensionality reduction attack. AGI is a larger vision that can attract more outstanding people and has stronger cohesion. So I have an organizational advantage, and I use that advantage to... it's a dimensionality reduction attack.

但如果你是一个商业公司,你的愿景就是服务好 C 端用户 ,那么是另外一个故事 。他有其他的优势,在产品上、用户服务上、流量上会有优势,但是在技术上没有优势。

But if you are a commercial company whose vision is to serve C-end users well, that's another story. You may have advantages in product, user service, and traffic, but not in technology.

现在比较有利的形势是,模型技术是最重要的。你要将模型做好,其他……对,然后这个可以解释我们之前发展的轨迹。 我们真的是选择 AGI,然后我根本没有想着说要做很多的用户。

The current favorable situation is that model technology is the most important. You need to build a good model, and everything else... yes, and this can explain our development trajectory. We truly chose AGI, and I never thought about having many users.

去年春节我们突然火的时候,那不在我们的剧本里面,我们完全没有想过那个东西。我们就是想把这个技术做好。但是那时候我发现,其实我们这个组织相对于完全商业化、以产品为目标的组织来讲,在人才跟组织上是有额外优势的。

When we suddenly became popular last Spring Festival, that was not in our script. We never thought about it. We just wanted to do this technology well. But at that time, I realized that compared to a fully commercialized, product-oriented organization, our organization has extra advantages in talent and organization.

这也很神奇。那时候 C 端大家都抢得头破血流,结果被一个没有去抢的人牵走了。这个也确实说明我前面说的逻辑是有道理的,确实是有人才和组 织上的优势。

This is also amazing. At that time, everyone in the C-end was fighting fiercely, but in the end, it was taken away by someone who didn't join the fight. This also confirms that the logic I mentioned earlier is correct; there is indeed an advantage in talent and organization.

这个人才优势不是说我 的人比他更聪明,而是这 些人才我怎么组织起来,怎么激励他,然后怎么 合作。这个东西是有优 势的。因为你把聪明的人聚在一起,并不是说他自然而然就能够合作,自然而然就能够非常有激情地去奔一个目标、去完成的,所以你需要一个愿景。

This talent advantage does not mean that my people are smarter than others, but rather how I organize these talents, how to motivate them, and how to collaborate. There is an advantage in this. Because gathering smart people together does not mean they can naturally cooperate, naturally have great passion to pursue a goal and achieve it. So you need a vision.

我前面说,我们在很多方面都要非常克制。但是我们的核心利益是什么?其实我们的核心利益只有一点:我们最大的核心利益是要保持团队 的稳定性。这是我们最大的核心利益,甚至可以认为是唯一的核心利益。

As I said earlier, we must be very restrained in many aspects. But what is our core interest? Actually, we have only one core interest: our biggest core interest is to maintain the stability of the team. This is our biggest core interest, and it can even be considered the only core interest.

只要我能够保持团队的稳定性,我一定能做成,一定能做成 AGI,就这么简单。只要大家都不 走,我们能够继续做,我就一定能……基本上没有什么风险。只是说早点晚一点,将遇到挫折;如果遇到挫折,大家都不走,那么我又可以继续。

As long as I can maintain the stability of the team, I will definitely succeed, I will definitely build AGI. It's that simple. As long as everyone stays, we can continue to work, and I will definitely... basically there is no risk. It's just a matter of sooner or later, encountering setbacks; if we encounter setbacks and everyone doesn't leave, then I can continue again.

钱肯定不是问题,资源不是问题,其他要素都是容易获得的。对我们来讲,只有一个核心利益,只有一个没法退让的:我们必须要保持团队的稳定性。

Money is certainly not a problem, resources are not a problem, and other factors are easy to obtain. For us, there is only one core interest, one thing that cannot be compromised: we must maintain the stability of the team.

这也是我们面临的一个非常大的挑战,或者说,我觉得是最大的风险。当然,这个风险随着我们近期的这一次融资,得到了比较大的解除。因为大家拿到的期权都还是比较多的,金额还是比较大的。

This is also a very big challenge we face, or I think it is the biggest risk. Of course, this risk has been largely alleviated with our recent financing round. Because everyone has received quite a lot of options, and the amounts are relatively large.

从团队稳定性来讲,只要最重要的一些员工、最老的员工能够稳定,那么其他人是不太会走的。其他人哪怕期权少一点、收入少一点,他也不会走。因为他不是全奔着钱来的,大家都希望在一个能够做成 AGI 的环境里面去做这个事情。

From the perspective of team stability, as long as the most important employees and the oldest employees can be stable, then others are unlikely to leave. Even if others have fewer options or lower income, they won't leave. Because they are not all coming for money; everyone hopes to work in an environment that can achieve AGI.

所以对人才来讲,还是有吸引力的。从历史上来看,我们的人才流动是比较少的,跟同行比,我们人才流动永远是比较少的。但是,这依然是我们最大的挑战,是唯一的挑战,可以这么讲。

So it is still attractive to talents. Historically, our talent turnover is relatively low, always lower compared to peers. However, this remains our biggest challenge, the only challenge, so to speak.

其他都 是时间问题,其他最多导致我们晚半年、晚一年,但是不会说做不出来。肯定是不缺钱,肯定是不缺资源,其实这些都是不缺的。

Everything else is a matter of time; at most, it might set us back half a year or a year, but it won't prevent us from achieving it. We definitely have no shortage of money or resources—in fact, we lack nothing.

所以我们现在做的很多事情,都是为了 保持团队的稳定性。除了这一点以外,其他我觉得我们都是可以不要的,都是可以克制的。我们一直非常克制,不愿意跟任何一家互联网大厂或者小厂成为对手。

So much of what we're doing now is aimed at maintaining team stability. Apart from that, everything else is dispensable and can be restrained. We've always been very restrained, unwilling to become adversaries with any major or minor internet companies.

我希望我能够给他赋能,或者希望我能够协助大家去做这个事情,希望能够帮助大家做这个事情。这也是我们前面说的,我们的商业意义的一部分 。前提是大家不要管…

I hope I can empower them, or assist everyone in doing this, and help everyone accomplish it. This is also part of our commercial significance, as we mentioned earlier. The premise is that everyone should not interfere...

在这个前提下,我们是很愿意协助、帮助任何人,甚至我们的竞争对手,包括阿里、 智谱、月之暗面,去做得更好。因为我们并不损失什么东西,我们本来也是开源的,开源也是没有把边界尽可能说清楚,对怎么做,希望你能够复现;你如果复现不了,你讲,我告诉你怎么复现。

Under this premise, we are very willing to assist and help anyone, even our competitors, including Alibaba, Zhipu AI, and Moonshot AI, to do better. Because we lose nothing; we are open-source by nature, and open-source does not clearly define boundaries. As for how to do it, we hope you can reproduce it. If you can't, just tell us, and we'll show you how.

这本来就是开源的一部分,不会因为你是竞争对手就怎么样。当然,如果是合作伙伴,我会做更多的 事情。但是在大的利益上面是没有冲突的。

This is inherently part of open-source; it doesn't change just because you're a competitor. Of course, if you're a partner, I'll do more. But there is no conflict on major interests.

在对外面打交道的时候,我们的态度是:我们只做 AGI 的主线。这就是我刚刚说的这个 GPT、CoT、Agent 等等,就只做主线。AI 领域很广泛,有很多东西我们觉得它不在这个主线上面,比如说 3D、视频生成,我觉得可能跟智能的主线没有太大的关系,我们不会去做。

When dealing with the outside world, our attitude is: we only focus on the AGI main line. This is what I just mentioned—GPT, CoT, Agent, etc.—we only work on the main line. The AI field is vast, and there are many things we don't consider to be on this main line, such as 3D and video generation. I think they don't have much to do with the main line of intelligence, so we won't pursue them.

还有一些,比如说世界模型,我觉得现在跟智能的上限也没有太大的关系,所以我们也不会去做。但是其他人做,我们也很乐意帮忙。有没有时间是一回事,但是利益上是没有冲突的。

There are also things like world models; I think they are not closely related to the upper limit of intelligence for now, so we won't pursue them either. But if others do, we are happy to help. Whether we have time is one thing, but there is no conflict of interest.

我们也希望这些 AI 技 术能够用到各种生产环 境里面去,能够提高社会的生产效率,能够帮助各行各业去提高生产效率。我们是非常有动力做这个事情的。

We also hope that these AI technologies can be applied in various production environments, to improve social productivity and help various industries enhance efficiency. We are very motivated to do this.

我有没有时间、有没有人手,或者我们的同学自己感不感兴趣,那是另外一回事。但是在利益上是没有冲突的,我们是希望能够达 到这个目的。并且我们认为,这个跟商业没有任何冲突,我该拿到的利益还是没有少。

Whether I have time, whether I have manpower, or whether our classmates are interested, that is another matter. But there is no conflict of interest; we hope to achieve this goal. Moreover, we believe this has no conflict with business; the benefits I am entitled to have not diminished.

我觉得我们之前秉持这种态度,其实我们并没有因此而少拿到任何东西,并没有因为我开源,并没有因为我们的善意或者说我对其他人提供帮助,而导致我们少拿了任何东西。比如说去年的 C 端用户,我们现在 C 端用户还是比较多的,也还是比较稳固的。

I think the attitude we have held before—we haven't actually received any less because of it. It's not because of open source, nor because of our goodwill or helping others, that we have received any less. For example, last year's C-end users; our C-end users are still quite numerous and stable.

我们今年的 B 端,我觉得也是比较乐观的。我并没有因为我们的善意而影响到我 的商业利益,完全没有影响,反而可能还有加分。这个看起来违反直觉,但它确实是这样的。或者我们反过来想,如果违背它,也是必然的,我能够拿到更多东西吗?没有。

As for our B-end this year, I think it is also quite optimistic. My goodwill has not affected my commercial interests at all; there is no impact, and it might even be a plus. This seems counterintuitive, but it is indeed the case. Or if we think the other way, must it necessarily yield more? No.

有个问题是,怎么理解世界模型和 AI 智能上限提高没有关系?我说的是在现在这个阶段, 这是我们的判断。我们有看到,这是我们自己的 AI 路线图,不是唯一的路线图。

One question is: how to understand that the world model is unrelated to the improvement of AI intelligence's upper limit? What I said is that at this stage, this is our judgment. We see that this is our own AI roadmap, not the only one.

从我们的理解、从我们的判断来看,当前最重要的是把 AI 训练做好。把 AI 训练做好,并不需要世界模型,甚至不用多模态。因为你把 AI 训练范围缩小一点,没有多模态,只是有一部分任务你做不了,但是不影响这个算法的成立。

From our understanding and judgment, the most important thing now is to do AI training well. Doing AI training well does not require a world model, or even multimodality. Because if you narrow down the scope of AI training, even without multimodality, you only cannot do certain tasks, but it does not affect the validity of the algorithm.

多模态最终还是要做的。现在重 要的是训练,然后下一步要解决的是持续学习的问题,再下一步要解决的是它自己问问题。但是这个路线图里面没有世界模型,没有视频生成。

Multimodality will eventually be necessary. Now the key is training; the next step is to solve the problem of continuous learning, and then the next is for it to ask questions on its own. But this roadmap does not include a world model or video generation.

视频生成一开始出 来的时候就很火,好像这是必须要做的,如果你不做,你就不是一个 AI 公司一样。所以我就很奇怪,这个其实你只要仔细去想一下,它跟智能的路线图 是没有什么关系的。

When video generation first emerged, it became very popular, as if it was something that had to be done—if you don't do it, you are not an AI company. So I find it strange; actually, if you think carefully, it has nothing to do with the roadmap of intelligence.

事实上,你也发现,那个视频生成一开始 Sora 出来之后,所有人都做,大公司、小公司都做。但是小公司后来都 把它砍掉了。它跟 智能的上限没有关系。

In fact, you also see that after Sora came out, everyone started doing video generation—big companies, small companies alike. But later, small companies cut it. It has nothing to do with the upper limit of intelligence.

但在商业上,它是个好生意,在商业上是个好生意。但是这跟智能没有什么关系。我们不会因为它是一个好商业而去做它,我们只会因为它是智能路线图上的东西,才会去做。

But commercially, it's a good business. Commercially, it's a good business. However, this has nothing to do with intelligence. We won't do it just because it's a good business; we will only do it because it's on the roadmap for intelligence.

视频生成那个,因为那个相对比较清楚,就拿它举例。然后世界模型,世界模型这个含义就没那么清楚,因为很多东西都可以说它是世界模型。

Take video generation for example, because that's relatively clear. As for world models, the meaning is not so clear, since many things can be called world models.

从我们的判断,世界模型和智能还不是现在这个阶段最重要的事情。最重要的才是 AI 训练,以及 AI 训练之后怎么解决持续学习。这是我们公司的判断, 当然每个公司它的判断是不一样的。

In our judgment, world models and intelligence are not the most important things at this stage. The most important is AI training, and then how to solve continuous learning after AI training. This is our company's judgment. Of course, each company has its own judgment.

我刚说的是,我们公司对公司来讲最重要的问题,就是人员的稳定性。从另外一个维度来讲,我们缺什么呢?我们跟美国的差距是什 么?其实差距只有一点,就是资源。

What I just said is that for our company, the most important issue is workforce stability. From another dimension, what are we lacking? What is the gap between us and the US? Actually, there is only one gap: resources.

我们没有那么多卡,我们的卡的数量还是比较少的。我们现在大概是两万张 H 等效算力,这其中大部分是刚到,最近一两个月刚到,可能还有很多机器还没有 来。

We don't have that many GPUs; our number of GPUs is relatively small. We currently have about 20,000 H-equivalent compute power, most of which have just arrived in the recent one or two months, and there may still be many machines that haven't arrived yet.

我们总的算力去年比较少,今年我们在非常激进地扩充这个算力。我们现在大概是两万张 H 等效,接下来几个月我们还会有大批量的机器买过来,基本上都是英伟达。

Our total compute power was relatively low last year. This year, we are very aggressively expanding it. We currently have about 20,000 H-equivalent, and in the next few months, we will purchase a large number of machines, basically all from NVIDIA.

我们需要多少卡?现在肯定是越多越好 。在我们能够承受的范围内 ,肯定是卡越多越好,这是毫无疑问的。所以我们现在策略是,在合理的价格里面,能买到多少卡就买多少卡。

How many GPUs do we need? Now, definitely the more the better. Within what we can afford, definitely the more the better, there's no doubt. So our current strategy is to buy as many GPUs as we can at a reasonable price.

如果说 这笔融资用完之后,我能 买多 少卡,我就买多少卡。花钱的速度不在计划里面,而是只要价格合理,有多少我都买。那么如果说我在半年之内就把钱花完了,我觉得是个好事。如果说我在半年之内把钱都花完,那这个太幸福了,这太理想了。

If after using up this round of financing, I can buy as many GPUs as I can, I will buy them. The speed of spending is not planned; as long as the price is reasonable, I will buy as many as I can. So if I spend all the money within half a year, I think that's a good thing. If I spend all the money within half a year, that would be too happy, too ideal.

实际上,要把这么多钱花完非常不容易,买不到那么多卡,很难买,而且价格也很高,也不能说花非常高的价格去买,还得确保这个价格是合理的。

In fact, it's extremely difficult to spend all this money. We can't buy that many cards—they're hard to acquire and very expensive. We can't just pay an exorbitant price; we have to ensure the price is reasonable.

如果我半年之内就把这个钱花完的话,可能是最理想的。因为我把钱变成英伟达卡,肯定是比放在银行里面好。放银行里面,我现在已经可能是存两个点,好像就要。但是买英伟达卡,十个月的社会成本。

If I could spend all this money within half a year, that would be ideal. Because turning the money into NVIDIA cards is definitely better than leaving it in the bank. In the bank, I might only get a two-point interest rate, but buying NVIDIA cards has a ten-month social cost.

所以肯定是你能买 到多少就买到多少。最先如果买了卡之后,后面我有很多空间,我通过提供服务或者我怎么样,我总是能够有现金流的。我有现金流我就可以活,我就不需要在我账上面放很多很多钱。

So we just buy as many as we can. Once we have the cards, we'll have a lot of room to generate cash flow by providing services or whatever. With cash flow, we can survive, and we won't need to keep a huge amount of money in our account.

所以说,我们只担心买不到那么多卡。如果能够把钱都变成卡的话,那我们会毫不犹豫把所有的钱都变成卡,并且在这里面我们是愿意付一定的溢价的。就是我们愿意付一定的溢价去把它变成卡,因为这太划算了。

Therefore, our only concern is not being able to buy enough cards. If we could convert all the money into cards, we would do it without hesitation, and we would be willing to pay a certain premium. Because it's such a great deal.

哪怕我们在付完 溢价之后,其实我们也很难实现这个目标。那么客观上来讲,如果说我今年能够花掉两百亿,那么就属于我们的采购部门业绩超级好了。

Even after paying the premium, it's still hard to achieve this goal. Objectively speaking, if I could spend 20 billion this year, that would mean our procurement department performed exceptionally well.

我们跟美国的差距主要在资源上 面,然后人上面差距不是很大的。人上面几乎没有差距,因为就是同一批人,可能是中国人。中国人出去的时候,有一些人留在国内,有些人留在国外,有些人去国外,他并没有说是聪明的人去国外,没有的。

The gap between us and the US is mainly in resources, not in people. There's almost no gap in talent because it's essentially the same group of people—Chinese. Some Chinese go abroad, some stay behind. It's not that the smarter ones go abroad; that's not the case.

其实它是比较随机的。最聪明的那些人,可能不是说一半多一点去国外的,但有一半少一点留在国内。国内人才是不缺的,而且我们的基数大,我们每年有那么多人的补充。

In fact, it's quite random. The smartest people—maybe not more than half go abroad, but less than half stay in China. We don't lack talent domestically, and our base is large, with many new additions every year.

人才不是瓶颈,资源是最大的瓶颈。资 源首先影响到人才培养,因为算力少,我们能做的实验机会比较少,所以我们的人才整体上比美国有差距。人才的差距,本质上也是因为算力的差距。

Talent is not the bottleneck; resources are the biggest bottleneck. Resources first affect talent cultivation, because with limited computing power, we have fewer opportunities to run experiments. So overall, our talent lags behind the US. The talent gap is essentially due to the computing power gap.

在现在最大的模型上,我们其实训不起的。我们哪怕把五百亿全花掉,其实也训不起。哪怕能堆起来也用不起。现在最大的模型,它激活大概是 800B;国内的话,我们还在几十 B 的这个规模,国内最大模型可能就几十 B 激活,那么差一个数量级。

In the largest models currently available, we actually can't afford to train them. Even if we spent all 50 billion, we still couldn't train them. Even if we could assemble the resources, we couldn't afford to use them. The largest model now has about 800B activated parameters; domestically, we are still at the scale of tens of billions. The largest model in China might have only tens of billions activated, which is an order of magnitude less.

如果我要训练跟 AI 同样大的模型,应该需要五万张 GB300,或者华为 950,二十万卡。这只是 训练,还没有考虑做研究。所以我们跟美国之间最大的差距是在资源上面。

If I wanted to train a model as large as AI's, I would need 50,000 GB300s, or Huawei 950s, 200,000 cards. That's just for training, not considering research. So the biggest gap between us and the United States is in resources.

我们现在的资源,以及我们今年之内、接下来几个月的资源,包括马上大资源,也只足够我们在 10B 激活的这个规模里面去做更多的实验。因为我在几十 B 激活的这个规模里面,还有很多实验要做,还有很多事情需要搞清楚的。我们应该离能够训 800B 的模型还比较远,时间还非常多,没有那么多卡。

Our current resources, as well as the resources we will have within this year and the next few months, including the large resources coming soon, are only enough for us to conduct more experiments at the scale of 10B activation. Because there are still many experiments to do and many things to figure out at the scale of tens of billions activation. We are still far from being able to train an 800B model; we still have a long way to go and not that many cards.

所以我们跟美国的区别,我认为就是资源 上的区别。我们可能会认为,我们看到的所有区别,包括人才的区别、模型能力的区别、应用的区别,可以认为都是因为算力资源上的区 别。

So the difference between us and the US, I think, is a difference in resources. We may believe that all the differences we see, including differences in talent, model capability, and applications, can be attributed to differences in computing resources.

算力资源 一方面是国内本身卡就买不到,另外一方面是我们的资本投入比美国要少。我们在资本投入层面上少了很多,基于人才的工资在这里面占的比重是很低的。你看他们开的薪水一个亿美金这种,但算起来,人才的薪水还是占比很少,大头还是算力。

On the one hand, domestic cards are simply unavailable; on the other hand, our capital investment is much less than that of the US. We have a lot less capital investment, and the proportion of talent salaries in this is very low. You see they pay salaries of 100 million US dollars, but when you calculate, the salary cost of talent is still a small part; the big part is computing power.

这个问题目前基本上无解 ,因为华为的产量也是有限的。因为我要训 800B 的话,我就得二十万张华为最新的卡,这只是训练,还没有考虑做研究。

This problem is basically unsolvable for now because Huawei's production capacity is limited. If I want to train an 800B model, I need 200,000 of Huawei's latest cards, and that's just for training, not considering research.

所以我们现在根本就不会考虑,到这么大的一个规模上去 跟美国竞争。我们现在输,还是在我们能够训练得起、能够用得起的规模上,就几十 B 激活的规模上要先把它做好。再等下一步有更多的资源的时 候,再把它做到一百五十 B、一百五十六 B,或者两百五十 B 激活的这个规模。

So we currently don't even consider competing with the US at such a large scale. Where we are losing now is at the scale we can afford to train and use—the scale of tens of billions activation—and we need to do it well first. Then, when more resources become available, we can expand to 150B, 156B, or 250B activation scale.

所以我们跟美国在这里面有一个目前看起来很难弥补的差距。你要硬要训那么大模型也是能训,但是你没法做充分的研究。就是你在训之前,没法做充分的研究。

So there is a gap between us and the US that seems difficult to bridge at present. If you insist on training such a large model, you can do it, but you cannot conduct sufficient research. That is, before training, you cannot do enough research.

还有个大家比较关心的问题是,所有大模型竞争,它终局的差距会体现在什么地方?就是当大模型最终差距会 体现在哪里?我觉得这个差距最终可能是没有太大差距的。

Another question that many people care about is, where will the ultimate gap lie among all the major models competing? That is, when it comes to the eventual gap between large models, where will it be? I think in the end, there may not be a significant gap.

最终的差距应该是三方面:一方面成本,一方面时间,一方面用户体验。除此以外,可能是没有什么差距的。

The final gap should be in three aspects: one is cost, another is time, and the third is user experience. Beyond that, there may be no gap.

成本比较好理解,你提供同样的服务、同样质量的服务,你能够以什么成本来提供。同样一件服务,不如说比亚迪的电池,在同等技术量的情况下,其他家是不是能够按照这个价格来提供,我觉得这个是 个较难的事情。不是那 么容 易做到的,肯定是壁垒。

Cost is easier to understand: providing the same service with the same quality, at what cost can you deliver it? Take BYD's batteries for example—under the same level of technology, can other companies provide them at the same price? I think that's a difficult thing. It's not easy to achieve; there are definitely barriers.

所以成本肯定是一个差异,我觉得可能成本是排在第一位的区别。然后第二个就是时间,你什么时候能够做到。你早几个月、晚几个月,它就不一样了。

So cost is definitely a differentiator. I think cost might be the primary difference. Then second is time—when you can achieve it. A few months earlier or later makes a difference.

第三个可能是体验,用户体验还是有些不一样的。这个用户中间还是会有一些用户粘性和用户壁垒的,但它可能不本质。本质还是第一个成本,第二个时间。你先做出来还是后做出来,你做到同样东西,是快点还是慢一点。

The third might be experience: user experience can still vary. There can be some user stickiness and user barriers among users, but it may not be essential. The essence is still the first: cost, and the second: time. Whether you do it first or later, whether you achieve the same thing faster or slower.

长期商业化路径和产品线进一步丰富后,怎么定价? 我们觉得现在最值得做、最值得花精力的,还是做 AGI。现在要把 AGI 往前再往前推,就是把智能的下限往上推,往前推。

After the long-term commercialization path and product line are further enriched, how should we price? We believe that the most worthwhile thing to do and invest effort in right now is to work on AGI. We need to push AGI further forward—that is, to raise the lower bound of intelligence and push it ahead.

这个应该是在现阶段,比做更多的产品线、考虑更多商业化路径更划算,或者说它是一个收益更大的路径。我觉得在未来的一段时间,以及过去的任 何一段时间,可能都是这样的。就是说,如果说我们

This should be, at the current stage, more worthwhile than developing more product lines or considering more commercialization paths. In other words, it's a path with greater returns. I think for the foreseeable future and indeed for any past period, this has been the case.

半年前讨论出来的商业化是什么?一定是我要去做广告,然后我要去做电商,我要 去在产品里面植入 电商, 然后要跟本地生活什么很大一体。肯定没用的,因为变化太快了。你这个产品,如果说你在前面 ,或者我们现 在这个阶 段,去花很多时间做商业化考虑、商业化路径,产品线它生命周期很短。

What was the commercialization we discussed six months ago? It was definitely about doing advertising, then e-commerce, embedding e-commerce into our product, and integrating with local services and such. It's certainly useless because things change too fast. If you spend a lot of time on commercialization considerations and paths at this stage, the product life cycle is very short.

我觉得没有到这个时候。这是我们的判断,或者至少前面的这些经验都是支持这个判断的。在过去三年,任何一个时候你来跟我说商业化路径、产品线,都是 浪费时间,因 为你并不 能够预测、不能够预见未来。你能够预见的是很少的。

I don't think we're there yet. That's our judgment, or at least our experience so far supports this judgment. At any point in the past three years, if you had come to me talking about a commercialization path or product line, it would have been a waste of time, because you cannot predict or foresee the future. What you can foresee is very little.

因为我看时间,不知道大家要中场休息一下吗?要先吃点东西吗?还是我们继续聊?大家没有意见,我就继续说。

Looking at the time, does anyone want to take a break? Maybe grab something to eat? Or should we keep going? No one objects, so I'll continue.

在做全球领先的 AGI 研究,并在适当时候进行商业化。我觉得我们一 直在做商业化 ,我认为 一直在做商业化,只是没有以商业化为目标。我们是以 AGI 为目标,但是我们一直在做商业化,所以我们才有 C 端的用户,才有 B 端的收入。

We are conducting world-leading AGI research and will commercialize it at the appropriate time. I think we have always been commercializing, just not with commercialization as the goal. Our goal is AGI, but we have always been commercializing, which is why we have C-end users and B-end revenue.

从历史经验来看,这个策略是成功的。然后我觉得,我们离彻底转向商业化的时间点,应该是很遥远的。所以 最大的还是技 术的延伸 ,然后做下一代的技术,更多地解决现在的问题。这些收益比,我个人觉得,就在现在能看到的未来,在任何一个时候你聚焦在产品上都太早了。

From historical experience, this strategy has been successful. I think the time for a complete shift to commercialization is still very far away. So the biggest thing is the extension of technology, then developing the next-generation technology to better solve current problems. The return on investment, in my opinion, focusing on products at any point in the foreseeable future is still too early.

所以这也是我们克制的一部分。我希望这些商业机会能够让其他人来做,我们希望这些商业机会、怎么用这 个 AI,是 整个社会 、我们所有的合作伙伴一起来做,大家一起来分享这个收益,而不是我想独吞,这不可能。而且我们没有那么多精力,我们的组织也没有那么多人去做这个事情。

So this is also part of our restraint. I hope these business opportunities can be taken by others. We hope that these business opportunities—how to use this AI—are a collaboration with the entire society and all our partners, sharing the benefits together. It's not about me wanting to monopolize, which is impossible. Besides, we don't have that much energy, and our organization doesn't have enough people to do that.

对于合作伙伴,其实我们这个融资是精心挑选的。具体怎 么合作的这个提议,但首先我觉得,利益是比较一致的,就是跟我们利益最一致的、对我们最没有敌意的,或者说最希望我们能够做成功的。不是所有人都希望我们能做成功的,因为我们还是损害了很多其他人的利益的。

As for partners, we selected this financing carefully. Regarding the proposal on how to cooperate, first of all, I think our interests are relatively aligned—those whose interests are most aligned with ours, who are least hostile to us, or who most want us to succeed. Not everyone wants us to succeed, because we have harmed the interests of many others.

公司重大战略、技术业务决定的流程和决策机制。我们公司整体上是建立在共识的基础上的,我并 不是说我一个人决定所 有事 情,而是我要寻求共识。我在公司内的权威以及在公司内的影响力,是建立在共识的基础上的。

The decision-making process for major company strategies, technologies, and business. Our company is built on consensus. I don't decide everything by myself; I seek consensus. My authority and influence within the company are based on consensus.

比如说我要做一个事情,我肯定 是先看我们大家的 共识是 什么,大家想不想做。然后有可能我会有一些引导或者倾向,但是这引导的作用是有限的,引导的作用是非常有限的,还是建立在一个共识的基础上。 这个决策机制其实 是一种 寻求共识的机制,这并不是说我能够推动一个什么事情,它一定是共识 ,我才能够推 得下去, 然后我才会去推。

For example, if I want to do something, I first look at what our consensus is—what everyone wants to do. Then I might guide or lean a certain way, but the effect of guidance is limited—very limited—and it is still based on consensus. This decision-making mechanism is actually a consensus-seeking mechanism. It's not that I can push something through; it must be a consensus before I can push it forward, and then I will push it.

目前主要精力基本上都在 DeepSeek 这边。我们休息五分钟。我快点写。大家可以开麦,给我一点反馈。我们这边没问题,要不然先休息五分钟。

Currently, most of our focus is on DeepSeek. Let's take a five-minute break. I'll write quickly. Everyone can unmute and give me some feedback. We're fine here, so let's take a five-minute break.

各家模型最终效果拉开差距,应该是一个综合上的。比较模型效果,肯定是得在相同的成本上来比较,这个才是有意 义的。因为你 比较两辆 车,也是同价位的车来比较。

The difference in final performance among various models should be comprehensive. Comparing model performance must be done under the same cost to be meaningful. Because comparing two cars, you compare cars of the same price.

Anthropic 跟现在超过 OpenAI,这是不是长期的?我觉得这不是长期的,这肯定是极端性的。OpenAI 和 Google,未来大概 率还是会 交替上升,应该是会交替上升。其实现在 Anthropic 在 Code Agent 上的优势没有那么大,其实并没有说它碾压 OpenAI。

Is Anthropic's current surpassing of OpenAI long-term? I don't think so. It's definitely extreme. OpenAI and Google will likely alternate in the future, they should alternate. Actually, Anthropic's advantage in Code Agent is not that big, it's not overwhelming OpenAI.

我们公司可能有一半的人,平时有一半的人觉得 Ope nAI 是更好的。其实 Anthropic 它有先发优势,但这先发优势应该很快就没了,并不是一个它能够长期占得住的优势。大家这三家都很厉害,这三家里面的效率是最高的,它花的成本、它花掉的、它烧掉的钱应该是最少的。

In our company, about half of the people usually think OpenAI is better. Actually, Anthropic has a first-mover advantage, but that advantage should soon disappear. It's not an advantage it can hold long-term. All three companies are very strong, and among them, the one with the highest efficiency spends the least cost, money burned.

在全球 AI 中主导分工的时 候,中国 公司很有可能扮演的一个角色还是产量最大。常理来讲,我们的产能最大,包括芯片,芯片可能我们的产能最大,我们的电力最多,所以我们的 AI 最终很有可能我们是三体之一。

When it comes to the division of labor in global AI, Chinese companies are likely to play the role of having the largest output. By common sense, our production capacity is the largest, including chips, chip production capacity may be the largest, and we have the most electricity, so our AI will likely become one of the three bodies.

中国人会把这产品做到最便宜,然后 再在效果上,毕竟国 外的商 品,现在很多商品中国产跟美国差并没有太大的区别。未来可能 AI 也是这样,但是中国产的 AI 可能价格会更便宜。这个便宜可能是系统性的低,就跟其他行业中国提供的服务更便宜可能是一样的。

Chinese people will make the product the cheapest, and then in terms of effectiveness, after all, for many foreign goods, there is not much difference between Chinese-made and American-made now. In the future, AI may be the same, but Chinese-made AI may be cheaper. This cheapness may be systematic, similar to how Chinese services are cheaper in other industries.

我要做事情的时候,我习惯的思考是:我现在这个时候做什么收益最大?如果说我觉得现在做产品收益最大,我觉得去做产品;如果说我觉得现在把 AGI 先实现收益最大,那我就是先去做 AGI。显然, 我觉得现在做产品不是 收益 最 大的。

When I want to do something, my habitual thinking is: what gives the greatest return at this moment? If I think making products gives the greatest return, I'll make products; if I think achieving AGI first gives the greatest return, then I'll go for AGI. Obviously, I think making products is not the greatest return now.

如果你是国产卡适配的问题,国产卡现在其实是有一个历史性的机遇的。因为之前国产卡适配有一个难题,叫生态不好。就买了卡之后,但是用不起来,它没有英伟达那个生态。所以说英伟达的护城河 是很强的。

If you are talking about the adaptation of domestic cards, domestic cards now actually have a historic opportunity. Because previously, the challenge for domestic card adaptation was poor ecosystem. After buying the cards, they couldn't be used effectively; they didn't have NVIDIA's ecosystem. So NVIDIA's moat is very strong.

但是这个事情在发生变化。英伟达 CUDA 的护城河在快速地被瓦解,快速被瓦解的原因可能是有三方面。一方面是现在有了 AI,然后有了 AI 之后,我要建立起这个生态比以前容易很多了 ,因为 AI 可 以写代 码。

However, this situation is changing. NVIDIA's CUDA moat is rapidly crumbling, and the reasons for this rapid erosion may be threefold. First, with the advent of AI, building this ecosystem has become much easier than before, because AI can write code.

我可以用 AI 来构建这个生态,就可以把跟英伟达一模一样的生态构建出来。第一个,因为有 AI;第二个,有一些新的技术。比如说,我们家出了一个技术,叫 TileLang,是一种高级语言 。

I can use AI to build this ecosystem, replicating an ecosystem identical to NVIDIA's. The first reason is AI; the second is the emergence of new technologies. For example, we have developed a technology called TileLang, which is a high-level language.

用这个高级语言来写 CUDA 的算子,可以很快地把英伟达的整套生态全部都写一遍,再结合 AI,这个看起来没有什么障碍。但是现在还没有完成,还没有做完,不过这个技 术路线看起来 是没有什 么障碍的。

Using this high-level language to write CUDA operators, it's possible to quickly rewrite NVIDIA's entire ecosystem. Combined with AI, this seems feasible without major obstacles. Although it's not yet complete or finished, this technical route appears to have no significant barriers.

还有一点,因为 CUDA,英伟达它是从游戏卡演变出来的,所以它在很多地方,游戏卡的设计跟游戏卡的设置是一脉相承的。CUDA 是兼容游戏卡的。以前因为 AI 计算是 一个很小的领 域,比游 戏卡的市场要小,所以这样是合理的。

Another point is that CUDA evolved from NVIDIA's gaming cards, so in many aspects, the design and settings of gaming cards are consistent. CUDA is compatible with gaming cards. Previously, AI computing was a niche field, smaller than the gaming card market, so this compatibility made sense.

但是现在计算卡的市场已经比游戏卡更大了,就没有理由这两个还需要耦合起来。现在的趋势是,以后就不再耦合了。那么专用芯片,不管是华为还是英伟达自己,以后都是专用芯片,都不是之前的这些东西了。

But now the computing card market is larger than the gaming card market, so there is no reason for the two to remain coupled. The current trend is that they will no longer be coupled. Going forward, dedicated chips—whether from Huawei or NVIDIA itself—will all be specialized, not like the previous ones.

在这个背景下,原来 英伟达生态的作用 就大幅 度减少了。因为对于一个专用芯片来讲,跟 CUDA 没有什么关系,它跟 CUDA 是不绑定的。或者说,这个芯片在设计的时候,就已经考虑到怎么建立这个生态了。

Against this backdrop, the role of the original NVIDIA ecosystem is greatly diminished. For a dedicated chip, it has little to do with CUDA; it is not tied to CUDA. Or rather, the chip's design has already taken into account how to build its own ecosystem.

有点复杂,但 anyway,国产 AI 芯片替代现在是有个历史性机会的。 我们认为,未来一 年之内 ,我们能够看到有一个事情被验证:国产芯片的生态完全没有问题。之前认为是有问题的,认为是用不起来、不好用,但是 未来我觉得一年之 内,我 们能够扭转这个认知,或者会用事实来扭转这些。

It's a bit complicated, but anyway, there is a historic opportunity for Chinese AI chip substitution. We believe that within the next year, we will see something validated: the ecosystem of domestic chips is perfectly fine. Previously, it was thought to be problematic—unusable or difficult to use—but I think within the next year, we can reverse this perception, or use facts to change it.

国产 AI 芯片的硬件和生态都没有问题,唯一有问题的是产能不够。国产卡适配这一点,没有障碍,英伟达挡不住。如果是在一个正常的商业环境里面,我能买到英伟达的卡,那么国产替代是比较难的;但是在英伟达的 卡买不到的情况下 ,所有 人都迫不得已,都要去做国产芯片。

The hardware and ecosystem of domestic AI chips are flawless; the only problem is insufficient production capacity. There is no obstacle to adapting domestic cards; NVIDIA cannot stop it. In a normal business environment where I can buy NVIDIA cards, domestic substitution is difficult; but when NVIDIA cards are unavailable, everyone has no choice but to adopt domestic chips.

在这个背景下,国产卡的适配是没有任何障碍的。国产卡建立起跟英伟达一样、甚至比英伟 达还好的生态 ,我觉得 没有障碍,但还需要时间。

In this context, there is no obstacle to the adaptation of domestic GPUs. I believe it is possible for domestic GPUs to establish an ecosystem that is on par with or even better than NVIDIA's, but it will take time.

我们现在主要是跟华为有合作。华为他们自己适配,但我们自己会参与这个生态,会深入参与到华为这个里面去。华为的问题还是产能不足。像华为给我们的大概是一万六千张卡的产 能,互联网大 厂可能是 十几万张,我们一万多张,我觉得这个比例也是比较……但这已经可能是华为就这么多产能了。

Currently, we are mainly collaborating with Huawei. Huawei handles their own adaptation, but we will participate in and deeply engage with their ecosystem. Huawei's issue remains insufficient production capacity. For example, Huawei provides us with around 16,000 GPUs, while major internet companies might get over 100,000 GPUs. Our share of over 10,000 GPUs seems relatively small, but this might be all the capacity Huawei has.

所以我们也没法指望在华为上面训后面那个更大的模型,或者说训练几百 B 参数激活的模型,有时就有说在今年。但是明年、后年说不定是有机会的。

So we cannot rely on Huawei to train our next larger model or models with hundreds of billions of activated parameters, which we had hoped to do this year. However, there might be opportunities next year or the year after.

华为卡适配,我们主要做的 工作是把它的高级语言编译器做好,把 TileLang 做好。把 TileLang 做好之后,问题可能就迎刃而解了。这个事情说起来比较复杂,但是我们在做,做完之后再来解释会清晰比较多。

Our main work in adapting Huawei's GPUs is to develop a high-level language compiler and refine TileLang. Once TileLang is well-developed, the problems may be easily solved. This is a complex matter, but we are working on it and will provide a clearer explanation after completion.

你可以理解,V3 训练的时候,它用的还是英伟达的卡,但是已经不用英伟达的生态了。 V3 用英伟 达的卡, 但是没有用英伟达的生态,而是我们先写一个高级编译器叫 TileLang,然后基于 TileLang 的生态来完成其他所有的事情,就已经几乎不依赖英伟达的生态了。

You can understand that during V3 training, it still used NVIDIA GPUs but no longer relied on NVIDIA's ecosystem. V3 used NVIDIA GPUs but not their ecosystem; instead, we first wrote a high-level compiler called TileLang and then completed everything else based on the TileLang ecosystem, achieving near-total independence from NVIDIA's ecosystem.

只要我把这一套东西,把这个过程重新在华为卡上做一遍,那么就完成了。我觉得这里可能是 一个历史性的使 命,即可 以完全扭转之前大家对国产卡生态不好的认知。现在有些天时地利基本上都全了,就差时间。我觉得一年之内,应该会有很多人在上面都有,或者说都明白这问题算是解决了,剩下就是产能的问题。

If we can replicate this process on Huawei's GPUs, the task will be accomplished. I believe this could be a historic mission that completely changes the previous perception of a poor ecosystem for domestic GPUs. Now, the timing and conditions are largely in place; only time is needed. I think within a year, many people will realize that this problem is solved, leaving only the issue of production capacity.

我对国产算力是比较 乐观的 。我觉得在这一点上,英伟达是在掘自己的坟墓。华为的超节点,华为的 950 超节点,在性能和价格上可以完全平替英伟达的 GB200、GB300。价格肯定要贵,但贵得有限。价格贵百分之五十、百分之一 百,贵百分之一百无所 谓, 贵百分之两百都无所谓。

I am quite optimistic about domestic computing power. In this regard, NVIDIA is digging its own grave. Huawei's super node, the Huawei 950 Super Node, can fully replace NVIDIA's GB200 and GB300 in terms of performance and price. The price will definitely be higher, but only by a limited margin. It is acceptable whether the price is 50%, 100%, or even 200% higher.

比如贵百分之一百,我觉得已经可以认为在价格上可以平替了。在任务上也是可以平替 的,所有 GB300 能 做的任务,华为超节点都能做,延时什么都一样。唯一的代价是,四张华为卡顶一张英伟达的卡,同时落后两年。

For example, a 100% price increase could already be considered a price parity in practice. Functionally, it is also a direct replacement; all tasks that GB300 can perform, Huawei's super node can handle equally well, with the same latency. The only trade-off is that four Huawei GPUs are equivalent to one NVIDIA GPU, and we are two years behind.

四张顶一张这个可以理解。落后两年的意思是,四张华为 950 能顶一张 GB300。落后两年是时间上落后两年。华为 950 超节点是今年 Q3 还是 Q4 的货,英 伟达 GB200 是两年前 Q 3 的货,差两年。

"Four cards covering one" is understandable. "Two years behind" means that four Huawei 950 can match one GB300. The two-year gap refers to a time lag. The Huawei 950 super node is a Q3 or Q4 release this year, while NVIDIA's GB200 was a Q3 release two years ago, so there is a two-year difference.

英伟达今年 Q3 可能已经有新一代了。所以我们跟美国在芯片上的差距,我认 为生态上以后不会再有差 距,但是在芯片上是四倍加两年。

NVIDIA may already have a new generation by Q3 this year. So in terms of the gap between us and the US in chips, I think there will no longer be a gap in ecosystem, but in chips it is "four times plus two years."

还有问题是,我们会不会向上游进行垂直整合?我希望不要。所以有个问题是,我们是不是会往上游应用进行垂直整合?我们希望不用这个,我希望是其 他人来做这个事情。我不 希望我把所有东西都吃了,我只要吃一块就好,我只要吃我最擅长那一块,或者我觉得我们自己认为最核心那一块,然后跟用户

Another question: whether we will vertically integrate upstream? I hope not. So the question is, whether we will vertically integrate into upstream applications? We hope not to do that; I hope others will do it. I don't want to take everything for myself; I just want one piece, the piece I am best at, or the core piece we believe is most important, and then work with users...

全文转写(下) Transcript (Part 2)

直接相关那一块,我觉得可能对于我们的很多产业伙伴来讲,他们比我们更关注类似……应该是别人的意义,不应该我把它诠释了。

For the directly relevant part, I think many of our industry partners are more concerned about it than we are... It should be someone else's interpretation, not for me to define.

未来我们会不会自建大型的集群?我觉得自建大型的集群是肯定要的,我们自己一直在做这个事情,我们所有的集群都是自己建的。但是未来要不要自研芯片,我觉得取决于这里的收益有多大,取决于收益有多大 。

Will we build large-scale clusters ourselves in the future? I think building large clusters is definitely necessary — we've been doing that all along, and all our clusters are self-built. But whether we'll develop our own chips in the future depends on the returns, on how much benefit there is.

Tesla、培育……这 个事情。假如说你是运营发电厂的,你并不一定需要去造发电机,对不对?发电设备可以是别人造的,只要它的价格合理,你为什么要自己 造?

Tesla, nurturing... this matter. If you operate a power plant, you don't necessarily need to build the generators, right? The power equipment can be made by others; as long as the price is reasonable, why would you build it yourself?

所以,我希望不用去做芯片。我希望能够以合理的价格买到芯片,这 样我就不用去做芯片。我 觉得这个很有可能是这样的,就是最近英伟达芯片的利润,不见得是可以……虽然……

So, I hope we don't have to make chips. I hope we can buy chips at reasonable prices, so we don't need to make them. I think it's very likely that the recent profit margins of Nvidia chips are not necessarily... although...

我们希望只做一块。我觉得 AI 这个事情很大,并不需要我……我只做一块。如果聚焦,并且我认为这里的生意利益已经足够大,就如果是 AI 时代会产生很多家万亿级别的公司,我觉得我们是其中一家。

We want to do just one part. I think AI is a huge matter, and we don't need to do everything—we only do one piece. If we stay focused, and I believe the business benefits here are already large enough, then in the AI era there will be many trillion-dollar companies, and I think we'll be one of them.

我们一直做其中一小块,我们是其中一家已经没有什么问题了,并不需要我……我并不是信心勃勃要做地方。我觉 得最可能是比较难变,且你真的想做别的话,你还是会有更多的主意,这个会……这是我们的态度。

We've always done a small part, and being one of them is already fine. We don't need to... I'm not overly ambitious to do everything. I think it's most likely hard to change, and if you really want to do something else, you'll have more ideas. That's our attitude.

譬如说,至少在 To B、To C 这两个业务上,目前能看到的,真的想做 To C 闭环的,反而我们做得好;真的想做 To B 闭环的,说不 定也没有我们做得好。你想要得越多……

For example, at least in the To B and To C businesses, from what we can see now, among those who truly want to achieve a closed loop in To C, we actually do it well; among those who want to close the loop in To B, perhaps none does better than us. The more you want...

然后还有,To B 这个可以多说几句。To B 业务的上限应该还是需求,在现在这一代 AGI 、AI 技术的背景下,To B 的需求应该是有限的。它会快速增长,但并不是一个无穷大的事情,最终还是受制于需 求,不是算力。

Also, I can say more about To B. The ceiling of To B business is still demand. In the context of the current generation of AGI and AI technology, the demand for To B should be limited. It will grow rapidly, but it's not an infinite thing; ultimately it is constrained by demand, not by computing power.

在当前技术的条件下,收入有多大,最终还是取决于需求。因为你这样想,我十个月就能收回成本,我肯定是如果有去处,我一块买,是因为没有那么多事情。对,需求应该会越来越大,如果随着技术持续有突破,这个 需求会越来越大。

Under current technological conditions, how much revenue you earn ultimately depends on demand. Because think about it—if I can recoup my cost in ten months, I would definitely buy more if there were somewhere to put it, but there just isn't that much to do. Yes, demand should keep growing, and if technology continues to break through, demand will grow even more.

多模态布局,我们一直在做。对产品来讲,它很重要;对 C 端用户产品来讲,它很重要。但是对智能的上限,它是一个组件,它不是主线本身。

We've been working on multimodal layout all along. It's important for products, especially for consumer-facing products. But for the upper limit of intelligence, it's a component, not the main thread itself.

但是它作为一个组件,多模态我们肯定会做,而且我们也在做。我们应该会上相关的模型,就是我们 V4、V4 的后续版本会支持原生的多模态。但是我们对多模态、对智能来讲,它是个组件,我们不把它当作智能本身,它 的主线跟搜索一样。

But as a component, we will definitely do multimodal—and we are doing it. We should release related models; our V4 and subsequent versions will support native multimodal. But we see multimodal as a component of intelligence, not intelligence itself. Its main line is like search.

搜索也是一个组件,多模态也可以,我们的理解它也是一个组件。Scaling,我们是信 Scaling,肯定是规模越大,效果越好,能够解锁更多的功能。阻止我们 Scaling 的其实就是算力,并不是我们不想 Scaling,是我们没有那么多的算力去做 这个 Scaling。

Search is also a component, and multimodal is similarly, in our understanding, a component. We believe in scaling; surely the larger the scale, the better the effect, unlocking more capabilities. What prevents us from scaling is compute power. It's not that we don't want to scale; it's that we don't have enough compute power to do so.

我们没有触摸到这个上限在哪里。我们训练这么大的模型,并不是因为我觉得这么大的模型就够了,而是我刚好有这么多资源。我是按照我的资源来算,我能够接受、能够训练的 模型是多大,是这样算出来的,并不是这个模型就够了。

We haven't touched the upper limit. We train such large models not because we think this size is sufficient, but because we happen to have these resources. We calculate based on our resources—what model size we can afford and train—that's how it's determined, not because this model is enough.

目前来看,模型收益还是非常明显的。我们还没有机会碰到这个 Scaling 的墙,这离我们还很远。硅谷在说 Scaling 到头的时候,那是对硅谷来说;对中国人来讲,我们离那个还很远,我们根本就没有 Scaling 到那个程度。

So far, the benefits of scaling models are still very obvious. We haven't had the chance to hit the scaling wall; it's far from us. When Silicon Valley talks about scaling hitting a plateau, that's true for Silicon Valley; for the Chinese, we are still far from that—we haven't scaled to that extent at all.

这个 Scaling 包括数据的 Scaling、模型规模的 Scaling,然后训练成本。我们离探索这个 Scaling 的上限还比较远,就是没有那么多算力。但我们也会不遗余力地去推进这个 Scaling 的上限。

Scaling includes scaling of data, model size, and training cost. We are still far from exploring the limits of scaling, simply because we lack compute power. But we will spare no effort to push the boundaries of scaling.

包括我们融资完之后,我们有更多的算力,可能训练更大的模型。时间内买到,就在越短时间内买到;如果能半年精力全花完,那么这是最好的,实际上做不到。

After we complete financing, we will have more compute power and can possibly train larger models. We will buy compute in as short a time as possible; if we could spend all our energy in half a year, that would be best, but in reality it's not achievable.

下一代模型的核心能力,我觉得它必须要有持续学习的能力,它才能叫下一代模型。在那之前,我们 能做的就是成本,然后效果做得更好,速度做得更快。但是要有大突破,它应该是具备持续学习的。

The core capability of the next-generation model, I think, must include continual learning for it to be called the next generation. Until then, all we can do is improve cost, effectiveness, and speed. But for a major breakthrough, it should possess continual learning.

然后还有一个问题,就是为什么好像我们特别在乎这个模型的计算效率?因为我确实发现,不是所有人都 很在乎这个模型的效率。一个 商业化公司来讲,没有动力去追求模型的效率,因为模型效率高一点……

Then there's another question: why do we seem to care so much about the computational efficiency of the model? Because I've indeed noticed that not everyone cares much about model efficiency. For a commercial company, there is no incentive to pursue model efficiency, because if the model is more efficient...

所以,它没有那么大的动力去追求模型效率要非常高。所以几个创业公司,他们自己并不会……你没有听 到他们说低成本是他们追 求的东西,因为那个不符合他们的利益。低成本了,你还赚什 么钱?低成本了,你就收不到什么钱了。

So, they don't have much incentive to pursue very high model efficiency. So several startups, they themselves won't... You haven't heard them say that low cost is something they pursue, because that doesn't align with their interests. If it's low cost, how can you make money? If it's low cost, you can't charge much.

相反,这是我们愿景的一部分。或者说,我们的愿景,我们的同学在乎这个成本,因为我们的同学都是普通人, 他知道用这个都是要花钱的。他能够有这个同理心:别人要花钱来用,那么别人如果便宜一点,别人能够接受的程度会更高。

On the contrary, this is part of our vision. Or rather, our vision—our colleagues care about this cost because they are ordinary people. They know that using this costs money. They can have this empathy: others have to pay to use it, so if it's cheaper for others, they will be more acceptable.

所以我们有很多同 学还是希望我们能够把 成本做到更低。但是如果你从商业角度 讲 的话,其实不会这么考虑。从商业来讲……

So many of our colleagues still hope that we can make the cost even lower. But from a business perspective, it wouldn't be considered this way. From a business standpoint...

对于不管是创业公 司来讲,还是大公司来讲,这个服务成本根本就不高。但是我觉得我们希望它是比较轻的,我希望它是一个负担得起的,特别是在中国算力紧缺的这样的背景下,是负担得起的,能在国产卡上用的。

For both startups and large companies, the service cost isn't actually high. But I think we want it to be lightweight, we want it to be affordable, especially in the context of China's computing power shortage, it should be affordable and runnable on domestic chips.

我觉得低成本首先是一个结果。我们的模型确实一直在模型架构上往一个更低成本的方向走,这跟我们的愿景有关系 。我们还有很多在算法上的方法,成本还可以往下走。

I think low cost is first of all a result. Our models have indeed been moving toward a lower-cost direction in model architecture, which is related to our vision. We also have many algorithmic methods that can further reduce costs.

成本往下走还有一个原因是,成本越低,我就越能训 练更大的模型,我就越能承担起更大的模型。在同样算力上,在算力有限的情况下,如果我的计算效率更高,我就能够承担起更大的模型。

Another reason for lowering cost is that the lower the cost, the larger models I can train, and the larger models I can afford. With the same computing power, under limited computing power, if my computational efficiency is higher, I can afford larger models.

对大公司来讲,他不一定会这么考虑。对大公司来讲,资源是可以加的,可以通过加资源来解决。 但是我们会优先考虑成本效率。

For big companies, they may not necessarily think this way. For big companies, resources can be added, and problems can be solved by adding more resources. But we would prioritize cost efficiency.

数据类模型的价值,数据这个范围比较广,数据应该几乎就等于模型的一半。为什么我会觉得,如果我想要那个,或者假如设 AI 能够占 GDP 的百分之二十,然后如果说我想要在这里面占 百分之五,绝对行不通。因为我肯定会被另外一个人打败,如 果另外一个人说我只要占百分之一,那么肯定就会被他打败。

The value of data-based models is broad. Data is a wide-ranging domain, and data should almost be half of the model. Why do I think so? If I wanted that, or suppose AI could account for 20% of GDP, and then if I wanted to take 5% of that, it absolutely would not work. Because I would definitely be defeated by another person. If that other person says they only need to take 1%, then I would surely be defeated by them.

如果说我的目标是,我要占 AI、全人类 GDP 的百分之五,理论上 算这个账还是成立的。你看 OpenAI,他算这个账好像是能算得过来的,理论上是没问题的。但是他有个问题,他会被另外一 个愿意只占百分之一的人打败。

If my goal is to take 5% of AI, that is, 5% of humanity's GDP, theoretically the math still works. Look at OpenAI; they seem to have calculated it, and theoretically it's fine. But they have a problem: they will be defeated by another person willing to take only 1%.

因为另外一个人说,我做得这么好,但是我只要拿全球 GDP 的百分之一就可以了,那么就会把他打败。这时候如果又出来另外一个人,说我只要百分之零点一就可以了,那么又会把前面的人给打败。

Because another person says, 'I do this well, but I only need to take 1% of global GDP,' and that will defeat him. Then if yet another person appears and says, 'I only need to take 0.1%,' that will defeat the previous one.

从宏观上来讲,不管这百分之几是从哪里拿,彼此之间没有区别,都是一样。拿得多的人会被拿得少的人打败。甚至你还不用真的拿得 多,愿景如果是拿得多的话,你就会被愿景是拿得少的人给打败。

From a macro perspective, it doesn't matter where these percentages are taken from; there is no difference between them—it's all the same. Those who take more will be defeated by those who take less. Even if you don't actually take more, if your vision is to take more, you will be defeated by someone whose vision is to take less.

其实大家都没有拿到钱,然后只是一个愿景。你愿景是拿得多,你就先输了,你就会面临着更大的困难 。这个世界就是这样。

In fact, no one has gotten the money yet; it's just a vision. If your vision is to take more, you lose first, and you will face greater difficulties. That's just how the world is.

OpenAI 从一开始觉得他真的能够垄断这个世界,但是实际上他会遇到很多很多挑战者。他会遇到挑战,他就不会那么轻松。美国会遇到过的挑战,那 么他在未来可能还会遇到中国的挑战,因为中国人愿意拿得更少,就可以给你提供这个服务。

OpenAI, from the very beginning, thought it could really monopolize the world, but in reality it will encounter many, many challengers. It will face challenges, and it won't be so easy. The U.S. has also faced challenges, and in the future, it may also face challenges from China, because the Chinese are willing to take even less and can provide this service to you.

在中国,也会有人愿意拿得更少一点。但最后这里会有个平衡,因为你拿得 太少之后,公司商业逻辑就不成立,就活不下去了。所以你拿得太少了,你活不下去;你拿太多了,你会被拿得少的人打败。

In China, there will also be people willing to take a little less. But in the end, there will be a balance, because if you take too little, the company's business logic won't hold, and it won't be able to survive. So if you take too little, you can't survive; if you take too much, you'll be defeated by those who take less.

所以对我们来讲,我们并不是利润要拿最多的钱,或者说算收益最大化的定价,而是只赚一个合理的收益。这是一个解释。

So for us, it's not about maximizing profits or setting prices to maximize returns—we aim to earn a reasonable profit. That's one explanation.

我是相信这个事的,我并不是去为这个事情找理由,因为没必要找理由。我本来就这么做的,我这么做肯定 是有理由的。这个理由可能不是非常惯常,但是我觉 得公司本来就不惯常。

I believe in this, and I'm not making up reasons for it, because there's no need to. I've always acted this way, and there must be a reason for it. The reason might not be conventional, but I think the company itself is unconventional.

我们公司的管理其实是两条线:一条线是从上到下,一条是从下到上。从下而上,就是每个人自己想做什么,自己做,没有人管他,没有 KPI。

Our company's management actually follows two lines: one from top to bottom, and the other from bottom to top. Bottom-up means each person does what they want to do, with no one supervising them and no KPIs.

从上到下,就是正式的,我们要集体做一个什么事情,需要全公司的人一起来配合。比如说我们要发 V4,那么就得分工,每个人得做一部分。那个是从上到下,然后从上到下这个我们叫做正式。

Top-down is formal: when we need to do something collectively, everyone in the company must cooperate. For example, to release V4, we have to divide the work, and each person does their part. That's top-down, and we call that formal.

一般我们希望这正式不要占用员工所有时间的一半,就正式不要超过一半。他还有一半的时间,是不被安排的,他想做什么就做什么。这是一个研究的范围,让他可以自己去探索,按照他觉得什么重要,他去探索什么,没有前置的要求。

Generally, we hope formal work doesn't take up more than half of employees' time—formal should stay under half. The other half is unscheduled; they can do whatever they want. That's a research area for them to explore freely, according to what they think is important, with no preconditions.

只要公司能够支持,公司算力能够支持他做,或者说他不需要算力,他需要做得很少,那么他根本就不用来协调。所以我们现在是这样的一个组织方式。

As long as the company can support it—whether in terms of computing power or if it requires minimal resources—then they don't even need to coordinate. So this is how our organization currently works.

有些人觉得我们是从上而下的,有些人觉得我们是从下而上的,我觉得两个都对。我的一个标准是,正式最好不要超过一半 。

Some people think we're top-down, others think we're bottom-up. I think both are right. My standard is that formal should ideally not exceed half.

我们一般也不太加班。加班有两个原因。第一个是,做研究是需要一个比较松弛的环境。你如果逼 得很紧,就没法做研究。因为既然就是要你自己有这个兴趣,你自己平时要去想这些问题,所以得是在一个比较松弛的环境里,才有可能能够探索。这是一个出于研究文化的需要。

We generally don't work overtime. There are two reasons for overtime. First, research requires a relatively relaxed environment. If you push too hard, you can't do research. Since you need to have genuine interest and think about these problems in your daily life, only in a relaxed environment can you explore. This stems from the culture of research.

第二个是,我们非常聚焦。 我们非常聚焦,就意味着我们要做的事情很少。那我就没那么多事情要做,我就不需要加班。

Second, we are very focused. Being very focused means we have very few things to do. So I don't have that many tasks, and I don't need to work overtime.

这个跟前面的克制是一脉相承的。因为我克制,所以很多时候我就不做了。那么我要做的事情少了,,每个人分到的工作就少了。你看到我们的产品很多都不完善的,我们也没有去补它。

This is consistent with the restraint I mentioned earlier. Because I am restrained, I often choose not to do things. With fewer things to do, each person's workload is reduced. You can see that many of our products are not fully polished, and we haven't rushed to fix them.

请各位投资人自由开麦交流吧。我稍微提醒一下,梁文锋讲了很多一些比较敏感的信息,请大家千万不要外传一些数字或一些情况,包括卡量什么的。同时也千万不要录屏对外做分享。很感谢各位, 请各位有问题的话,自由开麦交流。

Investors, please feel free to speak up. A quick reminder: Liang Wenfeng has shared some sensitive information. Please do not disclose any numbers or details, including information about GPU quantities. Also, please do not record your screens or share externally. Thank you very much. If you have questions, please feel free to speak.

杨哥,您能分享 更多关于持续学习什么时候带来突破的时间线吗?包括如果实现持续学习,还需要哪些架构、算法的创新,以及其他关键要素?

Mr. Yang, could you share more about the timeline for when breakthroughs in continual learning might happen? And if continual learning is achieved, what architectural and algorithmic innovations, as well as other key elements, would be needed?

人少,需要研究。现在全世界都在研究这个问题。或者说,对投资人来讲,现在投资人看到的最多的 是 AGE NT;但对于我们这 些研究的人来讲,现在看到更多的是学习,以及怎么解决学习这个问题。

With few people, we need to do research. The whole world is currently researching this problem. Or, from an investor's perspective, what investors see most now is AGENT; but for us researchers, what we see more of now is learning, and how to solve the problem of learning.

其实学习可能不是一项技术,它是一个问题。怎么解决这个问题,可能有很多种技术,不是 一项技术,它 不是一个东西,会是很多东西。或者说,AGI 是由很多东西组成的,需要模型,然后还需要 很多其他东西。

Actually, learning may not be a single technology; it is a problem. There might be many technologies to solve this problem. It's not one technology, not one thing—it will be many things. Or rather, AGI is composed of many things, requiring models and many other components.

梁总,谢谢。非常感谢今天这个机会。我首先非常想回应和感激一下,您最开始说的让我非常感动,也给我们很多启发。您提 到,这个团队带着最大的善意,希望能够在这个行业里面,对这个社会、人类智能的发展有所推动,做一点点贡献。

Mr. Liang, thank you. I really appreciate this opportunity today. First, I want to respond and express gratitude for what you said at the beginning—it deeply moved me and gave us a lot of inspiration. You mentioned that the team carries the greatest goodwill, hoping to push forward the development of society and human intelligence, and make a small contribution to this industry.

并且在这个里面,是带着这种使命感和愿景。我觉得这跟我们所服务公司的企业文化非常相近,就是“修己达人”。我深刻地理解了为什么您会带领团队做开源的事情。

And in doing so, you carry this sense of mission and vision. I feel this is very similar to the corporate culture of the company we serve—'cultivating oneself to benefit others.' I now deeply understand why you have led the team to do open-source work.

我会形象地感觉,像我们去做一个榕树 这样的小鸟天堂生态,利万 物而不争,但这样它就会被大家接纳 ,万物共生共存,最终就会无处不在。所以,我们会以这次投资,同样表达我们对这个使命和愿景的认可、支持和尊重。

I would intuitively feel that if we build an ecosystem like the 'Bird Paradise' of a banyan tree, benefiting all things without contending, then it will be embraced by everyone, with all things coexisting and ultimately becoming ubiquitous. Therefore, through this investment, we also express our recognition, support, and respect for this mission and vision.

同时,我们也希望能够在 这个产业未来,我们所擅长的一些领域等,去贡献一些力量。在这块也想跟您继续请教和探讨。比如在未来生态 的共建当中,现在开源之后,这个行业里有多少伙伴、人才、团队,能够比较好地复现咱们开源目前的 一些模型和成果?

At the same time, we also hope to contribute to the future of this industry in the areas where we excel. I'd like to continue seeking your advice and discussing this. For example, in the co-construction of the future ecosystem, after the open-source release, how many partners, talents, and teams in the industry can effectively reproduce the models and achievements that we have open-sourced so far?

未来在下一步想让这个生态进一步发展的时候,您感觉在哪几个方面,需要更多高质量的人才能够 衔接到我们的模型,把它复现?还是说现在 GPU 的算力相对有一些稀缺?大家未来会不会是一 种模型矩阵的方式?

In the next step, when we want to further develop this ecosystem, in what aspects do you think we need more high-quality talent to connect with our models and reproduce them? Or is GPU computing power relatively scarce right now? Will everyone adopt a model-matrix approach in the future?

比如咱们把大模型基模做得越来越好,各行各业的伙伴和团队去做一些模型矩阵里面的垂 直行业模型,或者一些应用模型。这块目前的发展怎么样?未来一步一步,两年、三年,您感觉 会长成一个什么样的生态?

For example, as we make the foundation models better and better, partners and teams from various industries will develop vertical industry models or application models within the model matrix. How is this developing now? Step by step over the next two or three years, what kind of ecosystem do you think it will grow into?

这是我第一个想跟您请教的。第二个,您刚才也跟很多伙伴分享了很多关于 AI 硬件方面的观察。像 AI 全球大公司,可能单体都会做千亿美元级别的投入,中国目前看起来在硬件算力上有些短板。

That's my first question. Secondly, you just shared many observations about AI hardware with our partners. Global AI giants may each invest at the scale of a hundred billion dollars. China currently seems to have some shortcomings in hardware computing power.

您感觉这个多长时间可以解决,并支撑咱们 AI 的发展,让 算力和硬件短板不给 AGI 拖后腿?您觉得这是不是中国人未来使命必达,我们肯定能做出来,只是时间和资金投入的问题?

How long do you think it will take to resolve this and support the development of our AI, so that computing power and hardware shortcomings do not hold back AGI? Do you think this is a mission that we Chinese are determined to accomplish? We can surely do it; it's just a matter of time and financial investment?

但同时,可能它会是两方面的。一方面,模型的进步会让模型智能化的提升,导致单一任务或者某些智能化 单体对硬件、对算力的消耗逐渐递减,不再需要 那么大算力的运算,因为模型的进步会让 它巧算。我不知道我理解得对不对。

At the same time, it might be twofold. On one hand, model improvements will enhance model intelligence, causing the consumption of hardware and computing power for a single task or certain intelligent entities to gradually decrease. They will no longer need such massive computation, because model improvements will enable more clever computation. I don't know if my understanding is correct.

另外一方面,硬件的技术进步会让算力的算能效能更强大。这会不会是两边相向而行的路径?现在如果是 在这个时点,用现在的 960 也好,还是 H200 也好,去做千亿美元级别的算力投入,您刚才提到说您是给它按三年摊销,那它的实际生命周期,您觉得技术迭代是四五年?

On the other hand, advances in hardware technology will make computing power more efficient. Could this be a path where the two sides move toward each other? If at this point in time, using the current 960 or H200 to invest in computing power at the scale of a hundred billion dollars, you mentioned earlier that you amortize it over three years. What do you think is its actual lifecycle? Is the technology iteration cycle four or five years?

或者直白地说,会不会现在算力中心 按现在的卡去建了万卡集群,可能三年之后它其实就是相对不那么先进的算力了?会不会有现阶段是 under construction、不够用,三年之后变成相对不那 么优质的算力有冗余的情况 ?我不知道会不会有这样一种现 象。以上两个问题请教您,谢谢。

Or, to put it bluntly, if computing centers build 10,000-card clusters with current cards, might they become relatively less advanced computing power in three years? Could there be a situation where at this stage they are under construction and insufficient, but in three years they become redundant as relatively lower-quality computing power? I don't know if such a phenomenon will occur. I would like to ask you these two questions, thank you.

我们 现在觉得,可能每个企业都面临的问题是人才不够。但我觉得这个人才 短缺会是阶段性的。我们在每个行业发展的初期,人才都是不够的。

We now think that the problem every enterprise may face is a shortage of talent. But I believe this talent shortage will be temporary. In the early stages of every industry, talent is always insufficient.

包括以前做网站,刚开始做网站的 时候,做网站的人很少,人才很缺。后来互联网要做服务端,人才也是很缺的。但是这种人才短缺都非常快会被解决, 也就两三年,因为会培养出大量的人。

For example, when we first started building websites, there were very few people doing it, and talent was scarce. Later, when the internet needed server-side development, talent was also scarce. But these talent shortages were resolved very quickly, within just two or three years, because a large number of people were trained.

AI 人才的短缺也是阶段性的,并且我们已经看到,大幅度被缓解了。因为 AI 人真的不缺,每个公司很快会把人培养出来,培养人是很快的。所以,AI 这个行业整体上 ,不管是生态、模型公司还是什么,人才都不缺。

The shortage of AI talent is also temporary, and we have already seen it greatly alleviated. Because AI talent is really not in short supply; every company can quickly train people, and training people is fast. So overall, in the AI industry, whether in the ecosystem, model companies, or whatever, there is no shortage of talent.

人才缺肯定是一个短期现象。历史上从来没有出现过长期缺某一类人的情况。我还记得十几年前说飞行员很缺,飞行员的培养周期很长,但也很快被解决 了 。

A talent shortage is definitely a short-term phenomenon. In history, there has never been a long-term shortage of a certain type of person. I remember more than a decade ago it was said that pilots were in short supply, and the pilot training cycle was long, but it was also quickly resolved.

所以大家不用担心缺人才的问题。以及国内现在做模型的公司有点太多了,还是太多了。美国可能就三家,中国做基模的东西太多了。

So there is no need to worry about talent shortages. Also, there are too many companies doing models in China now, really too many. The US may have only three, while China has too many doing foundation models.

最终一定是不需要那么多人去做基模的,一定会收敛。所以资源也是比较分散,某种程度上也比较浪费。就先于

In the end, it is certain that not that many people are needed to build foundation models; it will definitely converge. So resources are also quite dispersed, and to some extent, quite wasteful. So first,

每一家都要做同样的事情,但美国只要三家做,资源只集中在这三家。中国资源分得很散,每一家拿到的资源就更少。我觉得这个肯定是会收敛的,一定会,但这需要过程,最终一定会收敛。

不需要那么多家,因为现 在可能大家觉得做这个事情的利 润率非常高,所以一定 要自己做。但当他发现这个事情可能没有那么高利润的时候,可能就不做了。最近肯定是没有那么高利润的,我不相信有那么高利润, 因为这不符合客观规律。

We don't need that many companies, because right now people probably think the profit margin for this is very high, so they insist on doing it themselves. But when they discover that it may not be so profitable, they might stop. Certainly, there is no such high profit recently; I don't believe there is such high profit, because it doesn't conform to objective laws.

这意味着我们是处在一个阶段上:如果有一个非常高的利润率,这一定不符合客观规律。我们应该是一个合理的利润。所以这是产业的一个现状,我觉 得肯定会收敛。

This means we are at a stage where if there is an extremely high profit margin, it must not be in line with objective laws. We should have a reasonable profit. So this is the current state of the industry, and I think it will definitely converge.

就是大家做大模型的那一部分,其中不要说某一家独占,说“我要拿走全部利润”,这个肯 定不行。如果说每一家都只拿合理的利润,那么其实不需要那么多 人去做 大模型。中国最后 有个三四家竞争,竞争就很充 分了,价格绝对已经够 打价格战了。

As for everyone building large models, it's not about one company monopolizing and saying, 'I'm taking all the profits'—that definitely won't work. If each company only takes a reasonable profit, then there's actually no need for so many people to build large models. In the end, China will have three or four companies competing, which is already sufficient competition, and prices will definitely be low enough to trigger price wars.

至于生态上的,我没有什么太多的想法。我们希望能够扶持更多的人,但是我们并没有那么多的精力。我们是有这个意愿,并且不会有利益冲突,但是我们有没有去做是另外一回事。

As for the ecosystem, I don't have much to say. We hope to support more people, but we don't have that much energy. We have the willingness, and there won't be conflicts of interest, but whether we actually do it is another matter.

但至少这 里边是没有利益冲突的,我 们是希望合作共赢的。 首先,我绝对不认为大模型公司可以拿走大部分利润,这个不可能,因为这么多家大模型公司,现在差距不用那么大。

But at least there is no conflict of interest here; we hope for win-win cooperation. First, I absolutely don't think a large model company can take away most of the profits—that's impossible, because with so many large model companies, the gap isn't that big now.

差距 只 有两个东西:一个是时间,一个是成本。所以不至于 哪一家有暴利,我觉得不至于有暴利。 成本控制得好的人就多赚 一点,成本控制得差的人就少赚一点,仅仅此而已。

The gaps are only in two things: time and cost. Therefore, it's unlikely that any company will have windfall profits; I don't think there will be windfall profits. Those with good cost control will earn more, and those with poor cost control will earn less—that's all.

大家能……你相信以后肯定是有很多人可以……未来其实会……大家数据应用这些的循环迭代……

Everyone can... You believe that in the future there will definitely be many people who can... In fact, the future will... the cyclic iteration of data applications and so on...

现在可以听到吗?谢谢。对,感谢您的回答,也非常深刻地理 解和尊重您的这种行业里面的生态战略定位。比如说数据这一块,现在公开的数据,相信模型公司都已经可以有渠道获取,这个方法应该都不 成问题,只是时间和成本的问题。

Can you hear me now? Thank you. Yes, thank you for your answer. I also deeply understand and respect your ecosystem strategic positioning in the industry. For example, regarding data, for the currently public data, I believe model companies already have channels to obtain it. This method shouldn't be a problem; it's just a matter of time and cost.

那么后续比如说将到真正到 AGI 的时候,有可能大家一个设 想或者理想的状态是,模型可以自我迭代、自我学习,就是自己训练自己。那么 这一块的话,目前这个数据,您感觉仿真数据是不是可以用起来,还是说真实数据的质量最高?

Then later, for example, when we truly reach AGI, a possible vision or ideal state is that models can self-iterate, self-learn, that is, train themselves. In this regard, with the current data, do you think simulated data can be used, or does real data have the highest quality?

如果说还是需要来自于真实数据,那会不会是限制这个 AI 的智能还是在人类的过 往……这个 层面,因为它依 赖的是人 类真正曾经有过的真实数 据?还是说可以突破这个上限,通过模拟数据、仿真数据、创造数据等等的方式,让这个模型能力去超越人类过往的所有真实……

If real data is still needed, wouldn't that limit AI's intelligence to the scope of human history... because it depends on the real data that humans have actually had? Or can it break through this ceiling through simulated data, synthetic data, created data, etc., enabling the model's capabilities to surpass all of humanity's past reality?

我觉得是能超越的。我觉得有两点超越的,比如说围棋,AlphaGo 他下了一手人类从来没有见到过的棋。就是说,他肯定是在一定的范围内超越人类 的。

I think it can surpass. I think there are two aspects of surpassing. For example, in Go, AlphaGo made a move that humans had never seen before. That is, it definitely surpasses humans within a certain range.

但是他可能也有上限,他可能也是有局限性。但是这个局限性我们现在看不到。我们认为,所以笼 统地认为,它是可以基于人类已经有的、我们已经能说出来的知识上,予以超越的。

But it may also have an upper limit, and it may have limitations. However, we cannot see these limitations now. So we generally believe that it can surpass based on the knowledge that humans already have and have been able to articulate.

好了,我简单快速地重复一下。就是想请教,对于 AI Infra 这一块,未来相信算力现在大家都是千亿美元级别地在投入。那么这块的话,有可能我们相信中国人 未来在硬件上是使命必达,有一天可能会有高效率的算力,但有可能这个在实践的过程当中,目前还是一个 掣肘。

Okay, let me quickly recap. I'd like to ask, regarding AI Infra, in the future, computing power is something everyone is investing in at the hundred-billion-dollar level. In this regard, we may believe that the Chinese people will achieve their mission in hardware, and one day there may be highly efficient computing power, but in practice, it is still a constraint at present.

那么未来会不会 是两方面向下而行?一方面是模型的能力迭代之后,它其实对于算力从硬算变成巧算,所以单位模型或者任务对算力的要求和消耗会逐步地边际降低。另一方面,比如 说硬件像卡的能力提升,会让迭代速度越来越快,单卡效率提升。

In the future, will it move in two directions? On one hand, as model capabilities iterate, computing power shifts from brute-force to intelligent computation, so the per-model or per-task demand and consumption of computing power will gradually diminish at the margin. On the other hand, improvements in hardware, such as card capabilities, will make iteration faster and increase single-card efficiency.

那么这个在过程当中是怎么样的一个现象?会不会是说,现在去建了万卡集群,去买了 H200 或者 960,但是过个两三年,它就变 成了一个相对没有那么优质的算 力,相对又变成了一个陈旧的器件?

So what is the phenomenon in this process? Is it possible that if we build a 10,000-card cluster now and buy H200 or 960, in two or three years, it will become relatively less high-quality computing power and become an obsolete component?

英伟达的卡基本上你可以按照五年折旧。华为的卡最多按三年折 旧吧。华为 950 今年 能用还挺好的 ,明年用我觉得还可以,后面再用我觉得可能就真的太费电了。

NVIDIA cards can basically be depreciated over five years. Huawei cards should be depreciated over at most three years. Huawei 950 works well this year, it's still okay next year, but I think using it any longer would really consume too much electricity.

华为卡生命周期肯定是会短一点,因为它本来就已经比英伟达晚两年了。但是我觉得差距没那么大。

Huawei's card lifecycle will definitely be shorter, because it was already two years behind Nvidia. But I think the gap isn't that big.

如果说 B200 现在能买到多少,我觉得都划算。假如说对于腾讯来讲,阿里巴巴能买到的话,再看量;如果合理价格能买到,肯定都是划算的。但不是一个算成本的时候,相信买不到。

If you can buy B200s now, no matter how many, I think it's worth it. For Tencent, for instance, if Alibaba can get them, then it depends on quantity; if they can be bought at a reasonable price, it's definitely worthwhile. But this isn't the time to count costs; I believe you can't buy them.

明白。我们算力落后,这是一个事实。这个事实通过三方面来消解。第一方面是我们承受模型落后,我们只能够 用 比他们更小的模型,小多少 问题就是训练 。我们要承受一定的模型落后,以及模型的尺寸更小。

Understood. Our compute power is behind; that's a fact. This fact can be mitigated in three ways. First, we accept that our models are behind; we can only use smaller models than theirs. How much smaller is a matter of training. We have to accept a certain model gap, as well as smaller model sizes.

这落后有一个好处,落后意味着你有更多的时间,你有一定的技术,然后这样的话你就可以用巧妙的方法

Being behind has one advantage: it means you have more time, you have certain skills, and then you can use clever methods.

所以我们跟美国的差距可能是落后美国 12 个月,落后美国可能 12 到 18 个月,或者说 6 到 12 个月。反正简单说,就是落后美国两年,然后只用美国二十分之一的算力把这个事情做出来。

So the gap between us and the U.S. might be 12 months behind, maybe 12 to 18 months behind, or 6 to 12 months. In short, we're two years behind the U.S., yet we got it done using only one-twentieth of the U.S.'s compute.

这个叙事就是落后一到两年,但是只用它二十分之 一的算力。那么未来我们要把这个叙事改写,就是我们用它几分之一的算力,但是把这个时间缩得更短,缩到 6 个月、3 个月,我觉得这是一个目标。

That narrative is being one to two years behind, but using only one-twentieth of the compute. In the future, we want to rewrite that narrative: using a fraction of the compute, but shortening the time to 6 months or 3 months. I think that's a goal.

以及我们甚至可以在某一些方面超越他们。 但是在总体算力还是有数量级差距的情况下,全面超越是不现实的;但是在某一些重点、有取舍的地方,我们有一些地方超越可能是可以的。

And we might even surpass them in certain areas. But when the overall compute still has an order-of-magnitude gap, comprehensive surpassing is unrealistic; however, in some key areas with trade-offs, we might be able to surpass them in some places.

感谢刚才的分享。我这边有两个快速的技术问题。在刚才聊技术路线当中,提到了我们这一阶段 要解决的核心问题是 持续学习,也是目前国外研 究的热点,叫 Recursive Improvement。想请教一下,目前来看,从技术上最大的难点是什么?从您的视角来看,这个什么时候可以解决?这是第一个问题。

Thank you for sharing just now. I have two quick technical questions. In the earlier discussion of the technical roadmap, you mentioned that the core problem we need to solve at this stage is continual learning, which is also a hot research topic abroad, called Recursive Improvement. I'd like to ask: from a technical standpoint, what is the biggest challenge right now? From your perspective, when might this be solved? That's the first question.

第二个问题,同时您刚才提到,先解决持续学习,然后再去做智能。 我也想理解一下,您这么提 背后的技术根源是什么?是不是意味着解决完持续学习问题以后,DeepSeek 后续也会做通用智能?这两个问题请 教一下。

Second question, and you just mentioned solving continual learning first, and then pursuing intelligence. I'd also like to understand the technical rationale behind this. Does it mean that after solving the continual learning problem, DeepSeek will subsequently also work on general intelligence? I'd like to ask about these two questions.

技术问题其实 解释起来有点……它的难 点在于,我们现在还没有找到非常 work 的方法。全世界 都还没有找到好的方法,大家都在摸索。所以现在还在摸索阶段,就是不知道谁能摸到下一个解决这个问题的方法。

Explaining the technical problem is a bit... The difficulty is that we haven't yet found a method that really works. The whole world hasn't found a good method either; everyone is still exploring. So we're still in the exploratory stage, and it's unclear who will stumble upon the next solution to this problem.

现在还在探索阶段。我们有很多思路,有很多现在看起来有前途的一些想法,但是都还没有做通。对,这是第一个。

We are still in the exploration phase. We have many ideas, many that seem promising right now, but none have been fully realized yet. Yes, that's the first point.

第二个是,现在我们内部比较看重这 样一个叙事:训练我们的下一版模型, 我们希望它能够帮助我们 自己的开发。它能够提升 DeepSeek 的效率,我们的模型最首先是提高 DeepSeek 自己的工作 效率,让我们开发下一版模型的时候,它能够提供更多的帮助。

Secondly, internally we now place great emphasis on the narrative: when training our next model, we hope it can assist us in our own development. It can improve DeepSeek's efficiency; our model primarily increases DeepSeek's own work efficiency, so that when we develop the next version, it can provide more help.

或者说简单一点,我们做的模型,第一目标不是大家用得好用,而是我们自己用得好用。首先是对我们自己有用。对我们自己有用之后 ,我在开发下一版模型的时候就会更快。

Or to put it simply, the primary goal of our models is not for everyone to use well, but for ourselves to use well. First and foremost, it must be useful to us. Once it is useful to us, we can develop the next model faster.

我们内部很多人的想法是这样的: 首先要对我们自己有用,首先是给我 们自己用。然后这是实现 AGI 最快的方法。当我们自己好用,那意味着可能别人也 好用,但是首先得保证我们自己好用。

Many people inside the company think like this: first it needs to be useful to us, first for our own use. This is the fastest way to achieve AGI. When it works well for us, it may also work well for others, but we must ensure it works well for us first.

这个叙事有点奇怪,但是确实很多人就是这么想的。 而不是说用户用得好用,我是希望对我 们自己帮助更大,这样我们可以实现 AGI,会快很多。所以这个叙 事的逻辑是,来帮助我们实现 AGI。

This narrative is a bit odd, but indeed many people think this way. Rather than focusing on how well it works for users, I hope it can be more helpful to ourselves, so that we can achieve AGI much faster. So the logic of this narrative is that it helps us achieve AGI.

但首先是帮助我们实现。我们实现 AG I 需要这个帮助。现在是非常确定,我们确实需要人工智能来帮助我们实现 AGI。虽然说它还不是自主地工作,它还是只是跟人搭配,但是已经很有用了。

But first, it helps us achieve it. We need this help to achieve AGI. It is now very certain that we do indeed need AI to help us achieve AGI. Although it doesn't work autonomously yet and still requires human collaboration, it is already very useful.

第二个问题,就是刚才提到了,先解决持续学习的问题,然后再进入通用智能。这是您对后续 的一个预期吗?想 理解一下这背后的技 术根源是什么?为什么先要解决持续学习,然后再进入通用智能?您对这块后续的理解 。

The second question is about what was just mentioned: first solving the problem of continuous learning, and then moving on to AGI. Is this your expectation for the future? I'd like to understand the technical rationale behind it. Why do we need to solve continuous learning before entering AGI? What is your understanding of this going forward?

因为解决持续学习这个问题 ,可以大大加快我们的研发进度。如果我先解 决了持续学习这个问题, 那么通用智能这个问题就不在话下了。我有了 AI 的辅助,如果 AI 能够持续学习,它的能力应该是非常强的。

Because solving continuous learning can greatly accelerate our R&D progress. If we solve continuous learning first, then the AGI problem will be a piece of cake. With AI's assistance, if AI can learn continuously, its capabilities will be extremely powerful.

现在的 Agent 的能力受限,是因 为它不能持续学习,它不能有效地持续学习。如果说能够先把持续学习做完,那 AI 的能力是非常强的,它能够非常大地提升我们自己研究 的效率。

Current agents are limited in capability because they cannot learn continuously—or at least not effectively. If we can first achieve continuous learning, AI's capabilities will be very strong, and it can greatly improve the efficiency of our own research.

持续学习先做出来,通用智能可能就很容易了,用它来做就很容易。所以我说这是一个我们比较希望看到的结果,我们比 较省力, 我们就轻松。否则现在你要去人工做通用 智能, 它是一个比较累、 比较苦,是一个数据密集、人力密集的事情,性价比也不高。

Once continuous learning is achieved, AGI might become much easier—using it to build AGI would be straightforward. So I'd say this is the outcome we hope to see, as it would be much more labor-saving and easier for us. Otherwise, if we had to manually work on AGI now, it would be tiring and arduous—data-intensive and labor-intensive, with low cost-effectiveness.

小问题,线上的问题你看一下聊天群。觉得 AGI 还需要多久?国内硬 件在 这个时间能追上吗?

A small question—please check the chat group for the online questions. How much longer do you think AGI will take? Can domestic hardware catch up in that time?

华为 950,现在华为是给我们一万六千卡,这应该是可以公开说的。应该 是比互联网大厂少一个数量级。华为也只能给我们这么多,因为这个价格也不便宜。

Huawei 950—currently Huawei is providing us with 16,000 cards. This should be something we can disclose publicly. That's about an order of magnitude less than what internet giants have. Huawei can only give us this many, because the price is not cheap.

互联网大厂的追求会更大,对互联网大厂来讲,它更需要。对我们来讲,我们可以买一些不合规的卡。所以我们买华为 950 的目的,还是希望帮华为把这个生态做好。

Internet giants have even bigger aspirations and a greater need for them. For us, we can buy some non-compliant cards. So our purpose in purchasing Huawei 950 is actually to help Huawei improve its ecosystem.

一万六千卡的华为 950, 只相当于四千卡的 B 系列。所以说不是一个很大的量,意义不是很大。不够训一个下一代的模型,它只够训我们现在这一代模型,不够训下一代模型。但是可以让华为把这个先做好,就是关于华为 950。

Huawei's 16,000 cards are equivalent to only 4,000 cards of the B-series. So it's not a very large quantity, and its significance is limited. It's not enough to train a next-generation model; it's only enough to train our current generation, not the next. However, it can allow Huawei to get things done first—that's the point regarding Huawei 950.

我觉得在 AI 这个事业上面,在 AI 这个事情上,应 该国内一两年是能够做到跟国外差不多的,或者可能今年就能 做到平替国外的模型。在 AI 这个事情上,就现在的一个方式、现在这个范式,不是很难的事情,所以今年应该就能做到的。但它还不是 AGI。

I think in the AI industry, on this AI matter, domestic efforts should be able to match those abroad within a year or two, or perhaps this year we can already achieve models that are on par with foreign ones. In the AI field, given the current approach and paradigm, it's not very difficult, so we should be able to do it this year. But it's not AGI yet.

国产硬件在这个时间能追上吗?我觉得国产方面可能需要一个几年的时间。国内首先得解决生态的问题,因为生态是一个信心的问题。解决生态的问题,然后再解决产能的问题,我觉得应该是能 够逐步解决的。

Can domestic hardware catch up in this timeframe? I think it might take a few years for domestic products. First, we need to address the ecosystem issue, because the ecosystem is a matter of confidence. After solving the ecosystem problem, then we can tackle the capacity issue, which I believe can be resolved gradually.

我不太相信未来五年之后,我们还卡在产能的问题上。现在肯定是卡在产能问题上,今年、明年、后年,我觉得可能都还是卡在产能问题上,但五年之后,我觉得可能不一定,我还是比较乐观的。

I don't really believe that five years from now we'll still be stuck on the capacity issue. Right now we definitely are stuck on capacity; for this year, next year, and the year after, I think we might still be stuck, but five years later, I think that may no longer be the case. I'm rather optimistic.

然后第二个问题是,思考未来的组织架构,对人员的规划规模。首先,我们之前的组织架构是非常分散的,因为没有组织架构。但是我 们人在进一 步扩大之后,这些肯定是需要做出 一些改变的。

Next, the second question is about thinking about the future organizational structure and the scale of staffing plans. First, our previous organizational structure was very scattered because there was no structure. But as our team further expands, these aspects certainly need some changes.

现在只能说,这里会需要做很多改变,但现在很难一下子全部表达出来。最后应该还是要分不同的 部门,应该有些部门是我们需要建立起比较严谨的层级结构的。另外一些部门,可能还是会维持一个比较松散、比较扁平的结构。

For now, I can only say that many changes will be needed, but it's hard to articulate them all at once. In the end, we should still divide into different departments; for some departments we need to establish a rather rigorous hierarchical structure. Other departments may remain relatively loose and flat.

随着人员的增加,我们会做这个调整。应该是我们马上就得做这个调整,因为我已经在做这个调整。已 经不做这个调整的话,很多事情是没法推进下去的。确实有很多部门,是应该有组织架构的。

As the workforce grows, we will make this adjustment. In fact, we need to make it immediately, because I've already started doing it. If we don't make this adjustment, many things cannot move forward. Indeed, many departments should have an organizational structure.

然后大家还有问题是,CV 是先 于它的哪个版本是吧?我觉得我们现在,线上发了这个 GCV4 这个版本,还是比较粗糙的,然后能力方面还很多都需要时间。

Then you also asked about which version CV is ahead of, right? I think right now, the GCV4 version we've released online is still quite rough, and many aspects of its capabilities still need time.

我们一般,对我来讲,一个比较舒适的发版节 奏,大概是两三个月发一版。上次发版可能是四月底,那么下次发版可能是六月底,大概是这样子。如果没有意外的话,应该还是会一版比一版好的。

For me, a comfortable release cadence is generally about one version every two to three months. The last release was probably at the end of April, so the next one might be at the end of June, something like that. Barring any unexpected circumstances, it should get better with each version.

在 50B 激活的这个尺度上面,我感觉最终我们跟现在的这波开源不会 有太大差别。就在推理速度跟效 果上,我感觉可能就不会有太大的差别。

At the scale of 50B activation, I feel that in the end, we won't be too different from the current wave of open-source models. In terms of inference speed and effectiveness, I think there might not be a big difference.

但是跟它那个大尺寸的模型,就它一个没公开的模型 ,应该差距还是比较大的。那个差距应该是,我们这个激活参数应该是做不到的,这就肯定得是一个更大尺寸的模型,可能比如说是 1 50B。

But compared to its large-scale model, an unpublished model, the gap should be quite significant. That gap is something that our activated parameters cannot achieve; it would definitely require a larger model, perhaps like 150B.

那么 150 B,我觉得 按照我们现在训练的进度,今年,乐观的是今年年底可以开始去训练,至少是明年四……差距还是比较大的。对,这是跟 OCE 的差距。

As for 150B, I think based on our current training progress, optimistically we can start training by the end of this year, or at least by April next year... The gap is still significant. Yes, that's the gap with OCE.

你好,我其实有一个问题,就 是经常您说到这个 AGI,它实现的过程是一个渐进式的过程,而 不 是会有一个突变。所以我可以理解,这是一个没有临界点的过程吗?

Hello, I actually have a question. You often mention that AGI's realization is a gradual process, not something that happens suddenly. So can I understand that this is a process without a critical point?

它没有临界点,但是它是非线性的。我们现在比较相信一 个叙事是,AI 可以加速 AI 的研究,AI 可以加速 AI 的研究。就是说, 它不是线性的,因为你可以用 A I 来加速你自己的研究,所以它到后面可能是非线性的。

It has no critical point, but it is nonlinear. We currently believe in a narrative that AI can accelerate AI research. That is, it's not linear, because you can use AI to accelerate your own research, so it might become nonlinear later.

明白。所以当下您的这个,我理解前面的结论可能就是,语言模型继续 scaling 是够的,是足够的,做到这 个状态。

Understood. So currently, my understanding of your earlier conclusion is that continuing to scale language models is sufficient, enough to reach this state.

只能说语言模 型的 scaling,我现在没有看到有上限。我们现在的智力水平,或造成美国的智力水平,都没有看 到上限。

I can only say that for language model scaling, I currently see no upper limit. Our current level of intelligence, or rather the intelligence level in the US, has not seen an upper limit either.

明白。因为我 非常好奇,就是你说,对美国来讲,800B 激活这个模型,说比它能训练起,但它也用不起来,所以它只能训练出来,它也很难拿出来给大家用了,因为确实有点贵。

Understood. Because I'm very curious, you said that for the US, a model with 800B activated parameters, although it can be trained, cannot be used in practice. So it can only be trained, but it's hard to make it available for everyone because it's really expensive.

我之前其实很好 奇的一点就是,人类其实拥有语言 能力就是近十万年的事情,然后前面可能进化了三十七亿年。但在训练 AI 这个事情上,可能是可以,你说那个顺序是可以反过来的,但最终可能还是会进入到所谓,也不一定叫世界的模型,还是 要进入到物理模型或者具身的那部分,是吧?就是在这个上限之后。

One thing I've been curious about is that humans have had language ability for only the last hundred thousand years, while evolution may have gone on for 3.7 billion years before that. But in the case of training AI, perhaps that order can be reversed. Yet ultimately, it might still enter into what could be called a world model, or rather a physical model or embodied part, right? That is, after reaching this upper limit.

对,我觉得具身肯定还是要进入,最终具身。所以对我们公司来讲,自然而言,可能终点都是具身。因为对一个正常人来讲,他的需求并不是电脑,对不对?因为正常的人,他吃喝玩乐、衣食住行,他不需要电脑。

Yes, I think embodiment definitely needs to come in; ultimately, it will be embodied. So for our company, naturally, the endpoint might be embodiment. Because for an average person, what they need is not a computer, right? Normal people have needs like eating, drinking, having fun, clothing, food, shelter, and transportation; they don't need a computer.

他需要的是,所以他还是需要具 身 智能来解决具体人力的需求。 如果说目标是解约人力需求的话,所以具 身我觉得就是绕不过去的。

What they need is embodied intelligence to address specific labor needs. If the goal is to reduce human labor demands, then embodiment, I think, is unavoidable.

明白。所以阶段性地讲,假如说到达类似,也不一定是临界点,就是可以自进 化、可以 比较好的自进化,或者接近这个 与 SV 的这个上升的到那个时间点上,我很好奇会希 望这个状态的 AI 落地的第一个让是……

Understood. So stage-wise, let's say we reach something like—not necessarily a critical point—but the ability to self-evolve well, or near that point of ascent to something like AGI, I'm curious: at that point, what would be the first area you'd want this state of AI to be applied to?

可能跟现在不一样。我们是希望它能够,就是直接假如说没有具身,那么我们对 AGI 的定义,或者我们希望 AGI 能做什么呢?它能帮我迭代下一版模型,能帮我迭代下一版模型一样。

It might be different from now. We would hope that it can—directly, if without embodiment, then what would be our definition of AGI, or what would we hope AGI could do? It could help me iterate the next version of the model, just like that.

然后如果 有了具身之后,我们希望它做 的也是,让它来迭代下一版的具身,它 来做下一版机器人。

And then if we have embodiment, we would hope it would also be able to iterate the next version of embodiment—it would make the next version of robots.

我还是很好奇一点,就是因为我看之前 DeepSeek 的采访等等,应该是在一些重要的方向的 选择和研究的选择上面,品位和直觉是很重要的,而不是简单的工程优化。那如果之后 AI 可 以 自 进化的话,这个品位、taste、直觉 这些东西还重要,或者说重要的会是什么?

I'm still curious about one thing: from what I've seen in DeepSeek's interviews and so on, it seems that in major directional choices and research decisions, taste and intuition are very important, not just simple engineering optimization. So if AI can self-evolve later, are these qualities—taste, intuition—still important, or what would become important?

AI 现在不缺品位和直觉,它缺的是持续学习的能力。AI 的品位和直觉没 有问题。你让它写个文章,它的品位和直 觉,我觉得没有什么问题。

AI doesn't lack taste and intuition right now; it lacks the ability for continuous learning. AI's taste and intuition are fine. If you ask it to write an article, its taste and intuition, I think, are not a problem.

梁总,我屏幕上留了个问题,我给您再念一下。其实是想请教一下,continuous learning,就是持续学习,您也提到,很多研究员也提到是一个还没解决的研究问题。然后那个 coding agent,尤其是追上 M ILES,就是

Mr. Liang, I left a question on my screen, let me read it to you again. Actually, I wanted to ask about continuous learning, which you and many researchers have mentioned as an unsolved research problem. Then as for the coding agent, especially catching up to MILES, that is,

到这个 Office 水平、 MIS 水平是一个比较确定的目标。对 于一个还没解决的研究 Scaling 目标和一个比较确定的 Scaling 目标,您认为应该怎么分配研究资源,尤其 是研究员这块的人才资源,怎么样才能 达到最好的平衡和效果?

Reaching the Office level and MIS level is a relatively certain goal. For an unsolved research scaling goal and a relatively certain scaling goal, how do you think research resources should be allocated, especially the talent resources of researchers, to achieve the best balance and effect?

模型 Office 是 一个比较确定的目标,但是模型 MIS,我觉得 还不好说是确定的,只能说是一个目标。模型 Office,我觉得应该是比较确定的。

The model Office is a relatively certain goal, but the model MIS, I think, cannot yet be said to be certain; it can only be called a goal. The model Office, I feel, should be fairly certain.

不占资源。做朋友 任 意研究,他不消耗卡,只需要很少的卡,他需要的是想法;他 CoT 也不消耗人才资源, 因为不需要有人一直在那里做。他不 是 一个项目,而是需要有很多人都在那里想这个问题。

It takes up no resources. Pursuing any arbitrary research, it does not consume GPUs; it only requires very few GPUs. What it needs is ideas. Its CoT does not consume talent resources either, because no one needs to be there all the time. It is not a project; rather, it requires many people to think about this problem at the same time.

所以并不需要给它分配什么资源,因为它不需要资 源。你要训模型,要做模型、发模型,做 模型的效 率 实验才需要资源。刚才讲的,对人、对卡的资源消耗都不多。

So there is no need to allocate resources to it, because it doesn't need resources. It is when you train models, build models, release models, and conduct efficiency experiments on models that resources are needed. As I just said, it consumes very little in terms of both people and GPUs.

所以我们叫这个叫“摸奖”。门槛很低,谁都可以 去摸,但是谁能摸出什么来,这个可能我也不知道是看天赋还是看什么。

So we call this 'a lucky draw.' The threshold is very low; anyone can take a chance, but what anyone might get out of it probably depends on talent or something—I don't know.

所以这里并不需要我们去分配资源。只是说,我们跟其他公司不一样的地方,就是我们会花时间去讨论这个问题,会去想这个问题,然后把它当做一个重要的事情 。

So here there is no need for us to allocate resources. It's just that what sets us apart from other companies is that we spend time discussing this problem, thinking about it, and treating it as an important matter.

在公司里,它是一个重要的问题,是一个我们会花时间去思考 的问题,但是并不需要花很多资源去做。

Within the company, it is an important issue, one that we take time to think about, but it does not require a lot of resources to pursue.

然后下面我还看一个问题:大模型的幻觉问题比较影响用户的体验。幻觉问题也是有一个方法可以解决的,但是这是一个长命题。幻觉问题可以认为是一个可以通过更好的 Post-training 解决的,是一个能解、能够改善的问题。

Then I also want to look at another issue: the hallucination problem in large models significantly impacts user experience. There is a way to address hallucinations, but it's a long-term effort. Hallucinations can be seen as something that can be resolved through better post-training—it's a solvable and improvable issue.

只是大家没 有花很大的力气去做。或者说,对我来讲, 幻觉是一个问题,但我们归结可能是产品问题。我 们会去解决它,但是不是重点的问题。

It's just that not much effort has been devoted to it. Or, for me, hallucination is a problem, but we classify it as a product issue. We'll tackle it, but it's not a priority.

前面还有一个标数据的问题。我们在数据标注方面,这跟我们的资本 投 入有关。以我们这个资本投入的结构,支撑不起那么多高质量数据标注的成本 ,因为成本很高。

Earlier, there's also the data labeling issue. In terms of data annotation, it's related to our capital investment. Given our capital investment structure, we cannot afford the cost of high-quality data labeling at scale, because the cost is very high.

美国数据标 注的成本跟中国 数 据标注成本没有什么区别。中国去标数据并没有成本优势,尤其是标高端数据上并不会有成本优势,使得我 们很难投入去像美国这样标数据。这条 路在中国是很难的,因为标数据实在太贵了,不管是我们外标还是我们自己 标 ,都很难受。

The cost of data labeling in the US is no different from that in China. China has no cost advantage in labeling data, especially for high-end data, which makes it difficult for us to invest in labeling data the way the US does. This path is very difficult in China because labeling data is too expensive, whether we outsource or do it ourselves—it's tough.

所以现在基 本上是两条腿走路。并不是说我们完全不能标,而是因为标数 据有一些成本低,有一些成本高。 我们先标成本低的。

So now we basically have to balance two approaches. It's not that we can't label data at all; rather, labeling some data is cheaper, and other data is expensive. We label the cheaper ones first.

所以你 也可以认为,现在我们公司有一半的人在标数据。有一半的核心研究员,最重要的人,有一半在标数据。我们就集中在标数据。解决 AI 这个问题,在 现在这个阶段靠的就是标数据。你只 要看就是一个数据。

So you can also think of it as half of our company is now labeling data. Half of our core researchers—the most important people—are labeling data. We are focused on labeling. Solving AI at this stage relies on labeling data. You just have to look at the data.

第一个问题,您刚才讲,中国的模型比美国的模型在效率上肯定是强的。你讲的还有 一些其他方面,未来也有可能会比美国强。你觉得是哪些方面,我们 在智能或者在其他方面会比美国强?

First question: you just mentioned that Chinese models are definitely more efficient than American models. You also mentioned some other aspects where they might surpass the US in the future. Which aspects do you think we will be stronger in—intelligence or other areas?

我觉得在很多体验 方面,有可能我们是能比美国做得好的。即不说自己的体验,自己的产品体验觉 得挺好的,我觉得我们在用户体验方面不 一定会比美国差。

I think in many aspects of experience, we might be able to do better than the US. Not to mention our own experience—we think our product experience is pretty good. I believe we won't necessarily be worse than the US in user experience.

在产品方面,产品能力上不一定会比美国差 。成本应该也会比美国低,所以有可能中国还是会有竞争力的。

In terms of products, our product capabilities may not necessarily be inferior to those of the United States. Costs should also be lower than in the US, so it's possible that China will still be competitive.

其他方面,如果说有结构性的优势, 我觉得可能也没有。但是在成本和产品这两方面,我觉得确实是有一定的结构性优势的。

In other aspects, if you ask whether there are structural advantages, I think there may not be. But in terms of cost and products, I do think we have certain structural advantages.

成本这个很好理解,是因为他们都不用做,所以他们就不发展这个能力。他们肯定没有我们重视这个事情。我们 可 以把它当做一个非常重要的事情,但对于他们来讲,这个是不重要的。

The cost aspect is easy to understand. Because they don't have to do it, they don't develop that capability. They certainly don't value it as much as we do. We can treat it as a very important matter, but for them, it's not important.

产品也是,原生通过很多公司,产品能力还是可以的。所以我觉得这两方面可能是有结构性优势的。

The same goes for products. Through many native companies, the product capabilities are decent. So I think these two areas may have structural advantages.

好的。第二个问题,请教您,您刚刚讲到后训练,我们投入相对成本又高,像 Anthropic 和 OpenAI 都投入巨大的金 额。这次融资以后,您觉得我们会在后训练方面加大投入吗?

Okay. Second question, may I ask you: you just mentioned post-training. Our relative investment cost is high, and companies like Anthropic and OpenAI have invested enormous amounts. After this round of financing, do you think we will increase investment in post-training?

差距主要是在高质量数据的标注上,然后主要是在 AI 研究里面。肯定会加大投入,但是像高质量数据的标注,典型不是资本投入。

The gap mainly lies in the annotation of high-quality data, and also primarily in AI research. We will definitely increase investment, but high-quality data annotation is typically not a matter of capital investment.

高质量数据标注的瓶颈,我觉得是时间,就是需要时间。因为对 OpenAI 来讲、对国外来讲、对 Anthropic 来讲,他们都更早,然后资本更多,卡也更多 。

The bottleneck for high-quality data annotation, I think, is time. It requires time. Because for OpenAI, for overseas companies, for Anthropic, they started earlier, have more capital, and more GPUs.

这种情况下,我们国内可以认为是最近半年才开始做,所以在时间上,我觉得是需要更多时间的。这个跟资本投入关系不太大,因为哪怕没有更多的资本投入,原来的资本也够它以最快的速度在扩张了。

In this situation, we in China have only started doing this in the last six months or so. Therefore, in terms of time, I think we need more time. This is not closely related to capital investment, because even without additional capital, the existing capital is enough for it to expand at the fastest pace.

但是这个速度是有上限的,瓶颈不是说我能够马上有更多的人,并不卡在钱上,也不卡在卡上。但是它确实是在一个快速扩张的过程。

But this speed has an upper limit. The bottleneck is not that I can immediately have more people; it's not about money, nor about GPUs. But it is indeed in a process of rapid expansion.

所以我们觉得一年之内,高质量数据这个问题做得比较好,我觉得国内应该是可以预期的。我觉得面可能没有那么冷,但它确实是需要时间。

So we think that within a year, the issue of high-quality data can be handled quite well, and I think that can be expected domestically. I think the situation might not be that cold, but it does take time.

谢谢。然后第三个问题,就是我们看到 Anthropic 他们用自己的模型做自己的产品,推出了很多纵向的金融、法律……

Thank you. Then the third question is that we see Anthropic using its own models to build its own products, launching many vertical products such as finance, law...

我还没有想得很明白,未来我们国内的商业模式是怎么样的,或者说最顺利的路径是 怎样的,我们还没有到那个阶段。 国内情况和国外情况不一定一样,国内到时候是什么样,我觉得现在还比较难判断。

I haven't fully figured out what our domestic business model will be in the future, or what the smoothest path is. We haven't reached that stage yet. The domestic situation may not be the same as abroad. It's hard to judge what it will be like domestically at that time.

国内我们现在以目前的情况来看的话,我觉得最合理的做法应该是全力做通用的 Agent,其他的 Agent 优先级应该更低,包括金融、医生 这些 Agent。要先做 Co ding,因为 Coding Agent 能够做到很多,还有很多垂直的 Agent。现阶段我们觉得最重要的,应该还是 Coding Agent。

Domestically, given the current situation, I think the most reasonable approach is to fully focus on general-purpose agents, while other agents should have lower priority, including finance and doctor agents. We should first work on coding, because coding agents can do a lot, and there are many vertical agents as well. At this stage, we think the most important thing is still coding agents.

讲得很清楚,谢谢您。还有一个问题,我们也想请教您。其实我们做 DeepSee k,也很 佩服您,一直在以一个非常纯粹的科研 方式来做 DeepSeek。

That's very clear, thank you. Another question, we'd also like to ask you. Actually, when we do DeepSeek, we also admire you very much, as you have always been doing DeepSeek in a very pure scientific research manner.

但现在这个行业确实也走到了资本市场,走到了资本化这条路上。您又是一个非常负责任的人,无论是对小伙伴还是对投资人,您也都 是非常负责的。那么,在纯粹科研、纯粹做 AGI 的方向和资本市场之间的平衡 ,您未来肯定还要走资本市场,肯定还有公众股道,您觉得以后怎么 balance?这个您是怎么考虑的?

But now this industry has indeed moved into the capital market, onto the path of capitalization. You are also a very responsible person, both to your team and to investors. So, regarding the balance between pure research, purely doing AGI, and the capital market, in the future you will definitely have to go into the capital market, and there will be public shareholders. How do you plan to balance this? What are your considerations?

我现在觉得应该是能够做到的,就都要。假如说我今年能够有 几个亿美金的 B 端收入,再加上我们 C 端有用户,那么这本身就已经有一定的商业 基础。以明年我们 B 端有收入,如果这个需求可以再增大的话,公司离净利润已经不远了,可能就已经是净利润了。

I think it should be achievable, and I want both. For example, if this year we can have several hundred million dollars in B-end revenue, plus we have users on the C-end, then that itself already provides a certain commercial foundation. With B-end revenue next year, if this demand can increase further, the company will not be far from net profit, or might already be profitable.

可能就已经不是一个纯烧钱的阶段了,所以我感觉后面能做的、能操作的动作应该还是比较大的。或者说,最坏情况卖 API,可能都能够支撑一个上市公司。就如果说技术后面没有新的进步了,我们的技术就冻结在这里了,那么最后我们就全力卖 API,把这些服务做好,我觉 得也够的 。

It might already not be a stage of purely burning money, so I feel that the actions we can take and operate later should be relatively significant. Or to say, in the worst case, selling APIs might even be able to support a listed company. If there are no further technological advancements later, and our technology is frozen here, then ultimately we will focus on selling APIs and doing these services well; I think that would be sufficient.

所以我觉得还是有信心的,确实没有那么难。因为确实是在一个杠杆高的地方,又在一个发展非常快的领域,它可能就是没有那么难。只能讲,我们希望有更大的梦想,但是我们也有保底可以拿出来的业绩 。

So I still feel confident; it's indeed not that difficult. Because it's a place with high leverage and in a rapidly developing field, so it might just not be that hard. We can only say that we hope to have bigger dreams, but we also have bottom-line performance we can show.

感谢您分享。前面提到一些问题,想跟您再问一下。因为我觉得 DeepSeek 其实最大的和其他公司的区别在于,我们的组织跟其他人不一样,或者说组织形式不一样。

Thank you for sharing. Regarding some issues mentioned earlier, I'd like to ask you further. Because I think the biggest difference between DeepSeek and other companies is that our organization is different from others, or rather, the organizational form is different.

但组织形式既跟我们的目标相关,可能也要去考虑组织自身的边界和效率。我不知道,从宏观的角度考虑 ,我们这个组织形式会有一个好的学习对象吗?历史上可能 Bell Labs,还是一个什么样的形态可能是比较理想的?还是说我们自己觉得也没有理想的,更多还要靠我们 逐渐去探索?

But the organizational form is related to our goals, and we also need to consider the organization's own boundaries and efficiency. I don't know, from a macro perspective, is there a good model for our organizational form to learn from? Historically, maybe Bell Labs, or what kind of form might be ideal? Or do we ourselves think there is no ideal, and it still relies on us to explore gradually?

可能在一个新的时期,只能靠我们自己来做探索。因为我们的组织形式其实跟美国的几家公司肯定都不一样,对吧?三家公司自己也不一样,但他们至少会从一个商业公司的角度出发去 探索。

Perhaps in a new era, we can only rely on ourselves for exploration. Because our organizational form is definitely different from several American companies, right? The three companies themselves are different, but they at least explore from the perspective of a commercial company.

首先,我们没有模仿的对象。每一步都是我们从实际情况出发,实事求是,根据实际情况来做决策,找到我们应该怎么做。所以它是一个时代的产 物,或者说是现实情况的一个反应, 它并不是一个模仿的结果。

First, we have no object to imitate. Every step is based on our actual situation, seeking truth from facts, making decisions according to reality, and finding what we should do. So it is a product of the times, or a reflection of reality, not a result of imitation.

就是说,在这个情况下,我确实最优解可能就是这样,或者说我自己认为,我们自己 选的路就是这样。每一步我们肯定都是思考过的,肯定都是选过,反正选的结果就是这么选的 。它并不是因为看到谁这么选,我们这么 选,而是因为我们分析了利弊而这么选。

That is to say, under these circumstances, the optimal solution might indeed be like this, or I myself believe that the path we chose is like this. We must have thought through every step, and definitely made choices; anyway, the result of choices is this. It's not because we saw someone else choose this way, but because we analyzed the pros and cons and chose this way.

那么在未来可能也是一 样,我们并没有去模仿谁。我觉得我们跟 Bell Labs 还是不一样的,因为它明 确是不需要有商业化的,因为 它是个 ……但是我们明确是要有商业 化的。我们最终还是要能活下去,我们毕竟 是 一个公司,政府不会给我一分钱。

In the future, it may be the same; we are not imitating anyone. I think we are still different from Bell Labs, because it clearly didn't need commercialization, because it was a... but we clearly need commercialization. We ultimately need to survive; after all, we are a company, and the government won't give us a single penny.

所以我们可以有非常远大的使命,但我们归根结底是一个公司 ,我们要考虑怎么活着。所以 B 端对我们肯定是重要的,因为可能以后可能得靠它活着。只是说现在不重要,因为它现在是个成本线。

So we can have a very ambitious mission, but at the end of the day, we are a company, and we need to consider how to survive. So the B2B side is definitely important to us, because we may have to rely on it to survive in the future. It's just that it's not important now because it's currently a cost center.

所以我觉得这个还是不一样的,跟 Bell Labs 还是不一样的。我们本质上还是一个公司。历史上也有很多公司,它也有利润以外的追求,但并不能说它有利润以外的追求,它就不是个公司。

So I think this is different; it's different from Bell Labs. We are essentially a company. In history, there have been many companies that also had pursuits beyond profit, but that doesn't mean they weren't companies because they had pursuits beyond profit.

很多公司做得很伟大,因为它有一种利润以外的追求。那个追求最后不但没有影响到它的商业化,反而能让它商 业化得更好。我们本质上还是一个公司 ,只是说我 们在考虑赚哪些钱、什么时候 赚钱、赚多少 钱 、靠 什么赚钱,只是说我们有取舍。

Many companies achieve greatness because they have a pursuit beyond profit. That pursuit not only doesn't hinder their commercialization, but it actually helps them commercialize even better. We are essentially a company; it's just that we consider which money to earn, when to earn it, how much to earn, and what to rely on to earn it. It's just that we have trade-offs.

梁总,我有两个小问题,快速跟您请教一下。一个是说,您刚刚其实有提到 MILOS 的这个,可能不是一个很确定的目标,但是肯定会往这个方向去做。然后您也提到激活参数可能比如说……

Mr. Liang, I have two quick questions for you. First, you just mentioned MILOS; it might not be a very definite goal, but we will definitely move in that direction. And you also mentioned that the activation parameters might, for example...

下一代可能会是在 150 到 250 B 左右的一个情况。这种情况下,您觉得 150 到 250 B 是对标 O4.7,还是可能会对标到别的?这是第一个问题。

The next generation might be around 150 to 250 B. In this case, do you think 150 to 250 B is benchmarked against O4.7, or could it be benchmarked against something else? That's the first question.

第二个问题,您前面也提到了,我们现在在整个 inference 上面用了一些除了 CUDA 以外的编译语言。我理解原来咱们基于英伟达的生态 ,可能会用 PTX 这些比较多。现在用上 更多类似您刚提到 Tile Lang 这些的话,是不是会大幅降低我 们在 inference 上面的一些效率,或者短期降低效率 ?

Second question, you mentioned earlier that we are using some compilation languages other than CUDA for inference. I understand that originally, based on Nvidia's ecosystem, we might have used PTX and similar more. Now, if we use more of something like TileLang you just mentioned, will it significantly reduce our efficiency in inference, or reduce efficiency in the short term?

我不知道您会怎么看编译语言的这种变化带来的效率损失,还是说长期其实是一个可以补充的状态,是提高效率 ?

I wonder how you view the efficiency loss brought by this change in compilation languages, or is it actually a complementary state in the long run, improving efficiency?

对,是大幅提高效率。所以说这是一个机遇。相当于以前你是离不开 CUDA 的生态的,现在我们可以抛弃它的生态,用一个更简单的方法,就用 TileLang。 它是高级语言,写那个程序也是很快的,要写的代码量很小了 ,我可以把 它重写一遍。

Yes, it significantly improves efficiency. So this is an opportunity. It's like before, you couldn't leave the CUDA ecosystem; now we can abandon its ecosystem and use a simpler method, just use TileLang. It's a high-level language, and writing programs in it is very fast; the amount of code needed is very small, and I can rewrite it entirely.

明白。所以这两个都其实是一 个,像您提到的 AI 带来的比较大的机会 ,不是一个可能短期需要去弥补的 缺点 。

I see. So these two are actually one and the same—the significant opportunity that AI brings, as you mentioned—rather than a shortcoming that needs to be remedied in the short term.

对,它是技术发展的大机会。它不是 AI 的机会,因为我们还有一个项目,我们在用 AI 来写 TileLang。

Right, this is a major opportunity for technological development. It's not an opportunity for AI itself, because we have another project where we use AI to write TileLang.

现在所有 TileLang 都是人写的,但是它已经比原来写 CUDA 的要快很多了。

Currently, all TileLang is written by humans, but it is already much faster than writing CUDA originally.

互动版:图/公式 + 针对本篇提问 →