DeepSeek founder Liang Wenfeng in His Own Words: 64 Quotes from DeepSeek's Investor Call
Key takeaways:
Vision-driven: No KPIs, no written vision, even "no organization"—it runs on goodwill toward the world and an obsession with AGI. "We're just a bunch of very ordinary people."
Restraint as strategy: API pricing aims only to recover hardware costs in ten months. No fighting for consumer traffic, no chasing enterprise fads. "The more restrained you are, the more likely you'll pull it off."
Open source is the point: Even the strongest model will be open-sourced. Not worried about others competing. "I only worry they won't be able to deploy it."
AGI roadmap: CoT → Agents → continual learning → "singularity" → embodied intelligence. "After continual learning, it can develop the next version itself."
Compute reality: About 20,000 "H-equivalent" GPUs; 12–18 months behind the U.S.; chasing with one‑twentieth the compute. For domestic chips: "The ecosystem is fine; capacity is the bottleneck."
Endgame view: China may match foreign models this year, but "that still isn't AGI." "There are too many base-model companies in China; it will converge."
Vision, motivation, and culture
*"We don't have an organization; we're driven by a vision."*
When we started the company, our original intent wasn't "how much money will I make," or "go to the capital markets," or "IPO," or anything like that. That wasn't the intent. The first few dozen people never thought that way. If they did, they wouldn't have joined.
Overall, we're doing this with a lot of goodwill toward the world, and we think it's useful to humanity. It's something beyond money.
About twenty years ago, in management, the person I admired most was Jack Welch, former CEO of GE. Looking back now, most of what he said may no longer be right, but he got one thing right: the most important thing for a company is its vision. Managing a big company isn't about rules and procedures—it's about vision. What is vision? Vision isn't a slogan on the wall. Vision is how you do things, not what you say.
In fact, we don't really have an organization—we're driven by a vision, organized by a vision. That vision isn't even written down. We've never written anything.
We're not especially capable. We don't have more money than others, and it's not that our people are better than everyone else's. We're just a bunch of very ordinary people.
Our company is built on consensus. It's not that I decide everything. I seek consensus. My authority inside the company is built on consensus.
Our management has two lines: one top-down and one bottom-up. Bottom-up means everyone decides what they want to do and just does it. No one manages them. No KPIs. In general, we hope "formal work" doesn't take more than half of an employee's time. The other half is unallocated—they can do whatever they want.
We generally don't work overtime. First, research needs a relaxed environment; if you push too hard, you can't do research. Second, we're very focused, which means we do very few things.
CoT doesn't consume resources. For research, it doesn't burn GPUs—you only need a small number of cards. What you need is ideas. The barrier is low; anyone can try it. But who can actually get something out of it—I don't know if that's talent or what.
Open-source strategy
*"Even our strongest model will be open source, because I don't see what's good about closed source."*
Why do we insist on open source? Because the vision itself requires open source. Without that vision, you can't organize people.
Zhipu also open-sources, but their open source isn't the same as ours. Their open source has a "being chased" feeling—they don't see it as the point. For us, it is the point.
AI is big enough that in the end it may take up 10% of human society's GDP. One person can't monopolize it. You have to share it with others, or you won't survive.
I think we will open-source, and our strongest model will probably also be open-sourced, because I don't see what's good about closed source.
Even if you open-source the model and tell people everything, the barrier is still very high. It's still very hard for others to actually use it.
If we only make, say, a 6× profit, open source doesn't really hurt. But if you want a 100× profit, then open source does affect that.
We were open-source from the start. We won't treat you differently just because you're a competitor—that's part of open source. I'm not worried about others deploying our model to compete with us. I want them to be able to deploy it. I only worry they can't—some detail is wrong, results get worse, or costs get high.
Are the open-source models we provide the same as what we deploy ourselves? Yes, they're the same. We won't open-source a weaker model and then use a better one internally. Last year, basically everything on the consumer side was open source, and we didn't see a conflict with our consumer service.
Restraint and pricing
*"Restraint is a strategy. Ten months to recoup costs is a reasonable profit."*
If we want to make AI work in our hands, the first thing, I think, is restraint. You can't think "a huge share of human GDP should be mine." The more you think like that, the less you'll succeed.
For me, restraint is a strategy. Sometimes you give up something to gain other things.
For our API pricing, what we consider a reasonable profit is: buy a batch of equipment from the market, and recoup the cost in ten months. I think that's reasonable. It's not profit-maximizing. If you wanted to maximize profit, you'd price higher. In this price range, demand isn't elastic. Even if I raise the price by 50%, token usage doesn't change much.
Let me tell a story. For our DCP (Decoder Context Parallelism) model, at first we worried demand would be too high, so we set the price relatively high. People on the team weren't very happy. Later I lowered it—down to one quarter—and everyone was happy. When we cut the price, people in the company group chat were cheering.
Someone just commented that "ten months to break even is too high." True, there's still room to lower prices. There's room for optimization on the model side too, so there's still a lot of room overall.
At that cost level—ten months to break even—we can do it. Other companies can't. Alibaba or Tencent, without our optimizations, their costs should be several times higher.
If I cut prices further, demand won't increase, or will increase only slightly. At this price, everyone can afford it and is satisfied. Cutting prices won't bring the company more revenue, and it won't add more value to society. Even if it's cheaper, it doesn't increase "happiness" by much.
The lower the cost, the bigger the model I can train, and the more I can afford to take on larger models. With limited compute, if my compute efficiency is higher, I can take on larger models.
A commercial company has no incentive to pursue model efficiency. If efficiency improves… costs drop… then what do you earn? For us, it's part of the vision. Our teammates are ordinary people; they know using this costs money. They have empathy: others have to pay to use it, so if it's cheaper, it's easier to accept.
Commercialization
*"Both C-end and B-end are side products on our road to AGI."*
We do AGI to do AGI. We just happen to produce something we can commercialize, so we use it. That's different from other companies. Other companies build the model to serve consumer users or enterprise users. For us, both consumer and enterprise are side products on the road to AGI.
When we suddenly got popular around last year's Spring Festival, that wasn't in our script. We never expected it. We just wanted to make the technology good.
When users suddenly surged last year, we didn't try to keep them for monetization, or fight for those commercial interests. We didn't fight for users, we didn't try to make money—but we worked hard to serve them well.
This year it's quite possible our API/AI revenue could reach a few hundred million dollars in ARR, if demand continues expanding and if we can buy more GPUs. If AI revenue can reach a billion dollars, our cash flow can basically turn positive.
Over the past three years, at any point, if you came to me to talk about a commercialization roadmap or product lines, it would be a waste of time, because you can't predict the future.
If this year we can have a few hundred million dollars of enterprise revenue, and next year enterprise revenue continues, and demand keeps growing, we won't be far from net profit. In the worst case, just selling APIs can support a listed company. If there's no new technical progress later and our technology "freezes" here, then we'll fully focus on selling APIs and doing the services well—that's enough.
AGI tech roadmap
*"AI isn't lacking taste or intuition. It's lacking continual learning."*
You can think of AI progress as climbing steps. Last year's step was CoT. This year's step is Agents. Why steps? Because each next step builds on the previous one. Agents rely on CoT, and CoT relies on the previous step—the language model.
After Agents, we think the next problem to solve is continual learning: how to let the model keep learning continuously, rather than just giving it one strong training run. It should be able to learn over a long period like humans do.
After continual learning, we may reach a "singularity." It's when a model that can keep learning can already do everything humans can do. It can develop its own versions, do research by itself, and build the next version—build even more advanced AI models.
This "singularity" isn't really a point; it's also a gradual process. It may be a long, gradual change—not a sudden jump.
After that, I think you get embodied intelligence. Then it enters the real world: it can do housework, take care of the elderly, and so on.
We think this roadmap is the easiest. We can follow it without overtime. If the roadmap were reversed—for example, if you try to do embodied intelligence first—you'd suffer. It's hard labor.
AI isn't lacking taste or intuition. Its taste and intuition are fine. If you ask it to write an article, I think its taste and intuition are basically fine. What it's lacking is continual learning.
Everyone in the world is researching this. From an investor's perspective, what they see most is Agents. But for researchers like us, what we see more is learning, and how to solve the learning problem.
Internally we care about this narrative: for the next version of our model, we hope it can help our own development. It can improve DeepSeek's efficiency. Our model's first job is to raise DeepSeek's own work efficiency. First it has to be useful to us. That's the fastest path to AGI.
Compute, chips, and the China–U.S. gap
*"We're behind the U.S. mainly on resources. On people, almost no gap."*
Right now we have about 20,000 H-equivalent compute cards, and most of those arrived recently—within the past one or two months.
Our gap with the U.S. is mainly resources. On people, almost no gap, because it's basically the same group of people. Talent isn't the bottleneck; resources are the biggest bottleneck. The talent gap is essentially also a compute gap.
For the largest models today, we can't afford to train them. The biggest models have about 800B active parameters; domestically, we're still at tens of billions. If I want to train a model as large as the leading ones, I'd need 50,000 GB300s or Huawei 950s—200,000 cards. That's just training, not even research.
Within what we can afford, more cards is always better. At a reasonable price, we buy as many as we can. Ideally, if we can spend the money within six months, that's best. Turning money into Nvidia cards is definitely better than leaving it in the bank.
Our gap with the U.S. may be 12 months, maybe 12–18 months, or maybe 6–12 months. Simply put: about two years behind, but doing it with one-twentieth of the U.S. compute. In the future we want to rewrite that narrative: use a fraction of the compute, but shorten the time to 6 months, 3 months—maybe even surpass them in some areas.
Nvidia's CUDA moat is being eroded quickly. One reason is that AI can now write code, so I can use AI to build the ecosystem. Another reason is new technologies—for example TileLang (an open-source high-performance AI operator programming language developed by Peking University's School of Computer Science). Using a higher-level language to write CUDA operators, you can quickly rewrite Nvidia's whole ecosystem.
There's a historic opportunity for domestic AI chips to replace imports. We believe within the next year we'll see something validated: the ecosystem for domestic chips has no problem. People used to think there were problems—couldn't use them, not usable—but within a year, I think we can change that perception.
Our goal in buying Huawei 950 is still to help Huawei improve the ecosystem. Huawei 950 "super nodes" can substitute Nvidia GB200/GB300 in performance and price. The only tradeoff is: four Huawei cards equal one Nvidia card, and you're two years behind in time. Huawei 950 super nodes will ship in Q3 or Q4 this year; Nvidia GB200 shipped in Q3 two years ago.
You can depreciate Nvidia cards over five years. Huawei cards, at most three years. Huawei 950 is fine to use this year; next year should still be OK. After that, it may really become too power-hungry.
What DeepSeek won't do
*"Video generation, 3D, world models—these don't set the upper bound of intelligence."*
AI is a broad field. There are many things we think aren't on the main line—for example 3D and video generation. I don't think they're strongly related to the main line of intelligence, so we won't do them.
When video generation first came out, it was very hot—like you had to do it, or you weren't an AI company. But you can see what happened: after Sora, everyone did it—big companies, small companies. Then small companies cut it. It doesn't affect the ceiling of intelligence. Commercially, it's a good business, but it's not about intelligence. We won't do something just because it's a good business.
World models, I think, also don't have much to do with the upper bound of intelligence right now, so we won't do them. In our view, world models and intelligence aren't the most important thing at this stage.
Multimodal matters a lot for products, and for consumer products. But for the ceiling of intelligence, it's a component, not the main line itself. We will definitely do multimodal, and we're doing it. V4 and later versions will support native multimodal.
Data labeling
The cost of data labeling in the U.S. isn't really different from China. China doesn't have a cost advantage in labeling, especially for high-end labeling. That makes it hard for us to invest in labeling data the way the U.S. does. This path is hard in China because labeling is just too expensive.
Right now, half of the most important people in the company are labeling data.
The bottleneck isn't that we can't quickly hire more people. It's not about money, and not about cards either. But we are expanding fast. So we think within a year, it's reasonable to expect that the high-quality data problem in China can be handled much better.
Competition and endgame
*"There are too many base-model companies in China. It will converge."*
Our biggest core interest is keeping the team stable. That's our biggest core interest—maybe the only one. As long as I can keep the team stable, I'll be able to do AGI. Money isn't the problem, resources aren't the problem; everything else is easy to get.
We've always been very restrained. We don't want to become an opponent of any big internet company or small company. I'd rather help them do this.
Open source, goodwill, helping others—none of that caused us to lose anything. No impact at all. If anything, it may be a plus.
There are a bit too many model companies in China right now—still too many. The U.S. maybe has three. China has too many doing base models. Everyone is doing the same thing. Resources get spread thin, and each company gets less. I think it will definitely converge.
If my goal is to take 5% of global GDP from AI, I'll be defeated by someone willing to take only 1%. Because that person says: I do it well, but I only need 1% of global GDP. Then if another person comes along and says: I only need 0.1%, they beat the previous one. People who take more get beaten by people who take less. OpenAI at the beginning thought it could monopolize the world, but in reality it will face many challengers.
In AI, in China, in one or two years we can get to roughly the same level as abroad, or maybe even substitute foreign models this year. In the current paradigm, it's not that hard. But it's still not AGI.
Right now, returns from model scaling are still very obvious. We haven't had a chance to hit the "scaling wall." It's still far away. When Silicon Valley says scaling is over, that's for Silicon Valley. For China, we're still far from that—we haven't scaled to that level at all.
Source: https://mp.weixin.qq.com/s/HxvyxDGXusNRMzes3i-U7g


