DeepSeek's Liang Wenfeng Breaks His Silence
Inside China · Fred Gao · 2026-07-23
DeepSeek's Liang Wenfeng argues that in the pre-AGI race, deliberate restraint — capping margins, open-sourcing models, and refusing to maximize revenue or market share — is the strategy most likely to reach AGI first, because whoever is willing to extract the smallest slice of value wins.
Liang inverts the usual playbook: instead of racing to lock in users, revenue, and closed-source moats, DeepSeek optimizes for probability of reaching AGI and assumes AGI will monetize itself. The pricing math is deliberate — a roughly sixfold profit ceiling with a ten-month payback keeps competitors from undercutting deployment, making open source safe, whereas hundredfold margins would collapse that equilibrium. The claim, still contested, is that in a pre-revenue technology wave, willingness to extract less is a competitive weapon, not charity.
Liang frames AI not as a tool for monopoly but as a historical wave larger than any firm. The right response is to hold back, cap your own gains, and share openly — restraint is itself the strategy.
If your vision is to capture 5% of global GDP through AI, someone willing to take only 1% will defeat you, and someone willing to take 0.1% will defeat them. In the pre-revenue phase, ambition of extraction is itself a competitive disadvantage.
Liang frames giving up short-term revenue, users, and closed-source moats as an intentional strategy — trading immediate benefits for a higher probability of reaching AGI.
Liang's optimization function is not incremental share or revenue but the probability that DeepSeek gets to AGI — and he assumes AGI itself will provide enormous commercial value automatically.
At roughly sixfold profit (ten-month payback), third parties can't undercut DeepSeek on deployment cost, so open-sourcing the model doesn't threaten revenue — but chasing a hundredfold margin would break that equilibrium.
Open
- · Does the sixfold-margin equilibrium hold once frontier training costs rise further?
- · If AGI arrives, does it actually monetize itself automatically, or does distribution and lock-in still decide the winner?
- · Can a restraint-based strategy survive competitors and investors who demand higher extraction?
Pipeline
- source kind
- url
- generated by
- anthropic
- candidates
- 153 (selected 5)
- embeddings
- —
Sections
Candidate pool grouped by section. Selected candidates are bolded.
Considered candidates (148)
Below top-k · 148
- claimThe only real gap with the US is compute resourcesc 0.95
Talent is not the bottleneck — Chinese and US labs draw from the same pool, split roughly evenly. Every visible gap (talent, model capability, applications) ultimately reduces to a gap in compute.
- mechanismAI's sheer size makes monopoly self-defeatingc 0.90
Because AI could eventually be ten percent of global GDP, trying to monopolize it invites resistance that will block you outright. The market is too large for any one player to hold, so sharing is the only viable path.
- claimTen-month payback is the deliberate pricing anchorc 0.90
DeepSeek prices its API to recover cost in roughly ten months — equivalent to about a sixfold profit margin — which Liang treats as commercially sufficient rather than maximizing revenue.
- mechanismOpen weights don't transfer the deployment moatc 0.90
Even with open source, third parties struggle to reproduce DeepSeek's low deployment cost — the principles are known but the engineering, management, and organizational will to execute them are rare.
- claimContinuous learning is the one missing step between current AI and human replacementc 0.90
Current AI cannot replace employees because it cannot learn on the job the way a new hire does over a few months. Adding learning-to-learn would close that single remaining gap.
- claimWorld models and video generation are not on DeepSeek's intelligence roadmapc 0.90
DeepSeek's roadmap runs from training, to continuous learning, to models asking their own questions — with no place for world models or video generation. Multimodality will come eventually, but these hyped directions aren't on the critical path.
- claimNVIDIA's CUDA moat is rapidly disintegratingc 0.90
The ecosystem advantage that made domestic chips unusable is collapsing, opening a historic window for domestic AI chip substitution.
- mechanismSolve continuous learning first, and general intelligence becomes easyc 0.90
If AI could learn continuously, its capabilities would jump and it could dramatically accelerate DeepSeek's own research. Under that condition, general intelligence stops being a data- and manpower-heavy grind and becomes something the AI helps deliver.
- claimAI's taste is already fine — continuous learning is the real gapc 0.90
Current models don't lack taste or intuition; they lack the ability to keep learning. Ask one to write an article and the aesthetic judgment is already there.
- mechanismSesame seeds versus watermelonsc 0.85
Current C-end and B-end revenue are 'sesame seeds' in front of the 'watermelon' of AGI — worth casually picking up at low cost, but not worth stopping to fight over.
- claimAGI roadmap is clear — with a caveat about contextc 0.85
Current models already beat humans when given complete context and instructions; the remaining gap to AGI is that supplying that full context is impractical, and models can't yet learn continuously the way a new hire does over two months.
- mechanismAI progress is a staircase where each step reuses the lastc 0.85
Language models enabled chain-of-thought, and chain-of-thought enabled Agents. Each new capability layer is built on the previous one, so no step is wasted and the trajectory is traceable rather than mysterious.
- claimThe thing you most want, you can't get; the thing you don't care about comes easilyc 0.85
Chasing C-end traffic or B-end revenue directly is what everyone else does. DeepSeek pursues AGI and finds commercial success arrives as a byproduct they didn't chase.
- claimTeam stability is the only core interestc 0.85
The single non-negotiable thing is keeping the team together. Money, resources, and everything else are readily obtainable — if the team stays, AGI is guaranteed, only the timing is uncertain.
- mechanismOnly do things that raise the intelligence ceiling, not things that are good businessc 0.85
Video generation is a good business, but a good business is not a reason for DeepSeek to pursue something. The filter is whether a direction lifts the intelligence ceiling — commercial attractiveness alone doesn't qualify.
- claimThe endgame gap between large models will be smallc 0.85
Wenfeng expects the final gap between competing large models to be modest, reducing to three axes: cost, time-to-market, and user experience. Beyond those, capabilities converge.
- claimPushing AGI forward beats enriching product linesc 0.85
At this stage, raising the intelligence floor has higher returns than expanding products or building commercialization paths. This has been true for the past three years and will remain true for the foreseeable future.
- claimChinese AI will be systematically cheaper, like other Chinese industriesc 0.85
Just as Chinese-made goods have closed the quality gap with foreign versions while remaining cheaper, Chinese AI will likely offer comparable capability at systematically lower prices.
- mechanismTileLang as a high-level compiler decouples DeepSeek from CUDAc 0.85
DeepSeek built TileLang, a high-level language for writing CUDA operators, so they can rewrite NVIDIA's ecosystem quickly and port it to any chip. V3 already used NVIDIA cards without using NVIDIA's ecosystem.
- claimDeepSeek hasn't touched the scaling ceiling — compute is the only limitc 0.85
Model returns are still very obvious. The reason DeepSeek trains models of a certain size is simply because that's what its resources allow, not because further scaling wouldn't help. Silicon Valley's scaling-wall discourse doesn't apply.
- claimThe current narrative — two-year lag on one-twentieth the computec 0.85
China trails the US by roughly 12 to 18 months while using only one-twentieth of the compute. The goal now is to rewrite this by shrinking the time gap to three or six months at similar compute ratios.
- claimThe next model's first customer is DeepSeek itselfc 0.85
Internally, the priority isn't that users find the model useful — it's that DeepSeek's own team finds it useful for building the next model. If it helps them iterate faster, that's the fastest path to AGI.
- claimAGI's job is to iterate the next AGIc 0.85
DeepSeek's working definition of AGI is a model that can help iterate the next-version model. Once embodiment enters the picture, the same recursive definition applies — a robot whose job is to build the next-version robot.
- claimContinuous learning research doesn't need resources, it needs thoughtc 0.85
Free-form research on continuous learning consumes almost no cards and no dedicated headcount. What it needs is many people turning the problem over in their heads, not a project plan.
- evidenceHalf of DeepSeek's core researchers are labeling datac 0.85
Roughly half the company — including the most important people — is currently working on data annotation. Solving AI at this stage really does come down to data.
- claimDeepSeek has no model to imitate — not even Bell Labsc 0.85
Every step is decided from the actual situation, not copied from anyone. Bell Labs is not the template because Bell Labs didn't need to commercialize, and DeepSeek explicitly does — the government won't give a single cent.
- claimAGI is the real goal; commercial wins are byproductsc 0.80
For Liang, reaching AGI is the point. Business success, revenue, and market position are just natural side effects of moving the technology forward, not objectives in themselves.
- claimHigh-margin closed models are strategically weakc 0.80
American labs pursue scale and lock-in via closed models and high profits because they have the resources to. Liang argues this posture will inevitably lose to competitors content with reasonable margins.
- mechanismOpen source and commerce reconcile via slim-margin pricingc 0.80
Liang argues open source and business only conflict when you try to maximize profit. Charging just enough to recover hardware in about ten months lets open models coexist with a sustainable business and attracts partners to build on top.
- claimDeepSeek will stay on the AGI critical path and cede the restc 0.80
The company will concentrate on language models, chains of thought, agents, and continual learning, and will deliberately skip areas like video generation and world models — handing those opportunities to the wider ecosystem.
- claimDemand for API is inelastic below a certain pricec 0.80
Further price cuts wouldn't meaningfully increase demand because users already find current prices affordable, so cutting more would neither grow revenue nor add social value.
- claimNo visible benefit to closed sourcec 0.80
Liang says he simply cannot see any concrete advantage from keeping models closed — pointing to ByteDance's closed model as an example of a strategy with no discernible payoff.
- implicationContinuous learning triggers a self-iteration singularityc 0.80
Once a model can learn continuously, it can do everything humans can, including researching and building its own next version. That recursive self-improvement is what people mean by the singularity.
- mechanismStanding on high-ground tech is a dimensional-reduction strike on applicationsc 0.80
Building at the AGI frontier makes C-end and B-end products almost trivial to produce. Users who arrived last Spring Festival couldn't be driven away even when DeepSeek didn't want to maintain them.
- mechanismAGI as a vision is an organizational recruiting weaponc 0.80
A big vision coheres more excellent people and coordinates them once they arrive. Smart people don't naturally cooperate; a shared AGI mission is what makes them charge together, and that's the real advantage over commercially-oriented rivals.
- evidenceFrontier models are an order of magnitude larger than what China can trainc 0.80
The current largest models have around 800B active parameters; the largest domestic model is only several tens of B active. Even spending all 50 billion yuan, DeepSeek could not train at frontier scale.
- claimWithin a year, the 'domestic chips are unusable' perception will reversec 0.80
Liang predicts that within a year, facts will demonstrate that domestic chip hardware and ecosystem have no fundamental problems — the only remaining issue is production capacity.
- claimDeepSeek refuses to vertically integrate into products or chipsc 0.80
Liang wants to own one piece — the core model — and let others do applications and chip manufacturing. Trying to do everything is how you lose to the specialist who only wants one percent.
- claimCommercial companies have no incentive to pursue efficiencyc 0.80
For-profit AI companies don't chase low cost because low cost means less revenue — DeepSeek's efficiency obsession is a values-driven anomaly, not a business strategy.
- claimChina has too many base-model companies and will convergec 0.80
The US has roughly three companies building base models; China has too many, spreading resources thin. Convergence is inevitable — once firms discover the profit margin isn't as high as they assumed, most will stop.
- claimContinuous learning is the actual blockerc 0.80
Nobody in the world has found a method that works for continuous learning yet — everyone is still groping. DeepSeek has many promising ideas internally but none have been made to work.
- claimAGI arrives without a critical point but not linearlyc 0.80
There is no discrete threshold at which AGI switches on, but the trajectory is nonlinear because AI can be used to accelerate AI research. That recursive feedback is what bends the curve.
- claimChina's structural advantages are cost and product, not intelligencec 0.80
Chinese labs won't necessarily beat the US on intelligence, but on cost and product experience there are genuine structural advantages. US labs simply don't need to care about cost the way Chinese labs do.
- claimSelling API alone could support a public companyc 0.80
With several hundred million dollars of B-end revenue plus C-end users, DeepSeek is not far from net profit. In the worst case where technology froze today, going all-out on API sales would already be enough to sustain a public company.
- mechanismPrice cuts hurt competitors more than they help usersc 0.75
Slashing prices halves the ARR of anyone whose revenue depends on inference, so a cost-optimized player like DeepSeek can price at levels that make competitors' economics untenable.
- claimWanting competitors to succeed at deploying the modelc 0.75
Rather than fearing that Tencent or others will steal C-end users by hosting DeepSeek's open weights, Liang actively hopes they deploy it well — his only worry is that they'll do it badly.
- contextContinuous learning is the missing ingredientc 0.75
Liang identifies continuous learning — the way a human employee accumulates situational knowledge over weeks — as the specific missing capability separating current systems from human-replacement AGI.
- claimEach paradigm hits a ceiling that superhuman performance can't break throughc 0.75
CoT already surpassed top humans at math olympiad and programming but still couldn't reach AGI. Agents will similarly max out — solving everything they can and still failing to replace employees.
- claimEmbodied intelligence should come last, not firstc 0.75
The preferred roadmap is: continuous learning, then self-iteration, then embodied intelligence. Once the model can self-iterate, embodied intelligence doesn't need humans to build it — the model produces it.
- claim3D, video generation, and world models are off the intelligence main linec 0.75
The AGI main line is language models, CoT, Agents, and continuous learning. Things like 3D, video generation, and world models don't currently raise the intelligence ceiling, so DeepSeek won't work on them — though they're happy to help others who do.
- mechanismBuy as many cards as possible, as fast as possiblec 0.75
DeepSeek's procurement strategy is to convert all available capital into GPUs at any reasonable price, willing to pay a premium. Cash in the bank earns two points; a card gains ten months of social value.
- mechanismAI itself can rebuild the CUDA ecosystemc 0.75
Because AI can write code, building a full NVIDIA-equivalent ecosystem is dramatically easier than before — you can use AI to reconstruct the whole thing.
- claimThe China-US chip gap is fourfold plus two yearsc 0.75
Ecosystem parity will arrive, but on raw chips China lags by roughly four Huawei 950s per NVIDIA GB300, plus a two-year time lag in generations.
- claimContinuous learning is the defining feature of the next-generation modelc 0.75
Before that breakthrough, the work is on cost, effect, and speed. But calling something 'next-generation' requires continuous learning ability.
- mechanismLower cost lets you train larger models on the same computec 0.75
On limited compute, higher computational efficiency directly translates into the ability to afford larger models — so cost reduction is a lever for scaling, not just a customer benefit.
- mechanismCompute lag is absorbed by accepting model lagc 0.75
DeepSeek dissolves its compute disadvantage by accepting that it must use smaller models than US labs. The lag itself buys time, and that time gets converted into cleverer methods.
- implicationEmbodiment is the unavoidable endpointc 0.75
Ordinary humans don't need computers — they need eating, clothing, housing, and transportation solved. If the goal is to relieve human labor, embodied intelligence is where this has to land.
- mechanismOpen research as buying lottery ticketsc 0.75
The threshold to try is low — anyone can draw — but whether anyone pulls something out is unpredictable, maybe talent, maybe luck. So there's no point allocating dedicated resources; you just make it a topic the company takes seriously.
- claimThe post-training bottleneck is time, not moneyc 0.75
Closing the gap on high-quality annotation isn't gated by capital or cards — the existing capital already funds fastest-possible expansion. It's gated by time; domestic efforts really only started in the last half year.
- claimCoding Agent first, vertical Agents laterc 0.75
The most reasonable domestic strategy is to go all-out on a general Coding Agent and deprioritize vertical Agents like finance or medical. Coding covers a lot of ground and subsumes many verticals.
- claimLeaving CUDA via TileLang is an opportunity, not a costc 0.75
Moving off CUDA to a high-level language like TileLang substantially improves efficiency rather than degrading it. The code is much smaller, writing it is fast, and rewriting from scratch becomes feasible — losing 1–2% at the hardware level is acceptable.
- claimThe American lead is cyclical, not structuralc 0.70
Liang thinks OpenAI, Anthropic, and Google's lead won't last, and that Anthropic's coding-agent edge will fade. The gap comes down to compute, not talent — talent is distributed globally by chance.
- mechanismThe company runs on vision, not org structure or KPIsc 0.70
DeepSeek has no formal organization, no KPIs, no assessments, and no written vision statement. Coordination happens through a shared attitude toward the world, visible in how people act rather than what is posted on a wall.
- caveatAI breaks the historical open-source-vs-commerce tradeoffc 0.70
Open-sourcing traditional software used to destroy the business, since markets were only a few billion dollars. AI's addressable market is so much larger that giving away the model doesn't foreclose a viable business.
- claimRestraint is what let ordinary people beat better-resourced rivalsc 0.70
Liang insists DeepSeek had no money, no GPUs, no fame, and no special talent at founding. Restraint is the only explanation he can offer for why an ordinary group succeeded where richer, more famous labs did not.
- exampleNot chasing the Spring Festival user surgec 0.70
When DeepSeek suddenly got a flood of users last Spring Festival, they deliberately declined to retain, monetize, or convert them into a super-app play against ByteDance or Tencent.
- caveatThe open-source model is identical to the deployed onec 0.70
DeepSeek does not keep a stronger private variant — the model released to the community is the same one they serve — which Liang says was already shown last year to cause no C-end conflict.
- mechanismThe chosen roadmap is picked for laziness, not ambitionc 0.70
Each step in this order requires very little new work because you can use earlier technology to help build later technology. Reversing the order — starting with embodied intelligence — would be exhausting and bitter.
- implicationDeepSeek will stay at the tens-of-B active scale for nowc 0.70
Current resources only support experiments at the 10B-active scale, where there are still many things to figure out. Scaling to 150B–250B active waits until more compute arrives; 800B is far off.
- mechanismCompute constrains research, not just trainingc 0.70
Even if you could stack the cards to train an 800B model, you couldn't afford to do sufficient research beforehand. The compute gap therefore compounds into a research gap, which compounds into a talent-development gap.
- mechanismCost is the number-one long-run differentiatorc 0.70
Like BYD undercutting rivals on battery cost at equivalent technology, delivering the same service at lower cost becomes the primary moat. Time-to-market matters second; user experience creates some stickiness but isn't essential.
- claimDeepSeek commercializes, but not as the goalc 0.70
DeepSeek has always been commercializing — hence its C-end users and B-end revenue — but AGI is the goal and commercialization is a byproduct. The point of fully pivoting to commercialization is still very far off.
- mechanismCompute cards are decoupling from gaming cardsc 0.70
CUDA's design was tied to gaming cards because AI was once a smaller market. Now compute exceeds gaming, so dedicated AI chips — from Huawei and even NVIDIA — no longer need CUDA compatibility, further eroding NVIDIA's ecosystem lock-in.
- claimHuawei's supernode is already a price-level substitute for NVIDIA's GB300c 0.70
Even if Huawei's supernode is 100–200% more expensive per equivalent task, that's acceptable — it counts as a real substitute, and NVIDIA is 'digging its own grave.'
- mechanismVision itself is the competitive variable, not realized revenuec 0.70
No one has actually captured the money yet — it's all vision. So the company that visibly aims to take less has already won the framing war, because customers and challengers will route around anyone whose stated ambition is to take more.
- implicationOpenAI's monopoly assumption will not survive contact with cheaper challengersc 0.70
OpenAI began believing it could monopolize the field, but it will meet many challengers — including Chinese firms willing to take less and still provide the service. Monopoly pricing invites its own undercutting.
- caveatThere is a floor: take too little and the business can't survivec 0.70
The race downward has a limit. Take too little and the business logic collapses; take too much and someone taking less defeats you. DeepSeek's stated position is to earn a reasonable return, not a maximized one.
- claimFormal work is capped at roughly half of each person's timec 0.70
Liang's one management standard is that formal, coordinated work should not exceed half of employees' time. The other half is unassigned exploration — no prerequisites, driven by what the person thinks is important.
- mechanismRestraint reduces workload, which removes the need for overtimec 0.70
Because the company is restrained about what it takes on, the set of things to do stays small, so each person's assigned work stays small. Many products are visibly incomplete — that's the deliberate consequence of not filling everything in.
- implicationHigh profit margins in base models would violate objective lawc 0.70
A very high profit margin at this stage doesn't conform to how markets work — a reasonable profit is what's coming. Three or four Chinese firms competing is already enough to trigger a price war.
- claimLarge-model companies will not capture most of the profitc 0.70
With so many large-model firms and gaps between them narrowing, no one will earn windfall profits. The only real differentiators left are time and cost — firms that control cost earn a bit more, and that's the whole story.
- claimParity with foreign models this year, but that isn't AGIc 0.70
Under the current paradigm, Chinese models should substitute for foreign ones this year or within one to two years. That's a solvable engineering target — reaching AGI is a different problem that at minimum requires continuous learning.
- claimNo ceiling visible yet in language-model scalingc 0.70
At current intelligence levels — including the frontier US labs — no ceiling is yet visible in scaling language models. That's why the roadmap still runs through more scale rather than a pivot.
- contextChina has no cost advantage in data annotationc 0.70
Labeling costs in China are essentially the same as in the US, especially for high-end data. That removes the usual expected structural edge and makes matching US labeling investment genuinely hard.
- claimA pursuit beyond profit tends to help commercialization, not hurt itc 0.70
Many great companies had missions beyond profit and that pursuit ended up making them commercialize better. DeepSeek is still a company at bottom — the trade-off is only which money to earn, when, and how much.
- contextThe work feels easy because so much was given upc 0.65
From outside DeepSeek looks like it chose hard mode by focusing on research, but Liang says the team doesn't even need overtime — the ease comes from having refused to spread themselves across products, users, and monetization.
- mechanismThe API is a byproduct, not a productc 0.65
Exposing models through an API required no extra work — no sales, no customer service. It's just a stair-step on the AGI path that happened to be commercializable.
- claimRestraint on everything except the main linec 0.65
DeepSeek deliberately avoids becoming adversaries with anyone and stays willing to help competitors like Alibaba, Zhipu, and Moonshot. Open source already means no clear boundaries, and helping others hasn't cost them any commercial upside.
- evidenceCommercialization plans have consistently been a waste of timec 0.65
Every commercialization conversation over the past three years — ads, e-commerce embedding, local services — has turned out to be wasted effort because change is too fast and product lifecycles are too short. You cannot foresee enough to plan around it.
- implicationLeave commercial opportunities for partners to capturec 0.65
DeepSeek explicitly wants partners and society to build on top of its AI rather than swallowing all the returns itself. This is framed as restraint — the organization doesn't have the people or the appetite to do it all.
- implicationExport controls forced the domestic-chip transitionc 0.65
If NVIDIA cards were freely available, domestic substitution would be hard. Because they can't be bought, everyone is forced onto domestic chips — NVIDIA can't block this.
- claimDomestic hardware's bottleneck is confidence, then capacityc 0.65
The ecosystem problem is fundamentally a confidence problem, and only after that is solved does production capacity become the constraint. Five years out, being stuck on capacity seems unlikely.
- claimTeam stability matters more than top-tier talentc 0.60
Liang argues the biggest lever for reaching AGI is not recruiting the very best people but keeping a team stable over time. He describes DeepSeek as ordinary people doing extraordinary things.
- implicationChina will play the low-cost systematic role in AIc 0.60
Liang expects AI competition to mirror other manufacturing sectors, with China providing cheaper, more efficient services at scale because US firms lack the incentive to push costs to the limit.
- claimCUDA's moat is being eroded, opening a window for Chinese chipsc 0.60
Liang believes new techniques are chipping away at Nvidia's CUDA ecosystem wall. Building a fresh ecosystem around Chinese chips is a historic opportunity, and once fab capacity constraints ease, the compute gap will close.
- implicationOnly three or four base-model players will remainc 0.60
The number of foundation model companies will shrink to a handful that compete fully, and any of them that chase very high margins will be pushed out of the market.
- evidenceDeepSeek's pricing is deliberately below profit-maximizingc 0.60
Liang says user demand for tokens is inelastic in the current range — halving or doubling price barely changes consumption — so the ten-month payback pricing is a chosen restraint, not a market-forced level.
- evidenceAlibaba and Tencent's costs are several times higherc 0.60
Liang estimates that competitors like Alibaba and Tencent, lacking DeepSeek's optimization work, run at several times DeepSeek's inference cost — which is why the pricing bites them.
- contextCost optimization sits in a startup sweet spotc 0.60
Very small startups lack the resources to drive costs this low, and large companies can't organize around it — leaving a mid-sized firm like DeepSeek uniquely positioned.
- evidenceLast Spring Festival's breakout wasn't in the scriptc 0.60
DeepSeek's sudden popularity wasn't planned — they just wanted to make the technology well. That accident proved their non-commercial organization had real talent advantages against companies fighting bloody battles for C-end users.
- claimMultimodality isn't necessary for the algorithmic frontierc 0.60
Doing AI training well doesn't require world models or even multimodality. Narrowing scope loses a portion of tasks but doesn't affect the validity of the underlying algorithm.
- caveatEven with money, cards are hard to actually buyc 0.60
Spending the money is harder than raising it — cards are scarce, prices are high, and you can't pay unlimited premiums. Spending 20 billion yuan in a year would count as procurement performing exceptionally well.
- mechanismThe company runs on consensus, not top-down authorityc 0.60
Wenfeng's authority is built on consensus-seeking, not on decree — he cannot push something through unless the team already agrees. His guidance has limited effect; the decision mechanism is fundamentally a consensus mechanism.
- claimTo B revenue is capped by demand, not computec 0.60
Under current AI technology, enterprise demand is limited — it will grow, but ultimately how much revenue is possible depends on demand, not on how much compute you can throw at it.
- contextTwo management lines: top-down formal and bottom-up self-directedc 0.60
The company runs on two tracks. Formal top-down work coordinates shared deliverables like a V4 release, while a bottom-up track lets each person work on whatever they choose with no KPIs.
- mechanismResearch requires a relaxed environment, not overtimec 0.60
You can't do research if you're pushed tightly — real research needs people mulling problems in their spare time, which only happens when the environment is loose. Overtime culture is incompatible with the kind of exploration DeepSeek wants.
- claimLearning, not agents, is what researchers are actually working onc 0.60
Investors currently see agents as the frontier, but researchers see the open problem as learning — how to solve continuous learning. It isn't a single technology but a problem addressable by many.
- claimAI talent shortage is a phase, not a permanent conditionc 0.60
Every industry looks talent-starved in its first years. Website builders, server-side engineers, pilots — each shortage resolved within a few years as people were trained up, and AI is already visibly following the same curve.
- claimSynthetic data can push models beyond human data ceilingsc 0.60
Real data limits models to what humans have already produced, but that ceiling can be broken. AlphaGo's move that no human had seen shows models can surpass humans within a bounded domain — synthetic and simulated data extend that further.
- caveatComprehensive surpassing is unrealistic given the compute gapc 0.60
With an order-of-magnitude compute deficit, catching up across the board isn't feasible. Selective surpassing in a few carefully chosen areas is.
- evidence16,000 Huawei 950s equal only 4,000 B-series cardsc 0.60
The Huawei allocation is an order of magnitude below what internet giants have, and in raw terms equals only about 4,000 B-series cards. It's enough to train the current generation model but not the next.
- claimBuying Huawei cards is partly an ecosystem investmentc 0.60
DeepSeek could buy non-compliant NVIDIA chips instead, so the purpose of taking Huawei's 950s is largely to help Huawei build out its ecosystem. The near-term compute value is secondary.
- implicationARR of hundreds of millions is plausible but not the goalc 0.55
If demand keeps expanding and more GPUs can be acquired, API revenue could reach hundreds of millions or a billion dollars — enough to cover R&D — but Liang refuses to make this the priority.
- mechanismAnchoring senior employees anchors everyone elsec 0.55
If the oldest and most important employees are retained through generous options, junior staff won't leave either — because most people are here to work in an environment that can actually achieve AGI, not primarily for money.
- evidenceGoodwill and open source haven't cost commercial upsidec 0.55
Being open source and helping competitors hasn't reduced C-end retention or B-end growth — if anything it's been a bonus. The counterfactual of behaving worse wouldn't have yielded more.
- contextCommercial companies win in product, not technologyc 0.55
If your vision is to serve C-end users well, you'll have advantages in product, service, and traffic — but not in technology. In the current era where model technology is what matters most, that's the wrong axis.
- contextDeepSeek runs on roughly 20,000 H-equivalent GPUsc 0.55
Most of DeepSeek's current 20,000 H-equivalent cards arrived only in the last month or two, with more machines still en route. Compute has been scaling up aggressively this year after a lean prior year.
- contextUS capital spending, not $100M salaries, drives the resource gapc 0.55
Headline hundred-million-dollar salaries make talent look like the differentiator, but they are a small fraction of US AI capex. The bulk of investment is compute, and that's where China falls short.
- caveatHuawei alone cannot close the gapc 0.55
Training an 800B model would need 200,000 of Huawei's newest cards just for training, before research is considered. Huawei's output is limited, so the resource problem is currently unsolvable domestically.
- claimModel comparisons only make sense at equal costc 0.55
The right way to compare models is at the same cost, like comparing cars in the same price range. The gap between a well-made and poorly-made model shows up as a comprehensive difference, not at any single link.
- implicationChina's likely global role is as the largest AI producerc 0.55
With the largest production capacity — including for chips — and the most electricity, China is well-positioned to be the manufacturing base of global AI. This makes it plausibly one of the three major players.
- evidenceHuawei allocated DeepSeek only ~16,000 cardsc 0.55
Huawei gave DeepSeek roughly 16,000 cards of capacity while internet giants got over a hundred thousand — capacity, not capability, is the binding constraint.
- contextMultimodality and search are components, not the main linec 0.55
DeepSeek will support native multimodality in V4 and beyond, but treats it — like search — as a component of a product, not as intelligence itself.
- claimCurrent paradigm has a ceiling we cannot yet seec 0.55
The current approach can surpass articulated human knowledge, but it likely has limits that aren't visible from where we stand now. Whether real or simulated data breaks through is an open question with many methods being tried.
- evidenceAgents are limited today precisely because they can't learn continuouslyc 0.55
Current agent capabilities are capped by their inability to continuously learn from experience. That's why the continuous-learning problem sits at the front of the roadmap.
- claimHallucination is a product problem, not a core research problemc 0.55
Hallucination is solvable through better post-training and mostly hasn't gotten the effort it deserves. It matters for user experience but isn't a key open problem.
- implicationHigh-quality data gap could close within a yearc 0.55
Given the pace of expansion, doing high-quality data well domestically should be achievable inside a year. The outlook isn't as cold as it looks — it just needs time.
- contextThe founding intent was never IPO or maximum profitc 0.50
Liang says the first few dozen employees came precisely because the company wasn't chasing capital markets or profit maximization. Anyone with that goal would not have joined.
- exampleThe DDCP price cut that made the team cheerc 0.50
DeepSeek initially priced its DDCP model high fearing overload, but the team was unhappy. When Liang cut the price to a quarter, colleagues cheered — evidence that cheapness for users is the shared motivation.
- implicationReasonable margins will keep compressingc 0.50
Sixfold profit looks high but reflects current AI efficiency; Liang expects reasonable margins to drift down to three- or fourfold over time, though the absolute profit pool remains large.
- caveatThe singularity is really a gradual process, not a discontinuityc 0.50
Despite the label, the transition through self-iteration is expected to be a long gradual change rather than a mutation. The word 'singularity' is a habit of speech inherited from earlier prophets.
- evidenceAfter Sora, small companies quietly abandoned video generationc 0.50
When Sora appeared everyone piled in, but small companies later cut it because it had no relation to the intelligence ceiling. The market's own behavior validates treating video as commercial, not foundational.
- claimAnthropic's lead over OpenAI is temporaryc 0.50
OpenAI and Google will most likely keep rising alternately with Anthropic; Anthropic's Code Agent lead isn't crushing and its first-mover advantage should disappear soon. Among the three, one operates at notably higher efficiency and lower burn.
- exampleThe power-plant analogy for not making chipsc 0.50
If you run a power plant, you don't have to build your own generators — you buy them at a reasonable price. Same for DeepSeek and chips.
- contextLiang's decisions are driven by 'greatest return right now'c 0.50
Whether to build products or push for AGI is decided by which has the greatest return at the moment. Right now, doing products doesn't have the greatest return.
- claimBuying top NVIDIA chips would be worthwhile if it were possiblec 0.45
If Chinese firms could freely buy B200s at reasonable prices, it would clearly pay off. But the constraint isn't cost — the chips simply can't be bought.
- example150B model needed to close the gap with unreleased frontier modelsc 0.45
At the 50B-active scale, DeepSeek expects rough parity with the current open-source wave. But matching the unreleased large model from frontier US labs requires something more like 150B, with training starting optimistically by year-end.
- exampleUsing AI to write TileLangc 0.45
DeepSeek has an internal project applying AI to write TileLang kernels. Even the current human-written version is already much faster to produce than CUDA.
- contextA rare four-hour glimpse into DeepSeek's founderc 0.40
A transcript of a May 20 conversation between DeepSeek founder Liang Wenfeng and his investors surfaced publicly, offering a nearly four-hour window into his thinking on open source, US-China AI, and the road to AGI. Its authenticity cannot be independently verified.
- exampleZhipu's open source feels forced; DeepSeek's is sincerec 0.40
Liang contrasts DeepSeek with Zhipu, which also open-sources but does so reluctantly. For DeepSeek open source flows from its actual vision, which he says is what makes it hold up over time.
- implicationOpen source strengthens internal cohesionc 0.40
Liang cites employee sense of accomplishment and company cohesion — not just external goodwill — as a concrete internal payoff from choosing to give things away.
- contextThe industry's yearly reality: chatbots last year, B-end revenue this yearc 0.40
Everyone was chasing C-end traffic last year and is now scrambling for B-end revenue to stay 'at the table.' DeepSeek treats both waves as beside the point.
- caveatThe domestic-chip story isn't finished yetc 0.40
The TileLang-plus-AI path to full ecosystem parity looks unobstructed but is not yet complete — it still needs time and the Huawei port has to be done.
- contextEcosystem support: willingness without capacityc 0.40
DeepSeek is open to supporting the broader ecosystem and sees no interest conflict, but doesn't have the energy to actively build it out. Cooperation is welcome; commitment beyond that isn't promised.
- contextHuawei cards depreciate faster than NVIDIA'sc 0.40
NVIDIA cards can be depreciated over roughly five years, while Huawei cards last at most three because Huawei starts two years behind. The gap is real but not enormous.
- contextThe organization has to stop being purely flatc 0.40
DeepSeek's previous structure was flat because there was effectively no structure. As headcount grows, some departments need rigorous hierarchy while others stay loose — that reorganization is already underway.
- caveatModel MIS is a direction, not a definite goalc 0.40
Model Office scaling is treated as a relatively definite target. Model MIS is only a direction they'll work toward — not something to bet on as certain yet.
- contextThe company is deliberately unconventional and does not justify itselfc 0.30
Liang isn't looking for external justifications for how the company operates. The reasons exist but may not be conventional — and the company isn't conventional to begin with.
Janitor
Non-content spans (acknowledgements, references, footnotes, headers, boilerplate) are dropped before the decomposition runs.
- total spans
- 458
- kept
- 412
- dropped
- 46
- content · 412
- noise · 39
- header_footer · 4
- boilerplate · 2
- metadata · 1