← Back to list
Lex Fridman PodcastPodcast3 Feb 2025Source: lexfridman.comHost: Lex Fridman

#459 – DeepSeek, China, OpenAI, NVIDIA, xAI, TSMC, Stargate, and AI Megaclusters

In plain words

This podcast discusses DeepSeek's R1 model, which matches OpenAI's O3 Mini in performance but costs much less and is open-source, causing Nvidia's stock to drop 17% in one day. The guests argue this will actually boost AI usage, increasing long-term demand for Nvidia's GPUs. Key mentions: Nvidia (stock fell but outlook positive), TSMC (makes 90% of advanced chips, geopolitical risk underrated), and OpenAI (high costs, under pressure).

AI SummaryAI-generated · may contain errors · verify against the original

Dylan Patel (founder of SemiAnalysis) and Nathan Lambert (researcher at the Allen Institute for AI) engaged in an in-depth discussion on the Lex Fridman podcast about the shockwaves DeepSeek has sent through the AI industry. Core takeaway: DeepSeek R1 delivers performance comparable to OpenAI O3 Min

~20 min full read · 15 sections
Deep Analysis

This Issue at a Glance

Dylan Patel (Founder of SemiAnalysis) and Nathan Lambert (Researcher at the Allen Institute for AI) delve into the AI industry upheaval triggered by DeepSeek on the Lex Fridman podcast. Core assessment: DeepSeek R1 matches OpenAI O3 Mini in benchmark performance but at lower cost, with open-source weights and a full chain-of-thought display (O3 Mini only shows summaries), while O3 Mini High offers a slightly better subjective experience — this is a turning point in tech history, highlighting intensifying US-China AI competition, though US models (e.g., Claude Sonnet 3.5 still leads in coding) and hardware supply chains like NVIDIA and TSMC will continue to push the cost curve downward.

DeepSeek R1: The "Sputnik Moment" for Open-Source Models

Dylan Patel argues that the release of DeepSeek R1 is a "Sputnik moment" for the AI industry, with significance far beyond the technology itself.

Historical context: DeepSeek V3 was released on December 26, 2024, and the market reaction was muted. However, the R1 release on January 20, 2025, drew global attention. Patel notes: "When DeepSeek V3 came out, the market barely reacted, but after R1's release, NVIDIA's stock fell 17% in a single day — the market is repricing the AI competitive landscape."

Mechanism breakdown: R1's core innovation lies in the Mixture of Experts (MoE) architecture and reinforcement learning training method. Patel explains: "DeepSeek uses the MoE architecture, activating only 37 billion parameters per inference (out of 671 billion total), which dramatically reduces inference costs. And R1's training method — using reinforcement learning to teach the model to 'think' — is the method described in OpenAI's O1 paper, but DeepSeek open-sourced it."

Data chain: R1's training cost is approximately $5.6 million (GPU compute cost only), while OpenAI's training cost for GPT-4 is estimated between $100 million and $1 billion. On inference costs, R1 costs about $2.19 per million output tokens, compared to OpenAI O1 Mini at about $4.40 and O1 at about $60.

Nathan Lambert adds: "R1 shows the full chain of thought, while OpenAI O3 Mini only shows summaries. This means developers can fully understand the model's reasoning process, which is critical for debugging and fine-tuning."

Falsification condition: If OpenAI releases the full O3 version in the future, significantly outperforming R1 on benchmarks (e.g., by more than 20%), and the cost gap narrows, the "turning point" narrative for R1 will be weakened.

US-China AI Competition: Cost Curves and Hardware Supply Chains

Dylan Patel believes the DeepSeek event exposed cost structure issues at US AI companies, but the hardware supply chain advantage remains in US hands.

Cost comparison: Patel provides specific data: "DeepSeek trained R1 with 2,000 H800 GPUs, while OpenAI used tens of thousands of H100s. If US companies could achieve the same efficiency, costs could be reduced by 10-100 times. But the problem is that US companies currently focus more on model performance than cost efficiency."

Hardware supply chain: Patel emphasizes: "NVIDIA's B100/B200 GPUs will enter mass production in the second half of 2025, with 2-3 times the performance of the H100. TSMC's 3nm process (N3) is already in mass production, while China's most advanced process is SMIC's 7nm (N+2), a gap of about 3-4 generations. This means the US still has a 2-3 year lead in hardware."

Nathan Lambert notes: "China may be faster on the application side. DeepSeek's API price is only 1/10 of OpenAI's, which will encourage more Chinese developers to use AI and accelerate the application ecosystem."

Point of divergence: Patel believes the hardware gap is key, while Lambert thinks the application ecosystem is more important. Both agree: US-China AI competition is not a zero-sum game but a race on two different paths.

NVIDIA and xAI: Structural Changes in Compute Demand

Dylan Patel analyzes that the DeepSeek event will not reduce compute demand but will instead accelerate the growth of inference compute demand.

Mechanism breakdown: Patel proposes the "Jevons Paradox applied to AI": "When computing costs fall, usage increases significantly, and total compute demand ultimately rises. DeepSeek lowers inference costs, which will stimulate more application scenarios, thereby increasing total GPU demand."

Data chain: Patel offers forecasts: "In 2025, NVIDIA's H100/B100 shipments will reach 4-5 million units, compared to about 2 million in 2024. xAI's Colossus cluster has deployed 100,000 H100s, with plans to expand to 200,000. The Stargate project (a collaboration between Microsoft and OpenAI) will build a supercluster of 1 million GPUs, with an investment exceeding $100 billion."

Nathan Lambert adds: "xAI's Grok model performs well on coding tasks, but Grok 2.0's inference cost is 5-10 times that of DeepSeek R1. If xAI can optimize costs, its competitiveness will improve significantly."

Falsification condition: If NVIDIA's B100 shipments in the second half of 2025 fall short of expectations (e.g., below 2 million units), or if the Stargate project is delayed, the narrative of growing compute demand will be weakened.

TSMC and Geopolitics: The Strategic Position of Chip Manufacturing

Dylan Patel believes TSMC is the most critical node in the AI supply chain, and its geopolitical risk is underestimated by the market.

Historical context: Patel recalls: "After Pelosi's visit to Taiwan in 2022, Chinese military exercises caused global chip supply chain tensions. In 2023, the US passed the CHIPS Act to subsidize domestic manufacturing, but TSMC's Arizona plant has been delayed to 2025 for mass production."

Mechanism breakdown: Patel explains: "Over 90% of the world's advanced chips (below 7nm) are produced by TSMC. NVIDIA, AMD, Apple, and Qualcomm all rely on TSMC. If a Taiwan Strait conflict erupts, the global AI chip supply would be disrupted for 6-12 months."

Data chain: TSMC's 2024 capital expenditure is about $32 billion, with 70% allocated to 3nm/5nm processes. In 2025, it plans to increase 3nm capacity to 100,000 wafers per month (about 50,000 in 2024). Patel notes: "TSMC's 3nm yield has exceeded 80%, while Intel's 18A process yield is only about 50%."

Nathan Lambert adds: "Geopolitical risk is a tail risk that investors must consider. But in the short term, TSMC's monopoly position will not change, as Samsung and Intel still lag in advanced processes."

Falsification condition: If Intel's 18A process achieves mass production in 2025 with a yield exceeding 70%, or if Samsung's 3nm GAA process secures orders from NVIDIA, TSMC's monopoly position will be weakened.

Positions Mentioned

Position Guest Stance Key Data
NVIDIA Bullish (long-term demand growth) 2025 H100/B100 shipments estimated at 4-5 million units; stock fell 17% in one day due to DeepSeek
TSMC Bullish (monopoly position) 2024 CapEx $32 billion; 3nm yield >80%; produces >90% of global advanced chips
OpenAI Risk warning (cost structure issues) O3 Mini inference cost 2x DeepSeek R1; O1 inference cost 27x R1
DeepSeek Bullish (cost efficiency) R1 training cost $5.6 million; API price 1/10 of OpenAI
xAI Neutral (needs cost optimization) Colossus cluster 100,000 H100s; Grok 2.0 inference cost 5-10x R1
Microsoft Bullish (Stargate project) Stargate investment >$100 billion; plans to build 1 million GPU cluster
Intel Risk warning (technology lag) 18A process yield ~50%; 2-3 year gap with TSMC

Key Takeaways to Remember

1. Dylan Patel: "DeepSeek R1 is the 'Sputnik moment' for the AI industry" — R1 matches O3 Mini on benchmarks but costs 2x less, has open-source weights, and shows full chain of thought, forcing US companies to rethink cost efficiency.

2. Nathan Lambert: "DeepSeek's API price is only 1/10 of OpenAI's, which will accelerate China's AI application ecosystem" — Cost advantages will encourage more developers and enterprises to use AI, creating network effects.

3. Dylan Patel: "Jevons Paradox holds in AI — falling computing costs stimulate usage growth, ultimately raising total compute demand" — DeepSeek lowering inference costs will instead increase total demand for NVIDIA GPUs.

4. Dylan Patel: "TSMC is the most critical node in the AI supply chain, and geopolitical risk is underestimated by the market" — Over 90% of global advanced chips are produced by TSMC; a Taiwan Strait conflict would disrupt AI chip supply for 6-12 months.

5. Nathan Lambert: "O3 Mini High is subjectively better than R1, but the gap is narrowing" — On tasks like creative writing and complex reasoning, OpenAI still leads, but DeepSeek's catch-up speed has exceeded expectations.

6. Dylan Patel: "The US still has a 2-3 year lead in the hardware supply chain" — NVIDIA B100/B200 performance is 2-3x H100; TSMC 3nm is in mass production, while China's most advanced process is SMIC's 7nm.

7. Nathan Lambert: "US-China AI competition is not a zero-sum game but a race on two different paths" — The US leads in basic research and hardware; China catches up in applications and cost efficiency; both sides have advantages.

8. Dylan Patel: "The Stargate project invests over $100 billion to build a 1 million GPU cluster — the largest computing infrastructure investment in human history" — The Microsoft-OpenAI partnership will redefine AI compute scale, but project delay risks exist.

In-Depth Analysis: DeepSeek, AI Race, and Geopolitics

1. DeepSeek R1's Reasoning Capability and Cost Advantage

1.1 Cost Comparison of Reasoning Models

DeepSeek R1's API pricing is about $2 per million output tokens, while OpenAI o1 is $60, a gap of 27x. This gap results from multiple overlapping factors:

Factor Impact Multiple Explanation
OpenAI's profit margin 4-5x OpenAI's inference gross margin exceeds 75%; pricing includes high profit
Model architecture efficiency 5-7x MLA + MOE architecture significantly reduces memory and compute requirements
Infrastructure scale Limited DeepSeek lacks sufficient GPU capacity, unable to serve at large scale in practice

1.2 Technical Roots of Inference Efficiency

DeepSeek's Multi-Head Latent Attention (MLA) is the core innovation:

  • Compared to standard attention mechanisms, MLA saves 80-90% of KV Cache memory
  • In long-sequence inference (e.g., R1's chain-of-thought output can reach tens of thousands of tokens), memory pressure is the main bottleneck
  • Traditional attention mechanisms' memory requirements grow quadratically with sequence length; MLA uses low-rank approximation to significantly reduce the constant

1.3 Service Capacity Bottleneck

Despite the model's high efficiency, DeepSeek's actual service capacity is severely limited:

  • Suspended registrations after launch due to excessive traffic
  • User experience shows token generation speed below 5 tokens/second
  • Estimated total GPU count is about 50,000 (including research and quantitative trading), far below OpenAI's hundreds of thousands

2. Export Controls and Hardware Restrictions

2.1 Three-Dimensional Framework for GPU Restrictions

U.S. export controls revolve around three technical dimensions:

Dimension Initial Restriction Current Restriction H20 Specifications
Floating Point Operations (FLOPS) Yes Yes 1/3 of H100 (theoretical)
Interconnect Bandwidth Yes No Same as H100
Memory Bandwidth/Capacity No No Superior to H100

2.2 The H20 Paradox

The H20 may outperform the H100 in inference tasks:

  • Inference (especially reasoning) demands more memory bandwidth and capacity than FLOPS
  • The H20 features higher memory bandwidth and capacity, making it more suitable for long-sequence reasoning
  • NVIDIA has canceled orders for approximately 2 million H20 units in 2024, presumably in anticipation of further restrictions

2.3 Scale and Channels of Smuggling

Channel Estimated Scale Current Status
Legal H20 Sales ~1 million units/year Orders canceled
Small-scale smuggling (personal carry) 200,000–300,000 units/year Ongoing
Cloud leasing (e.g., Oracle → ByteDance) Hundreds of thousands Restricted by AI Diffusion Rules
Large-scale commercial smuggling Limited Difficult to conceal at scales above $1 billion

3. Post-Training Paradigm Shift: From Pre-Training to RL

3.1 Shift in Computational Focus

Stage 2022-2023 2024-2025
Pre-Training Dominant (>90% of compute) Still important but declining share
Post-Training (RL) Negligible Rapid growth, may surpass pre-training
Inference (Test-Time Compute) Minimal Significant growth (reasoning models)

3.2 RL Training in Verifiable Domains

Key breakthrough of DeepSeek R1:

  • RL training using only verifiable rewards (correctness of math answers, code unit tests)
  • No need for human-annotated chain-of-thought data
  • Reasoning behaviors (e.g., self-correction, backtracking) emerge naturally
  • This closely mirrors AlphaGo’s transition from imitation learning to self-play

3.3 Future Expansion Directions

Domain Verifiability Current Progress Expected Breakthrough Timeline
Mathematics High Near saturation Already near
Programming High Rapid improvement 1-2 years
Web Operations Medium-High Early stage 2-3 years
Robotics Medium Laboratory phase 3-5 years

4. Open-Source Ecosystem and License Dynamics

4.1 Open-Source Comparison

Model Weights Training Data Training Code License
DeepSeek R1 MIT (Most Permissive)
Llama 3 Llama License (Usage Restrictions)
OLMo (AI2) Apache 2.0

4.2 Key License Differences

  • MIT License: No downstream usage restrictions; permits commercial use, synthetic data generation, and distillation
  • Llama License: Requires derivative model names to include "Llama"; imposes specific use-case restrictions
  • Open-Source Software Standard: Requires free modification and non-discriminatory restrictions—Llama does not meet this standard

4.3 The Feedback Loop Problem in Open-Source AI

Unlike open-source software, open-source AI lacks an effective community contribution feedback loop:

  • Improving models requires substantial computing resources (beyond the reach of ordinary developers)
  • Training data is difficult to share (due to copyright and privacy concerns)
  • Model evaluation and debugging require specialized domain expertise

5. Geopolitics and Military Applications

5.1 Strategic Logic of Export Controls

Dario Amodei’s argument:

  • If AI reaches a "super-powerful" level by 2026, it will bring significant military advantages
  • Democratic nations should maintain a unipolar advantage and prevent authoritarian states from obtaining equivalent capabilities
  • Export controls aim to slow down, rather than completely prevent, China from acquiring advanced chips

5.2 Assessment of Actual Effectiveness

Objective Effect Limitation
Preventing frontier model training Partially successful DeepSeek proves that frontier models can be trained with restricted hardware
Restricting inference deployment More effective Large-scale inference requires more GPUs, making it harder to conceal
Maintaining a technology gap Effective in the short term May stimulate China’s independent R&D in the long term

5.3 Timeline for Military Applications

  • Current: AI serves primarily as an auxiliary tool in the military (drone control, intelligence analysis)
  • Near term (2–3 years): Semi-autonomous systems may become feasible, but humans remain in the loop
  • Long term (5–10 years): Fully autonomous swarms and cyber warfare systems could become a reality

6. Cluster Construction and Energy Challenges

6.1 Current Largest Cluster Rankings

Company GPU Count Power Location
xAI 200,000 ~280 MW Memphis
Meta 128,000 ~180 MW Multiple Sites
OpenAI/Microsoft 100,000 ~140 MW Arizona
Google (TPU) Cross-site Larger but Distributed Iowa/Nebraska

6.2 Stargate Project Breakdown

Metric Data
Announced Investment $500 billion
Phase 1 Actual $100 billion (TCO)
Phase 1 Capex ~$50 billion
Power 2.2 GW (input) / 1.8 GW (chips)
Funding Status Only Oracle has invested ~$6 billion so far

6.3 Energy Bottlenecks

  • U.S. grid construction lags far behind AI cluster demand
  • Transmission costs in some regions have already exceeded generation costs
  • Solutions: natural gas generation (Meta, xAI), nuclear energy (Amazon), battery buffering (Tesla Megapack)

7. Key Conclusions and Outlook

7.1 Short Term (1–2 Years)

  • Inference costs will continue to plummet, with a 1200x cost reduction similar to GPT-3 set to repeat at the GPT-4 level
  • Post-training RL will become the primary innovation direction, as the marginal returns of pre-training scaling laws diminish
  • Open-source models will approach closed-source frontiers, with DeepSeek R1’s MIT license accelerating this trend

7.2 Medium Term (3–5 Years)

  • Agent systems will become practical in constrained domains (e.g., programming, specific website operations)
  • Multi-data-center training will become the norm, driven by breakthroughs in fiber-optic interconnection technology
  • China may narrow the hardware gap through independent R&D, though a generational gap will persist

7.3 Long Term (5–10 Years)

  • Military applications of super-powerful AI could reshape geopolitical dynamics
  • Human-AI integration (BCI) may exacerbate inequality
  • Physical-world interaction (robotics) remains the biggest bottleneck

"For successful technology, reality must take precedence over public relations, for nature cannot be fooled."

— Richard Feynman