← Back to list
Colossus (Invest Like the Best / Business Breakdowns)Podcast30 Sep 2025Source: joincolossus.comHost: Patrick O'Shaughnessy

Dylan Patel - Inside the Trillion-Dollar AI Buildout - [Invest Like the Best, EP.442]

In plain words

This piece breaks down the trillion-dollar AI buildout: money, power, and chips are reshaping the industry. Dylan Patel says the AI race is just in its first inning, with post-training and reinforcement learning driving future compute demand, not pre-training. Key holdings: OpenAI ($16B ARR but signed a $300B Oracle deal, high risk), NVIDIA (75% gross margin, effectively cutting prices via equity investments), and Oracle (could make $100B profit if OpenAI succeeds). He warns that traditional SaaS models are collapsing under AI's high cost of goods sold, since every token costs money.

AI SummaryAI-generated · may contain errors · verify against the original

Dylan Patel, on the podcast Invest Like the Best, offered a deep dive into the physical realities of trillion-dollar AI infrastructure construction. His core argument is that the AI race is still in its early stages, and the compute demands for post-training and reinforcement learning (RL) are far f

~15 min full read · 9 sections
Deep Analysis

This Issue at a Glance

Dylan Patel, founder and CEO of SemiAnalysis, is renowned for tracking data center construction via satellite imagery and mapping the flow of hundreds of billions in capital. The central theme of this issue is that the physical realities of AI infrastructure—from electricity and labor to supply chain bottlenecks—are reshaping the entire industry's power dynamics and investment logic. The most impactful judgment in the entire piece: Patel believes the AI race is still in the early stages of "the first inning," with compute demand for post-training and reinforcement learning (RL) far from reaching its ceiling, while the traditional SaaS economic model is collapsing under the weight of AI's high cost of goods sold (COGS).


I. The "Infinite Money Loop" of OpenAI-NVIDIA-Oracle: Power Dynamics and Capital Structure

Dylan Patel argues that the three-party transaction between OpenAI, NVIDIA, and Oracle is essentially "the highest-stakes gamble in capitalist history," with the core logic being: whoever can provide a balance sheet to front the capital expenditure (Capex) for AI infrastructure holds the power.

Patel breaks down the mechanics of this deal: OpenAI signed a five-year contract with Oracle worth approximately $300 billion. Oracle purchases GPUs from NVIDIA, while NVIDIA makes an equity investment of roughly $100 billion in OpenAI. Specifically, for a 1-gigawatt (GW) data center, construction costs amount to about $50 billion, of which NVIDIA's chips account for roughly $35 billion. NVIDIA's gross margin is 75%, translating to approximately $30 billion in gross profit—half of which flows back to OpenAI in the form of equity.

"It's like the Spider-Man pointing meme—OpenAI pays Oracle, Oracle pays NVIDIA, and NVIDIA pays OpenAI," Patel notes. He points out that this structure stems from a fundamental contradiction: OpenAI's demand for computing power is "infinite," yet its annualized revenue is only about $16 billion (by end of 2025), while its committed spending runs into the hundreds of billions. Therefore, it must find allies willing to "front the capital first and collect rents later."

Key Data Chain:

  • Annual rent for a 1 GW data center: $10–15 billion, totaling $50–75 billion over five years
  • OpenAI's current ARR: approximately $16 billion (estimated by end of 2025)
  • Total contract value signed by Oracle: approximately $300 billion
  • NVIDIA's equity investment: approximately $100 billion (corresponding to 10 GW capacity)

Patel's Assessment: If OpenAI succeeds, Oracle will reap $100 billion in pure cash profit; if it fails, Oracle will be saddled with massive debt. Meanwhile, NVIDIA, through this structure, "effectively cuts prices without lowering prices," while securing upfront cash and equity.


2. Tokenomics: The Fundamental Contradiction Between Exponential Growth in Inference Demand and Model Economics

Patel proposes a "Tokenomics" framework, arguing that the core of the AI economy is "the economics per token"—the growth in inference demand (doubling every two months) far outpaces the pace of hardware deployment, forcing model providers into a difficult trade-off between "bigger and slower" versus "smaller and faster."

"GPT-5 is not much larger than GPT-4 because OpenAI faces a fundamental dilemma: if the model is too large, no one can afford it or use it quickly enough." Patel points out that GPT-4.5 is actually smarter, but "no one can bear its cost and speed," so the vast majority of Anthropic's revenue comes from Claude Sonnet (a smaller model) rather than Opus (a larger model).

The contradiction between inference demand and hardware supply:

  • Inference token demand: doubling every two months
  • Hardware deployment speed: far below this growth rate
  • Solution: Model costs must continue to decline (costs from GPT-3 to GPT-4 fell by 2,000 times; DeepSeek further reduced them by 500–600 times)

Patel's reasoning: OpenAI chose to keep GPT-5 at a similar scale to GPT-4o in order to maximize the number of users it can serve, rather than pursuing extreme intelligence per single inference. This means that the consumer experience of "larger models" will only be realized once the hardware (inference latency) bottleneck is broken.

Falsifiable condition: If the growth rate of inference token demand slows to doubling every 6–12 months, or if the pace of model cost reduction fails to match demand growth, the current capital expenditure trajectory will face risk.


3. Post-Training and Reinforcement Learning: The True "First Inning"

Patel believes that pre-training in the text domain has entered the "late innings," but post-training and reinforcement learning (RL) have only just "thrown the first pitch"—this is the primary source of future compute demand.

"Inning" Assessment by Stage:

Stage Inning Key Bottleneck
Text Pre-training Late Innings Internet text data is largely exhausted
Non-text Pre-training (Video/Audio/Image) Early Innings Data is abundant but processing costs are extremely high
Post-training/RL First pitch just thrown Requires building a large number of "environments" to generate training data

Patel draws an analogy with an infant calibrating its senses: "A baby puts its hand in its mouth because the tongue is the most sensitive organ—it is calibrating its sense of touch. Models have not yet achieved this. We are still far from true general intelligence."

Diversity of Environments:

  • Simulated e-commerce platforms (e.g., a "fake Amazon") to train models in comparing and purchasing products
  • Math puzzles (progressing from simple to complex)
  • Medical case diagnosis
  • Code debugging and data cleaning

Key Data: Currently, there are approximately 40 startups in the Bay Area building these "environments," but Patel believes "whether they can succeed remains questionable."

Extrapolation: The compute demand from RL will "absorb most of the computational resources," as models must learn through trial and error, with each trial requiring substantial inference and training compute. Falsifiable Signal: If RL methods fail to significantly improve model performance on real-world tasks (e.g., complex manipulation, multi-step reasoning) within 12–18 months, the current expectations for RL compute investment may be too high.


4. Hardware Bottlenecks and Supply Chain Realities: Electricity, Labor, and "Diesel Truck Engines"

Patel argues that the biggest constraint on AI infrastructure is not the chips themselves, but electricity supply, labor shortages, and supply chain bottlenecks—physical realities that are giving rise to "crazy" solutions.

Current State of Electricity and Labor:

  • U.S. data centers currently account for 3-4% of national electricity consumption (of which AI accounts for roughly half)
  • Electrician wages have doubled, especially for "mobile electricians" who can relocate for work
  • Some companies are using diesel truck engines in parallel as emergency power sources, due to insufficient supply of gas turbines
  • Elon Musk is sourcing electrical equipment from Poland for shipment to the U.S., as domestic supply chains cannot meet demand

Supply Chain Bottlenecks:

  • GE Vernova and Mitsubishi are doubling gas turbine production capacity
  • The transformer supply chain is "completely sold out," and companies are beginning to consider imports from China
  • Different data center builders (Google, Vantage, EdgeConneX, QTS) have varying supply chains, leading to "all sorts of strange problems"

Key Data Points:

  • Total capital expenditure (including GPUs) for a single 500 MW data center: approximately $25 billion
  • The 2 GW data center OpenAI plans to build would consume as much electricity as the entire city of Philadelphia
  • Texas's ERCOT and PJM grids have introduced new rules allowing for 24-72 hours advance notice to cut half of a large load's power

Patel's Assessment: "We are relearning how to build power systems. This isn't because demand is too high, but because we haven't built new power infrastructure in 40 years."


5. US-China AI Competition: Two Distinct Philosophies of Capital Allocation

Patel argues that the US must win the AI race to maintain global hegemony, or else face economic recession and social instability; meanwhile, China is employing its "historically typical long-term strategy"—leveraging state capital for massive investment in the semiconductor supply chain in pursuit of self-sufficiency.

US Perspective:

  • "Without AI, the US may no longer be the world's dominant power by the end of this century"
  • AI must "significantly accelerate GDP growth," otherwise "once the pie starts being divided, it's over"
  • US capital primarily flows toward: building the largest data centers, training the best models

China Perspective:

  • Over the past decade, has invested $400–500 billion into the semiconductor ecosystem (via state-owned enterprises, tax policies, land subsidies, the "Big Fund," etc.)
  • Pursues supply chain security through "self-sufficiency," rather than short-term commercial returns
  • ByteDance may be the second or third largest GPU user globally (behind only OpenAI and Meta)
  • China's construction speed far exceeds that of the US: "If China wants to build a 10 GW data center, it could be completed within a few years; the US cannot build one in the short term"

Key Comparison:

Dimension US China
Capital Allocation Focus Building data centers, training models Building semiconductor supply chain, self-sufficiency
Construction Speed Slow (constrained by labor, regulation) Fast (government coordination, centralized decision-making)
Chip Capability 3–4 years ahead Catching up (Huawei, etc.)
Talent Global talent attraction Dense local talent pool (roughly half of the world's engineers are of Chinese descent)

Patel's Warning: "If the US pushes China into a corner, they will fight back. But if the US does not apply pressure, China will gradually overtake through long-term accumulation."


6. The Collapse of the SaaS Economic Model: The "High COGS Dilemma" in the AI Era

Patel argues that AI is fundamentally reshaping software economics—the traditional SaaS model of "low COGS, high customer acquisition cost" will be upended, as AI introduces high cost of goods sold (COGS) while simultaneously lowering barriers to entry for competitors.

Traditional SaaS Model:

  • R&D costs are relatively fixed
  • COGS is very low
  • Customer acquisition cost (CAC) is high
  • Profitable once reaching critical scale

SaaS Model in the AI Era:

  • COGS rises significantly (each token incurs computational costs)
  • CAC does not decline (sales remain difficult)
  • Barriers to entry for competitors drop substantially ("you can write code yourself with AI")

Patel's Reasoning:

1. AI reduces software development costs, meaning more enterprises can "build rather than buy"

2. AI SaaS products have high COGS, compressing profit margins

3. Result: The market becomes more fragmented, making it harder for individual SaaS companies to achieve "escape velocity"

4. "Large companies that have already scaled (such as YouTube) will do well, but new entrants will face severe challenges"

Historical Analogy: The Chinese SaaS market is relatively small because Chinese software developer costs are only 1/5 of those in the U.S., making enterprises more inclined to build in-house. AI is now extending this "build-in-house economy" globally.


Mentioned Positions

Position Analyst Stance Key Data
OpenAI Bullish but with risk warnings ARR ~$16B (end of 2025); 800M users; signed ~$300B Oracle contract
NVIDIA Bullish Gross margin 75%; effectively lowering prices via equity investments; ~$35B Capex in a 1GW data center
Oracle Bullish (conditional) Signed ~$300B contract; could generate $100B in profit if OpenAI succeeds
Anthropic More bullish (preferred over OpenAI) Revenue grew from <$1B to $7-8B; primarily from code-related products
Meta Strongly bullish Has "full stack" (hardware + models + recommendation systems + capital); glasses could become the next-generation human-computer interface
Google Turned bullish (previously bearish) Lowest token cost (vertically integrated TPU); waking up and investing aggressively
AMD Personally bullish, business "mediocre" Author's first multi-bagger stock; but sees difficulty competing with NVIDIA
xAI Risk warning Faces funding challenges; needs to find a business model (currently relies mainly on "Anime" chatbot)
ByteDance Not explicitly stated Could be the second or third largest GPU user globally
TSMC Not explicitly stated Author believes "if you believe in Taiwan risk, you cannot invest in any US tech giant"
CoreWeave Bullish (conditional) Signed ~$19B contract with Microsoft; credit risk in its contract with OpenAI
Nebius Bullish Signed ~$19B contract with Microsoft
Periodic Labs Bullish (author's personal investment) Uses RL methods to accelerate materials science/chemistry discovery
Cursor Neutral (power dynamics analysis) Revenue near $1B; most revenue flows to Anthropic; but owns user data and embedded models

Judgments Worth Remembering

1. Patel believes "Tokenomics" is the core framework for understanding the AI economy — the value creation and cost structure per token determine profit distribution across the entire industry chain. Inference demand doubles every two months, but hardware growth is far from keeping pace, so model costs must continue to decline.

2. "Post-training and reinforcement learning have only thrown the first pitch" — Patel draws an analogy to an infant putting its hand in its mouth to calibrate touch: models have not yet learned to understand the physical world through trial and error. The compute demand from RL will "absorb most of the computing resources."

3. "The traditional SaaS economic model is collapsing" — AI brings high COGS while lowering the barrier to entry for competitors ("you can use AI to write your own code"), leading to market fragmentation and making it harder for new entrants to achieve escape velocity.

4. "Electrician wages have already doubled, and companies are starting to use diesel truck engines in parallel as backup power" — The biggest bottleneck for AI infrastructure is not chips, but power supply and labor shortages. The U.S. is "relearning how to build an electrical system."

5. "If the U.S. does not win the AI race, it may no longer be the world's superpower by the end of this century" — Patel argues that without AI-accelerated GDP growth, the U.S. will "fall apart" due to debt, income inequality, and social fragmentation.

6. "ML research is like semiconductor manufacturing — there are thousands of knobs to tune, and you can't try them all" — Patel uses the "process knowledge" in semiconductor fabrication as an analogy for intuition and experience in AI research, emphasizing that "knowing what to try" is more important than "how many times you've tried."

7. "NVIDIA is doing something unprecedented — using its balance sheet to win the competition" — Through equity investments and demand guarantees, NVIDIA is effectively "cutting prices without cutting prices" while locking in upfront cash.

8. "Meta may be the only company with a full stack for the next-generation human-machine interface" — Patel believes that from hardware (glasses), models, recommendation systems, to capital, Meta is best positioned to dominate the new interaction paradigm of "you say it, AI does it."