← Back to list
Colossus (Invest Like the Best / Business Breakdowns)Podcast27 Aug 2024Source: joincolossus.comHost: Patrick O'Shaughnessy

Gavin Baker - AI, Semiconductors, and the Robotic Frontier - [Invest Like the Best, EP.385]

In plain words

This piece argues the biggest AI investment opportunity lies in the efficiency gap: GPU performance has improved 50x in 5 years, but supporting systems only 4-5x. Gavin Baker is bullish on NVIDIA (tracked since 2000, CUDA moat), Tesla (FSD 12.3 huge leap, 100-1000x more data than rivals), and Apple (smartphones as key AI inference nodes, Apple can charge a 'toll'). He warns most AI apps are thin 'LLM wrappers' with weak moats; real winners may emerge in 5-6 years.

AI SummaryAI-generated · may contain errors · verify against the original

Gavin Baker (Managing Partner and CIO of Atreides Management) delved into the frontiers of AI, semiconductors, and robotics during the program. Core views: Generative AI follows scaling laws, but infrastructure faces challenges, with enormous future demand for data centers. Baker emphasizes that imp

~14 min full read · 10 sections
Deep Analysis

Gavin Baker - AI, Semiconductors & Robotics Frontier - [Invest Like the Best, EP.385]

At a Glance

Gavin Baker (Managing Partner and CIO of Atreides Management) is one of the few investors who has tracked NVIDIA since 2000, with 25 years of experience in the semiconductor industry. The core theme of this episode: The efficiency gap in AI infrastructure—GPU performance has increased 50x, but supporting systems have only improved 4-5x—represents the biggest investment opportunity today. Baker proposes a complete AI efficiency equation (MAMF × SFU × checkpoint frequency × PUE) and argues that whoever gains an edge in "effective compute per dollar of capex per watt of power" will dominate the next generation of model competition.


Theme 1: Tech Giants Enter a "Digital God" Arms Race, ROI Is Not a Decision Factor

Gavin Baker argues that the core players among the Mag7 (Meta, Google, Microsoft, XAI) have abandoned traditional ROI frameworks for AI spending, because they believe the value of creating the first Artificial General Intelligence (AGI) is in the tens of trillions to hundreds of trillions of dollars.

Baker points out that the founders or controlling figures of these companies (Zuckerberg holds super-voting rights, Larry Page's influence within Google, Satya Nadella's decision-making power at Microsoft) all view losing this race as an existential threat to the company. "Larry Page has reportedly said multiple times inside Google, 'I would rather go bankrupt than lose this race.'" This conviction stems from their firm belief that scaling laws will continue to hold—model quality improves as compute increases.

Key inflection point: XAI built a 100,000-GPU coherent training cluster in Memphis, the first globally to reach this scale. Baker believes Elon Musk redesigned the data center architecture, enabling 100,000-GPU coherence even using Hopper (rather than the next-generation Blackwell). "This means we could see the first GPT-4.5-level model within 6–9 months." With Blackwell's arrival, cluster scale could expand to 300,000 GPUs, potentially giving rise to GPT-5/6-level models.

Readers should note: Baker is a long-term investor in these tech companies, and his narrative of an "inevitable arms race" also serves his own portfolio logic. He acknowledges that "if scaling laws are disproven, it would be a disaster for the entire data center infrastructure."


Theme 2: The AI Efficiency Equation — MAMF × SFU × Checkpoint Frequency × PUE

Baker proposes a unified efficiency equation for measuring AI lab competitiveness, arguing this is a more important differentiator than model quality itself.

The core issue: current theoretical GPU utilization (MFU) stands at only 35–40%, meaning over 60% of compute power is wasted. Baker decomposes MFU into:

1. MAMF (Maximum Achievable Matrix Multiplication Floating-point): Measures single-GPU software efficiency. NVIDIA GPUs can reach 83% under CUDA optimization, while AMD's Instinct MI 300 initially achieved only 25–30%, now improved to 60%. "This quantifies the CUDA moat — AMD's MI 250 is actually a decent chip, but no one uses it because MAMF might be only 10%."

2. SFU (System Floating-point Utilization): Measures the efficiency of supporting systems such as networking, storage, and memory. "GPUs have become 50x faster over the past 5 years, but the rest of the data center has only improved 4–5x. That's why MFU is so low — GPUs spend most of their time waiting."

3. Checkpoint Frequency: Due to frequent GPU failures (melting, optical link interruptions, switch outages), training clusters must frequently save model states. More reliable clusters require fewer checkpoints.

4. PUE (Power Usage Effectiveness): Electricity costs are becoming a key variable. "There are only about 3 places in the US that can reliably supply 1 GW of nuclear power, and electricity prices there could be 10x the average."

Ultimate goal: Effective compute per dollar of capex per watt (exaflops/second per dollar of capex per watt). Baker believes that design decisions across different labs will lead to efficiency differences exceeding 100%, directly determining whether the cost of GPT-7/8 is $500 billion or $1 trillion.

Investment Implications: Invest in the weakest links of the efficiency chain — currently SFU (networking, storage, memory technologies) and PUE (cooling, power management).


Theme 3: Decentralization of Inference – The Superphone and Apple’s “Toll Booth” Model

Baker predicts that inference computing will undergo a rebalancing from the cloud to edge devices, with smartphones becoming the key node for AI inference, and Apple occupying the most advantageous monopoly position.

Core logic: Computing has gone through cycles of centralization, decentralization, and re-centralization. Training must take place in large data centers, but inference follows the principle of “local first, cloud on demand.” “Local inference is free; cloud inference burns GPU hours.” Therefore, DRAM capacity on smartphones will replace flash storage as the most closely watched hardware metric for consumers — it determines the parameter scale of models that can run locally, and thus the IQ of the AI assistant.

Baker outlines Apple’s business model: when local models are insufficient, queries are routed to the cloud via a “router.” Apple will charge model providers a “toll,” much like Google pays to be the default search engine on iOS. “Eventually, you might pay $60 a month for cloud-based superintelligence, and $1,000 for hyperintelligence. Humans with higher-IQ AI will gain a massive competitive advantage.”

Falsification conditions: If improvements in model efficiency drive cloud inference costs toward zero, or if a new edge computing paradigm emerges (e.g., BCI brain-computer interfaces), this thesis may not hold.


Theme 4: The Application Layer Dilemma — LLM Wrappers and the Collapse of SaaS Metrics

Baker argues that investing in the AI application layer is currently extremely difficult, as the traditional SaaS growth curve has been disrupted by AI-native companies, yet most lack a sustainable moat.

Key data: Chase Coleman (Tiger Global) notes that ChatGPT is to AI what Netscape Navigator was to the internet (1994). Companies founded within two years of the internet's birth now account for less than 1% of the global internet market capitalization. The vast majority of giants were founded five to six years later. "All the companies VCs invested in from 1995 to 1997 looked like they were riding the wave, but the truly great companies emerged later."

Baker cites Eric Vishria of Benchmark, who observes that AI-native companies are "blowing up" traditional SaaS metrics — reaching $30 million ARR from zero in nine months with higher cash flow efficiency. However, they are essentially "thin wrappers around LLMs." "They aren't competing for software budgets; they are competing for labor budgets." The question is: how to build defensibility around these wrappers? Baker suggests possible paths include exceptional sales execution, deep integration, building RAG systems for small businesses, and composite AI systems (using small models for simple queries and routing complex queries to large models).

Key risk: If these companies fine-tune, they become locked into a specific generation of models, losing the ability to benefit from the rapid iteration of foundation models.


Theme 5: Robotics – The Underestimated "Blue-Collar Labor Revolution"

Baker argues that robotics (especially humanoid robots) could have a greater impact on the world within five years than the automation of white-collar jobs, and that Tesla's FSD is the first truly world-changing robot.

The inflection point of FSD: Tesla's FSD version 12.3 (fully based on deep learning, eliminating nearly all human-coded rules) delivered "the sum of all progress from the past 10 years." Version 12.5, running on AI4 hardware, represents another step-change. Baker estimates that the current compute level of FSD is equivalent to GPT-2, while Tesla is building a cluster of over 50,000 H100/H200 GPUs in Austin. "They will soon jump from GPT-2 to GPT-4.5-level compute, which means a 100x improvement."

Key data advantage: Tesla possesses a vision training dataset based on actual driving miles, with a scale 100 to 1,000 times that of its closest competitor, Waymo. "It's as if they have YouTube, all Meta platforms, the open internet, and X, while everyone else is stuck with Yahoo."

The logic behind humanoid robots: Both Elon Musk and Jensen Huang have publicly stated that humanoid robots will prevail because "the world is optimized for humans." A general-purpose robot can perform any task a human can, and after mass production, its cost advantage becomes significant. "To all non-humanoid robotics startups, I wish you venture returns on par with Lycos or MySpace, but none of you will become Google."

Biggest uncertainty: Whether synthetic video data is effective. "We know synthetic text data works, but we don't know if synthetic video data works. No one knows." If synthetic video data proves viable, the competitive landscape for FSD could be completely transformed.


Theme 6: Leadership — The Common Traits of Jensen, Elon, and Lisa Su

Baker believes the commonalities among these three CEOs—focusing on the most critical issues, a love for hearing bad news, and mission-driven leadership—set them apart from the vast majority of corporate leaders.

Specific traits:

  • No Fixed Schedule: Jensen told Baker, "I have no fixed schedule, no standing meetings. I find the company's most important problem, move my desk to that area, and marshal the best resources."
  • Love for Bad News: Baker quotes his father, a bankruptcy lawyer: "The common thread in all bankrupt companies is that the CEO doesn't like hearing bad news." Both Elon and Jensen demand that bad news reach them immediately.
  • No Hierarchy: Regardless of where a problem originates in the company, they work directly with the person who knows it best. "If the best model analyst is 24 years old, he sits right next to Jamie Dimon."
  • Mission-Driven: From Jensen's "making photorealistic virtual worlds a reality" to Elon's "making the world sustainable," the mission attracts top talent. "The world's brightest minds spent 20 years making it slightly more likely for someone to click a link. But if you have a mission, you get better employees."

Mentioned Positions

Position Guest Stance Key Data
NVIDIA Bullish (long-term hold) Coverage since 2000; GPU performance improved 50x in 5 years; CUDA MAMF at 83%
AMD Bullish (Lisa Su's leadership) Instinct MI 300 MAMF rose from 25% to 60%; only 20 days of cash when Lisa Su took over
Tesla Bullish (FSD & Optimus) FSD 12.3 described as "the sum of 10 years of progress"; Austin cluster >50K H100/H200; AI4 hardware
Apple Bullish (inference node monopoly) Will establish a "router" charging model; local models + on-demand cloud
Meta Bullish (AI ad ROI) Revenue re-accelerated to 5x due to AI ad targeting; open-sourced Llama but retains data
Google Neutral to slightly positive YouTube + Knowledge Graph + Maps data advantage; Larry Page's "rather go bankrupt" stance
XAI Bullish (data center innovation) First 100K GPU coherent cluster; Elon redesigned data center architecture
OpenAI Neutral (position not disclosed) Achieves internet-scale distribution via Microsoft
Anthropic Neutral (position not disclosed) Achieves distribution via Amazon
Waymo Neutral (competitive observation) Geofencing approach; LIDAR route; will "brute-force" the data problem
Mistral Neutral (technical recognition) Fewer parameters but evaluation results close to GPT-4 level

Judgments Worth Remembering

1. "These companies' ROIC actually improved after CapEx increased—because they are replacing human labor with GPU hours. The AI ROI debate is intellectually absurd." (Gavin Baker) — Meta rebounded from an 80% decline to a 5x gain driven by AI ad targeting, with revenue re-accelerating; this is the clearest case of AI ROI.

2. "GPUs have become 50x faster over the past five years, but the rest of the data center has only improved 4-5x. That's why MFU is only 35-40%." (Gavin Baker) — The weakest links in the efficiency chain (networking, storage, memory) represent the biggest investment opportunities.

3. "If you believe in scaling laws, you believe AI has extremely high marginal costs. This is the complete opposite of the traditional zero marginal cost of software." (Gavin Baker) — This explains why infrastructure efficiency will become the most critical competitive dimension for model companies.

4. "Chase Coleman's statistic: ChatGPT is to AI what Netscape Navigator was to the internet. Companies founded within two years of the internet's birth currently account for less than 1% of global internet market cap." (Gavin Baker) — Investors should remain highly humble about application-layer investments; the true winners may not emerge for 5-6 years.

5. "Tesla FSD version 12.3 represents the sum of all progress over the past 10 years. Current compute power is equivalent to GPT-2, and they will jump from GPT-2 to GPT-4.5, implying a 100x improvement." (Gavin Baker) — FSD is following a faster scaling law than GPT because there is more room to catch up.

6. "Tesla's visual training dataset is 100-1000x larger than Waymo's, the second-largest. It's like they have YouTube, all of Meta's platforms, and the open internet, while others only have Yahoo." (Gavin Baker) — The data flywheel is Tesla's deepest moat in autonomous driving.

7. "Humanoid robots will be to robotics what GPT was to AI—a general-purpose breakthrough. The world is optimized for humans, so general-purpose robots will ultimately win." (Gavin Baker) — Both Elon and Jensen have publicly agreed with this judgment; non-humanoid robotics startups will find it hard to become the "Google" of the space.

8. "Jensen said: 'I have no fixed schedule, no standing meetings. I find the company's most important problem and move my desk to that area.' Elon is the same—regardless of the company, he solves the most critical problem." (Gavin Baker) — "Focus on the most important problem + love bad news + no hierarchy" is a common trait among top CEOs.

9. Baker's AI efficiency equation: MAMF (software efficiency) × SFU (system efficiency) × checkpoint frequency × PUE (power efficiency) = effective compute per dollar per watt. — This is the core framework for evaluating AI lab competitiveness and a roadmap for finding investment opportunities.

10. "Synthetic text data works, but no one knows if synthetic video data works. If it does, the competitive landscape for FSD could completely change." (Gavin Baker) — This is the biggest source of uncertainty in Baker's optimistic view on Tesla's autonomous driving.