← Back to list
Colossus (Invest Like the Best / Business Breakdowns)Podcast9 Jul 2024Source: joincolossus.comHost: Patrick O'Shaughnessy

Martin Casado - Entering Uncharted AI Territory - [Invest Like the Best, EP.381]

In plain words

This episode explains the economics of AI: making content is now nearly free (e.g., turning a photo into a Pixar-style character costs $0.0001 vs. $200), but physical-world AI like self-driving is still uneconomical after $75 billion invested. Casado (a16z partner) argues AI exploits human-created data, not true understanding, so creative/chat AI works but robotics doesn't. He likes Ideogram (AI image generation, backed by a16z) and Cursor (AI coding tool that changes how code is written), and flags Tesla's self-driving as risky.

AI SummaryAI-generated · may contain errors · verify against the original

Martin Casado (a16z partner) discusses the opportunities and challenges brought by the AI revolution in this episode of Invest Like the Best. Core thesis: AI is entering uncharted territory and will profoundly reshape creativity, policy-making, agentic systems, and real-world data structures. Key co

~12 min full read · 10 sections
Deep Analysis

Martin Casado - Entering Uncharted AI Territory - [Invest Like the Best, EP.381]

At a Glance

In this episode, Martin Casado (a16z Partner) explores the core judgments of the AI revolution: The economic impact of AI on content creation is a foregone conclusion, but fundamental questions remain about the economic viability of robotics/physical-world AI. Casado proposes a key framework—AI models essentially "exploit" structured data created by humans (accumulated over 3,000 years) rather than truly understanding the physical world; this explains why AI for creativity, companionship, and language reasoning is already commercially viable, while physical-world AI such as autonomous driving remains uneconomical despite $75 billion in investment.


Theme 1: Three Waves of Marginal Cost Approaching Zero — AI as the Third "Zero Marginal Cost" Revolution

Casado argues that AI is driving the marginal cost of content creation toward zero, and its economic impact will be comparable to that of the microchip (which zeroed out computing costs) and the internet (which zeroed out distribution costs).

Historical Context:

  • First wave (1940s): ENIAC was 5,000 times faster than human calculation (a 4-order-of-magnitude improvement), zeroing out the marginal cost of computing → giving rise to the entire computing industry (IBM, DEC, HP)
  • Second wave (1980s–90s): The internet zeroed out the marginal cost of distribution → the internet boom
  • Third wave (present): AI zeroes out the marginal cost of content creation. Example: Converting a real human face into a Pixar-style character — a traditional designer would charge $200, while the AI inference cost is only $0.0001 — again a 4-order-of-magnitude gap

Data Points:

  • AAA game production costs can reach over $500 million, yet "there is no aspect of a game that cannot be created using models (music, sound effects, textures, 3D meshes, characters, cutscenes, images)," with theoretical COGS dropping to around $10
  • For every $1 of revenue generated by current AI models, approximately 1/3 comes from "emotional connection" applications (virtual friends, companions, interactive games)

Implications: Casado emphasizes that "when COGS decline, it is not a zero-sum game — we will create far more content." However, he warns that the supply cycle may become misaligned with the demand cycle (analogous to the overcapacity of internet fiber optics), potentially leading to a "great winter."


Theme 2: Human Data as "Structured Energy Reserves" — The Economic Dilemma of Physical World AI

Casado proposes a core framework: the success of current AI models hinges on their training using structured data already created by humans, rather than raw physical world data. This explains why creative AI is commercially viable, while robotics and autonomous driving remain uneconomical.

Mechanism Breakdown:

  • The human brain's "neocortex" (prefrontal cortex, approximately 250,000 years old) is responsible for creativity and linguistic reasoning — this part is relatively inefficient and easily surpassed by computers.
  • The human brain's "old brain" (navigating 3D space, approximately 4 million years old) is extremely efficient — the human brain consumes only 20 watts, while an autonomous vehicle's GPU system can reach 1.3 kilowatts.
  • Key finding: Research by Rohan (AGI House) shows that the higher the gzip compression ratio of a data corpus, the better the model's learning performance — indicating that models are merely "exploiting" the structure already embedded in human data.

Data Chain:

  • The industry has invested $75 billion in autonomous driving, yet unit economics remain more expensive than Uber.
  • The language and image data accumulated by humans over 3,000 years has been "largely exhausted."
  • Synthetic data is effective in axiomatic systems (chess, arithmetic, simulable spaces), but for language and creativity, "I don't even know what that means."

Implications: Casado argues that "we don't know if we can obtain enough data from merely observing the world through cameras to train economically viable models — you are competing against a brain that has evolved over 4 million years." He suggests that the "bitter lesson" may fail here — "everything in computer science is an S-curve; exponential growth always plateaus eventually."


Theme 3: The Moat of AI Companies — From "Model Competition" to "Traditional Barriers"

Casado argues that models themselves do not possess a durable competitive advantage, because distillation creates "abnormal economies of scale" — after a new model is released, competitors can catch up quickly at low cost. The true moat comes from traditional business elements.

Model Company Lifecycle:

1. Release a model → poor performance → reset to zero

2. Release a good-enough model → user surge ("sawtooth wave")

3. Users engage → grow bored → churn

4. Must release a new model → repeat the cycle

5. Must build a traditional moat during this process

Types of Effective Moats:

  • Integration Moat (API stickiness): OpenAI's API is widely integrated, but LLMs have "anti-stickiness" — unpredictability forces customers to prepare for switching
  • Community Moat: Sharing/viral spread/community for creative companies
  • Workflow Moat: User interface and operational habits independent of the model
  • Brand/Style Differentiation: Markets fragment rather than consolidate during expansion — image models already exhibit a "bookstore effect": Pico (anime), Midjourney (gritty fantasy realism), Ideogram (cartoon text), Adobe (photorealistic)

Casado emphasizes: "If you release a text model, you'd better have a good traditional business defense plan. Relying solely on the 'best model' is not enough to endure."


Theme 4: Regulatory Debate – The "Baptists and Bootleggers" Framework

Casado draws a parallel between the AI safety debate and the "Baptists and Bootleggers" economic phenomenon from the U.S. Prohibition era—where true believers (Baptists) and profiteering speculators (Bootleggers) jointly drive regulation.

Historical analogy (Internet era):

  • Early 2000s: Similar panic rhetoric emerged, such as "Digital Pearl Harbor" and "the end of sovereignty"
  • At that time, there were specific threat evidence (e.g., the Morris worm) and clear shifts in the security landscape (asymmetric warfare)
  • Conclusion: Do not hinder innovation; open source is beneficial and is the way to maintain U.S. leadership
  • Key security contributions came from academia and open source (SELinux, Sourcefire, Mandiant)

Current AI regulatory issues:

  • Lack of specific threat evidence comparable to the Internet era—arguments such as "recursively self-improving AI" and "biological weapons" have been effectively refuted
  • California's SB-1047 bill (which Casado calls a "terrible California bill") originates from a Berkeley professor and his students—academia has shifted from being past advocates of open source to being "fully complicit"
  • Casado believes "big tech has deeply embedded itself in government and academia, leaving too few neutral voices"

Implications: Casado calls on the "silent majority" (academia, small tech companies) to speak up. "If you don't work at a big company and won't benefit from regulation, wake up and get involved—don't let us lose our leadership through regulation."


Theme 5: Agents – The Most Divisive Topic Within the Team

Casado personally holds a "bearish" view on agents, but acknowledges that most of the team disagrees with him. He defines an agent as "a system with a control loop and state—given a high-level goal, it autonomously executes multiple steps and maintains state."

Technical Bottlenecks:

  • State Problem: Unsolved—"feels like an afterthought, not part of the reasoning architecture"
  • Control Loop: Unsolved convergence issue—"don't know how to handle the non-determinism of these models"
  • Planning Capability: Unsolved
  • Memory Problem: Unsolved

Similar Divergence on AI Coding:

  • One view: Coding is over, just tell AI what you want
  • Casado's view: 90% of coding time is maintenance (understanding and readability), humans still need to know everything; low-code/no-code never eliminated programmers, and AI may be the same
  • But he is bullish on AI-native IDEs like Cursor—"changes how coding is done, but does not replace programmers"

Inference: Casado believes "the breakthroughs in these large generative models may not necessarily solve the planning/agent problem. I would very much like to be proven wrong, but that's my gut feeling."


Theme 6: Infrastructure Innovation — The "New Frontier" of Data and Hardware

Casado argues that AI is driving comprehensive innovation across the entire infrastructure stack, from data management to custom chips.

On the data side:

  • Unstructured data (video, images, music) is converging with structured data
  • "We still don't know how to efficiently process large-scale unstructured data pools" — cross-video queries, workflow integration, and data extraction remain unsolved problems
  • Once existing data is exhausted, "nobody truly knows how to obtain more data without spending a fortune"

On the hardware side:

  • AI workloads are fundamentally different from cloud workloads — the communication patterns and network architecture requirements for training and inference vary significantly
  • The economics of custom ASICs have shifted: with a single training run costing $1 billion, a custom ASIC costs around $200 million; if it improves efficiency by 50%, $200 million yields $500 million in savings — "for the first time in a long while, it makes economic sense to design custom chips for a single project"
  • Network architecture innovation: the largest wave of innovation since the advent of data centers

Casado's assessment: "The entire infrastructure stack will be upended. The silicon space is still in a very early stage."


Mentioned Positions

Position Guest Sentiment Key Data
Ideogram (Image Generation) Bullish (a16z has invested) Former Google Imagen team, pioneers of diffusion models
Character AI Neutral observation (as a user behavior case) Founded by Noam Shazeer, one of the authors of the Transformer
Cursor (AI-native IDE) Bullish Changes the way of coding but does not replace developers
OpenAI Neutral (has integration moat) API widely integrated, but LLMs exhibit "anti-stickiness"
Meta (Llama) Neutral (open-source but lacks enterprise sales engine) Open-source model, but no enterprise-level GTM
Mistral Neutral (open-source + enterprise customer engagement) Engages in both open-source and enterprise solutions
Anthropic Neutral Pushing frontiers, may find an independent niche
NVIDIA Not explicitly stated (mentioned only as background) GPU power consumption reaches 1.3 kW (autonomous driving scenario)
Tesla (Autonomous Driving) Risk warning Autonomous driving is "notoriously difficult"; still uneconomical after $75 billion industry investment

Judgments Worth Remembering

1. "AI models essentially exploit structured data created by humans, rather than understanding the physical world" (Casado) — The 3,000 years of human-accumulated language/image data is a "structured energy reserve"; once depleted, model progress will slow significantly; synthetic data works in axiomatic systems, but for language and creativity, "I don't even know what that means."

2. "Distillation creates anomalous economies of scale — after a new model is released, competitors can catch up quickly at low cost" (Casado) — This means the "best model" is not a durable moat; model companies must build traditional business barriers (API stickiness, communities, workflows, brands) within a "sawtooth wave" cycle.

3. "Currently, for every $1 of revenue AI generates, about 1/3 comes from emotional connection applications" (Casado) — Companionship/virtual friends/interactive games are "a new modality computers have never solved," not traditional enterprise automation; Casado sees this as a precursor to changes in the future "access layer."

4. "After $75 billion invested in autonomous driving, unit economics are still more expensive than Uber" (Casado) — AI in the physical world competes against the human brain evolved over 4 million years (20 watts vs. GPU 1.3 kilowatts), while creative/language AI competes against the prefrontal cortex with only 250,000 years of history — this is the root cause of the economic disparity.

5. "The economics of custom ASICs have changed — a $1 billion training run, spending $200 million on a custom chip that improves efficiency by 50% yields a net gain of $300 million" (Casado) — "For the first time in a long while, it makes economic sense to customize chips for a single project," and the entire infrastructure stack will be disrupted.

6. "Market expansion leads to fragmentation, not consolidation — a 'bookstore effect' has already emerged in image models" (Casado) — Pico (anime), Midjourney (fantasy realism), Ideogram (cartoon text), Adobe (photorealistic) each find independent niches, rather than "only one winner."

7. "The 'Baptists and bootleggers' framework in AI regulation debates — true believers and speculators profiting from it jointly drive regulation" (Casado) — In the internet era, there was concrete evidence of threats (Morris worm) and clear changes in security posture, ultimately concluding "don't hinder innovation"; current AI regulation lacks similar evidence, but academia has shifted from open-source advocates to "full complicity."

8. "Three unresolved problems for agent systems: state, control loop convergence, and non-deterministic handling" (Casado) — Personally bearish, but most of the team disagrees; similar to the divide on AI programming — "low-code/no-code never eliminated programmers, and AI may be the same."