← Back to list
Colossus (Invest Like the Best / Business Breakdowns)Podcast31 Mar 2026Source: colossus.comHost: Patrick O'Shaughnessy

Sergey Levine - Building LLMs for the Physical World - [Invest Like the Best, EP.465]

In plain words

This interview is about the future of robots. UC Berkeley professor Sergey Levine argues that building a general-purpose robot brain (a foundation model) that can control any machine is better than making specialized robots for single tasks. His company Physical Intelligence is doing this, and its model can already open doors and wash pans, but struggles with tasks like changing a baby's diaper. He is optimistic about this general approach. Key mentions: Physical Intelligence (model improving fast), Boston Dynamics (cool demos but no real use), Tesla (its data-collection model is a good example to follow).

AI SummaryAI-generated · may contain errors · verify against the original

UC Berkeley professor and co-founder of Physical Intelligence, Sergey Levine, argues that building a general-purpose robot foundation model is the right path, as learning across robots, environments, and tasks is more scalable than developing narrow-domain experts. His company is developing a robot

~9 min full read · 6 sections
Deep Analysis

Here is the English translation of your analysis of the Sergey Levine interview, following all specified rules.

At a Glance

Sergey Levine, a Professor at UC Berkeley and co-founder of Physical Intelligence, argues that building a general-purpose robot foundation model is the correct path. Learning across robots, environments, and tasks is more scalable than developing narrow-domain experts. His company is developing a robot foundation model capable of controlling any physical system to perform any task, completing new tasks without direct training. Levine points out that everyday human actions remain the hardest problem in robotics, and that human trust and acceptance are as important as technological breakthroughs in determining when robots will integrate into daily life.

The Generality Bet: Why the "Harder" Path Might Be "Easier"

Sergey Levine believes that pursuing full generality, rather than building specialized robots for specific tasks, may be the simpler path in the long run. This draws on the experience of LLMs: building dedicated systems for specific tasks like machine translation was ultimately surpassed by general-purpose language models that could leverage vast amounts of web data.

  • Core Mechanism: By learning across multiple robots, environments, and tasks, a general model builds a "world understanding" of physical interaction. This understanding allows it to quickly "draw analogies" like a human when facing new tasks, without needing to be retrained from scratch. Levine emphasizes: "If you can leverage data from many sources, many applications, many robots, then you can have a model with physical understanding. Building new applications on this platform becomes much easier."
  • Data Challenge & Strategy: Unlike LLMs, which have access to internet-scale data, the robotics field lacks ready-made large-scale datasets. Levine's strategy is not to pre-quantify the required data volume, but to "make the system useful enough that it can go out into the world and collect more data on its own," forming a data flywheel similar to Tesla's. The key is reaching a critical threshold of "usefulness."
  • Unique Insight: Levine acknowledges that the generality path does not excel in demonstrations. A specialized robot can perfectly execute a single task in a controlled environment, which looks cooler. The value of a general model lies in "doing mundane things that any human can do, but doing them in any situation." This generalization capability is difficult to showcase in a single video.

Technological Breakthrough: From "Muscle Memory" to "Common Sense Reasoning"

Levine points out that the most surprising progress in robotics is that models have far exceeded expectations in dexterity and cross-morphology generalization, while the current biggest bottleneck has shifted from low-level physical control to mid-level semantic reasoning.

  • Dexterity & Generalization: Physical Intelligence's models, without special design, can perform highly dexterous manipulations simply by adding more data, and can adapt to different robot morphologies (e.g., multi-fingered hands, different degrees of freedom). The model can even automatically adapt without being told its own morphology.
  • The Science of "Common Sense": Levine defines "common sense" as applying semantic knowledge learned from other domains to solve a current physical task, as opposed to "muscle memory." It is implemented through "Chain of Thought": the robot enters a scene, first thinks about the task (e.g., "clean the kitchen"), then executes actions. This unlocks the prior knowledge the model gained from pre-training on internet data, allowing it to handle never-before-seen edge cases.
  • "Coach-Style" Improvement: A key finding is that when a model fails in a long-horizon task (e.g., cleaning an entire kitchen), simply adding high-level semantic instructions (e.g., "put the plate in the cupboard") as "coach-style" corrections can improve its generalization ability. This means the bottleneck has shifted from "can the robot do it?" to "can the robot understand the scene and choose the correct steps?" — a problem that can be supervised at low cost through language.

Future Vision: From "Tool" to "Platform," From "Cool" to "Useful"

Levine believes the value of a general-purpose robot foundation model lies not in creating a "humanoid Terminator," but in becoming a "platform" that inspires countless robot applications of various forms, much like the personal computer sparked the Cambrian explosion of the software ecosystem.

  • Morphology Innovation: Levine is open to, but not exclusively focused on, "humanoid robots." He believes a general model can adapt to various forms, from "a construction crew of 1,000 quadcopters" to "micro-surgical robots." The key is lowering the barrier to entry, allowing more people to "assemble a robot in their garage, load the foundation model, and make it move."
  • The "Cool" vs. "Useful" Trade-off: Physical Intelligence's strategy is to "be as cool as possible within the constraints of being useful." They choose tasks like "making espresso" or "folding laundry" that are both challenging and push the boundaries of technology. Levine specifically mentions the concept of a "Robot Olympics," which would compete on tasks like "opening a door" or "cleaning a greasy frying pan" — simple for humans but extremely hard for robots — rather than parkour or backflips.
  • The Last Tasks to Be Conquered: Levine believes that tasks involving intimate human interaction, requiring high empathy and fine physical control, such as "changing a baby's diaper" and "elderly care," will be the hardest for robots to master. This perfectly embodies Moravec's Paradox — what is easiest for humans is hardest for machines.

Position Moves

Position Guest's Stance Key Data
Physical Intelligence Bullish (Co-founder) The model can already complete almost all tasks in the "Robot Olympics" except "turning a shirt inside out" and "peeling an orange."
Boston Dynamics Appreciative (Technical level) Its Atlas robot is praised for its agility, which is "very human-like and very un-human-like," but Levine also notes it "has been doing cool demos for a long time without doing anything useful for customers."
Tesla (Autonomous Driving) Positive Analogy Its data collection flywheel (collecting data while humans drive) is a model the robotics field hopes to replicate.
Roomba (iRobot) Historical Reference Mentioned as "the best-selling consumer robot of all time."

Judgments Worth Remembering

1. The Generality Bet (Sergey Levine): Pursuing full generality may be simpler in the long run than developing narrow-domain experts, because a general model can build a "world understanding" from broader data, allowing it to adapt quickly to new tasks.

  • Support: Analogous to LLMs replacing dedicated systems like machine translation; Physical Intelligence's model can clean a new kitchen without being retrained for each one.

2. "Coach-Style" Improvement (Sergey Levine): When a robot fails, simply providing high-level language instructions as "coaching" can improve its generalization, meaning the bottleneck has shifted from physical control to semantic understanding.

  • Support: Physical Intelligence found that adding semantic labels (rather than more action data) for failure scenarios improved model performance, making "talking to the robot" an effective improvement method.

3. The "Robot Olympics" (Sergey Levine, citing Benji Holson): The true measure of a robot's capability should be tasks like "opening a door" or "cleaning a greasy frying pan" — simple for humans but extremely hard for machines — rather than parkour or backflips.

  • Support: Physical Intelligence's general model can complete all but two items on this list (due to hardware limitations), proving the power of the generalist approach.

4. The Last Tasks to Be Conquered (Sergey Levine): Tasks involving intimate human interaction and fine physical control, such as "changing a baby's diaper" and "elderly care," will be the hardest for robots to master.

  • Support: This is the ultimate expression of Moravec's Paradox — what humans are best at (interacting with people, fine manipulation) is hardest for machines, with extremely low tolerance for error.

5. Data Flywheel > Data Scale (Sergey Levine): The key is not knowing in advance how much data is needed, but making the system useful enough that it can enter the real world and collect more data on its own, forming a self-reinforcing flywheel.

  • Support: Analogous to Tesla's autonomous driving system, whose data collection capability far exceeds its needs; the robotics field should aim to reach this critical point of "usefulness."

6. The "Cool" vs. "Useful" Trade-off (Sergey Levine): In robotics, the coolest demos (e.g., backflips) are often not the most useful, while the most useful capabilities (e.g., generalized cleaning) look mundane.

  • Support: Boston Dynamics' demos are cool but haven't translated into real customer value; Physical Intelligence's strategy is to pursue "cool" within the constraints of being "useful."

7. Falling Hardware Costs as a Key Catalyst (Sergey Levine): Robot hardware costs have plummeted over the past decade (from ~$400,000 for the PR2 to ~$3,000 for a robotic arm today), making general-purpose robots feasible.

  • Support: Low-cost, low-precision hardware combined with powerful learning algorithms can compensate for the lack of precision, dramatically lowering the deployment barrier.

8. Optimistic Researcher, Pessimistic Entrepreneur (Sergey Levine): In robotics, his optimism about the technology's prospects is higher than most senior researchers, but lower than many robotics entrepreneurs.

  • Support: He sees many "puzzle pieces" that can fit together, but acknowledges the long history of failures in robotics, requiring caution.