← Back to list
Lex Fridman PodcastPodcast20 May 2020Source: lexfridman.comHost: Lex Fridman

#97 – Sertac Karaman: Robots That Fly and Robots That Drive

In plain words

MIT professor Sertac Karaman says small consumer drones are easier to make than self-driving cars, but scaling up delivery drones is much harder due to safety and regulation. He's positive on his own company Optimus Ride, which uses remote human monitors for multiple vehicles in closed areas like campuses, not full autonomy. He also thinks lidar (a laser-based sensor) may become so cheap that it's silly not to use it alongside cameras.

AI SummaryAI-generated · may contain errors · verify against the original

At a Glance

MIT professor and co-founder of Optimus Ride, Sertac Karaman, discussed the technical challenges of flying robots and driving robots on the Lex Fridman podcast. The core argument: current consumer-grade autonomous drones are easier to achieve than autonomous driving, but large-scale deployment of flying robots (e.g., logistics transport) is far more difficult than autonomous driving. Key conclusion: over the past 50 years, most deployed robots have operated in isolated environments or confined spaces; a truly large-scale, consumer-facing autonomous flight system has yet to emerge and is expected to be resolved only after large-scale deployment of autonomous driving. Karaman emphasizes that scaling is the core challenge, involving complex issues such as safety, regulation, and system reliability.

~15 min full read · 8 sections
Deep Analysis

Theme 1: Autonomous Flight vs. Autonomous Driving — The Key to the Difficulty Reversal Lies in "Scale"

Karaman believes that autonomous flight is easier to achieve than autonomous driving in consumer drone scenarios (such as aerial photography), but once it involves large-scale logistics, transportation, and other applications, the difficulty will surpass that of autonomous driving.

  • Historical Context: Over the past 50 years, robots deployed by humans (in factories, warehouses, and on Mars) have either operated in isolated environments or confined spaces, rarely interacting directly with humans. The real challenge is placing robots into environments where humans live their daily lives.
  • Mechanism Breakdown: Large-scale deployment implies "density"—thousands of autonomous vehicles in a city, so dense that "when you see one, you look around and see another." This density has never been achieved. Ground vehicles can mitigate risk by "domesticating the environment" (e.g., designating dedicated lanes, setting speed limits), whereas aerial vehicles are far more complex.
  • Extrapolation: Karaman predicts that this density will first be realized by autonomous cars, and only then by flying robots. He gives a timeline of "after large-scale deployment of autonomous driving."

> "We really haven't yet seen any kind of machine at massive scale, large scale being deployed and flown. And I think that's going to be after we kind of resolve some of the large scale deployments of autonomous driving."


Theme 2: Simulation and Perception — From "Simulating Cameras" to "Simulating Human Behavior"

Karaman points out that the core bottleneck of simulation technology has shifted from "simulating physics" to "simulating human behavior," with the latter being the most difficult challenge to overcome.

  • Historical Context: Early simulations could easily handle "interoceptive sensors" (e.g., inertial measurement units) and "dynamics" (e.g., car rolling, aircraft flight), but simulating "exteroceptive sensors" (cameras, radar) has always been difficult. In recent years, camera simulation has approached a "tipping point," enabling realistic scene rendering.
  • Mechanism Breakdown: Karaman proposes a "friend/mother test" — showing a friend or mother two images (one rendered, one real), they would struggle to distinguish them unless a human is present. This is because the human brain has evolved to be extremely sensitive to recognizing humans, while being relatively less discerning of "man-made environments." Simulating human behavior requires mathematical modeling, but human actions are difficult to describe with rules.
  • Implications: This bottleneck also affects autonomous driving. The AI revolution of 2012-2013 solved the problem of "knowing where other objects are," but "predicting what other objects will do next" remains an unsolved challenge. Karaman believes the solution may lie in "iterated learning," which involves accumulating data through extensive experimentation (including failures).

> "We're still missing what everybody else is going to do next. You want to know where you are, you want to know what everybody else is, and then you want to predict what other people are going to do. That last bit has been a real challenge."


Theme 3: Gaming and Ethics — How Autonomous Vehicles "Coexist" with Humans

Karaman emphasizes that the interaction between autonomous vehicles and humans is not only a technical issue but also a social and ethical one, involving a trade-off between "efficiency" and "sustainability (livability)."

  • Mechanism Breakdown: When an empty vehicle (with no driver) travels on the road, humans perceive it as an "object" rather than a "person," and may be more inclined to "abuse" it (e.g., cutting it off, failing to yield). Karaman's research involves "social value orientation," which assesses whether another driver is aggressive or conservative based on observed driving behavior, and adjusts one's own strategy accordingly.
  • Data Chain: He notes that if a vehicle is too aggressive, it can improve transportation efficiency (reducing delays, increasing capacity), but it lowers "livability"—people do not want to live in an environment surrounded by aggressive robots. Conversely, if it is too conservative, efficiency declines.
  • Extrapolation: Karaman argues that as autonomy increases, this trade-off will become "explicit," and society must openly debate "whether we value efficiency or sustainability more." He specifically points out that testing autonomous vehicles itself involves a trade-off between "innovation and risk"—the risk is small but non-zero, and the public needs to be fully informed.

> "If robots are going around being aggressive, you don't want to live in that environment. However, if you're not being aggressive, then you're probably taking up some delays in transportation. So you're always balancing that."


Theme 4: Optimus Ride’s Strategy — A Pragmatic Path from “Geofencing” to “Human-Machine Collaboration”

Karaman outlined Optimus Ride’s differentiated strategy: focusing on closed/semi-closed environments with “transportation scarcity” (such as campuses and naval shipyards), and achieving early deployment through “human-machine collaboration” (one person monitoring multiple vehicles) rather than full autonomy.

  • Mechanism Breakdown: Optimus Ride does not pursue “full self-driving.” Instead, it uses humans as “AV controllers” (similar to air traffic controllers) to remotely monitor multiple vehicles. The vehicles themselves can safely travel at low speeds (e.g., 5 mph), but to increase to 25 mph, human intervention is required for complex scenarios (such as left turns at intersections). Karaman emphasized that this intervention is for “efficiency” rather than “safety” — the vehicles should be safe under all circumstances.
  • Data Chain: He estimated that within a 2-mile by 2-mile area, parking alone could save tens of millions of dollars (in a typical area) to billions of dollars (in downtown New York). Optimus Ride aims to replace traditional 20–30 seat shuttles with 4–6 seat small vehicles, as smaller vehicles offer greater flexibility, shorter wait times, and require no app — passengers can board directly, and the vehicle identifies them to provide destination options.
  • Inference: Karaman believes that such “unexpected” applications (e.g., campus shuttles, warehouse robots) may become widespread before “fully autonomous robotaxis,” gradually acclimating the public to coexisting with autonomous vehicles, ultimately making the arrival of “full autonomy” feel natural rather than abrupt.

> “We want to go from that to 10 people operate 50 vehicles. How do we do that? The help shouldn't be for safety. Help should be for efficiency. Vehicles should be safe no matter what.”


Theme 5: LiDAR vs. Cameras – Not "Either/Or," but "When to Use"

Karaman is open to the notion that "LiDAR is a crutch," but believes the future is more likely to be dominated by "sensor fusion," and that LiDAR may become a "no-brainer" option as costs decline.

  • Mechanism Breakdown: Karaman acknowledges that a camera-only system is theoretically feasible and could be realized within the next 20 years. However, LiDAR is widely used today because it is "easier to build"—simply "mount the LiDAR on the vehicle, write a few lines of code, and press a button to demonstrate." A camera-only system requires more complex computer vision technology.
  • Extrapolation: He proposes two possible paths: 1) LiDAR becomes so cheap in the future that "not using it would seem foolish"; 2) camera-only systems require more powerful computing resources, while LiDAR can reduce computational complexity (thereby lowering hardware costs). He predicts that early deployed autonomous vehicles will use "low-resolution or short-range solid-state LiDAR" (such as MEMS scanning types), as these are low-cost and technologically mature.

> "There will be a time when you can only use cameras and you'll be fine. At that time, it's very possible that you find the LiDAR system as another robustifier, or it's so affordable that it's stupid not to just put it there." (Meaning: There will come a time when cameras alone are sufficient. But by then, LiDAR may be so cheap that not adding it would be foolish.)


Theme 6: Drone Racing and "High-Throughput Computing" – Bridging the Gap from Competition to Safety

Karaman views drone racing as a testing ground for "high-throughput computing," a technology that could in the future enhance the safety of autonomous driving (e.g., by avoiding accidents).

  • Mechanism Breakdown: Human visual processing speed is approximately 100Hz, while robotic cameras can achieve 1kHz (10x slow motion). At the moment of an accident (e.g., veering off the road during a high-speed lane change), if a vehicle has sufficiently fast perception and computing capabilities, it can take over control and come to a safe stop. This requires a "redesign from the chip level"—integrating computing units directly next to pixels to bypass the bandwidth bottleneck of copper wires.
  • Data Chain: Karaman notes that current CPU clock speeds are limited by the speed of light (at a 3GHz clock, light travels only a few inches within the chip), and transistor sizes are approaching the quantum tunneling limit. Therefore, dedicated chips need to be designed for specific tasks (e.g., high-speed camera data processing).
  • Extrapolation: The AlphaPilot Challenge (hosted by Lockheed Martin) is precisely aimed at advancing this technology. Karaman argues that machines have already surpassed humans in "repeatability" and "precision" (e.g., flying the same course ten times in a row—machines remain stable while humans tire), but still lag in "strategy." He predicts that machines may soon beat humans on "moderately complex" courses, but "fully autonomous aggressive flight" will still take years.

> "We end up building systems that see things at a kilohertz, like a human eye would barely hit a hundred hertz. Imagine things that see stuff in slow motion, like 10X slow motion. That will be very useful."


Mentioned Positions

Position Guest Sentiment Key Data
Waymo Neutral (viewed as a long-term research project) No specific data provided
Tesla Slightly positive (acknowledges its product and iteration strategy) No specific data provided
Optimus Ride Bullish (from the founder's perspective) Target: 10 operators managing 50 vehicles; a 2-mile × 2-mile area could save tens of millions to billions of dollars in parking costs
NVIDIA Neutral (viewed as a hardware partner) No specific data provided
Lockheed Martin (AlphaPilot) Neutral (viewed as a technology testing ground) No specific data provided

Judgments Worth Remembering

1. “Scalability is the true bottleneck for autonomous flight, not technical feasibility” (Karaman): Consumer drones are already viable, but large-scale logistics/transportation must first solve the scalability problem of autonomous driving, as ground environments are easier to “tame.”

2. “Simulating human behavior is far harder than simulating the physical world” (Karaman): Camera simulation is nearing an inflection point, but simulating human behavior (e.g., pedestrian gestures, driver intent) remains an “AI-complete” level challenge that cannot be solved through rule-based programming.

3. “The interaction between autonomous vehicles and humans is essentially a trade-off between ‘efficiency and livability’” (Karaman): Aggressive driving improves efficiency but reduces environmental livability; society must openly discuss this trade-off.

4. “Optimus Ride’s strategy is ‘human-machine collaboration’ rather than ‘full autonomy’” (Karaman): One human monitors multiple vehicles, with human intervention aimed at efficiency (not safety); the vehicle should remain safe in all circumstances.

5. “LiDAR is not a crutch; in the future, it may become a ‘why not use it’ option as costs decline” (Karaman): Pure camera systems are feasible, but LiDAR can reduce computational complexity; early deployments will use low-cost solid-state LiDAR.

6. “Drone racing is a testbed for ‘high-throughput computing,’ which could later be used to prevent traffic accidents” (Karaman): 1kHz perception combined with dedicated chips can take over a vehicle at the moment of an accident, with the technology originating from racing scenarios.

7. “The Bellman equation is the most beautiful equation in decision-making—it simultaneously reveals optimality and computational complexity” (Karaman): Theoretically optimal, but the “curse of dimensionality” can make its computation exceed the number of atoms in the universe in extreme cases; however, in practice, average cases are often feasible.

8. “Technology predictions suffer from a ‘time dilation’ effect—distant things are said to be near, and near things are said to be distant” (Karaman): People often “compress” all technological progress into the next three years, leading optimists to say “by year-end” and realists to say “in five years,” but the actual timeline is hard to predict.