AI expert Gary Marcus argues that deep learning (AI that learns from data) has big limits: it finds correlations but lacks common sense, like 'a bottle holds water' or 'people drink when thirsty.' He wants to combine deep learning with symbolic AI (which uses logic rules). Marcus is cautious about AI progress, saying it's gradual, not a sudden singularity. Key examples: DeepMind (lost $530M in 2018, its tech fails on language), GPT-2 (writes fluently but fails on simple logic), and Google Duplex (only handles simple tasks like booking haircuts, not general).
Gary Marcus, in his discussion on the Lex Fridman podcast, explored the necessity of integrating deep learning with symbolic AI. He pointed out that current deep learning suffers from fundamental limitations such as low data efficiency and a lack of common-sense reasoning, making it incapable of ach
Gary Marcus is a Professor Emeritus at New York University and founder of Robust.AI, who has long taken a critical perspective on the limitations of deep learning. The core theme of this issue: Current deep learning faces fundamental limitations in data efficiency, common sense reasoning, and causal understanding that cannot be resolved simply by scaling up computational power. A hybrid system that integrates symbolic AI with deep learning must be built. Marcus's core judgment: Deep learning excels at perception and classification, but it cannot grasp common sense such as "a bottle holds water" or "a person drinks water because they are thirsty"—and this is the "rate-limiting step" toward achieving general intelligence.
Marcus argues that the so-called "technological singularity" will not be a sudden shift at a single point in time, but rather a gradual, multidimensional evolution.
Inference: Marcus believes that physical common sense (e.g., a key opens a lock) may be easier for machines to master than psychological common sense (e.g., understanding others' motives), because robots can acquire physical data through experimentation, whereas psychological experiments are constrained by ethical limitations.
Marcus systematically critiques 10 challenges of current deep learning architectures, with the core issue being that systems learn correlations rather than causality and abstract concepts.
Data Support: Marcus mentions that DeepMind posted a loss of $530 million in 2018, and despite massive investments in applying similar technologies to language, "productivity is far lower."
Marcus explicitly opposes the extreme paths of "pure deep learning" or "pure symbolic AI," advocating for hybrid systems that combine the strengths of both.
Implication: Marcus argues that symbolic AI may need its own "ImageNet moment"—a breakthrough demonstration that captures the world's imagination. The Allen Institute for AI (AI2) is building a commonsense reasoning benchmark dataset, which could serve as a catalyst.
Marcus argues that the way human children learn—an innate knowledge framework combined with experiential learning—provides key design principles for AI.
Falsification Condition: If a future system can robustly predict "which containers will leak and which will not" simply by watching videos, Marcus says, "I would be delighted to see it."
Marcus argues that the fundamental reason current AI is untrustworthy is that it only processes correlations without understanding abstract concepts such as "harm" or "intent."
Extrapolation: Marcus argues that if AI could pass a "comprehension test"—for example, after watching Spartacus, answering "Why did everyone stand up and say 'I am Spartacus'?"—he would acknowledge that AI has made genuine progress.
| Position | Analyst View | Key Data |
|---|---|---|
| DeepMind | Risk Warning | 2018 loss of $530 million; poor results when applying AlphaGo technology to language tasks |
| GPT-2 | Risk Warning | Generates fluent text but lacks consistent concept understanding; fails on reasoning tasks like "DAX is DAX" |
| Google Duplex | Risk Warning | Only handles barber shop and restaurant reservations, and business hours inquiries; not general-purpose |
| CYC | Neutral (Historical Case) | 30 years of manually coded common sense, unsuccessful; Marcus argues this should not discredit the symbolic AI approach |
| AlphaGo/AlphaZero | Neutral (Hybrid System Case) | Deep learning + Monte Carlo tree search; not pure deep learning |
1. Marcus: Intelligence is a multidimensional variable, not a single IQ score. Machines far surpass humans in mathematics and games, but lag behind a five-year-old child in language comprehension—progress across different dimensions is asynchronous, and the singularity will not be a single moment.
2. Marcus: Deep learning excels at perceptual classification but cannot grasp common sense like "a bottle holds water." This is the "rate-limiting step" toward general intelligence, because reading comprehension, physical reasoning, and social interaction all depend on such background knowledge.
3. Marcus: Neural networks learn correlations, not algebraic operations on variables. A 1998 experiment showed that a network could learn all even numbers but failed to generalize to odd numbers—21 years later, GPT-2 similarly failed on "DAX is DAX" type reasoning, with the underlying architecture unchanged.
4. Marcus: AlphaGo is a hybrid system, not pure deep learning. It incorporates Monte Carlo tree search (a classical AI technique), and in the future, people will "relabel hybrid systems as deep learning."
5. Marcus: Humans are "a low bar that is easy to surpass." In agreement with Kahneman: machines should learn the flexible reasoning at which humans excel, while avoiding human flaws such as motivated reasoning, confirmation bias, and poor memory.
6. Marcus: The prerequisite for trustworthy AI is "deep understanding," not "deep learning." You cannot align values with a system that only processes correlations—understanding "harm" requires abstract concepts, which deep learning handles poorly.
7. Marcus: Evolution is cumulative—once a "good idea" emerges, the gene pool spreads it rapidly. A newborn antelope can descend a mountain within hours of birth, indicating that organisms carry a vast amount of innate knowledge; AI should draw on cognitive science rather than starting from scratch.
8. Marcus: The six-question test for evaluating AI reports—the core is "ask to see a demo." If Sundar Pichai says a system can converse like a human, you should ask, "Can I try it? How general is it?"