← Back to list
Lex Fridman PodcastPodcast7 Mar 2024Source: lexfridman.comHost: Lex Fridman

#416 – Yann Lecun: Meta AI, Open Source, Limits of LLMs, AGI & the Future of AI

In plain words

This is from a podcast with Meta's chief AI scientist Yann LeCun. He argues that current large language models (like ChatGPT) can't achieve human-level intelligence because they just predict the next word without understanding the real world. He supports Meta's open-source LLaMA models, saying open-source is the only way to prevent AI from being controlled by a few companies. He criticizes Google's Gemini for generating a 'black Nazi soldier' due to over-correction of bias. He envisions future AI as 'objective-driven', planning before acting rather than generating directly.

AI SummaryAI-generated · may contain errors · verify against the original

At a Glance Meta's Chief AI Scientist and Turing Award winner Yann LeCun discussed open-source AI, the limitations of LLMs, and the future of AGI on the Lex Fridman podcast. Key views: Meta firmly supports the development of open-source AI, having already open-sourced LLAMA 2 and soon releasing LLAM

~10 min full read · 8 sections
Deep Analysis

As a third-party independent analyst, I have analyzed the transcript of this podcast. The following is the interpretation report.

At a Glance

The guest in this episode is Yann LeCun, Chief AI Scientist at Meta and Turing Award laureate. He systematically elaborated on the limitations of the current AI path represented by LLMs and proposed a blueprint for future AI development centered on non-generative, Joint Embedding Predictive Architecture (JEPA). LeCun's core judgment is that current mainstream autoregressive large language models (LLMs) cannot achieve human-level intelligence because they lack an understanding of the physical world, persistent memory, reasoning, and planning capabilities; a true breakthrough requires abandoning generative models and shifting to "objective-driven AI" that predicts and plans in an abstract representation space.

Topic Sections

1. Limitations of LLMs: Data Scarcity and "System 1" Reactions

Yann LeCun argues that large models trained solely on language cannot gain a deep understanding of the world, with the root cause being the insufficient "bandwidth" of training data.

  • Data Volume Comparison: LeCun points out that the amount of data a 4-year-old child receives through visual senses is approximately 10^15 bytes, while all public text used to train an LLM (about 2×10^13 bytes) is far less. This proves that most human knowledge comes from observation and interaction with the physical world, not language.
  • "System 1" Operation: LeCun likens the operation of LLMs to human "System 1" (subconscious, automatic reactions). It merely predicts the next word based on preceding context. "It doesn't think about the answer; it retrieves the answer" (meaning: it does not plan the response but retrieves and outputs from accumulated knowledge). This "token-by-token" generation method results in extremely primitive reasoning capabilities because the computation per token is constant, unlike humans who can invest more thought into complex problems.
  • Root Cause of Hallucinations: LeCun believes hallucinations are an inherent flaw of LLMs. Since each prediction has a non-zero error probability, errors accumulate exponentially as the generated sequence lengthens ("probability drift"). While fine-tuning can cover common questions, the system is highly prone to producing nonsensical outputs when faced with "long-tail" prompts not seen in the training set.
2. The Path Forward: Embracing JEPA and "Objective-Driven AI"

LeCun proposes that the key to Advanced Machine Intelligence (AMI) lies in abandoning generative models and shifting to Joint Embedding Predictive Architecture (JEPA) and objective-driven AI.

  • Core Idea of JEPA: Unlike LLMs that attempt to "generate" the complete input at the pixel or token level, JEPA only learns to predict in an abstract representation space. By comparing a complete input with its corrupted/transformed version, it trains a predictor to forecast the representation of the former (complete) from the representation of the latter (corrupted). This forces the system to extract only "predictable," high-level abstract information while ignoring unpredictable details like rustling leaves, thereby building an abstract understanding of the world.
  • From Generation to Planning: LeCun argues that true intelligence requires "System 2" (deliberate planning). The blueprint he proposes is objective-driven AI: the system has an internal "energy function" to evaluate whether an answer/action is compatible with the current question/state. During reasoning, instead of directly generating an answer, the system searches in the abstract representation space through an optimization process (e.g., gradient descent) to find an "optimal idea" that minimizes this energy function, which is then decoded into language or action. This process is akin to "thinking before speaking," and the computation can be adjusted based on problem complexity.
  • Abandoning Reinforcement Learning (RL): LeCun advocates minimizing the use of RL due to its extremely poor sample efficiency. A better path is: first learn a world model through observation (e.g., video), then use this model for Model Predictive Control (MPC) to plan actions. RL is only used to fine-tune the world model or objective function when planning results do not match expectations.
3. Open Source is the Only Answer to AI Bias and Power Concentration

LeCun firmly believes that facing the unavoidable bias and risk of power concentration in AI systems, open source is the only solution.

  • Bias is Unavoidable: LeCun points out that any AI system will have bias because "bias is in the eye of the beholder." Trying to fine-tune a single system to please everyone is impossible and may lead to factual errors, like Gemini generating "Black Nazi soldiers."
  • Danger of Power Concentration: He warns that in the future, every interaction with the digital world will be mediated by AI assistants. If these systems are controlled by only a few West Coast companies, it poses a significant threat to democracy, cultural, and linguistic diversity. "We cannot have the digital food of all citizens controlled by three West Coast companies."
  • Open Source Enables Diversity: Open-source foundation models (e.g., LLaMA) allow any organization (governments, enterprises, communities) to fine-tune based on their own data, culture, and values, thereby generating diverse AI systems. For example, the French government would not accept its citizens' information being controlled by a US company; India needs AI that can speak 22 official languages; Africa needs AI that can provide medical information in local languages. Only an open-source platform can support such a diverse AI ecosystem.
4. Rebuttal of AI "Doomsday" Theories: Gradual Development, Not a Catastrophic Event

LeCun strongly opposes the view that AI could spiral out of control and destroy humanity, arguing it is based on a series of false assumptions.

  • Not an "Event": AGI will not be suddenly invented one day, like in science fiction movies. It is a gradual process; we will first have "cat or parrot-level" intelligence and learn how to set safety guardrails for it along the way.
  • "Good AI" vs. "Bad AI": Even if a "rogue AI" emerges, there will be more, more powerful "good AIs" to counter it. In the future, everyone will have their own AI assistant, which will automatically block malicious information from hostile AIs, much like spam filters.
  • "Dominance Drive" is Not Inevitable: LeCun believes AI systems will not inherently have a drive for dominance. This desire is "hardcoded" in specific social species (like humans, chimpanzees) and is not a necessary byproduct of intelligence. Humans have a strong incentive to create "obedient" AI.
  • Analogy to the Turbojet Engine: AI safety is not something that can be solved with a single mathematical proof. It is more like the design of a turbojet engine, which became so reliable through decades of gradual iteration, improvement, and testing. "You need to build better AI systems; they will be safer because they are designed to be more useful and controllable."

Position Moves

Position Guest's Stance Key Data
Meta (LLaMA 2/3) Bullish (Core of open-source strategy) LLaMA 2 has millions of downloads; LLaMA 3 is coming soon, will be larger, better, and multimodal.
Google (Gemini 1.5) Risk Warning (As a negative example) Factual errors like generating "Black Nazi soldiers" due to excessive "de-biasing."
OpenAI (GPT-4) Neutral (As a typical LLM representative) Used as the primary example for discussing LLM limitations.
DeepMind Neutral (As a peer) Mentioned for related work on BYOL (non-contrastive learning) and world models.
Tesla (Optimus) Neutral (Industry observation) Believed to have "re-energized the entire industry" around humanoid robots.
Boston Dynamics Neutral (Industry observation) Their robotics technology relies heavily on hand-crafted dynamic models, not general intelligence.
Figure AI, Unitree Neutral (Industry observation) Mentioned as emerging humanoid robotics companies.

Judgments Worth Remembering

1. LLMs are "System 1," not a path to AGI (Yann LeCun). Support: They lack understanding of the physical world, persistent memory, reasoning, and planning capabilities; their operation is akin to human subconscious reactions, not deliberate thought.

2. The "bandwidth" of language data is far lower than that of sensory data (Yann LeCun). Support: The amount of data a 4-year-old receives through vision (10^15 bytes) is 50 times the total text used to train an LLM (2×10^13 bytes), proving most knowledge comes from observing the physical world.

3. The core of JEPA is "predicting in an abstract space," not "generating pixels" (Yann LeCun). Support: By predicting the representation of a corrupted input rather than the complete input, the system is forced to learn high-level, predictable abstract information, thereby building a world model. This is key to solving the "Moravec's paradox."

4. The blueprint for future AI is "objective-driven AI," which plans through an optimization process (Yann LeCun). Support: During reasoning, the system searches for the abstract representation of the best answer by optimizing an "energy function," rather than directly generating tokens. This allows computation to be adjusted based on problem complexity and naturally supports adding "safety guardrails" as optimization objectives.

5. AI bias cannot be eliminated; open source is the only solution (Yann LeCun). Support: Bias is in the eye of the beholder; no single system can please everyone. Only by open-sourcing foundation models and allowing groups with different cultures and values to fine-tune them can a diverse AI ecosystem be created, avoiding power concentration.

6. AI "doomsday" theories are wrong because they assume AGI is an "event" (Yann LeCun). Support: AGI will be a gradual development; we will first have "cat-level" intelligence and learn how to control it along the way. The future will be a struggle between "good AI" and "bad AI," not between humans and a single super AI.

7. AI safety should be analogous to the reliability design of a turbojet engine (Yann LeCun). Support: Safety is not achieved through a single mathematical proof but through decades of gradual design, testing, and iteration. Better AI systems are inherently safer and more controllable systems.

8. The ultimate value of AI is to "make humans smarter," like the printing press (Yann LeCun). Support: AI will serve as an intelligent assistant for everyone, amplifying human intelligence. This is similar to how the printing press popularized knowledge, ultimately leading to the Enlightenment. Although the process involves growing pains, its long-term impact is positive.