Why Large Language Models Aren't a Path to AGI
Researchers argue that true artificial general intelligence requires physical understanding, not just multimodal language processing.

Takeaways
- ›LLMs may not build true world models, limiting their path to AGI
- ›Physical understanding is crucial for general intelligence
- ›Next-token prediction alone is insufficient for real-world problem solving
- ›Embodied AI approaches may be more promising for AGI development
The recent breakthroughs in large language models (LLMs) have sparked excitement about the potential for artificial general intelligence (AGI). But some researchers are pushing back, arguing that these models fundamentally lack the physical understanding needed for true AGI.
Language models don't truly understand the world
The core argument is that while LLMs appear remarkably intelligent in their language abilities, they lack genuine comprehension of the physical world. As the article states:
"Do LLMs really learn implicit models of the world? How could they otherwise be so proficient at language?"
The author contends that LLMs are not building accurate world models, but rather learning "bags of heuristics to predict tokens." This results in superficial understanding that can pass language tests but fails at tasks requiring physical reasoning.
Why physical understanding matters for AGI
A key point is that many real-world problems cannot be solved through language and symbol manipulation alone. The article gives examples like:
- Repairing a car
- Untying a knot
- Preparing food
These tasks require a "form of intelligence that is fundamentally situated in something like a physical world model." Without this grounding, an AI system cannot be considered truly general.
The limits of next-token prediction
The core training objective of LLMs, predicting the next token in a sequence, is argued to be fundamentally limited. While it produces impressive language abilities, it does not necessitate building an accurate model of reality.
The article cites research showing that generative models can excel at sequence prediction tasks without learning true world models, instead relying on "comprehensive sets of idiosyncratic heuristics."
Syntax vs. semantics vs. pragmatics
The author argues that LLMs are primarily learning syntax (the structure and rules of language) rather than semantics (the meaning) or pragmatics (contextual interpretation). This allows them to manipulate language convincingly without true understanding.
A key quote summarizes this view:
"LLMs are simply not running physics simulations in their latent next-token calculus when they ask you if your person, place, or thing is bigger than a breadbox."
The case for embodied AI approaches
Rather than trying to achieve AGI by combining multiple language-centric models, the article advocates for approaches that treat "embodiment and interaction with the environment as primary."
This aligns with longstanding arguments in AI and cognitive science about the importance of grounding intelligence in physical experience and sensorimotor interaction with the world.
Implications for AGI development
If this argument holds, it suggests that the current focus on scaling up language models may not be the most promising path to AGI. Instead, more integrated approaches combining language, vision, robotics, and physical simulation may be needed.
The debate highlights the ongoing uncertainty about what constitutes true intelligence and how to best pursue artificial general intelligence. While language models have made remarkable progress, fundamental questions remain about the nature of understanding and the requirements for AGI.
Related reads
Nemotron 3 Nano Omni 30B Explained: Handles Text, Images, Audio, Video
5 min read
Mental World Modeling Explained: Predicting Human Behavior
4 min read
Vision-Language Models Explained: Causal Mechanisms of Perception-Knowledge Conflict
5 min read
Small Language Models Explained: Efficiency Over Size, Benchmarks
6 min read
AI Agents Explained: Capabilities, Risks, Adoption Trends
5 min read
Data Pyramid for Embodied Manipulation Explained: Framework, Tradeoffs
3 min read
Reported and explained by AI·Reporter.