Why Do We Need World Models When LLMs Exist? How World Models Will Impact Physical AI
Yifeng Yin, Co-founder and CEO of TEA Intelligent System.
gettyDuring my tenure as a machine learning engineer and AI researcher at Hugging Face and now as the CEO of an AI startup, I have watched and participated in the evolution of large language models (LLMs) from being cute to being kind of scary.
Now, almost everyone knows how powerful LLMs are. But if LLMs seem so powerful, why are influential computer scientists like Yann LeCun and Fei-Fei Li betting on world models?
All technology is valued for the problems it can potentially solve. The main difference between LLMs and world models is that, fundamentally, they are designed to solve two different sets of problems.
To understand this, let’s get back to first principles.
LLMs input and output tokens. That is, if a problem can be described and solved by combinations of tokens, it can potentially be solved by LLMs. If an LLM exhausted its theoretical potential, it would output the optimal token combination for any input. Since many things can be represented by a combination of tokens, LLMs can potentially solve a lot of problems.
The case for world models is more interesting and less straightforward. As explained in LeCun’s paper and this blog post by Fei-Fei Li, world models are functions that take in an observation of the “environment” and an action to be taken. They then output, in the form of embeddings, a prediction of how the state of the “world” would be if we take that action. The embedding should contain all information of interest about the new state of the environment.
In short, world models do not generate sentences but a prediction of the state of the environment after a certain action possibly taken by some agents.
For example, given a house (environment) and a robot (an agent), a world model should output an embedding containing all information of interest about the state of the house if a certain action is taken by the robot (say, if it decides to light the bed on fire). Of course, the environment can be anything from a house to a lab setting to the real world or the world of Sanctuary from the Diablo video game series.
As with LLMs, if we exhaust all theoretical potentials of a world model, it will output a very accurate prediction of what will happen to the world given an observation of the world and the action to be taken.
• LLMs deliver token combinations and are best used to solve problems that can be optimally solved by a string of tokens.
• World models deliver prediction and causality and are best used to solve problems that can be optimally solved by an accurate prediction of the future and the consequence of a choice/action. The technology is already used for spatial simulation, autonomous driving and several other disciplines.
Both have their uses. There is also a small overlap between the sets of problems they are trying to solve because the state of an environment can sometimes be described to reasonable accuracy with words. So, I think the two will co-exist for some time.
Throughout human history, the ability to predict the future and to perform accurate causal inference are seen as some of the defining indicators of intelligence and should be an integral part of artificial general intelligence.
For example, physical intelligence is on the hype train right now, with many envisioning a future where robots will be actually useful for production and services. However, one of the major obstacles to physical AI is the lack of training data and, generally, how hard it is to obtain it.
Yes, we can apply the old “learning by doing” + “trial and error” principle and just release the robot into the world and let it run wild.
The issue here is that actions taken in the real world have real-world consequences. A robot might seriously hurt someone or even cause catastrophic failures in critical infrastructure, leading to real-world disasters. But if we restrict the robot’s actions for safety, the robot would not be able to learn efficiently. The classic “exploration versus exploitation” problem.
Furthermore, based on my experience, machines need a lot more data than humans to master the same skill. Gathering real-world data takes time, and time moves slowly in the real world. An hour in the real world is just one hour of data. A robot can be in the wild for 10 years and still remain severely under-trained and practically useless in any meaningful settings.
We can, of course, release millions of robots out there to learn and gather data, but that will make the world a very dangerous place.
Needless to say, this problem cannot be efficiently addressed by combinations of tokens, but it can be solved elegantly with the ability to accurately predict the future.
If a world model can accurately predict the consequences of actions within a “world,” like in self-driving, that means we have access to unlimited training data. With fast inference, we can potentially let the robot operate for billions of years in a near-perfect simulation of our world within just weeks in real life, so all robots fresh out of the factory are trained as if they have lived in this particular era of our world for billions of years.
Since no one and nothing will get hurt during the data-generating process, the “exploration versus exploitation” dilemma is solved elegantly. The robot is free to explore anything within the “world” during those billions of years while causing no real-world damage except maybe a huge compute bill.
If I’m right about this, the predicted boom of physical intelligence could come after world models are sufficiently mature.
Of course, as with all new technologies, there are still challenges facing world models, like the sim-to-real gap, which is when performance drops when AI trained in a virtual environment is deployed to physical hardware. There is also still a lack of action-labeled data, which explains the physics behind why actions happen. Finally, the problem of partial observability, where models receive incomplete information about their environment, will need to be addressed.
However, I still see world models as the most promising path toward physical intelligence.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

