The Shift Beyond LLMs: How AI Researchers Are Chasing the Next Frontier of World Models
After eight years studying large language models (LLMs)—the artificial intelligence technology powering popular chatbots like ChatGPT and Claude—computer scientist Louis Castricato began to feel he had hit an unbreakable dead end for fundamental progress. “We have essentially moved past the era of groundbreaking foundational LLM research,” Castricato explained. “Today, the space is almost entirely focused on building new applications, not advancing core knowledge.”
That frustration led him to walk away from his doctoral program at Brown University and launch a new startup called Overworld. True to its name, the company chases a far more ambitious goal than text-based AI: building artificial intelligence that can actually understand and navigate the physical world, not just process language.
Right now, AI chatbots remain an enormous profit opportunity, with investors pouring trillions of dollars into top developers like OpenAI and Anthropic on the bet that the technology will keep delivering outsized returns. But a growing cohort of AI innovators are turning their attention to what they see as the field’s next frontier: “world models,” systems designed to teach AI—and eventually robots—how to interact and respond to real physical environments.
The movement counts some of the industry’s most prominent leaders among its backers, including Fei-Fei Li, often called the “Godmother of AI.” Li describes world models as “one of the most important and most overloaded terms in AI today.”
At the core of world model research is a simple, provocative argument: AI can never reach true general intelligence if it only understands text pulled from books and the internet. To be genuinely smart, it also needs to understand the physical context around it. “Where language models learn the statistical structure of text, world models learn the statistical structure of space and time: how light falls on a surface, how a garden looks from an angle no camera has ever captured, how objects respond to force and follow the laws of physics,” Li, founder of San Francisco-based startup World Labs, wrote in a recent essay.
Another leading advocate for the approach is AI pioneer Yann LeCun, who stepped down as Meta’s chief AI scientist last year to launch the Paris-based Advanced Machine Intelligence Labs. “World model is quickly becoming a buzzword,” LeCun noted on a recent episode of the Unsupervised Learning podcast. For him, a world model is fundamentally a tool that lets an AI agent “predict the consequences of its own actions.”
There is no single universal definition of a world model, and framing often shifts based on what type of technology a researcher is building, from autonomous robots to more immersive interactive video games. Training LLMs on the totality of humanity’s written text, news, and visual media has already produced AI assistants that are reshaping knowledge work and many creative fields. But proponents of world models argue that current generative AI models, which work by iteratively predicting the next word or pixel to generate new content, have fundamental limitations.
Chatbots cannot physically pick up a coffee mug, points out Martial Hebert, dean of computer science at Carnegie Mellon University. “There’s all the geometry of the world, the dynamic of how I move my hand, the physical interaction of the contact with the cup,” Hebert explained. “This is much more complex than just predicting the next word in a sentence.”
For Hebert, who has spent more than 40 years researching robotics, world models offer a faster, cheaper path to building “physical AI”—another fast-growing buzzword in the tech sector. “Some people may have different definitions, but physical and embodied AI are kind of the evolution of what we used to call robotics,” Hebert said in an interview. The same AI advances that made chatbots so capable can also be used to build AI with a broad enough understanding of its environment to act as a functional robot brain, he argues.
He draws a parallel to how the human body works: “In your body and spinal cord you have a very general model of how to balance, how to walk around, and you can adapt to your knee hurting in the morning, so you now walk a little differently. You don’t need to think about that. You have a general model somewhere in your nervous system and brain that allows your body to adapt very quickly.”
Smarter autonomous robots are not the only end goal for world model research. Castricato launched Overworld last year, and the small Rhode Island-based startup is building interactive video game worlds where dynamic environments—for example, a spooky forest—can shift and respond as a player’s virtual character moves through the space and interacts with objects around them. “There’s no other world model where you can just walk through doors or where you can interact with a detailed environment like this,” Castricato said in an interview. “We optimize for interaction above anything else.”
While near-term commercial use cases are not as immediately obvious as AI coding assistants or chatbot tools, world model developers are already drawing significant interest from venture capitalists. Steve Jang, co-founder and managing partner at Kindred Ventures, says his firm has invested in Overworld and multiple other world model-focused startups, including Causal Labs, which builds AI models for weather forecasting, and Extropic, which designs specialized computer chips optimized for running world models. “I think that the future is many different types of models with many different philosophies and architectures,” Jang said. “I don’t think that it’ll be one large, dense model to rule them all.”
In her recent essay, Li set out to create a “taxonomy of world models” to cut through confusion around the many competing visions for the technology. “A video model that produces gorgeous but physically impossible flames, a language model improvising a playable game, and a physics engine that faithfully simulates combustion all go by the same name,” she wrote. She split world models into three distinct categories:
Renderers: The most commercially viable today, these prioritize high visual fidelity for virtual worlds but are not accurate enough to train functional robots.
Simulators: These build virtual training environments that accurately replicate the physical structure and rules of the real world.
Planners: These systems focus on predicting what an AI agent or robot should do next when navigating an unstructured real-world environment.
“A robot that can plan is a robot that can work, and the entire industry is racing to be the one that gets there first,” Li wrote.
The Shift Beyond LLMs: How AI Researchers Are Chasing the Next Frontier of World Models