Physical AI: World Models and Embodied Intelligence
Keywords:
Physical AI, World Models, Embodied Intelligence, Vision-Language-Action Models, Robot Foundation Models, Sim-to-Real, Reinforcement Learning, Generative SimulationAbstract
Physical AI names a shift in artificial intelligence away from purely digital tasks and toward systems that perceive, reason about, and act in the physical world. Between 2024 and 2026 this shift accelerated as robot foundation models, learned simulators, and embodied agents moved from isolated demonstrations to shared datasets, open policies, and reusable world models. This review organizes the area around two technical pillars. The first is world models: neural networks that learn predictive models of environment dynamics and support planning or behavior learning through imagined rollouts, from the early recurrent world model of Ha and Schmidhuber through the Dreamer family and joint-embedding predictive architectures. The second is generalist action policies, especially vision-language-action models such as RT-1, RT-2, Octo, and OpenVLA trained on cross-embodiment corpora like Open X-Embodiment. We connect these pillars to the simulation and data infrastructure that feeds them, including GPU-accelerated physics, domain randomization, and video generation framed as world simulation. We summarize benchmarks and evaluation practice, survey humanoid and manipulation applications, and discuss open challenges: data scarcity, the sim-to-real gap, safety, generalization, long-horizon control, and reproducible evaluation. All quantitative figures in the tables and charts are representative values collated from disparate reports and are intended to illustrate trends rather than to support controlled comparison. We close with directions toward neuro-symbolic integration and foundation world models.



