July 2, 2026
Innovator Coffee EP-38 World Models: The Missing Layer Between AI and the Physical World
<p>Welcome to the Innovator Coffee, a podcast that bridges the gap between people and the world of AI and innovation. Follow us to explore the top AI products, ecosystem insights, and the emerging trends.</p><p><br></p><p><strong>Guest Bios</strong></p><p><strong>Dr. ChongKai Gao</strong></p><p>Dr. ChongKai Gao is a Visiting Ph.D. Student at Stanford University in Dr. Feifei Li’s lab, where he conducts research on robotic manipulation, visual planning, and world models. His work focuses on enabling robots to reason about future physical states, perform long-horizon task planning, and bridge AI perception with real-world decision making.</p><p><br></p><p><strong>Dr. JunFan Zhu</strong></p><p>Dr. JunFan Zhu is the organizer of the San Francisco Robotics & World Model Reading Club, bringing together researchers from leading AI labs to discuss frontier advances in robotics, embodied AI, and world models. His interests span world model architectures, evaluation, tactile intelligence, VLA systems, and the future roadmap of physical AI.</p><p><br></p><p><strong>Episode Description</strong></p><p><strong>Large Language Models (LLMs)</strong> have transformed how AI understands language. But what happens when AI must understand and interact with the physical world?</p><p><br></p><p>In this episode of <strong>Innovator Coffee</strong>, Stanford researcher <strong>Dr. ChongKai Gao</strong> and robotics researcher <strong>Dr. JunFan Zhu</strong> explain <strong>World Models</strong>, one of the fastest-growing areas in <strong>AI, Robotics, and Embodied AI</strong>. We explore how world models differ from LLMs, their relationship with <strong>Vision-Language-Action (VLA)</strong> models and <strong>Spatial Intelligence</strong>, and why companies like <strong>NVIDIA, Meta, and Google DeepMind</strong> are investing heavily in this field.</p><p><br></p><p>We also discuss the biggest challenges in <strong>Physical AI</strong>, from data and simulation to evaluation, and what it will take for robotics to reach its own "ChatGPT moment."</p><p><br></p><p><strong>Timestamps</strong></p><p><strong>00:00 – 09:20 | What Is a World Model?</strong></p><ul><li>World Models vs. Large Language Models</li><li>Why predicting the future matters more than predicting the next token</li></ul><p><strong>09:21 – 21:40 | Mapping the World Model Landscape</strong></p><ul><li>Five major technical routes</li><li>JEPA, Dreamer, VLA, diffusion models, and hybrid architectures</li></ul><p><strong>21:40 – 29:45 | Spatial Intelligence vs. Decision Making</strong></p><ul><li>Fei-Fei Li's Spatial Intelligence</li><li>Why robotics may need action before perfect 3D reconstruction</li></ul><p><strong>29:45 – 38:30 | Visual Planning for Robot Manipulation</strong></p><ul><li>Why robots need to "imagine" before acting</li><li>Visual planning versus language planning</li></ul><p><strong>38:30 – 46:00 | Research Bottlenecks and Missing Pieces</strong></p><ul><li>Why deployment is much harder than demos</li><li>Tactile sensing, evaluation, calibration, and the "missing layer" between intelligence and capability</li></ul><p><strong>46:00 – 55:00 | Where Will World Models Create Real Business Value?</strong></p><ul><li>Games, autonomous driving, warehouse automation, and robotics</li><li>Which applications may commercialize first?</li></ul><p><strong>55:00 – 01:10:30 | Robotics' ChatGPT Moment</strong></p><ul><li>Why robotics doesn't have a scaling law yet</li><li>Data flywheels, simulation, evaluation, and the future of embodied AI</li></ul><p><strong>01:10:30 – End | Betting on the Future of Physical Intelligence</strong></p><ul><li>Why leading researchers are investing in world models</li><li>The biggest open questions and what founders, researchers, and investors should watch next</li></ul><p><br></p><p>Tom Kong</p><p>*Stanford EE alumni,</p><p>*Founder@ Stanford AGI Adventist Community (10K+ members so far from top VC, Engineers, startups from Silicon Valley )</p><p>*AI Lecturer, a serial entrepreneur in media and data. Advisor @ techtimes.com and heyboss.ai</p><p>*AI deployment for 8 years, with NLP and recent LLMs (RAG, Agent, Diffusion)</p><p><br></p><p>Wickey Wang</p><p>*IT Security Compliance Leader & University Faculty</p><p>*VC advisor and Angel Investor with cybersecurity and AI focus</p><p>*GAI Security book co-author</p><p>Co-founderQuestions, Suggestions, Feedback and Comments? You can find us in LinkedIn:</p><p>https://www.linkedin.com/in/wickey-wang-cisa-six-sigma-green-belt-2aaa913/</p><p>https://www.linkedin.com/in/thomaskong/</p><p><br></p>