World models emerge as distinct paradigm from LLMs, learning from sensory grounding not text
World models learn directly from video, audio, and sensory interaction like humans do, enabling common sense and physical reasoning that text-only LLMs fundamentally lack, representing a ne…