Robert: Multimodality and multilinguality work; abstraction, reasoning, and world models still fail
Current LLMs excel at fusing text/image/audio and handling mixed languages, but lack true logical reasoning, contextual awareness, and persistent world models — autonomous agents remain scaffolded by human-designed workflows.