Alibaba on Tuesday launched the Qwen Robot Suite, its first comprehensive series of large models for embodied intelligence, developed by its AI research unit Tongyi Lab. The suite divides robot intelligence into three interconnected layers: Qwen-RobotNav, a vision-language navigation model that unifies instruction following, goal navigation, object tracking, and autonomous driving into a single framework; Qwen-RobotWorld, a video world model that lets machines predict how physical scenes will evolve before acting, spanning manipulation, driving, and navigation contexts; and Qwen-RobotManip, a generalist vision-language-action (VLA) model built on the Qwen3.5-4B architecture and trained on a corpus of more than 38,100 hours assembled entirely from open-source data. All three provide language-first interfaces and can be composed via standard Qwen model calls. Alibaba said the suite has entered pilot testing with selected Alibaba Cloud enterprise clients.
Alongside the three models, Alibaba disclosed Qwen-RobotClaw, an internal agent framework that enables Qwen vision-language models to invoke the Robot Suite components as tools for physical-world execution, while managing the context and memory required for sessions of up to 20 minutes — allowing sustained long-horizon planning beyond frame-by-frame visual reaction. The launch extends the Qwen model family, which already spans text, vision, code, audio, and video, into the physical world, and positions it as a potential general substrate for robotic applications. It also marks Alibaba’s formal entry into the “embodied AI” race alongside rivals including Google DeepMind, Figure AI, and ByteDance, all competing to move AI models from digital interfaces into machines that can perceive and act in real environments.