Alibaba has unveiled its first ‘is’ models for robots

Chinese internet giant Alibaba has unveiled its first series of artificial intelligence models for robots.

According to the company’s statement, the Qwen-Robot Suite comprises three main models.

The Qwen-RobotManip model has been trained on 38,000 hours of video and is built on the ‘vision-language-action’ (VLA) principle, which enables the robot to perform tasks based on visual information and text instructions.

Qwen-RobotNav has been trained on 15.6 million data samples in the areas of trajectory planning and vision-language reasoning, and acts as a navigator.

Finally, Qwen-RobotWorld is a so-called video world model that predicts trajectories based on current observations. It has been trained on 8.6 million video-text pairs and can synthesise video data to train robots and help them simulate a trajectory before performing an action.

According to the company, the versatility of the Qwen-Robot Suite, which is already being tested by selected Alibaba clients, enables robots to dynamically perceive information, reason and act in real time. Thanks to these models, industrial manipulators, delivery robots and robot dogs can function seamlessly in unfamiliar environments and interact with new objects, whilst strictly adhering to the laws of physics and carrying out commands in simple human language.

Traditional robots based on large multimodal models often become disoriented in unfamiliar environments and struggle to process new instructions, as they are unable to dynamically translate verbal commands into physical actions.

- Реклама -