Maxwell smart: new Chinese embodied model tops global AI ranking on physical tasks

A new embodied AI model known as Maxwell – developed by the Chinese Academy of Sciences’ Institute of Artificial Intelligence for Industries – has shot to first place on the Meta-World benchmark for robot physical tasks.
Maxwell scored 91.9 – the highest ever recorded on the simulation benchmark.
Meta-World was set up by researchers from Stanford University, UC Berkeley and other institutions to evaluate how well robots perform 50 everyday physical tasks. They range from grasping and carrying to opening doors and using drawers.
00:54
Robots waltz on stage with humans in China’s WorldSkills competition
The list draws submissions from leading teams worldwide, including Physical Intelligence, Google DeepMind, Alibaba, Meituan, Carnegie Mellon University, MIT, the University of Cambridge and the Toyota Research Institute.
Of the latest submissions, from July, second-placed FabriVLA – developed by Shenzhen-based Youibot – scored 90. SUREFlow, developed by researchers at Kyungpook National University in South Korea, was third with a score of 88.3.
This means the Chinese large models are surpassing their American rivals on learning to act – beyond thinking and talking, they are entering the real physical world.
Robots can already perform precise operations on assembly lines. But even the most precise work is often limited to fixed settings and preset commands. For robots to truly enter daily life, complete tasks on their own and adapt to new behaviours, the key is for them to understand the physical world by themselves.
To complete a task in Meta-World, a model must understand the spatial relationship between objects and targets, carry out a sequence of actions such as approaching, grasping and moving, and adjust its actions based on contact.

Beyond the score, according to a release on the institute’s website, Maxwell was able to perform more than 200 tasks without additional fine-tuning. Those included moving objects, adjusting positions, and operating switches and containers.
In demonstrations, Maxwell could move a nut onto a target post, use a handle to open cabinet doors and drawers, and grasp objects and carry them across containers.
The team also tested the model in LIBERO – a simulation environment that puts more emphasis on continuous operations.
There, it showed the ability to connect multiple steps in different scenarios: turning on a stove and putting a kettle on it; placing a cup in a microwave; putting a book into a storage box; and putting chocolate into a drawer. Its success rate was 99.1 per cent.
Maxwell is also designed for edge deployment – where the data is processed close to the user or data source, rather than through centralised cloud deployment.
It has 1 billion parameters and supports single-card inference, meaning it also has relatively low hardware requirements.
The achievement of the CAS institute is not isolated – many Chinese companies have recently made progress in this field.
They include Beijing-based DeepCybo, which built physical foundation model PhysBrain. Earlier this month, the robot was assessed against 28 benchmarks and achieved an overall score of 72.5. Against other open-source entries, PhysBrain topped 14 of the evaluations and came second in 10 of them.
It was just edged out by the best-performing closed-source model – OpenAI’s latest GPT-6 Astra – which had a score of 73.3.

Chinese company Unitree Robotics also open-sourced a foundation model on September 10. It has 6 billion parameters and can drive robots to complete more than 60 whole-body mobile and tabletop tasks.
They include organising shoe racks, cleaning a kitchen, arranging tableware, doing laundry, making beds and taking food out of a microwave – with the idea that robots could be trained as butlers.
Once focused on making joints and motors, Unitree Robotics was previously seen as a hardware company. But public information shows that after going public, the company plans to invest nearly half of its IPO proceeds in research and development for intelligent robot models.
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.