China has launched a national framework for sharing AI training datasets specifically designed for embodied intelligence -- the field of AI that controls physical robots, autonomous vehicles, and other systems that interact directly with the real world. The initiative, announced at the World Intelligence Expo in Tianjin on May 30, aims to accelerate the country's push into physical AI by solving one of the field's most persistent bottlenecks: the scarcity of high-quality training data for real-world interaction.

The framework establishes standardized protocols for collecting, annotating, and sharing datasets across categories including robotic manipulation, locomotion, object recognition in unstructured environments, and human-robot interaction. Participating institutions include Tsinghua University, the Chinese Academy of Sciences, Baidu, and several robotics startups.

"Embodied AI requires orders of magnitude more diverse data than language models," said Professor Zhang Wei, director of the China Institute of AI and Robotics, who helped design the framework. "No single organization can collect enough manipulation data, locomotion data, and scene understanding data to build truly capable systems. National coordination is the only path to scale."

Why Data Sharing Matters for Physical AI

Training a language model requires text, which exists abundantly on the internet. Training a robot to pick up an arbitrary object from an arbitrary surface requires physical interaction data -- recordings of successful and failed grasps, force measurements, visual observations from multiple angles, and contextual information about the environment. This data is expensive to collect, difficult to standardize, and often proprietary.

The Chinese framework addresses this by creating shared data repositories that participating organizations can both contribute to and draw from. Contributors earn credits based on the volume and quality of their data, which they can spend to access datasets from other participants. A central governance body reviews data quality and enforces annotation standards.

The approach mirrors successful data-sharing initiatives in other domains, including genomics and climate science, where collaborative data infrastructure has accelerated research beyond what any individual institution could achieve.

The 150-Robot Showcase

The embodied intelligence pavilion at the World Intelligence Expo provided a tangible demonstration of why China is investing in this infrastructure. More than 80 enterprises showcased nearly 150 robots, from industrial manipulators capable of precision assembly to consumer robots designed for elderly care and childcare.

Several exhibitors demonstrated robots trained using shared datasets from early pilot programs of the national framework. A manipulation system from RoboSense, for example, showed the ability to handle dozens of different household objects despite being trained primarily in a warehouse environment -- a level of generalization that its developers attributed to the diversity of training data available through the sharing program.

"Before the framework, we were training our robots on data we collected ourselves in our own lab," said RoboSense CTO Liu Chen. "Now we have access to manipulation data from 15 different facilities with different objects, lighting conditions, and table surfaces. The improvement in generalization is dramatic."

Geopolitical Dimensions

The dataset cooperation framework reflects China's broader strategy of leveraging coordinated national effort to compete in AI domains where it sees potential advantages. While the United States leads in large language models, China has identified embodied intelligence as a field where its manufacturing base, robotics industry, and government coordination capabilities could provide an edge.

The initiative also raises questions about data governance and international competitiveness. If Chinese robotics companies have access to significantly larger and more diverse training datasets than their Western counterparts, the resulting capability gap could be difficult to close.

What to Watch

The first major test of the framework will come in the second half of 2026, when participating institutions are expected to publish benchmark results comparing models trained on shared datasets versus proprietary data. If the results show meaningful improvements in robotic generalization, expect other countries to explore similar cooperative models for physical AI training data.

"Embodied AI requires orders of magnitude more diverse data than language models. National coordination is the only path to scale."
— Zhang Wei, Director, China Institute of AI and Robotics
150
Robots showcased
80+
Participating companies
15
Data-sharing facilities