China Unveils Massive Plan to Build AI Training Datasets Amid Data Shortage Concerns
China is facing a critical data shortage that could impact its AI ambitions, prompting the National Data Administration to unveil a massive plan to build AI training datasets. The draft plan aims to expand the supply, circulation, and commercialization of high-quality training datasets across nine core industrial sectors by 2028.
The plan covers emerging fields like autonomous driving and healthcare, as well as finance and transportation. It's not just about text data; the initiative explicitly calls for multimodal datasets covering text, code, images, audio, and video.
China is investing approximately $295 billion into AI-focused data centers over the next five years, with a target deadline of 2028 to have a comprehensive ecosystem of multimodal datasets ready for commercial use. The plan also emphasizes data governance, management standards, and rights frameworks around training data.