![]() |
| Infographic. / Courtesy of Board of Audit and Inspection |
The Board of Audit and Inspection (BAI) revealed through an audit that public artificial intelligence (AI) training datasets created with a 1.6 trillion KRW government investment to support the AI industry have been poorly managed, including building data similar to other agencies and producing unusable information.
The BAI disclosed the results of its audit on the status of AI industry promotion on August 12. As of November last year, the Ministry of Science and ICT (MSIT) built 908 types of training data and made them available on AI Hub. Additionally, 26 organizations including the Ministry of Food and Drug Safety, the Seoul Metropolitan Government, and Korea Expressway Corporation are providing 313 types of data directly usable for AI companies on their respective portals.
According to the BAI, agencies were building duplicate data without checking plans from other organizations. Notably, despite MSIT spending 7.38 billion KRW to build training data for four image categories (wild animals, domestic waste, concrete cracks, and eggs), five organizations including the Seoul Metropolitan Government and Korea Expressway Corporation spent an additional 610 million KRW to build similar data.
In the case of wild animal training data meant for preventing roadkill, 25,000 images out of the data built by Korea Expressway Corporation in 2022 (12 types, 60,000 images)—representing 42%—were confirmed to be similar to data already built by MSIT in 2021.
Quality verification of data was also neglected. Among the top 20 organizations in terms of data building volume, 9 permitted deliveries without undergoing external quality verification during project execution.
The BAI stated that the road driving CCTV training dataset built by Korea Expressway Corporation in 2025 lacked vehicle location details in its annotation information, which is essential information added during data processing so machines can recognize learning targets, rendering it unusable for training autonomous driving technologies.
For large-scale waste training data built by the Seoul Metropolitan Government in 2020, 538 out of 2,303 images (23.4%) had file names mismatched with actual images, such as labeling images of boxes or tables as bags, posing a risk of errors during AI training.
The BAI notified MSIT to introduce a pre-review system to assess the potential alternative use of existing data, share and adjust building plans across agencies in advance, and establish quality management standards.
Choi Min-jun
1
2
3
4
5
6
7