TY - GEN
T1 - Data Cleaning About Student Information Based on Massive Open Online Course System
AU - Yin, Shengjun
AU - Yi, Yaling
AU - Wang, Hongzhi
N1 - Publisher Copyright:
© 2020, Springer Nature Singapore Pte Ltd.
PY - 2020
Y1 - 2020
N2 - Recently, Massive Open Online Courses (MOOCs) is a major way of online learning for millions of people around the world, which generates a large amount of data in the meantime. However, due to errors produced from collecting, system, and so on, these data have various inconsistencies and missing values. In order to support accurate analysis, this paper studies the data cleaning technology for online open curriculum system, including missing value-time filling for time series, and rule-based input error correction. The data cleaning algorithm designed in this paper is divided into six parts: pre-processing, missing data processing, format and content error processing, logical error processing, irrelevant data processing and correlation analysis. This paper designs and implements missing-value-filling algorithm based on time series in the missing data processing part. According to the large number of descriptive variables existing in the format and content error processing module, it proposed one-based and separability-based criteria Hot+J3+PCA. The online course data cleaning algorithm was analyzed in detail on algorithm design, implementation and testing. After a lot of rigorous testing, the function of each module performs normally, and the cleaning performance of the algorithm is of expectation.
AB - Recently, Massive Open Online Courses (MOOCs) is a major way of online learning for millions of people around the world, which generates a large amount of data in the meantime. However, due to errors produced from collecting, system, and so on, these data have various inconsistencies and missing values. In order to support accurate analysis, this paper studies the data cleaning technology for online open curriculum system, including missing value-time filling for time series, and rule-based input error correction. The data cleaning algorithm designed in this paper is divided into six parts: pre-processing, missing data processing, format and content error processing, logical error processing, irrelevant data processing and correlation analysis. This paper designs and implements missing-value-filling algorithm based on time series in the missing data processing part. According to the large number of descriptive variables existing in the format and content error processing module, it proposed one-based and separability-based criteria Hot+J3+PCA. The online course data cleaning algorithm was analyzed in detail on algorithm design, implementation and testing. After a lot of rigorous testing, the function of each module performs normally, and the cleaning performance of the algorithm is of expectation.
KW - Data cleaning
KW - Dimension reduction
KW - Intermittent missing
KW - MOOC
KW - Time series
UR - https://www.scopus.com/pages/publications/85090031754
U2 - 10.1007/978-981-15-7981-3_3
DO - 10.1007/978-981-15-7981-3_3
M3 - 会议稿件
AN - SCOPUS:85090031754
SN - 9789811579806
T3 - Communications in Computer and Information Science
SP - 33
EP - 43
BT - Data Science - 6th International Conference of Pioneering Computer Scientists, Engineers and Educators, ICPCSEE 2020, Proceedings
A2 - Zeng, Jianchao
A2 - Jing, Weipeng
A2 - Song, Xianhua
A2 - Lu, Zeguang
PB - Springer
T2 - 6th International Conference of Pioneering Computer Scientists, Engineers and Educators, ICPCSEE 2020
Y2 - 18 September 2020 through 21 September 2020
ER -