TY - GEN
T1 - Efficient High Utility Itemset Mining on Massive Data
AU - Han, Xixian
AU - Wang, Kangao
AU - Wan, Xiaolong
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - In learning analytics and computer support for intelligent tutoring, high utility itemset mining (HUIM) is an interesting operation to find the sets of items which have high utilities. HUIM normally is considered to be more difficult than the traditional frequent itemset mining (FIM), since the utility of the itemset does not hold anti-monotone property. This paper analyzes the inefficiency of existing algorithms for HUIM on massive data. To address this, we introduce a novel algorithm named P2H, which utilizes a prefix-partitioning approach. P2H organizes the transaction table into memory-resident partitions. Each partition groups transactions that share a common initial item, enabling efficient discovery of high-utility itemsets. P2H handles the prefix-based partitions sequentially and independently. By the pre-computed data structure, P2H can skip most of the partitions and report the required high utility itemsets. For the partitions to be processed further, a set enumeration tree is proposed to compute the results quickly. Besides, a novel pruning strategy utilizing full suffix utilities is introduced to effectively reduce the exploration space. Extensive evaluation on both synthetic and real-world datasets demonstrates that P2H substantially outperforms state-of-the-art methods.
AB - In learning analytics and computer support for intelligent tutoring, high utility itemset mining (HUIM) is an interesting operation to find the sets of items which have high utilities. HUIM normally is considered to be more difficult than the traditional frequent itemset mining (FIM), since the utility of the itemset does not hold anti-monotone property. This paper analyzes the inefficiency of existing algorithms for HUIM on massive data. To address this, we introduce a novel algorithm named P2H, which utilizes a prefix-partitioning approach. P2H organizes the transaction table into memory-resident partitions. Each partition groups transactions that share a common initial item, enabling efficient discovery of high-utility itemsets. P2H handles the prefix-based partitions sequentially and independently. By the pre-computed data structure, P2H can skip most of the partitions and report the required high utility itemsets. For the partitions to be processed further, a set enumeration tree is proposed to compute the results quickly. Besides, a novel pruning strategy utilizing full suffix utilities is introduced to effectively reduce the exploration space. Extensive evaluation on both synthetic and real-world datasets demonstrates that P2H substantially outperforms state-of-the-art methods.
KW - Computer Support for Intelligent Tutoring
KW - HUIM
KW - Learning Analytics
KW - Massive data
KW - prefix-based partitioning
KW - subtree pruning
UR - https://www.scopus.com/pages/publications/105041619206
U2 - 10.1007/978-981-92-0042-9_24
DO - 10.1007/978-981-92-0042-9_24
M3 - 会议稿件
AN - SCOPUS:105041619206
SN - 9789819200412
T3 - Lecture Notes in Computer Science
SP - 327
EP - 342
BT - Learning Technologies and Systems - 24th International Conference on Web-based Learning, lCWL 2025 and 10th International Symposium on Emerging Technologies for Education, SETE 2025, Revised Selected Papers
A2 - Fernández-Manjón, Baltasar
A2 - Mendes, António José
A2 - Temperini, Marco
A2 - Kubincová, Zuzana
A2 - Spaniol, Marc
A2 - Xu, Guandong
A2 - Popescu, Elvira
A2 - Hao, Tianyong
A2 - Wang, Xiangmeng
A2 - He, Shuning
PB - Springer Science and Business Media Deutschland GmbH
T2 - 24th International Conference on Web-based Learning, ICWL 2025 and 10th International Symposium on Emerging Technologies for Education, SETE 2025
Y2 - 30 November 2025 through 3 December 2025
ER -