Skip to main navigation Skip to search Skip to main content

Efficient High Utility Itemset Mining on Massive Data

  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In learning analytics and computer support for intelligent tutoring, high utility itemset mining (HUIM) is an interesting operation to find the sets of items which have high utilities. HUIM normally is considered to be more difficult than the traditional frequent itemset mining (FIM), since the utility of the itemset does not hold anti-monotone property. This paper analyzes the inefficiency of existing algorithms for HUIM on massive data. To address this, we introduce a novel algorithm named P2H, which utilizes a prefix-partitioning approach. P2H organizes the transaction table into memory-resident partitions. Each partition groups transactions that share a common initial item, enabling efficient discovery of high-utility itemsets. P2H handles the prefix-based partitions sequentially and independently. By the pre-computed data structure, P2H can skip most of the partitions and report the required high utility itemsets. For the partitions to be processed further, a set enumeration tree is proposed to compute the results quickly. Besides, a novel pruning strategy utilizing full suffix utilities is introduced to effectively reduce the exploration space. Extensive evaluation on both synthetic and real-world datasets demonstrates that P2H substantially outperforms state-of-the-art methods.

Original languageEnglish
Title of host publicationLearning Technologies and Systems - 24th International Conference on Web-based Learning, lCWL 2025 and 10th International Symposium on Emerging Technologies for Education, SETE 2025, Revised Selected Papers
EditorsBaltasar Fernández-Manjón, António José Mendes, Marco Temperini, Zuzana Kubincová, Marc Spaniol, Guandong Xu, Elvira Popescu, Tianyong Hao, Xiangmeng Wang, Shuning He
PublisherSpringer Science and Business Media Deutschland GmbH
Pages327-342
Number of pages16
ISBN (Print)9789819200412
DOIs
StatePublished - 2026
Event24th International Conference on Web-based Learning, ICWL 2025 and 10th International Symposium on Emerging Technologies for Education, SETE 2025 - Hongkong, China
Duration: 30 Nov 20253 Dec 2025

Publication series

NameLecture Notes in Computer Science
Volume16425 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference24th International Conference on Web-based Learning, ICWL 2025 and 10th International Symposium on Emerging Technologies for Education, SETE 2025
Country/TerritoryChina
CityHongkong
Period30/11/253/12/25

Keywords

  • Computer Support for Intelligent Tutoring
  • HUIM
  • Learning Analytics
  • Massive data
  • prefix-based partitioning
  • subtree pruning

Fingerprint

Dive into the research topics of 'Efficient High Utility Itemset Mining on Massive Data'. Together they form a unique fingerprint.

Cite this