Skip to main navigation Skip to search Skip to main content

Relevant subspace search for outlier detection on massive data

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Outlier detection is an important task in the field of data mining, which has wide applications in real life. Since subspace outlier detection attracts widespread attention for its ability to effectively identify outliers in high-dimensional data, a series of subspace search algorithms for outlier detection are proposed to identify subspaces that reveal outliers. However, it is found that the existing algorithms often lose their effectiveness when processing massive data due to large memory consumption. Additionally, they struggle to achieve high-quality results because of the lack of strong pruning. In this paper, a novel algorithm RSLSH (Relevant Subspace Search based on Locality Sensitive Hashing) is presented to mine relevant subspaces for outlier detection on massive data. RSLSH applies a pre-computed disk-resident attribute list structure and LSH slice set to divide the huge database into small partitions. An adaptive relevance attribute acquisition method using LSH slices is employed to significantly reduce the search space for candidate subspaces. Within the depth-first subspace generation approach, a memory-resident attribute list structure is designed to reduce the I/O cost when calculating subspace correlations. Moreover, an update strategy for the memory-resident structure based on MFU is devised to enhance data reusability. The extensive experimental results show that RSLSH can discover relevant subspaces for outlier detection on massive data efficiently.

Original languageEnglish
Article number128870
JournalExpert Systems with Applications
Volume296
DOIs
StatePublished - 15 Jan 2026
Externally publishedYes

Keywords

  • 0000
  • 1111
  • Algorithm
  • Locality-sensitive hashing
  • Massive data
  • Outlier detection
  • Subspace search

Fingerprint

Dive into the research topics of 'Relevant subspace search for outlier detection on massive data'. Together they form a unique fingerprint.

Cite this