Skip to main navigation Skip to search Skip to main content

Biclustering: An application of dual topic models

  • Daniel Rugeles
  • , Kaiqi Zhao
  • , Cong Gao
  • , Manoranjan Dash
  • , Shonali Krishnaswamy
  • Nanyang Technological University
  • Agency for Science, Technology and Research, Singapore

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Biclustering is a data mining technique that allows simultaneous clustering of two variables. A common biclustering task for categorical variables is to find 'heavy' biclusters, i.e., biclusters with high co-occurrence values. Although algorithms have been proposed to extract heavy biclusters, they provide little information about relative importance of each bicluster, as well as importance of the variables for each bicluster. To address these problems, there have been attempts to apply mixture models using information theory or Bayesian method. Although they are able to rank the biclusters and the variables for each bicluster, they do not target at extracting heavy biclusters. Furthermore, these models constrain the search for biclusters in such a way that every cell in the matrix must participate in some bicluster. We attempt to alleviate these limitations using dual topic models. First of all, we develop a generalized LDA topic model that extracts dual topics, i.e., topics in opposite directions - row- and column-topics. To obtain better topics, it applies mutual reinforcement, i.e., considering column-topics while constructing row-topics, and vice versa. Heavy biclusters, the high co-occurred relationship, are extracted using thresholds. We show that our model Dual Topic to Biclusters (DT2B) is effective in extracting heavy biclusters by experimenting over a simulated data, a text corpus (NIPS author-document) and a microarray gene expression data. Results show that biclusters extracted by DT2B are better.

Original languageEnglish
Title of host publicationProceedings of the 17th SIAM International Conference on Data Mining, SDM 2017
EditorsNitesh Chawla, Wei Wang
PublisherSociety for Industrial and Applied Mathematics Publications
Pages453-461
Number of pages9
ISBN (Electronic)9781611974874
DOIs
StatePublished - 2017
Externally publishedYes
Event17th SIAM International Conference on Data Mining, SDM 2017 - Houston, United States
Duration: 27 Apr 201729 Apr 2017

Publication series

NameProceedings of the 17th SIAM International Conference on Data Mining, SDM 2017

Conference

Conference17th SIAM International Conference on Data Mining, SDM 2017
Country/TerritoryUnited States
CityHouston
Period27/04/1729/04/17

Fingerprint

Dive into the research topics of 'Biclustering: An application of dual topic models'. Together they form a unique fingerprint.

Cite this