Skip to main navigation Skip to search Skip to main content

Improved blog clustering through automated weighting of text block

  • Harbin Institute of Technology Shenzhen
  • The University of Hong Kong

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In this paper, a new clustering algorithm is proposed for blog data clustering. Considering the structure information of text blocks in blog data, we group the features of blog data into three groups and extend the Ic-means clustering algorithm to automatically calculate a weight for each feature group in the clustering process. We introduce a new objective function with group weight variables and present the Lagrangian method to derive the formula to calculate the group weights. This formula is added as a new step in the standard kmeans iterative clustering process to automatically compute the group weights according to the distribution of features, This new process guarantees the convergency of the clustering process to a local optimal solution. The experimental results have shown that this new algorithm performed better than k-means without group feature weighting on different blog data sets.

Original languageEnglish
Title of host publicationProceedings of the 2009 International Conference on Machine Learning and Cybernetics
PublisherIEEE Computer Society
Pages1586-1591
Number of pages6
ISBN (Print)9781424437030
DOIs
StatePublished - 2009
Externally publishedYes
Event8th International Conference on Machine Learning and Cybernetics, ICMLC 2009 - Baoding, China
Duration: 12 Jul 200915 Jul 2009

Publication series

NameProceedings of the 2009 International Conference on Machine Learning and Cybernetics
Volume3

Conference

Conference8th International Conference on Machine Learning and Cybernetics, ICMLC 2009
Country/TerritoryChina
CityBaoding
Period12/07/0915/07/09

Keywords

  • Auto-weighted
  • Blog
  • Clustering
  • Web mining

Fingerprint

Dive into the research topics of 'Improved blog clustering through automated weighting of text block'. Together they form a unique fingerprint.

Cite this