Abstract
Nowadays, the fast advance of internet technology has brought two challenges. The first one is explosion of information. The second one is new information appears rapidly. Obviously, clustering is a good solution to help users analyze information automatically, whereas traditional clustering algorithms are only suitable for small-scale and stable text collection. In order to solve this problem, a novel clustering algorithm based on vector compression particularly for large-scale text collection (LDVC) and its incremental version (I-LDVC) are proposed in this paper. LDVC selects related features to compress feature sets. Iterative training idea of self- organizing-mapping (SOM) is also imported in it to optimize selection approach. Besides, when novel texts appear, its incremental version (I-LDVC) can select small samples from original texts to alter neuron model to perform incremental clustering. In order to prevent it from over fitting to new added texts, I-LDVC adjusts the weights of samples along with training process. Experimental results demonstrate that LDVC has better performance and lower time complexity on large-scale text collection, and I-LDVC can cluster unstable text collection very well.
| Original language | English |
|---|---|
| Pages (from-to) | 136-147 |
| Number of pages | 12 |
| Journal | Information Technology and Control |
| Volume | 45 |
| Issue number | 2 |
| DOIs | |
| State | Published - 2016 |
| Externally published | Yes |
Keywords
- Incremental clustering
- Neuron model
- Self-organizing-mapping
- Vector compression
Fingerprint
Dive into the research topics of 'A novel clustering algorithm and its incremental version for large-scale text collection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver