Abstract
In this paper, we consider the problem of in-depth document analysis. In particular, we propose a novel document analysis method, named multidimensional latent semantic analysis (MDLSA), which enables us to mine local information efficiently from a document with respect to term associations and spatial distributions. MDLSA works by first partitioning each document into paragraphs and building a term affinity graph, which represents the frequency of term cooccurrence in a paragraph. We then conduct a 2-D principal component analysis to achieve an optimal semantic mapping. This analysis involves finding the leading eigenvectors of the sample covariance matrix of a training set to characterize the lower dimensional semantic space. A hybrid document similarity measure is designed to further improve the performance of this framework. Our algorithm is examined in two document applications: retrieval and classification. Experimental results demonstrate that the proposed technique outperforms current algorithms with respect to accuracy and computational efficiency.
| Original language | English |
|---|---|
| Article number | 6670128 |
| Pages (from-to) | 1625-1640 |
| Number of pages | 16 |
| Journal | IEEE Transactions on Cybernetics |
| Volume | 43 |
| Issue number | 6 |
| DOIs | |
| State | Published - Dec 2013 |
| Externally published | Yes |
Keywords
- Dimensionality reduction
- Multidimensional
- Principle component analysis (PCA)
- Semantic analysis
- Term association
Fingerprint
Dive into the research topics of 'Multidimensional latent semantic analysis using term spatial information'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver