Conference proceeding
TOFA: Trace Oriented Feature Analysis in Text Categorization
2008 Eighth IEEE International Conference on Data Mining, pp.668-677
12/2008
DOI: 10.1109/ICDM.2008.67
Abstract
Dimension reduction for large-scale text data is attracting much attention lately due to the rapid growth of World Wide Web. We can consider dimension reduction algorithms in two categories: feature extraction and feature selection. An important problem remains: it has been difficult to integrate these two algorithm categories into a single framework, making it difficult to reap the benefit of both. In this paper, we formulate the two algorithm categories through a unified optimization framework. Under this framework, we develop a novel feature selection algorithm called Trace Oriented Feature Analysis (TOFA). The novel objective function of TOFA is a unified framework that integrates many prominent feature extraction algorithms such as unsupervised Principal Component Analysis and supervised Maximum Margin Criterion are special cases of it. Thus TOFA can process not only supervised problem but also unsupervised and semi-supervised problems. Experimental results on real text datasets demonstrate the effectiveness and efficiency of TOFA.
Details
- Title: Subtitle
- TOFA: Trace Oriented Feature Analysis in Text Categorization
- Creators
- J Yan - Microsoft Res. Asia, Sigma Center, BeijingNing Liu - Microsoft Res. Asia, Sigma Center, BeijingQiang Yang - Dept. of Comput. Sci., Hong Kong Univ. of Sci. & Technol., KowloonWeiguo Fan - Dept. of Comput. Sci., Virginia Polytech. Inst. & State Univ., Blacksburg, VAZheng Chen - Microsoft Res. Asia, Sigma Center, Beijing
- Resource Type
- Conference proceeding
- Publication Details
- 2008 Eighth IEEE International Conference on Data Mining, pp.668-677
- Publisher
- IEEE
- DOI
- 10.1109/ICDM.2008.67
- ISSN
- 1550-4786
- eISSN
- 2374-8486
- Language
- English
- Date published
- 12/2008
- Academic Unit
- Business Analytics
- Record Identifier
- 9984083854502771
Metrics
14 Record Views