Text Representation Based on Key Terms of Document for Text Categorization

Jieming Yang; Zhiying Liu; Zhaoyang Qu

Text Representation Based on Key Terms of Document for Text Categorization

원문정보

Jieming Yang, Zhiying Liu, Zhaoyang Qu

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.4 2016.04 pp.1-22 SCOPUS

피인용수 : 0건 (자료제공 : 네이버학술정보)

초록

영어

The text representation, “bag of words” or vector space model, is widely used by most of the classifiers in text categorization. All the documents fed into the classifier are represented as a vector in the vector space, which consists of all the terms extracted from training set. Due to the characteristics of high dimensionality, feature selection algorithm is usually used to reduce the dimensionality of the vector space. Through feature selection, each document is represented by some representative terms extracted from the training set. Although the classification results based on this document representation methodare better, it is inevitable that some documents may contain few even none representative terms, and these documents must be misclassified. In this paper, we proposed a new text representation method, KT-of-DOC, which represents one document using some key terms extracted from this document. We selected key terms of each document based on six feature selection algorithms, Improved Gini Index (GINI), Information Gain (IG), Mutual Information (MI), Odds Ratio (OR), Ambiguity Measure (AM) and DIA association factor (DIA), respectively, and evaluated the performance of two classifiers, Support Vector Machines (SVM) and K-Nearest Neighbors (KNN), on three benchmark collections, 20-Newsgroups, Reuters-21578 and WebKB. The results show that the proposed representation method can significantly improve the performance of classifier.

Abstract
1. Introduction
2. Related Work
3. Algorithm Description
  3.1. Problem and Motivation
  3.2. Algorithm Implement
  3.3 Complexity Analysis
4. Experiment Setup
  4.1 Feature-Selection Algorithms
  4.2. Data Sets
  4.3 Classifiers
  4.4. Performance Measures
5. Results
  5.1 Results of Algorithm for SVM
  5.2 Results of Algorithm for KNN
6. Discussions
7. Conclusion
Acknowledgment
References

키워드

저자정보

Jieming Yang College of Information Engineering, Northeast Dianli University, Jilin, Jilin, China
Zhiying Liu College of Information Engineering, Northeast Dianli University, Jilin, Jilin, China
Zhaoyang Qu College of Information Engineering, Northeast Dianli University, Jilin, Jilin, China

참고문헌

자료제공 : 네이버학술정보

함께 이용한 논문

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

0개의 논문이 장바구니에 담겼습니다.

earticle