earticle

논문검색

WordNet-based Hybrid VSM for Document Classification

초록

영어

Many text classifications depend on statistical term measures or synsets to implement document representation. Such document representations ignore the lexical semantic contents or relations of terms, leading to losing the distilled mutual information. This work proposed a synthetic document representation method, WordNet-based hybrid VSM, to solve the problem. This method constructed a data structure of semantic-element information to characterize lexical semantic contents, and support disambiguation of word stems. As a template, lexical semantic vector consisting of lexical semantic contents was built in the lexical semantic space of corpus, and lexical semantic relations are marked on the vector. Then, it connects with special term vector to form the eigenvector in hybrid VSM. Applying algorithm NWKNN, on text corpus Reuter-21578 and its adjusted version, the experiments show that the eigenvector performs F1 measure better than document representations based on TF-IDF.

목차

Abstract
 1. Introduction
 2. Related Work
 3. Proposed Program
  3.1 The Motivation and Theoretical Analysis
  3.2 Hybrid VSM of Text Corpus
  3.3 Algorithm NWKNN
 4. Experiment and Result
  4.1 Experiment Setup
  4.2 The Results
 5. Conclusion
 References

저자정보

  • Luda Wang Xiangnan University, Chenzhou, China, School of Information Science and Engineering, Central South University, Changsha, China
  • Peng Zhang Xiangnan University, Chenzhou, China
  • Shouping Gao Xiangnan University, Chenzhou, China

참고문헌

자료제공 : 네이버학술정보

    함께 이용한 논문

      ※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

      0개의 논문이 장바구니에 담겼습니다.