WordNet-based Hybrid VSM for Document Classification

Luda Wang; Peng Zhang; Shouping Gao

WordNet-based Hybrid VSM for Document Classification

원문정보

Luda Wang, Peng Zhang, Shouping Gao

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.1 2016.01 pp.185-200 SCOPUS

피인용수 : 0건 (자료제공 : 네이버학술정보)

초록

영어

Many text classifications depend on statistical term measures or synsets to implement document representation. Such document representations ignore the lexical semantic contents or relations of terms, leading to losing the distilled mutual information. This work proposed a synthetic document representation method, WordNet-based hybrid VSM, to solve the problem. This method constructed a data structure of semantic-element information to characterize lexical semantic contents, and support disambiguation of word stems. As a template, lexical semantic vector consisting of lexical semantic contents was built in the lexical semantic space of corpus, and lexical semantic relations are marked on the vector. Then, it connects with special term vector to form the eigenvector in hybrid VSM. Applying algorithm NWKNN, on text corpus Reuter-21578 and its adjusted version, the experiments show that the eigenvector performs F1 measure better than document representations based on TF-IDF.

키워드

저자정보

Luda Wang Xiangnan University, Chenzhou, China, School of Information Science and Engineering, Central South University, Changsha, China
Peng Zhang Xiangnan University, Chenzhou, China
Shouping Gao Xiangnan University, Chenzhou, China

참고문헌

자료제공 : 네이버학술정보

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

0개의 논문이 장바구니에 담겼습니다.

earticle