Classifying Unsolicited Bulk Email (UBE) using Python Machine Learning Techniques

Sabah Mohammed; Osama Mohammed; Jinan Fiaidhi; Simon Fong; Tai hoon Kim

Classifying Unsolicited Bulk Email (UBE) using Python Machine Learning Techniques

원문정보

Sabah Mohammed, Osama Mohammed, Jinan Fiaidhi, Simon Fong, Tai hoon Kim

보안공학연구지원센터(IJHIT) International Journal of Hybrid Information Technology Vol.6 No.1 2013.01 pp.43-56

피인용수 : 0건 (자료제공 : 네이버학술정보)

초록

영어

Email has become one of the fastest and most economical forms of communication. However, the increase of email users has resulted in the dramatic increase of spam emails during the past few years. As spammers always try to find a way to evade existing filters, new filters need to be developed to catch spam. Generally, the main tool for email filtering is based on text classification. A classifier then is a system that classifies incoming messages as spam or legitimate (ham) using classification methods. The most important methods of classification utilize machine learning techniques. There are a plethora of options when it comes to deciding how to add a machine learning component to a python email classification. This article describes an approach for spam filtering using Python where the interesting spam or ham words (spam-ham lexicon) are filtered first from the training dataset and then this lexicon is used to generate the training and testing tables that are used by variety of data mining algorithms. Our experimentation using one dataset reveals the affectivity of the Naïve Bayes and the SVM classifiers for spam filtering.

키워드

저자정보

Sabah Mohammed Lakehead University
Osama Mohammed Lakehead University
Jinan Fiaidhi Lakehead University
Simon Fong University of Macau
Tai hoon Kim Konkuk University

참고문헌

자료제공 : 네이버학술정보

함께 이용한 논문

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

0개의 논문이 장바구니에 담겼습니다.

earticle