Clustering Large Scale Data Set Based on Distributed Local Affinity Propagation on Spark

Wei Lu; Peng Cao

Clustering Large Scale Data Set Based on Distributed Local Affinity Propagation on Spark

원문정보

Wei Lu, Peng Cao

보안공학연구지원센터(IJDTA) International Journal of Database Theory and Application Vol.9 No.10 2016.10 pp.241-250 SCOPUS

피인용수 : 0건 (자료제공 : 네이버학술정보)

초록

영어

Affinity Propagation (AP) is a new clustering method to cluster data set efficiently. In this paper, a Distributed method of Local Affinity Propagation (DLAP) is proposed to solve hardware bottleneck and time-consuming problem. DLAP refines AP by reducing the calculating data scale in each iteration and keeps a high quality clustering result. The method is implemented on Apache Spark distributed computation framework. Depending on high iteration efficiency on Spark, the method has an impressive result in time complexity. Experiments are conducted on two-dimensional data to show that the time cost of LAP on single machine is better than the two methods, FSAP and FAP, meanwhile the result of DLAP on Spark is better than that on Hadoop.

키워드

저자정보

Wei Lu School of Software Engineering, Beijing Jiaotong University, Beijing, China
Peng Cao School of Software Engineering, Beijing Jiaotong University, Beijing, China

참고문헌

자료제공 : 네이버학술정보

함께 이용한 논문

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

0개의 논문이 장바구니에 담겼습니다.

earticle