원문정보
초록
영어
As an open source cloud storage scheme, HDFS is used by more and more large enterprises and researchers, and is actually applied to many cloud computing systems to deal with huge amounts of data. HDFS has many advantages, but there are some problems such as NameNode single point of failure, small file problem, hot issues, etc. For HDFS hot issues, this paper proposes a dynamic Replication mechanism of HDFS hot file based on cloud storage(HDFS-DRM). The mechanism includes a Replication of the dynamic adjustment mechanism and adding, deleting duplicate node selection mechanism in two parts, by increasing the NameNode, BlockMap parameters, it records the number of reading requests of each file in a certain period of time to decide whether to increase or decrease the number of copies. The mechanism presents a replica placement method based on stage historical information and node load and selects the appropriate node to add or delete copies of documents to improve the utilization efficiency of the data node storage space effectively. Experimental results show that, HDFS - DRM in hot files case, compared to native HDFS file system access latency is significantly reduced, HDFS-DRM can solve the hot issues successfully.
목차
1. Introduction
2. Framework of Dynamic Replication mechanism system based on HDFS
2.1. Increasing access_log and Timer1 in NameNode
2.2. Adding a Separate Dynamic Replica Controller, its Main Functions
3. The Improvement of the BlockMap
4. The Node's Selection of Adding and Deleting Replication
4.1. The Selection of Node that Increase the Replication
4.2. The Selection of Node that Delete Replication
5. Experiment and Analysis
6. Conclusion
Acknowledgement
References