Compression Techniques Applied to DNA Data of Various Species

Vilas Machhi; Maulika S Patel

Compression Techniques Applied to DNA Data of Various Species

원문정보

Vilas Machhi, Maulika S Patel

보안공학연구지원센터(IJBSBT) International Journal of Bio-Science and Bio-Technology Vol.8 No.3 2016.06 pp.45-52 SCOPUS

피인용수 : 0건 (자료제공 : 네이버학술정보)

초록

영어

DNA sequences comprise of sequentially linked nucleotides, A, C, G and T. As a result of the genome projects, a significant amount of DNA sequences of various species are deposited in various databases. Human DNA contains about 3 billion base pairs. The number of genes within the DNA is 20,000 to 25,000. For storing DNA data of a single person, we require approximately 10 CD – ROMs. This amounts to huge data storage costs, subsequently making the use of these data such as analysis and retrieval quite challenging. DNA sequence analysis is useful in diverse areas such as forensics, medical research, pharmacy, agriculture etc. It is very necessary to address the storage issue of these exponentially growing data. In this paper we have implemented 4 different algorithms for DNA data compression: LZW (Lampel-ziv-Welch) algorithm, run length encoding algorithm, Arithmetic coding and Substitution method. The compression results on these algorithms are presented and compared on DNA sequence data of 10 different species.

키워드

저자정보

Vilas Machhi Alumnus, G H Patel College of Engineering & Technology Vallabh Vidyangar, Gujarat, India
Maulika S Patel G H Patel College of Engineering & Technology Vallabh Vidyangar, Gujarat, India

참고문헌

자료제공 : 네이버학술정보

함께 이용한 논문

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

0개의 논문이 장바구니에 담겼습니다.

earticle