earticle

논문검색

Global Approach for Script Identification using Wavelet Packet Based Features

초록

영어

In a multi script environment, an archive of documents having the text regions printed in different scripts is in practice. For automatic processing of such documents through Optical Character Recognition (OCR), it is necessary to identify different script regions of the document. In this paper, a novel texture-based approach is presented to identify the script type of the collection of documents printed in seven scripts, to categorize them for further processing. The South Indian documents printed in the seven scripts - Kannada, Tamil, Telugu, Malayalam, Urdu, Hindi and English are considered here The document images are decomposed through the Wavelet Packet Decomposition using the Haar basis function up to level two. Gray level co-occurrence matrix is constructed for the coefficient sub bands of the wavelet transform. The Haralick texture features are extracted from the co-occurrence matrix and then used in the identification of the script of a machine printed document. Experimentation conducted involved 2100 text images for learning and 1400 text images for testing. Script classification performance is analyzed using the K-nearest neighbor classifier. The average success rate is found to be 99.68%.

목차

Abstract
 1. Introduction
 2. Wavelet Packet Transform (WPT)
 3. Data Collection
 4. Preprocessing
 5. The proposed Model
  5.1 Feature Extraction
  5.2 Classification
 6. Experimental Results
 7. Conclusion
 References

저자정보

  • M.C. Padma PES College of Engineering,
  • P. A. Vijaya Malnad College of Engineering

참고문헌

자료제공 : 네이버학술정보

    함께 이용한 논문

      ※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

      0개의 논문이 장바구니에 담겼습니다.