Object detection in financial reporting documents for subsequent recognition

Petr Sokerin; Alla Volkova; Kirill Kushnarev

Object detection in financial reporting documents for subsequent recognition

원문정보

Petr Sokerin, Alla Volkova, Kirill Kushnarev

국제인공지능학회(구 한국인터넷방송통신학회) The International Journal of Advanced Smart Convergence Volume 10 Number 1 2021.03 pp.1-11 KCI 등재

피인용수 : 0건 (자료제공 : 네이버학술정보)

초록

영어

Document page segmentation is an important step in building a quality optical character recognition module. The study examined already existing work on the topic of page segmentation and focused on the development of a segmentation model that has greater functional significance for application in an organization, as well as broad capabilities for managing the quality of the model. The main problems of document segmentation were highlighted, which include a complex background of intersecting objects. As classes for detection, not only classic text, table and figure were selected, but also additional types, such as signature, logo and table without borders (or with partially missing borders). This made it possible to pose a non-trivial task of detecting nonstandard document elements. The authors compared existing neural network architectures for object detection based on published research data. The most suitable architecture was RetinaNet. To ensure the possibility of quality control of the model, a method based on neural network modeling using the RetinaNet architecture is proposed. During the study, several models were built, the quality of which was assessed on the test sample using the Mean average Precision metric. The best result among the constructed algorithms was shown by a model that includes four neural networks: the focus of the first neural network on detecting tables and tables without borders, the second - seals and signatures, the third - pictures and logos, and the fourth - text. As a result of the analysis, it was revealed that the approach based on four neural networks showed the best results in accordance with the objectives of the study on the test sample in the context of most classes of detection. The method proposed in the article can be used to recognize other objects. A promising direction in which the analysis can be continued is the segmentation of tables; the areas of the table that differ in function will act as classes: heading, cell with a name, cell with data, empty cell.

키워드

저자정보

Petr Sokerin Research laboratory «Monetary policy research and financial market analysis», Plekhanov Russian University of Economics, Russia
Alla Volkova KPMG Taxes and Consulting, LLC, Russia
Kirill Kushnarev Research laboratory «Monetary policy research and financial market analysis», Plekhanov Russian University of Economics, Russia

참고문헌

자료제공 : 네이버학술정보

함께 이용한 논문

※ 원문제공기관과의 협약기간이 종료되어 열람이 제한될 수 있습니다.

0개의 논문이 장바구니에 담겼습니다.

earticle