A deep learning framework for robust text extraction from heterogeneous image sources

Dayananda Kodala Jayaram, Puttegowda Devegowda

Abstract


Text extraction from an image is characterized by various challenges that result in various evolving solutions at current times. Existing solutions using artificial intelligence (AI) do have some promising outcomes, but also have their own shortcomings associated with accuracy and efficiency. The manuscript introduces a simplified and yet innovative AI framework towards extracting text by unifying character region-aware text detection. The model uses a convolutional neural network (CNN) along with transformer-centric optical character recognition (OCR) towards better performance on scene text recognition. Analysis is carried out on the standard and benchmarked International Conference on Document Analysis and Recognition 2015 (ICDAR 2015) dataset to find that the proposed model accomplished 94.7% accuracy and 0.82 seconds per image of response time, which are much significant improvements in performance, while observed in contrast to existing AI models, proving the model to deliver high-performance and cost-effective outcomes.

Keywords


Artificial intelligence; Convolutional neural network; Deep learning; Optical character recognition; Scene text recognition; Transformer

Full Text:

PDF


DOI: http://doi.org/10.11591/ijai.v15.i5.pp4724-4732

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Dayananda Kodala Jayaram, Puttegowda Devegowda

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

IAES International Journal of Artificial Intelligence (IJ-AI)
ISSN/e-ISSN 2089-4872/2252-8938 
This journal is published by the Institute of Advanced Engineering and Science (IAES).

View IJAI Stats