Indian sign language video generation using attention-enhanced generative adversarial network

Prachi Pramod Waghmare, Ashwini Mangesh Deshpande

Abstract


Sign language (SL) is the primary mode of communication for the Deaf signers. Despite advancements in deep learning, SL recognition, translation, and video generation face challenges like blurriness and inconsistencies. This research proposes a novel sign language translation (SLT) approach for Indian sign language (ISL) using an attention-driven generative adversarial network (GAN). The preprocessing pipeline includes video frame extraction, skeletal joint coordinate detection via OpenPose, and dynamic time warping (DTW) for pose data refinement. The squeeze and excitation (SE) attention mechanism enhances 2D convolutional layers, allowing the generator to focus on relevant skeletal pose sequences. A motion discriminator refines motion authenticity. Performance evaluation using two SL datasets demonstrates significant improvements in structural similarity index measure (SSIM), peak signal-to-noise ratio (PSNR), and temporal consistency metric (TCM) scores, achieving 99.60 (%) as SSIM, 31.10 dB as PSNR, and 0.9111 as TCM. The proposed model outperforms standard GAN and dynamic GAN in SL video generation.

Keywords


Deep learning; Generative adversarial network; Indian sign language; Sign language translation; Squeeze and excitation attention mechanism; Video generation

Full Text:

PDF


DOI: http://doi.org/10.11591/ijai.v15.i4.pp3672-3682

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Prachi Pramod Waghmare, Ashwini Mangesh Deshpande

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

IAES International Journal of Artificial Intelligence (IJ-AI)
ISSN/e-ISSN 2089-4872/2252-8938 
This journal is published by the Institute of Advanced Engineering and Science (IAES).

View IJAI Stats