">

SIGNSIGHT: A Unified Real-Time Vision Framework for Object Detection and Sign Language Recognition



EOI: 10.11242/viva-tech.01.09.48

Download Full Text here



Citation

Nanaware Karan, Pawar Samiksha, Khamkar Ayush, Prof.Minakshi Gaonkar, Rane Srushti, " SIGNSIGHT: A REVIEW OF AUTOMATED SIGN LANGUAGE RECOGNITION AND THE DEVELOPMENT OF A REAL-TIME DETECTION FRAMEWORK", VIVA-IJRI Volume 1, Issue 8, Article 1, pp. 1-7, 2025. Published by Artificial Intelligence And Machine Learning Engineering Department, VIVA Institute of Technology, Virar, India.

Abstract

Communication barriers faced by hearing- and speech-impaired individuals remain a significant challenge in real-world interactions. At the same time, real-time object detection systems have gained widespread importance in intelligent surveillance and human–computer interaction applications. This paper presents SIGNSIGHT, a unified real-time vision framework that integrates object detection and sign language recognition within a single platform. The proposed system employs a pretrained YOLOv4 model for real-time object detection and localization, enabling accurate identification of surrounding objects. For sign language recognition, MediaPipe is utilized to extract 21 hand landmark keypoints, which are processed using a convolutional neural network (CNN) for classification of 27 static and dynamic gestures. The parallel architecture allows simultaneous execution of both modules, providing seamless real-time performance. Experimental evaluation demonstrates reliable detection accuracy and efficient frame processing suitable for practical deployment. By combining accessibility-focused gesture recognition with general-purpose object detection, the proposed framework offers a scalable and assistive AI-based vision solution. The system contributes toward enhancing inclusive communication and advancing real-time multimodal human–machine interaction.

Keywords

- Artificial media detection, Deep learning models, Multimodal content analysis, Synthetic data identification, Video, and text verification..

References

  1. T. Starner and A. Pentland, “Real-time American Sign Language recognition from video using hidden Markov models,” Proc. Int. Symp. Computer Vision, 1995, pp. 265–270.
  2. S. Mitra and T. Acharya, “Gesture recognition: A survey,” IEEE Trans. Systems, Man, and Cybernetics, vol. 37, no. 3, pp. 311–324,2007.
  3. R. Yang and S. Sarkar, “Gesture recognition using hidden Markov models from fragmented observations,” Computer Vision and Image Understanding, vol. 108, no. 1–2, pp. 41–53, 2007.
  4. J. Shotton et al., “Real-time human pose recognition in parts from single depth images,” Proc. IEEE CVPR, 2011, pp. 1297–1304.
  5. K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv:1409.1556, 2014.
  6. A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet classification with deep convolutional neural networks,” Communications of the ACM, vol. 60, no. 6, pp. 84–90, 2017.
  7. Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, pp. 436–444, 2015.
  8. P. Molchanov et al., “Hand gesture recognition with 3D convolutional neural networks,” Proc. IEEE CVPR, 2015.
  9. A. Graves, A. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” Proc. IEEE ICASSP, 2013.
  10. F. Chollet, Deep Learning with Python, Manning Publications, 2017.
  11. C. Szegedy et al., “Going deeper with convolutions,” Proc. IEEE CVPR, 2015.
  12. Z. Zhang, “Microsoft Kinect sensor and its effect,” IEEE Multimedia, vol. 19, no. 2, pp. 4–10, 2012.
  13. T. Baltrušaitis, C. Ahuja, and L. Morency, “Multimodal machine learning: A survey,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423–443, 2019.
  14. S. Banerjee et al., “Indian sign language recognition using deep learning,” Procedia Computer Science, vol. 167, pp. 2035–2044, 2020.
  15. M. Abavisani, H. Patel, and V. M. Patel, “Improving the performance of unimodal dynamic hand-gesture recognition with multimodal training,” Proc. IEEE CVPR, 2019.
  16. A. Vaswani et al., “Attention is all you need,” Advances in Neural Information Processing Systems, 2017.
  17. C. Lugaresi et al., “MediaPipe: A framework for building perception pipelines,” arXiv:1906.08172, 2019.
  18. J. Redmon et al., “You Only Look Once: Unified, real-time object detection,” Proc. IEEE CVPR, 2016.
  19. A. Bochkovskiy, C.-Y. Wang, and H.-Y. M. Liao, “YOLOv4: Optimal speed and accuracy of object detection,” arXiv:2004.10934, 2020.
  20. G. Bradski, “The OpenCV library,” Dr. Dobb’s Journal of Software Tools, 2000.
  21. M. Abduallah et al., “Real-time sign language recognition system,” IEEE Access, vol. 9, pp. 132–145, 2021.
  22. D. W. Kim et al., “Vision-based hand gesture recognition using CNN,” Applied Sciences, vol. 8, no. 8, 2018.
  23. A. Jain et al., “Real-time sign language translation system,” Lecture Notes in Networks and Systems, Springer, 2021.
  24. S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.
  25. O. Koller, “Deep learning for sign language recognition,” Lecture Notes in Computer Science, Springer, 2020.