SIGNSIGHT: A Unified Real-Time Vision Framework for Object Detection and Sign Language Recognition
EOI: 10.11242/viva-tech.01.09.48
Citation
Nanaware Karan, Pawar Samiksha, Khamkar Ayush, Prof.Minakshi Gaonkar, Rane Srushti, " SIGNSIGHT: A REVIEW OF AUTOMATED SIGN LANGUAGE RECOGNITION AND THE DEVELOPMENT OF A REAL-TIME DETECTION FRAMEWORK", VIVA-IJRI Volume 1, Issue 8, Article 1, pp. 1-7, 2025. Published by Artificial Intelligence And Machine Learning Engineering Department, VIVA Institute of Technology, Virar, India.
Abstract
Communication barriers faced by hearing- and speech-impaired individuals remain a significant challenge in real-world interactions. At the same time, real-time object detection systems have gained widespread importance in intelligent surveillance and human–computer interaction applications. This paper presents SIGNSIGHT, a unified real-time vision framework that integrates object detection and sign language recognition within a single platform. The proposed system employs a pretrained YOLOv4 model for real-time object detection and localization, enabling accurate identification of surrounding objects. For sign language recognition, MediaPipe is utilized to extract 21 hand landmark keypoints, which are processed using a convolutional neural network (CNN) for classification of 27 static and dynamic gestures. The parallel architecture allows simultaneous execution of both modules, providing seamless real-time performance. Experimental evaluation demonstrates reliable detection accuracy and efficient frame processing suitable for practical deployment. By combining accessibility-focused gesture recognition with general-purpose object detection, the proposed framework offers a scalable and assistive AI-based vision solution. The system contributes toward enhancing inclusive communication and advancing real-time multimodal human–machine interaction.
Keywords
- Artificial media detection, Deep learning models, Multimodal content analysis, Synthetic data identification, Video, and text verification..
References
- T. Starner and A. Pentland, “Real-time American Sign Language recognition from video using hidden Markov models,” Proc. Int. Symp. Computer Vision, 1995, pp. 265–270.
- S. Mitra and T. Acharya, “Gesture recognition: A survey,” IEEE Trans. Systems, Man, and Cybernetics, vol. 37, no. 3, pp. 311–324,2007.
- R. Yang and S. Sarkar, “Gesture recognition using hidden Markov models from fragmented observations,” Computer Vision and Image Understanding, vol. 108, no. 1–2, pp. 41–53, 2007.
- J. Shotton et al., “Real-time human pose recognition in parts from single depth images,” Proc. IEEE CVPR, 2011, pp. 1297–1304.
- K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv:1409.1556, 2014.
- A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet classification with deep convolutional neural networks,” Communications of the ACM, vol. 60, no. 6, pp. 84–90, 2017.
- Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, vol. 521, pp. 436–444, 2015.
- P. Molchanov et al., “Hand gesture recognition with 3D convolutional neural networks,” Proc. IEEE CVPR, 2015.
- A. Graves, A. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” Proc. IEEE ICASSP, 2013.
- F. Chollet, Deep Learning with Python, Manning Publications, 2017.
- C. Szegedy et al., “Going deeper with convolutions,” Proc. IEEE CVPR, 2015.
- Z. Zhang, “Microsoft Kinect sensor and its effect,” IEEE Multimedia, vol. 19, no. 2, pp. 4–10, 2012.
- T. Baltrušaitis, C. Ahuja, and L. Morency, “Multimodal machine learning: A survey,” IEEE Trans. Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423–443, 2019.
- S. Banerjee et al., “Indian sign language recognition using deep learning,” Procedia Computer Science, vol. 167, pp. 2035–2044, 2020.
- M. Abavisani, H. Patel, and V. M. Patel, “Improving the performance of unimodal dynamic hand-gesture recognition with multimodal training,” Proc. IEEE CVPR, 2019.
- A. Vaswani et al., “Attention is all you need,” Advances in Neural Information Processing Systems, 2017.
- C. Lugaresi et al., “MediaPipe: A framework for building perception pipelines,” arXiv:1906.08172, 2019.
- J. Redmon et al., “You Only Look Once: Unified, real-time object detection,” Proc. IEEE CVPR, 2016.
- A. Bochkovskiy, C.-Y. Wang, and H.-Y. M. Liao, “YOLOv4: Optimal speed and accuracy of object detection,” arXiv:2004.10934, 2020.
- G. Bradski, “The OpenCV library,” Dr. Dobb’s Journal of Software Tools, 2000.
- M. Abduallah et al., “Real-time sign language recognition system,” IEEE Access, vol. 9, pp. 132–145, 2021.
- D. W. Kim et al., “Vision-based hand gesture recognition using CNN,” Applied Sciences, vol. 8, no. 8, 2018.
- A. Jain et al., “Real-time sign language translation system,” Lecture Notes in Networks and Systems, Springer, 2021.
- S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997.
- O. Koller, “Deep learning for sign language recognition,” Lecture Notes in Computer Science, Springer, 2020.
