Original Research (Published On: 14-Aug-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.64331Milia Habib, Teddy Nohra, Charbel Srour, Mohammad Kanaan and Zaher Merhi
Adv. Artif. Intell. Mach. Learn., 6 (4):5983-5994
1. Milia Habib: Department of Computer & Communications Engineering, Lebanese International University
2. Teddy Nohra: Department of Computer & Communications EngineeringLebanese International UniversityBeirut, Lebanon
3. Charbel Srour: Department of Computer & Communications Engineering Lebanese International University Beirut, Lebanon
4. Mohammad Kanaan: Department of Electrical Engineering Lebanese International University Beirut, Lebanon
5. Zaher Merhi: Department of Computer & Communications Engineering Lebanese International University Beirut, Lebanon
DOI: 10.54364/AAIML.2026.64331
Article History: Received on: 29-Apr-26, Accepted on: 07-Aug-26, Published on: 14-Aug-26
Corresponding Author: Milia Habib
Email: milia.habib@liu.edu.lb
Citation: Milia Habib et al. A Real-Time Deep Learning-Based Sign Language Recognition System for Words, Alphabets, and Numbers. Advances in Artificial Intelligence and Machine Learning. 2026;6(4):331. https://dx.doi.org/10.54364/AAIML.2026.64331
Abstract
This paper presents a sign language recognition system based
on deep learning and computer vision. It aims to support communication between
deaf and hard-of-hearing individuals and the community. The proposed system
translates hand gestures into textual output through real-time image and video
processing. It supports three recognition modes, including number, alphabet,
and word recognition. For alphabet and digit recognition, MobileNetV2 is
adopted due to its lightweight architecture and suitability for real-time
deployment. To improve recognition accuracy and robustness in practical
scenarios, a custom-collected dataset is combined with publicly available
American Sign Language (ASL) datasets. The resulting dataset covers static
alphabet gestures (excluding the motion-based letters J and Z) and digit gestures
corresponding to the numbers 0–9. For word recognition, the pretrained Inflated
3D ConvNet (I3D) model is adopted to process short video clips and learn both
spatial and temporal gesture features. Additionally, OpenCV is employed for
video processing and real-time frame capture, while DroidCam is used to capture
live input. For the MobileNetV2 model, a Region of Interest (ROI) is applied
before prediction. Experimental results from real-time testing demonstrate high
confidence scores for alphabet and number recognition, achieving 93.43% and
92.8%, respectively, at 30 FPS. In addition, the word recognition module
achieves moderate performance, with Top-1 and Top-5 accuracies of 39.2% and
68.0%, respectively.
Statistics
Article Views: 250
PDF Downloads: 10
