ISSN :2582-9793

A Real-Time Deep Learning-Based Sign Language Recognition System for Words, Alphabets, and Numbers

Original Research (Published On: 14-Aug-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.64331

Milia Habib, Teddy Nohra, Charbel Srour, Mohammad Kanaan and Zaher Merhi

Adv. Artif. Intell. Mach. Learn., 6 (4):5983-5994

1. Milia Habib: Department of Computer & Communications Engineering, Lebanese International University

2. Teddy Nohra: Department of Computer & Communications EngineeringLebanese International UniversityBeirut, Lebanon

3. Charbel Srour: Department of Computer & Communications Engineering Lebanese International University Beirut, Lebanon

4. Mohammad Kanaan: Department of Electrical Engineering Lebanese International University Beirut, Lebanon

5. Zaher Merhi: Department of Computer & Communications Engineering Lebanese International University Beirut, Lebanon

Download PDF Here

DOI: 10.54364/AAIML.2026.64331

Article History: Received on: 29-Apr-26, Accepted on: 07-Aug-26, Published on: 14-Aug-26

Corresponding Author: Milia Habib

Email: milia.habib@liu.edu.lb

Citation: Milia Habib et al. A Real-Time Deep Learning-Based Sign Language Recognition System for Words, Alphabets, and Numbers. Advances in Artificial Intelligence and Machine Learning. 2026;6(4):331. https://dx.doi.org/10.54364/AAIML.2026.64331


Abstract

This paper presents a sign language recognition system based on deep learning and computer vision. It aims to support communication between deaf and hard-of-hearing individuals and the community. The proposed system translates hand gestures into textual output through real-time image and video processing. It supports three recognition modes, including number, alphabet, and word recognition. For alphabet and digit recognition, MobileNetV2 is adopted due to its lightweight architecture and suitability for real-time deployment. To improve recognition accuracy and robustness in practical scenarios, a custom-collected dataset is combined with publicly available American Sign Language (ASL) datasets. The resulting dataset covers static alphabet gestures (excluding the motion-based letters J and Z) and digit gestures corresponding to the numbers 0–9. For word recognition, the pretrained Inflated 3D ConvNet (I3D) model is adopted to process short video clips and learn both spatial and temporal gesture features. Additionally, OpenCV is employed for video processing and real-time frame capture, while DroidCam is used to capture live input. For the MobileNetV2 model, a Region of Interest (ROI) is applied before prediction. Experimental results from real-time testing demonstrate high confidence scores for alphabet and number recognition, achieving 93.43% and 92.8%, respectively, at 30 FPS. In addition, the word recognition module achieves moderate performance, with Top-1 and Top-5 accuracies of 39.2% and 68.0%, respectively.


Statistics

Article Views: 250
PDF Downloads: 10