ISSN :2582-9793

Low-Resource Adaptation of Whisper and MMS for Lebanese Arabic Automatic Speech Recognition

Original Research (Published On: 11-Sep-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.65343

Zaher Merhi, Elias Elia, Jennifer Takla, Milia Habib, Mohamad Kannan and Rabih Rammal

Adv. Artif. Intell. Mach. Learn., - (-):-

1. Elias Elia: Department of Computer & Communications Engineering Lebanese International University Beirut, Lebanon

2. Jennifer Takla: Department of Computer & Communications Engineering Lebanese International University Beirut, Lebanon

3. Milia Habib: Department of Computer & Communications Engineering Lebanese International University Beirut, Lebanon

4. Mohamad Kannan: Department of Electrical Engineering Lebanese International University Beirut, Lebanon

5. Rabih Rammal: Department of Electrical Engineering Lebanese International University Beirut, Lebanon

6. Zaher Merhi: Lebanese International University, Beirut , Lebanon

Download PDF Here

DOI: 10.54364/AAIML.2026.65343

Article History: Received on: 10-May-26, Accepted on: 04-Sep-26, Published on: 11-Sep-26

Corresponding Author: Zaher Merhi

Email: zaher.merhi@liu.edu.lb

Citation: Elias Elias, et al. Low-Resource Adaptation of Whisper and MMS for Lebanese Arabic Automatic Speech Recognition. Advances in Artificial Intelligence and Machine Learning. 2026. (Ahead of Print) https://dx.doi.org/10.54364/AAIML.2026.65343


Abstract

Arabic automatic speech recognition has grown significantly in recent years, but performance remains limited for under-resourced dialects such as Lebanese Arabic, especially in real conversational scenarios where speakers switch between Lebanese Arabic, English, and French. This paper presents a low-resource dialect adaptation study on Lebanese speech transcription for meeting-oriented applications. Due to the lack of a clear open-source Lebanese speech corpus with aligned reference text, a custom dataset was created of approximately 1100 manually transcribed audio clips extracted from Lebanese podcast recordings. This paper investigates the effect of fine-tuning multiple pre-trained speech recognition models on this dataset, focusing on Whisper Small, Whisper Medium, Whisper Large, and Massively Multilingual Speech (MMS). The study compares baseline and adapted performance to determine how model size and multilingual pretraining affect dialectal generalization. The results show that fine-tuning improved Whisper Small from 44.11% to 40.36% WER, Whisper Medium from 44.37% to 35.32% WER, and Whisper Large from 37.20% to 33.83% WER. MMS also improved from 76.92% to 66.94% WER in the same low-resource setting. These findings highlight the difficulty of low-resource dialect adaptation, particularly in the presence of code-switching and limited training data, and show that model size alone does not guarantee better adaptation. This study provides an initial Lebanese dialect speech resource and offers useful insights for developing ASR systems for Lebanese meeting transcription and meeting understanding tasks.


Statistics

Article Views: 14
PDF Downloads: 1