Original Research (Published On: 10-Oct-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.65351Deepak Kumar and Monika Khatkar
Adv. Artif. Intell. Mach. Learn., - (-):-
1. Deepak Kumar: K R Mangalam University- Gurgaon
2. Monika Khatkar: Assistant Professor- School Of Engineering and TechnologyK R Mangalam University- Gurgaon
DOI: 10.54364/AAIML.2026.65351
Article History: Received on: 15-Jun-26, Accepted on: 03-Oct-26, Published on: 10-Oct-26
Corresponding Author: Deepak Kumar
Email: javadevdeepak@gmail.com
Citation: Deepak Kumar and Monika Khatkar. An Intelligent Speech Analytics Framework for Depression Detection: Comparative Assessment of Machine Learning Architectures. Advances in Artificial Intelligence and Machine Learning. 2026. (Ahead of Print) https://dx.doi.org/10.54364/AAIML.2026.65351
Abstract
ABSTRACT:
Depression affects emotional, cognitive, and behavioral
functioning and it is a pretty ubiquitous mental health (MH) disorder. It is
often described as subjective time consuming, and kind of expert-driven, so
clinical interviews plus self reported questionnaires are still the traditional
methods, for diagnosing.. Automated systems can detect depressed symptoms utilising
voice biomarkers with the aid of AI and speech signal processing (SSP). Emotional and psychological states can greatly affect the
variation in pitch, energy, pace, and spectrum characteristics of speech,
rendering it a potential tool to assess depression. This research is designed to develop an independent
depression detection system (DDS) using speech and voice (S&V) input and
evaluate ML and DL algorithms on depression and non-DD. The audio files should be used to get the data, labels
should be extracted from the audio files, features should be taken from
preprocessed data, the model (Mod.) should be trained with the extracted
features, and the performance of the Mod. should be evaluated. Librosa extracts Mel Frequency Cepstral Coefficients
(MFCCs), Chroma features, and Mel Spectrogram (MS) features from speech and is
able to gather them in a hybrid feature vector. We split the data straight into training (80%) and test
(20%) sets in a normal random fashion. The
accuracy, precision, recall for the MLP Mod. were observed as 98% , 99% and 99%
respectively, while the ROC-AUC score came out to 0.96 in the trials. Overall
the results propose that S&V signals can be leveraged for depression
identification and to strengthen automated MH assessments. Also, screening for
depression plus clinical decision support with speech AI appears to be
efficient, non-invasive alongside scalable.
Statistics
Article Views: 29
PDF Downloads: 4
