ISSN :2582-9793

Performance Evaluation of DenseNet-121 and Bi-LSTM with Attention Mechanism based Vocal Signal Separation

Original Research (Published On: 03-Oct-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.65348

Parveen Lehana and Arfana Chowdhary

Adv. Artif. Intell. Mach. Learn., - (-):-

1. Parveen Lehana: Department of Electronics University of Jammu India

2. Arfana Chowdhary: DSP Laboratory, Department of Electronics, University of Jammu Jammu-180006, India

Download PDF Here

DOI: 10.54364/AAIML.2026.65348

Article History: Received on: 06-Feb-26, Accepted on: 26-Sep-26, Published on: 03-Oct-26

Corresponding Author: Parveen Lehana

Email: pklehana@gmail.com

Citation: Arfana Chowdhary and Parveen Kumar Lehana. Performance Evaluation of DenseNet-121 and Bi-LSTM with Attention Mechanism based Vocal Signal Separation. Advances in Artificial Intelligence and Machine Learning. 2026. (Ahead of Print) https://dx.doi.org/10.54364/AAIML.2026.65348


Abstract

Audio source separation is the process of separating a mixed audio signal into its individual sources. This task remains difficult because of significant overlap in time and frequency. The presence of noise and fluctuating signal characteristics further deteriorate the task. Traditional signal processing methods rely on strong statistical assumptions, which often limit their effectiveness in real-world acoustic signal processing. To overcome some of these challenges, this paper presents a supervised attention-based deep learning approach for separating single-channel audio signals. This approach directly predicts the individual source magnitude spectrograms without employing traditional time-frequency masking methods. The approach is based on fixed-length segments of audio signals, which are represented at a constant rate and mapped to magnitude spectrograms using short-time Fourier transform (STFT). A structured preprocessing employed in the algorithm ensure normalisation, cropping, and temporal alignment, maintains constant input data size during training and testing phases. The separation approach employs a DenseNet-121 network that uses hierarchical spectral feature extraction. Thereby enabling the effective extraction of complex harmonic and formant characteristics. To predict time-frequency regions containing salient perceptual significance, circular spatial attention mechanism has been employed for enabling the network to concentrate on significant acoustic elements. The temporal characteristics and contextual parameters from spectrum frames are captured using a bidirectional long short-term memory (Bi-LSTM) network and multi-head self-attention. The combined feature representations are processed through fully connected layers of the DenseNet-121 network to estimate the magnitude spectrograms of the constituent sources. The setup allows for end-to-end learning of nonlinear source mappings. The network is trained using scale-invariant signal-to-distortion ratio (SI-SDR) as loss function and the trained network is evaluated using objective metrics, such as short-time objective intelligibility (STOI) and perceptual evaluation of speech quality (PESQ). The experimental results demonstrate satisfactory performance of separation, intelligibility, and perceptual quality across different mixing ratios. The approach establishes a low complexity, scalable, and interpretable solution to realistic problems in source separation. 


Statistics

Article Views: 1
PDF Downloads: 0