Original Research (Published On: 18-Aug-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.64334Ahmed A. Hashim and Inas Ismael Imran
Adv. Artif. Intell. Mach. Learn., 6 (4):6033-6049
1. Inas Ismael Imran: Informatics Institute for Postgraduate Studies, University of Information Technology and Communication (UoITC), Baghdad, Iraq,
2. Ahmed A. Hashim: Department of Business Information Technology, business Informatics college, University of Information Technology and Communication (UoITC), Baghdad, 10053, Iraq
DOI: 10.54364/AAIML.2026.64334
Article History: Received on: 03-Apr-26, Accepted on: 11-Aug-26, Published on: 18-Aug-26
Corresponding Author: Ahmed A. Hashim
Email: dr.ahmed.hashim@uoitc.edu.iq
Citation: Inas Ismael Imran and Ahmed A. Hashim. A Text-Guided Cross-Modal Diffusion Framework With AttentionBased Hand-Drawn Sketches for Face Synthesis. Advances in Artificial Intelligence and Machine Learning. 2026;6(4):334. https://dx.doi.org/10.54364/AAIML.2026.64334
Abstract
Face sketch-to-photo synthesis plays a key role in computer vision
with applications in law enforcement, digital entertainment, and human–computer
interaction. Existing generative adversarial network-based methods typically
face mode collapse, training instability, and poor performance across different
sketching styles. This study introduces Sketch-to-Face, a cross-modal
diffusion-based model that leverages Stable Diffusion to generate
photorealistic faces using sketch-based inpainting. The proposed approach
comprises three components: a Sketch Encoder with multiresolution attention
that produces CLIP-compatible embeddings from grayscale sketches, Cross-Modal
Fusion module employing bidirectional attention to bridge sketch spatial
features with text semantic features, and Automatic Mask Generator with
learnable refinement for adaptive inpainting guidance. A low-rank adaptation
was applied to fine-tune the U-Net, reducing trainable parameters to 20.5 M
while retaining the 860M-parameter backbone .Experiments on the Person Face
Sketches dataset (21K+ pairs) show
stable convergence and are evaluated using SSIM, LPIPS, FID, identity
similarity, runtime, and GPU memory. The results demonstrate a favorable
quality–efficiency trade-off compared with text-and-sketch diffusion baselines,
while qualitative results indicate plausible facial structure and visual
realism
Statistics
Article Views: 279
PDF Downloads: 8
