ISSN :2582-9793

A Text-Guided Cross-Modal Diffusion Framework with Attention-Based Hand-Drawn Sketches for Face Synthesis

Original Research (Published On: 18-Aug-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.64334

Ahmed A. Hashim and Inas Ismael Imran

Adv. Artif. Intell. Mach. Learn., 6 (4):6033-6049

1. Inas Ismael Imran: Informatics Institute for Postgraduate Studies, University of Information Technology and Communication (UoITC), Baghdad, Iraq,

2. Ahmed A. Hashim: Department of Business Information Technology, business Informatics college, University of Information Technology and Communication (UoITC), Baghdad, 10053, Iraq

Download PDF Here

DOI: 10.54364/AAIML.2026.64334

Article History: Received on: 03-Apr-26, Accepted on: 11-Aug-26, Published on: 18-Aug-26

Corresponding Author: Ahmed A. Hashim

Email: dr.ahmed.hashim@uoitc.edu.iq

Citation: Inas Ismael Imran and Ahmed A. Hashim. A Text-Guided Cross-Modal Diffusion Framework With AttentionBased Hand-Drawn Sketches for Face Synthesis. Advances in Artificial Intelligence and Machine Learning. 2026;6(4):334. https://dx.doi.org/10.54364/AAIML.2026.64334


Abstract

Face sketch-to-photo synthesis plays a key role in computer vision with applications in law enforcement, digital entertainment, and human–computer interaction. Existing generative adversarial network-based methods typically face mode collapse, training instability, and poor performance across different sketching styles. This study introduces Sketch-to-Face, a cross-modal diffusion-based model that leverages Stable Diffusion to generate photorealistic faces using sketch-based inpainting. The proposed approach comprises three components: a Sketch Encoder with multiresolution attention that produces CLIP-compatible embeddings from grayscale sketches, Cross-Modal Fusion module employing bidirectional attention to bridge sketch spatial features with text semantic features, and Automatic Mask Generator with learnable refinement for adaptive inpainting guidance. A low-rank adaptation was applied to fine-tune the U-Net, reducing trainable parameters to 20.5 M while retaining the 860M-parameter backbone .Experiments on the Person Face Sketches dataset  (21K+ pairs) show stable convergence and are evaluated using SSIM, LPIPS, FID, identity similarity, runtime, and GPU memory. The results demonstrate a favorable quality–efficiency trade-off compared with text-and-sketch diffusion baselines, while qualitative results indicate plausible facial structure and visual realism


Statistics

Article Views: 279
PDF Downloads: 8