ISSN :2582-9793

Model Compression Techniques for Scalable Deployment of Generative AI in Automotive Systems

Review Article (Published On: 15-Aug-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.64332

Arunachalam Thirunavukkarasu, Domenik Helms, Christian Ojeda-Bernal, Sven Mantowsky, Syed Bukhari and Robert Buecs

Adv. Artif. Intell. Mach. Learn., 6 (4):5995-6013

1. Arunachalam Thirunavukkarasu: German Aerospace Center (DLR)

2. Domenik Helms: German Aerospace Center DLR

3. Christian Ojeda-Bernal: Artificial Intelligence Engineer bei Valeo Schalter und Sensoren GmbH

4. Sven Mantowsky: AI engineer at ZF

5. Syed Bukhari: Chief AI Engineer inn the AI Lab Saarbrücken, ZF Friedrichshafen AG

6. Robert Buecs: Senior Engineer Embedded AI - Behavior Prediction & Planning

Download PDF Here

DOI: 10.54364/AAIML.2026.64332

Article History: Received on: 26-Apr-26, Accepted on: 08-Aug-26, Published on: 15-Aug-26

Corresponding Author: Arunachalam Thirunavukkarasu

Email: arunachalam.thirunavukkarasu@dlr.de

Citation: Arunachalam Thirunavukkarasu, et al. Model Compression Techniques for Scalable Deployment of Generative AI in Automotive Systems. Advances in Artificial Intelligence and Machine Learning. 2026;6(4):332. https://dx.doi.org/10.54364/AAIML.2026.64332


Abstract

The deployment of generative AI models within automotive systems is fundamentally constrained by the limited computational resources, memory capacity, and energy budgets of in-vehicle hardware. In a companion article \cite{part1}, we examined the scalability challenges, current applications, and evolving E/E architectures that frame this problem. The present paper builds on that foundation by providing an in-depth survey of model compression and optimization techniques that enable the deployment of large neural networks on resource-constrained automotive platforms. We first present an overview of deployment strategies including edge computing, hardware selection, optimized inference frameworks, AI compilers, dedicated accelerators, and knowledge distillation. We then conduct a detailed review of three core compression methodologies: pruning (both with and without retraining), quantization (post-training and quantization-aware approaches for both vision and language models), and low-rank tensor decomposition (including Canonical Polyadic, Tucker, Tensor Train, and Tensor Ring methods, as well as Neural Architecture Search-based compression). For each technique, we analyze the underlying principles, review state-of-the-art methods, and discuss their applicability to automotive deployment scenarios. A comparative analysis highlights the trade-offs among compression ratio, inference speedup, accuracy retention, and hardware compatibility. Our findings indicate that combining multiple compression strategies offers the most promising path toward practical, real-time generative AI deployment in vehicles, and we identify open research directions for the automotive AI community.


Statistics

Article Views: 494
PDF Downloads: 5