Original Research (Published On: 31-Aug-2023 )
DOI : https://doi.org/10.54364/AAIML.2023.1181Wenping Wang, Tong Chen, Sicong Liu, Zhiran Chen, Wenyan Hu, Dachi Chen, Yuanxin Wang, Qi Lyu and Cindy X. Le
Adv. Artif. Intell. Mach. Learn., 3 (3):1369–1388
1. Wenping Wang: Individual Researcher
2. Tong Chen: Google Inc, 1600 Amphitheatre Parkway, Mountain View, CA, USA 94043
3. Sicong Liu: Amazon Inc, 410 Terry Ave N, Seattle 98109, WA, USA
4. Zhiran Chen: Otter.ai, Inc., 800 W El Camino Real, Suite 170, Mountain View, CA 94040
5. Wenyan Hu: Meta Platforms, 1 Hacker Way Melon Park, CA, USA 94025
6. Dachi Chen: Meta Platforms, 1 Hacker Way Melon Park, CA, USA 94025
7. Yuanxin Wang: Carnegie Mellon University Pennsylvania, USA.
8. Qi Lyu: Michigan State University, 426 Auditorium Road, East Lansing, MI 48824
9. Cindy X. Le: Google Inc, 1600 Amphitheatre Parkway, Mountain View, CA, USA 94043
DOI: 10.54364/AAIML.2023.1181
Article History: Received on: 04-Jun-23, Accepted on: 23-Aug-23, Published on: 31-Aug-23
Corresponding Author: Wenping Wang
Email: wenpingw@alumni.cmu.edu
Citation: Tong Chen, et al. Faster, Stronger, and More Interpretable: Massive Transformer Architectures for Vision-Language Tasks. 2023;3(3):81.
Abstract
Multi-layered transformer architectures have lately dominated the domain of vision-language tasks. However, massive transformer architectures can often be inaccessible to many researchers due to their sheer model sizes, and they are often treated as black boxes with poor interpretability. In this paper, we examine the weaknesses of such architectures and propose our own solutions. In particular, we select one of the state-of-the-art models called Oscar \cite{li2020Oscar} and apply distilling techniques and attention visualization to address the aforementioned issues. Moreover, we attempt to improve the overall effectiveness of the Oscar model by making its inferred object tags more useful. We show with detailed experimentation that we can both improve the performance of vision-language tasks and make them more transparent and accessible to all researchers. We discuss the findings with detailed analysis, including the effects of tags and confidence, the training behavior of distillation, and point out future directions in the end.
Statistics
Article Views: 2902
PDF Downloads: 28
