From Algorithm to System: Integrated Design of Compiler and Toolchain for Large Model Inference Optimization
Shengyi Gao
European Journal of AI, Computing & Informatics, 2025, vol. 1, issue 4, 21-33
Abstract:
With the rapid growth of deep learning model scales, especially large models such as Transformers and GPT, efficient inference has become a critical challenge due to increasing computational and memory demands. This paper proposes an integrated optimization framework that unifies algorithmic simplifications, compiler transformations, and system-level scheduling to enhance large model inference performance. By tightly coupling quantization, pruning, operator fusion, memory reuse, and automated heterogeneous hardware scheduling, the framework achieves significant improvements in computation reduction, memory efficiency, and parallel execution. Theoretical analysis and design considerations demonstrate the framework's potential for predictable performance gains and scalability across diverse hardware platforms. Future work will focus on extending hardware support, distributed inference, and adaptive optimization strategies. This integrated approach lays a foundation for efficient, scalable, and accurate large model deployment in practical AI applications.
Keywords: large model inference; algorithm optimization; compiler optimization; toolchain scheduling; heterogeneous hardware (search for similar items in EconPapers)
Date: 2025
References: Add references at CitEc
Citations:
Downloads: (external link)
https://pinnaclepubs.com/index.php/EJACI/article/view/385/386 (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:dba:ejacia:v:1:y:2025:i:4:p:21-33
Access Statistics for this article
More articles in European Journal of AI, Computing & Informatics from Pinnacle Academic Press
Bibliographic data for series maintained by Joseph Clark ().