A Compiler-Aware Framework for Partitioned Neural Network Inference on FPGA DPUs
DOI:
https://doi.org/10.64552/wipiec.v12i2.144Keywords:
FPGA, Distributed Inference, Deep Learning, DPU, Vitis AI, XIRAbstract
The increasing adoption of Artificial Intelligence is driving the development of larger and more accurate neu-ral networks. However, their high computational cost leads to significant inference latency, especially on resource-constrained embedded platforms such as FPGA-based systems. To address this issue, AMD introduced the Deep Learning Processing Unit, an accelerator integrated into the Vitis AI toolchain for efficient execution of quantized neural networks. Nevertheless, models with many parameters can be difficult to deploy on a single board due to latency or capacity constraints. In these scenarios, splitting a model into separately compilable fragments becomes useful for more flexible deployment. This process is not immediate within the Vitis AI flow. Our analysis shows that manually partitioning a network into independently compiled sub-networks can remove compiler optimizations. In particular, losing the global view of the quantized XIR graph can prevent the toolchain from preserving DPU mapping, moving accelerable operations to the CPU and causing severe performance degradation. Based on this observation, we analyze the Vitis AI compiler and propose an XIR-level splitting framework that generates independently compilable .xmodel fragments while preserving the context required for DPU mapping. The approach keeps the workflow high-level, without requiring advanced hardware design expertise or low-level design changes. Experimental results on CNN and ConvViT-based models show that naive splitting can introduce slowdowns up to ×249 on CNNs and ×542 on ConvViT variants. The proposed framework restores correct DPU mapping by addressing boundary-context loss and incomplete dependency collection, bringing latency back to the expected range for hardware-accelerated execution.
References
F. Buccellato, E. Vacca, S. Azimi, C. De Sio, and L. Sterpone, “From De-tection to Intervention: An End-to-End System for Recognizing the “Sig-nal for Help” Gesture in Real-Time,” Intelligent Systems with Applica-tions, vol. 26, Art. no. 200536, 2025, doi: 10.1016/j.iswa.2025.200536.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Y. Ng, “Large Scale Distributed Deep Networks,” in Advances in Neural Information Processing Systems, vol. 25, 2012, pp. 1232–1240.
W. Jiang, E. H.-M. Sha, X. Zhang, L. Yang, Q. Zhuge, Y. Shi, and J. Hu, “Achieving Super-Linear Speedup across Multi-FPGA for Real-Time DNN Inference,” ACM Transactions on Embedded Computing Systems, vol. 18, no. 5s, pp. 67:1–67:23, Oct. 2019, doi: 10.1145/3358186.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, M. X. Chen, D. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, and Z. Chen, “GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism,” in Advances in Neural Information Processing Systems, vol. 32, 2019.
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro, “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism,” arXiv preprint arXiv:1909.08053, 2019.
C. Zhang, P. Li, G. Sun, Y. Guan, B. Xiao, and J. Cong, “Optimizing FPGA-Based Accelerator Design for Deep Convolutional Neural Net-works,” in Proc. 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA), Monterey, CA, USA, Feb. 2015,
pp. 161–170, doi: 10.1145/2684746.2689060.
G. Cora, E. Vacca, C. De Sio, S. Azimi, and L. Sterpone, “Fast SEU Detection and Recovery in FPGA-Based AI Accelerators,” ACM Transactions on Reconfigurable Technology and Systems, 2026, doi: 10.1145/3806052.
G. Cora, D. Rizzieri, C. De Sio, S. Azimi, and L. Sterpone, “In-Hardware Fault-Tolerance Controller for Multi-FPGA Clustered Architectures,” IEEE Transactions on Computers, 2026, doi: 10.1109/TC.2026.3709831.
Y. Zhu, Z. He, W. Jiang, K. Zeng, J. Zhou, and G. Alonso, “Distributed Recommendation Inference on FPGA Clusters,” in Proc. 31st Inter-national Conference on Field-Programmable Logic and Applications (FPL), 2021, pp. 279–285.
Y. Umuroglu, N. J. Fraser, G. Gambardella, M. Blott, P. H. W. Leong, M. Jahre, and K. A. Vissers, “FINN: A Framework for Fast, Scalable Binarized Neural Network Inference,” in Proc. ACM/SIGDA Interna-tional Symposium on Field-Programmable Gate Arrays (FPGA), 2017, pp. 65–74.
H. Sharma, J. Park, D. Mahajan, E. Amaro, J. K. Kim, C. Shao, A. Mishra, and H. Esmaeilzadeh, “From High-Level Deep Neural Models to FPGAs,” in Proc. 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2016, pp. 17:1–17:12.
S. I. Venieris and C.-S. Bouganis, “fpgaConvNet: A Toolflow for Map-ping Diverse Convolutional Neural Networks on Embedded FPGAs,” arXiv preprint arXiv:1711.08740, 2017.
AMD, “Deep Learning Processing Unit,” in Vitis AI User Guide (UG1414), Version 3.5, Sep. 28, 2023. [Online]. Available: https://docs. amd.com/r/en-US/ug1414-vitis-ai
AMD, “Vitis AI User Guide (UG1414),” Version 2.5, 2022. [Online]. Available: https://docs.amd.com/r/2.5-English/ug1414-vitis-ai
F. Buccellato, C. De Sio, S. Azimi, and L. Sterpone, “On-Hardware Re-silience Analysis of DPU-Accelerated CNNs on FPGA-Based Systems,” in Proc. 28th Euromicro Conference on Digital System Design (DSD), Salerno, Italy, 2025, pp. 34–41, doi: 10.1109/DSD67783.2025.00017.
F. Buccellato, C. De Sio, S. Azimi, and L. Sterpone, “Hardware-Aware Runtime Detection of Soft-Error Anomalies in DPU-Accelerated Neural Networks,” in Proc. IEEE International On-Line Testing Symposium (IOLTS), 2026.
AMD, “Xilinx Intermediate Representation,” in Vitis AI User Guide (UG1414), Version 3.5, Sep. 28, 2023. [Online]. Available: https://docs. amd.com/r/en-US/ug1414-vitis-ai/XIR
A. B. Yoo, M. A. Jette, and M. Grondona, “SLURM: Simple Linux Utility for Resource Management,” in Job Scheduling Strate-gies for Parallel Processing, D. Feitelson, L. Rudolph, and U. Schwiegelshohn, Eds. Berlin, Heidelberg: Springer, 2003, pp. 44–60, doi: 10.1007/109689873.
A. B. Kahn, “Topological Sorting of Large Networks,” Commun. ACM, vol. 5, no. 11, pp. 558–562, 1962, doi: 10.1145/368996.369025.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Federico Buccellato, Luca Mannini, Corrado De Sio

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
License Terms:
Except where otherwise noted, content on this website is lincesed under a Creative Commons Attribution Non-Commercial License (CC BY NC)
![]()
Use, distribution and reproduction in any medium, provided the original work is properly cited and is not used for commercial purposes, is permitted.
Copyright to any article published by WiPiEC retained by the author(s). Authors grant WiPiEC Journal a license to publish the article and identify itself as the original publisher. Authors also grant any third party the right to use the article freely as long as it is not used for commercial purposes and its original authors, citation details, and publisher are identified, in accordance with CC BY NC license. Fore more information on license terms, click here.