Autotuned Distribution of Multi-DNN Workloads on Multi-Accelerator SoCs

Authors

  • Federico Nicolás Peccia FZI Research Center for Information Technology, University of Tübingen
  • Avik Bhatnagar FZI Research Center for Information Technology, University of Tübingen
  • Oliver Bringmann FZI Research Center for Information Technology, University of Tübingen

DOI:

https://doi.org/10.64552/wipiec.v12i2.134

Keywords:

Multi-tenant DNN, FPGA, RTOS, SoC

Abstract

Many machine learning applications require heterogeneous Deep Neural Networks (DNNs) to work collaboratively. Although several works have focused on how to serve these systems using cloud solutions, less attention has been paid to the edge scenarios. Particularly, coordinating heterogeneous AI workloads across a custom application-specific System-on-Chip (SoC) with multiple accelerators presents significant challenges. This work first demonstrates how to implement such a baseline system using a SoC generator framework, performs an ablation study prototyping different versions on an FPGA, details how an RTOS can be used to achieve model parallelism on multiple accelerators, and identifies gaps and limitations by executing a multi-DNN autonomous driving application. To improve the utilization of the system and increase throughput, we propose a method to distribute the execution of individual layers across accelerators. Instead of partitioning all layers in the same manner and statically allocating them at compile time, we select the ideal partitioning for each one during compilation using an autotuning process, and then dynamically assign them to the available accelerators during runtime. We analyze the variability in layer execution in a system with multiple accelerators and use this information to guide the runtime allocation of partitions. We demonstrate that our method achieves a mean 29 % and 40 % improvement in accelerator utilization and throughput over the model parallelism baseline, and a 10 % and 9 % improvement over a round-robin runtime distribution of partitions.

References

A. Mousavian, D. Anguelov, J. Flynn, and J. Kosecka, “3d bounding box estimation using deep learning and geometry,” CoRR, vol. abs/1612.00496, 2016. [Online]. Available: http://arxiv.org/abs/1612.00496

C.-J. Wu, D. Brooks, K. Chen, D. Chen et al., “Machine learning at facebook: Understanding inference at the edge,” in 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2019, pp. 331–344.

C. Wang, Y. Bai, and D. Sun, “CD-MSA: Cooperative and Deadline-Aware Scheduling for Efficient Multi-Tenancy on DNN Accelerators,”IEEE Transactions on Parallel and Distributed Systems, vol. 34, no. 7,pp. 2091–2106, Jul. 2023.

C. Gao, Y. Wang, C. Liu, M. Wang et al., “Layer-Puzzle: Allocating and Scheduling Multi-task on Multi-core NPUs by Using Layer Heterogeneity,” in 2023 Design, Automation & Test in Europe Conference & Exhibition (DATE). Antwerp, Belgium: IEEE, Apr. 2023, pp. 1–6.

Z. Zhao, N. Ling, N. Guan, and G. Xing, “Miriam: Exploiting elastic kernels for real-time multi-dnn inference on edge gpu,” in Proceedings of the 21st ACM Conference on Embedded Networked Sensor Systems, ser. SenSys ’23. New York, NY, USA: Association for Computing Machinery, 2024, p. 97–110. [Online]. Available: https://doi.org/10.1145/3625687.3625789

A. Amid, D. Biancolin, A. Gonzalez, D. Grubb et al., “Chipyard: Integrated design, simulation, and implementation framework for custom socs,” IEEE Micro, vol. 40, no. 4, pp. 10–21, 2020.

H. Genc, S. Kim, A. Amid, A. Haj-Ali et al., “Gemmini: Enabling systematic deep-learning architecture evaluation via full-stack integration,” in Proceedings of the 58th Annual Design Automation Conference (DAC), 2021.

Linux foundation. zephyr [Online]. Available: https://www.zephyrproject.org/[Accessed: 16 April 2025].

T. Chen, L. Zheng, E. Q. Yan, Z. Jiang et al., “Learning to optimize tensor programs,” CoRR, vol. abs/1805.08166, 2018. [Online]. Available: http://arxiv.org/abs/1805.08166

E. Baek, D. Kwon, and J. Kim, “A multi-neural network acceleration architecture,” in Proceedings of the ACM/IEEE 47th Annual International Symposium on Computer Architecture, ser. ISCA ’20. Virtual Event: IEEE Press, Sep. 2020, pp. 940–953.

Y. H. Oh, S. Kim, Y. Jin, S. Son et al., “Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise Scheduling,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), Feb. 2021, pp. 584–597.

Y. Choi and M. Rhu, “PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing Units,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). San Diego, CA, USA: IEEE, Feb. 2020, pp. 220–233.

S. Ghodrati, B. H. Ahn, J. Kyung Kim, S. Kinzer et al., “Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural Networks,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), Oct. 2020, pp. 681–697.

H. Kwon, L. Lai, M. Pellauer, T. Krishna et al., “Heterogeneous Dataflow Accelerators for Multi-DNN Workloads,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), Feb. 2021, pp. 71–83.

S.-C. Kao and T. Krishna, “MAGMA: An Optimization Framework for Mapping Multiple DNNs on Multiple Accelerator Cores,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA). Seoul, Korea, Republic of: IEEE, Apr. 2022, pp. 814–830.

X. Zhang, C. Hao, P. Zhou, A. Jones, and J. Hu, “H2H: Heterogeneous Model to Heterogeneous System Mapping with Computation and Communication Awareness,” Apr. 2022.

S. Karl, A. Symons, N. Fasfous, and M. Verhelst, “Genetic Algorithm-based Framework for Layer-Fused Scheduling of Multiple DNNs on Multi-core Systems,” in 2023 Design, Automation & Test in Europe Conference & Exhibition (DATE). Antwerp, Belgium: IEEE, Apr. 2023, pp. 1–6.

S. Zheng, S. Chen, and Y. Liang, “Memory and Computation Coordinated Mapping of DNNs onto Complex Heterogeneous SoC,” in 2023 60th ACM/IEEE Design Automation Conference (DAC). San Francisco, CA, USA: IEEE, Jul. 2023, pp. 1–6.

J. Yang, H. Zheng, and A. Louri, “Versa-DNN: A Versatile Architecture Enabling High-Performance and Energy-Efficient Multi-DNN Acceleration,” IEEE Transactions on Parallel and Distributed Systems, vol. 35, no. 2, pp. 349–361, Feb. 2024.

M. Amine Hamdi, F. Daghero, G. Maria Sarda, J. Van Delm et al., “Match: Model-aware tvm-based compilation for heterogeneous edge devices,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 44, no. 10, pp. 3844–3857, 2025.

M. Kirschner, F. Peccia, F. Th¨ommes, V. P. Betancourt et al., “Characterization of execution time variability in fpga-based ai-accelerators,” in 2023 IEEE Intl Conf on Dependable, Autonomic and Secure Computing, Intl Conf on Pervasive Intelligence and Computing, Intl Conf on Cloud and Big Data Computing, Intl Conf on Cyber Science and Technology Congress (DASC/PiCom/CBDCom/CyberSciTech), 2023, pp. 1–8.

F. N. Peccia and O. Bringmann, “Integration of a systolic array based hardware accelerator into a dnn operator auto-tuning framework,” in Proceedings of the 2023 Workshop on Compilers, Deployment, and Tooling for Edge AI, ser. CODAI ’23. New York, NY, USA: Association for Computing Machinery, 2024, p. 21–26. [Online]. Available: https://doi.org/10.1145/3615338.3618130

C. Banbury, V. J. Reddi, P. Torelli, J. Holleman et al., “Mlperf tiny benchmark,” Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2021.

F. N. Peccia, S. Pavlitska, T. Fleck, and O. Bringmann, “Efficient edge AI: Deploying convolutional neural networks on fpga with the gemmini accelerator,” in 2024 27th Euromicro Conference on Digital System Design (DSD), 2024, pp. 418–426.

E. Baccour, N. Mhaisen, A. A. Abdellatif, A. Erbad et al., “Pervasive AI for IoT applications: A survey on resource-efficient distributed artificial intelligence,” IEEE Communications Surveys & Tutorials, vol. 24, no. 4, pp. 2366–2418, 2022.

Downloads

Published

2026-08-25

How to Cite

Peccia, F. N. ., Bhatnagar, A., & Bringmann, O. (2026). Autotuned Distribution of Multi-DNN Workloads on Multi-Accelerator SoCs. WiPiEC Journal - Works in Progress in Embedded Computing Journal, 12(2), 8. https://doi.org/10.64552/wipiec.v12i2.134