Dolunay: Architectural Support for Independent Thread Scheduling in a RISC-V SIMT Accelerator

Authors

  • Ahmet Zahit Can Computer Engineering, A. I. and Data Engineering Yıldız Technical University
  • Erkan Uslu Computer Engineering, Yıldız Technical University

DOI:

https://doi.org/10.64552/wipiec.v12i2.137

Keywords:

Independent Thread Scheduling, SIMT, GPGPU, RISC-V, Forward-progress guarantee, Barrier synchronization, Cooperative multitasking

Abstract

Single-Instruction Multiple-Thread (SIMT) architectures have revolutionized data-parallel computing by providing a high-throughput abstraction that simplifies vector management. However, traditional stack-based SIMT models do not support intra-warp synchronization primitives such as mutexes and spin-locks. This work introduces Dolunay, a RISC-V-based Independent Thread Scheduling (ITS) SIMT accelerator. By only adding three custom instructions, Dolunay employs a cooperative multitasking model and explicit synchronization barriers at the hardware-level, and provides the forward-progress guarantees necessary to implement starvation-free algorithms. We evaluate Dolunay on the Cmod A7-35T FPGA module and demonstrate its ability to correctly execute kernels that deadlock on traditional stack-based architectures while still achieving parallel execution for conventional compute-heavy kernels.

Author Biographies

Ahmet Zahit Can, Computer Engineering, A. I. and Data Engineering Yıldız Technical University

Computer Engineering, A. I. and Data Engineering, Yıldız Technical University, İstanbul, Türkiye.

Erkan Uslu, Computer Engineering, Yıldız Technical University

Computer Engineering, Yıldız Technical University, İstanbul, Türkiye.

References

E. Lindholm, J. Nickolls, S. Oberman, and J. Montrym, “Nvidia tesla: A unified graphics and computing architecture,” Micro, IEEE, vol. 28, pp. 39 – 55, 04 2008.

J. Nickolls and W. Dally, “The gpu computing era,” Micro, IEEE, vol. 30, pp. 56 – 69, 05 2010.

J. Nickolls, I. Buck, M. Garland, and K. Skadron, “Scalable parallel programming with cuda,” Queue , vol. 6, pp. 40–53, 03 2008.

J. Owens, M. Houston, D. Luebke, S. Green, J. Stone, and J. Phillips, “Gpu computing,” Proceedings of the IEEE , vol. 96, pp. 879–899, 05 2008.

NVIDIA Corporation, “Nvidia tesla v100 gpu architecture: The world’s most advanced data center gpu,” https://images.nvidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper.pdf, accessed: 2026-03-28.

Z. Jia, M. Maggioni, B. Staiger, and D. P. Scarpazza, “Dissecting the nvidia volta gpu architecture via microbenchmarking,” 2018. [Online]. Available: https://arxiv.org/abs/1804.06826

F. Elsabbagh, B. Tine, P. Roshan, E. Lyons, E. Kim, D. E. Shim, L. Zhu, S. K. Lim, and H. kim, “Vortex: Opencl compatible risc-v gpgpu,” 2020. [Online]. Available: https://arxiv.org/abs/2002.12151

S. Machetti, P. D. Schiavone, L. Orlandic, D. Huang, D. Kasap, G. Ansaloni, and D. Atienza, “e-gpu: An open-source and configurable risc-v graphic processing unit for tinyai applications,” 2026. [Online]. Available: https://arxiv.org/abs/2505.08421

C. Collange, “Simty: generalized SIMT execution on RISC-V,” in First Workshop on Computer Architecture Research with RISC-V , ser. First Workshop on Computer Architecture Research with RISC-V, vol. 6, Boston, United States, Oct. 2017, p. 6. [Online]. Available: https://inria.hal.science/hal-01622208

M. Naylor, A. Joannou, A. Markettos, P. Metzger, S. Moore, and T. Jones, “Advanced dynamic scalarisation for risc-v gpgpus,” 11 2024, pp. 260–267.

W. Fung, I. Sham, G. Yuan, and T. Aamodt, “Dynamic warp formation and scheduling for efficient gpu control flow,” 01 2008, pp. 407–420.

A. Tino, C. Collange, and A. Seznec, “Simt-x: Extending single- instruction multi-threading to out-of-order cores,” ACM Trans. Archit. Code Optim. , vol. 17, no. 2, May 2020. [Online]. Available: https://doi.org/10.1145/3392032

M. Nemirovsky and D. M. Tullsen, Fine-Grain Multithreading . Cham: Springer International Publishing, 2013, pp. 25–31. [Online]. Available: https://doi.org/10.1007/978-3-031-01738-4_4

J. Liu, Q. Li, Y. Luo, H. Zhang, J. Lu, S. Fan, J. Ye, Y. Liu, X. Liu, Y. Yang, Z. Ye, Y. Zeng, A. Shen, R. Huang, W. Cong, X. Zou, and M. Gao, “Titan-i: An open-source, high performance risc-v vector core,” in Proceedings of the 58th IEEE/ACM International Symposium on Microarchitecture , ser. MICRO ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 675–690. [Online]. Available: https://doi.org/10.1145/3725843.3756059

B. Tine, K. P. Yalamarthy, F. Elsabbagh, and K. Hyesoon, “Vortex: Extending the risc-v isa for gpgpu and 3d-graphics,” in MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , ser. MICRO ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 754–766. [Online]. Available: https://doi.org/10.1145/3466752.3480128

Downloads

Published

2026-08-25

How to Cite

Can, A. Z., & Erkan Uslu, E. (2026). Dolunay: Architectural Support for Independent Thread Scheduling in a RISC-V SIMT Accelerator. WiPiEC Journal - Works in Progress in Embedded Computing Journal, 12(2), 8. https://doi.org/10.64552/wipiec.v12i2.137