Instruction Flow Amplifier: A Novel Architecture for In-Order Processor Front-End Optimization

Authors

  • João Antônio Temochko Andre Department of Electrical Engineering, Centro Universitário FEI
  • Alexandre Beletti Ferreira Department of Informatics, Federal Institute of Sao Paulo

DOI:

https://doi.org/10.64552/wipiec.v12i2.139

Keywords:

Computer Architecture, RISC-V, ARM, Front-End, Gem5, In-Order Processor cores

Abstract

This paper proposes the Instruction Flow Amplifier (IFA), a novel architecture designed to mitigate frontend bottlenecks in in-order processor cores. Using the Gem5 simulator, we evaluate Cycles Per Instruction (CPI), Instructions Per Cycle (IPC), branch prediction accuracy, and instruction cache misses to demonstrate front-end bottlenecks and quantify the performance gains achieved by the proposed architecture. Unlike traditional trace caches that primarily focus on instruction storage, the proposed IFA integrates a lightweight Dependency Analysis Unit (DAU) to enable pseudo-out-of-order dispatch within an energy-efficient in-order infrastructure. The experimental results demonstrate that the IFA acts as an effective instruction supply accelerator. In workloads dominated by tight loops (MatMult), the proposed architecture effectively saturates the pipeline, achieving a 28% IPC improvement. In control-heavy workloads (WikiSort), the IFA mitigates fetch latency through aggressive buffering, improving IPC by 13.4% with negligible impact on branch prediction accuracy.

Author Biographies

João Antônio Temochko Andre, Department of Electrical Engineering, Centro Universitário FEI

Department of Electrical Engineering, Centro Universitário FEI, São Bernardo do Campo, Brazil.

Alexandre Beletti Ferreira, Department of Informatics, Federal Institute of Sao Paulo

Department of Informatics, Federal Institute of Sao Paulo, Sao Paulo, Brazil.

References

C. Kaynak, B. Grot, and B. Falsafi, “Confluence: Unified instruction supply for scale-out servers,” in Proceedings of the 48th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO-48), 2015.

M. Ferdman, T. F. Wenisch, A. Ailamaki, B. Falsafi, and A. Moshovos, “Temporal instruction fetch streaming,” in Proceedings of the 41st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO-41), 2008.

M. B. Breughe, S. Eyerman, and L. Eeckhout, “Mechanistic analytical modeling of superscalar in-order processor performance,” ACM Transactions on Architecture and Code Optimization (TACO), vol. 11, no. 4, pp. 50:1–50:26, 2014.

C. Arul Rathi, G. Rajakumar, T. Ananth Kumar, and T. S. Arun Samuel, “Design and development of an efficient branch predictor for an in-order risc-v processor,” Journal of Nano- and Electronic Physics, vol. 12, no. 5, pp. 05 021–1–05 021–4, 2020.

G. Reinman, B. Calder, and T. Austin, “Fetch directed instruction prefetching,” in Proceedings of the 26th Annual International Symposium on Computer Architecture (ISCA), ser. ISCA ’99. IEEE Computer Society, 1999, pp. 16–27.

W. A. Wulf and S. A. McKee, “Hitting the memory wall: implications of the obvious,” ACM SIGARCH Computer Architecture News, vol. 23, no. 1, pp. 20–24, 1995.

J.-L. Baer and T.-F. Chen, “Effective hardware-based data prefetching for high-performance processors,” IEEE Transactions on Computers, vol. 44, no. 5, pp. 609–623, 1995.

G. A. Chacon, “Processor memory system design for performance and security,” Dissertation, Texas A&M University, College Station, TX, 2023.

O. J. Santana, A. Falcón, A. Ramirez, and M. Valero, “Dia: A complexity-effective decoding architecture,” IEEE Transactions on Computers, vol. 58, no. 4, pp. 449–462, 2009.

G. Reinman, B. Calder, and T. Austin, “Optimizations enabled by a decoupled front-end architecture,” IEEE Transactions on Computers, vol. 50, no. 4, pp. 338–355, 2001.

P. S. Oberoi, “Out-of-order front-ends,” Preliminary Proposal, University of Wisconsin-Madison, Madison, WI, 2001.

O. Mutlu, J. Stark, C. Wilkerson, and Y. N. Patt, “Runahead execution: An effective alternative to large instruction windows,” IEEE Micro, vol. 23, no. 6, pp. 20–25, 2003.

R. Shioya, M. Goshima, and H. Ando, “A front-end execution architecture for high energy efficiency,” in Proceedings of the 47th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2014, pp. 419–431.

N. Binkert et al., “The gem5 simulator,” in ACM SIGARCH Computer Architecture News, vol. 39, no. 2, 2011, pp. 1–7.

M. R. Guthaus, J. S. Ringenberg, D. Ernst, T. M. Austin, T. Mudge, and R. B. Brown, “Mibench: A free, commercially representative embedded benchmark suite,” in Proceedings of the 4th Annual IEEE International Workshop on Workload Characterization (WWC-4). IEEE, 2001, pp. 3–14.

J. L. Hennessy and D. A. Patterson, Computer Architecture: A Quantitative Approach, 5th ed. Morgan Kaufmann, 2012.

G. M. Amdahl, “Validity of the single processor approach to achieving large scale computing capabilities,” in Proceedings of the April 18-20, 1967, spring joint computer conference. ACM, 1967, pp. 483–485.

S. Mashimo, A. Fujita, R. Matsuo, S. Akaki, A. Fukuda, T. Koizumi, J. Kadomoto, H. Irie, M. Goshima, K. Inoue, and R. Shioya, “An open source fpga-optimized out-of-order risc-v soft processor,” in International Conference on Field-Programmable Technology (ICFPT), 2019.

M. McKeown, J. Balkind, and D. Wentzlaff, “Execution drafting: Energy efficiency through computation deduplication,” in Proceedings of the 47th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2014, pp. 432–444.

R. Huerta et al., “Socgpu: Simple out-of-order core for gpgpus,” in Proceedings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2023.

A. Miyoshi, C. Lefurgy, E. Hensbergen, R. Rajamony, and R. Rajwar, “Critical power slope: understanding the runtime-leakage tradeoff in deep submicron technologies,” in Proceedings of the 16th international conference on Supercomputing, 2002, pp. 35–44.

A. Bardizbanyan, M. Själander, and P. Larsson-Edefors, “Reconfigurable instruction decoding for a wide-control-word processor,” in 2011 IEEE International Symposium on Parallel and Distributed Processing Workshops and Phd Forum (IPDPSW), 2011.

A. Djupdal, M. Själander, M. Jahre, S. Aunet, and T. Ytterdal, “Optimizing energy efficiency in subthreshold risc-v cores,” in ResearchGate Preprint, 2020.

Downloads

Published

2026-08-25

How to Cite

Temochko Andre, J. A., & Beletti Ferreira, A. (2026). Instruction Flow Amplifier: A Novel Architecture for In-Order Processor Front-End Optimization. WiPiEC Journal - Works in Progress in Embedded Computing Journal, 12(2), 7. https://doi.org/10.64552/wipiec.v12i2.139