Performance Evaluation of Cache Replacement Policies for Multi-Core Processor Architectures in Data-Intensive Computing Workloads

Authors

  • Mr. T. Narasimhulu BITS (A), Vizag, 530048 Author
  • Dr. B. Poorna Satyannarayana BITS (A), Vizag, 530048 Author

Keywords:

Cache replacement policy, Multi-core processors, Data-intensive computing, Cache hierarchy, Last-level cache

Abstract

The swift expansion of data-heavy applications has markedly heightened the performance requirements on contemporary multi-core processor designs. Applications like large-scale analytics, graph processing, machine learning, database management, scientific computing, and cloud services generate a significant amount of memory traffic and often work with data sets that surpass the capacity of on-chip caches. As a result, efficient cache management has become a crucial factor in determining overall system performance. Among the various cache management strategies, cache replacement policies are particularly vital as they determine which cache block should be removed to make room for new data in a full cache. Traditional policies such as First-In-First-Out (FIFO), Least Recently Used (LRU), and random replacement are still widely adopted due to their straightforward concepts and low implementation complexity. However, workloads that are data-intensive and run on multi-core systems often display characteristics like irregular memory access patterns, high cache contention, inter-core interference, data sharing, streaming accesses, and rapidly changing reuse behavior, which can diminish the effectiveness of conventional replacement strategies.

This study offers a thorough performance assessment of cache replacement policies for multi-core processor architectures under typical data-intensive workloads. The research contrasts traditional policies, including FIFO, LRU, Pseudo-LRU (PLRU), and Random replacement, with more adaptive and reuse-aware strategies such as Segmented LRU (SLRU), Re-Reference Interval Prediction (RRIP)-based replacement, and Dynamic RRIP (DRRIP). The proposed methodology utilizes a controlled multi-core simulation framework where processor cores, private cache levels, and a shared last-level cache are set up to execute workload categories with varying memory behaviors. Performance metrics include cache hit rate, miss rate, average memory access time, instructions per cycle, execution time, memory traffic, and inter-core fairness. Realistic experimental data are generated across different core counts and workload classes to illustrate how replacement decisions affect processor efficiency.

The results reveal that no single replacement policy is universally superior. LRU is effective for workloads with strong temporal locality but struggles when streaming or scan-intensive accesses fill the cache. FIFO and Random policies are simpler to implement but generally result in higher miss rates for locality-sensitive workloads. PLRU offers a practical approximation to LRU but may select suboptimal victims under complex access patterns. RRIP and DRRIP show better adaptability, especially for mixed and data-intensive workloads with both high-reuse and low-reuse cache blocks. The findings indicate that adaptive reuse-aware replacement can enhance average last-level cache hit rates by about 8–17% compared to baseline FIFO and Random policies and can decrease average memory access time and execution cycles in high-contention multi-core environments.

The study highlights a significant research gap in the joint evaluation of replacement policies across diverse workloads, increasing core counts, shared-cache contention, and performance fairness. It proposes a systematic evaluation framework that connects cache behavior with workload characteristics and multi-core interference. The research concludes that future cache systems should increasingly adopt adaptive, workload-aware, and hardware-efficient replacement mechanisms. Future research opportunities include machine-learning-guided replacement, phase-aware adaptation, quality-of-service-aware cache management, energy-sensitive replacement, chiplet-based cache hierarchies, and replacement mechanisms optimized for heterogeneous CPU, GPU, and AI accelerator architectures.

Author Biographies

  • Mr. T. Narasimhulu, BITS (A), Vizag, 530048

    Department of CSE

  • Dr. B. Poorna Satyannarayana, BITS (A), Vizag, 530048

    Department of CSE

References

1. Belady, L. A. (1966). A study of replacement algorithms for a virtual-storage computer. IBM Systems Journal, 5(2), 78–101.

2. Mattson, R. L., Gecsei, J., Slutz, D. R., & Traiger, I. L. (1970). Evaluation techniques for storage hierarchies. IBM Systems Journal, 9(2), 78–117.

3. Smith, A. J. (1982). Cache memories. ACM Computing Surveys, 14(3), 473–530.

4. Jouppi, N. P. (1990). Improving direct-mapped cache performance by the addition of a small fully-associative cache and prefetch buffers. In Proceedings of the 17th Annual International Symposium on Computer Architecture.

5. Kessler, R. E. (1991). The Alpha 21264 microprocessor. IEEE Micro, 19(2), 24–36.

6. Hennessy, J. L., & Patterson, D. A. (2019). Computer Architecture: A Quantitative Approach (6th ed.). Morgan Kaufmann.

7. Qureshi, M. K., Jaleel, A., Patt, Y. N., Steely, S. C., & Emer, J. (2007). Adaptive insertion policies for high performance caching. In Proceedings of the 34th Annual International Symposium on Computer Architecture.

8. Jaleel, A., Theobald, K. B., Steely, S. C., & Emer, J. (2010). High performance cache replacement using re-reference interval prediction (RRIP). In Proceedings of the 37th Annual International Symposium on Computer Architecture.

9. Qureshi, M. K., & Patt, Y. N. (2006). Utility-based cache partitioning: A low-overhead, high-performance, runtime mechanism to partition shared caches. In Proceedings of the 39th Annual IEEE/ACM International Symposium on Microarchitecture.

10. Qureshi, M. K., & Patt, Y. N. (2007). Utility-based cache partitioning: A low-overhead, high-performance, runtime mechanism to partition shared caches. IEEE Transactions on Parallel and Distributed Systems, 20(9), 1266–1277.

graph

Published

2026-07-26

How to Cite

Performance Evaluation of Cache Replacement Policies for Multi-Core Processor Architectures in Data-Intensive Computing Workloads. (2026). International Journal of Computer Engineering and Embedded Technologies, 1(1), 1-37. https://ijceet.com/index.php/ijceet/article/view/performance-evaluation-cache-replacement-policies-multi-core-pro