A Reconfigurable Computer Architecture Framework for Accelerating Parallel Matrix Computation in General-Purpose Processors
Keywords:
Reconfigurable computing, parallel matrix computation, computer architecture, general-purpose processor, coarse-grained reconfigurable architecture, CGRA, matrix multiplication, processing elementsAbstract
Matrix computation is a core and computationally demanding category of operations in today's computing landscape. Key operations such as matrix multiplication, matrix-vector multiplication, convolutional transformations, and factorizations are integral to fields like scientific computing, artificial intelligence, machine learning, computer vision, engineering simulation, graph analytics, signal processing, and high-performance computing. While modern general-purpose processors are equipped with multicore execution, SIMD/vector extensions, multilevel cache hierarchies, and multithreading capabilities, their static hardware design often fails to fully leverage the structural nuances of matrix workloads. Variations in matrix size, dimensionality, sparsity, data locality, precision needs, and parallelism patterns can result in suboptimal use of traditional execution resources. Recent advancements in coarse-grained reconfigurable architectures (CGRAs), FPGA-based accelerators, heterogeneous processors, and matrix-specific computing engines highlight that reconfigurability can effectively balance the flexibility of software with the efficiency of specialized hardware. Nonetheless, many current solutions are tailored to specific applications, necessitate distinct accelerator platforms, concentrate mainly on either dense or sparse computations, or offer limited integration with general-purpose processor execution environments. Recent reviews also pinpoint unresolved issues in reconfigurable matrix acceleration, such as memory access, workload imbalance, resource utilization, programmability, and adaptive design-space selection [1]–[4].
This study introduces a Reconfigurable Parallel Matrix Computing Framework (RPMCF), conceived as a closely integrated architectural extension to enhance matrix-oriented workloads on general-purpose processors. The framework integrates a host processor with a reconfigurable processing-element array, a configurable interconnection network, hierarchical local memory, an adaptive task scheduler, a matrix profiler, a runtime configuration manager, and a software programming interface. Rather than viewing matrix acceleration as a static fixed-function operation, the proposed framework dynamically adjusts the organization of processing elements, tiling strategy, execution mode, dataflow, memory allocation, and precision configuration based on workload characteristics. Both dense and sparse matrix workloads are analyzed prior to execution and assigned to different computational configurations, such as block-parallel, SIMD-style, pipelined reduction, and sparse-aware execution modes.
The research employs a design-and-evaluation approach that includes workload characterization, architectural modeling, RTL-based functional modeling, parallel execution simulation, benchmark testing, and comparative performance analysis. A practical prototype evaluation is conducted using representative square and rectangular dense matrices, along with structured sparse matrices, to evaluate execution time, throughput, speedup, processing-element utilization, memory efficiency, and estimated energy efficiency. The experimental findings suggest that the proposed framework enhances matrix-processing performance by dynamically aligning hardware resources with workload structure. For dense workloads, the architecture gains mainly from tiled data reuse, parallel multiply-accumulate pipelines, and minimized synchronization overhead. In the case of sparse and irregular workloads, adaptive work distribution and sparse-aware scheduling help reduce processing-element idle time.
The findings reveal that reconfigurable execution offers a viable compromise between fully general-purpose processors and highly specialized accelerators. In the tested configurations, the proposed framework achieves significant speedup over a baseline sequential processor model and noticeable improvement over a fixed parallel architecture, while retaining configurability across various matrix shapes and computational characteristics. The study provides a unified architectural model, adaptive mapping methodology, implementation workflow, and performance analysis for incorporating reconfigurable matrix acceleration into general-purpose computing systems. This framework is particularly pertinent to future heterogeneous processors where adaptable accelerators function alongside traditional CPU cores.
References
[1] Z. Hassan, A. Ometov, E. S. Lohan, and J. Nurmi, “Coarse-grained reconfigurable architectures for radio baseband processing: A survey,” Journal of Systems Architecture, vol. 154, Art. no. 103243, 2024. doi: 10.1016/j.sysarc.2024.103243.
[2] L. Liu, J. Zhu, Z. Li, Y. Lu, Y. Deng, J. Han, and S. Yin, “A survey of coarse-grained reconfigurable architecture and design: Taxonomy, challenges, and applications,” ACM Computing Surveys, vol. 52, no. 6, pp. 1–39, 2020. doi: 10.1145/3357375.
[3] Y. Liu, R. Chen, S. Li, J. Yang, S. Li, and B. da Silva, “FPGA-based sparse matrix multiplication accelerators: From state-of-the-art to future opportunities,” ACM Transactions on Reconfigurable Technology and Systems, vol. 17, no. 4, Art. no. 59, 2024. doi: 10.1145/3687480.
[4] A. S. Baroughi, M. B. Rajashekar, A. R. Baranwal, and Z. Fang, “HiSpMM: High performance high bandwidth sparse-dense matrix multiplication on HBM-equipped FPGAs,” ACM Transactions on Reconfigurable Technology and Systems, vol. 19, no. 1, Art. no. 10, 2026. doi: 10.1145/3774327.
[5] S. Kiefer, “Griddle: A novel hardware based matrix multiplier architecture,” M.S. thesis, California Polytechnic State University, San Luis Obispo, CA, USA, 2025.
[6] V. Isaac–Chassande, A. Evans, Y. Durand, and F. Rousseau, “Dedicated hardware accelerators for processing of sparse matrices and vectors: A survey,” ACM Transactions on Architecture and Code Optimization, vol. 21, no. 2, Art. no. 27, 2024. doi: 10.1145/3640542.
[7] K. Compton and S. Hauck, “Reconfigurable computing: A survey of systems and software,” ACM Computing Surveys, vol. 34, no. 2, pp. 171–210, 2002.
[8] R. Hartenstein, “A decade of reconfigurable computing: A visionary retrospective,” in Proceedings of the Conference on Design, Automation and Test in Europe, 2001, pp. 642–649.
[9] B. Mei, S. Vernalde, D. Verkest, H. De Man, and R. Lauwereins, “DRESC: A retargetable compiler for coarse-grained reconfigurable architectures,” in Proceedings of the IEEE International Conference on Field-Programmable Technology, 2002.
[10] H. Singh, M.-H. Lee, G. Lu, F. J. Kurdahi, N. Bagherzadeh, and E. M. C. Filho, “MorphoSys: An integrated reconfigurable system for data-parallel and computation-intensive applications,” IEEE Transactions on Computers, vol. 49, no. 5, pp. 465–481, 2000.
[11] J. M. P. Cardoso et al., “CGRAs: Architectures and compilation,” IEEE Micro, vol. 38, no. 2, pp. 58–65, 2018.
[12] K. Sankaralingam et al., “TRIPS: A polymorphous architecture for exploiting ILP, TLP, and DLP,” ACM Transactions on Architecture and Code Optimization, vol. 1, no. 1, pp. 62–85, 2004.
[13] J. W. Lee, M. C. Ng, and K. Asanović, “Vector-thread architecture,” in Proceedings of the International Symposium on Computer Architecture, 2006.
[14] J. Demmel, L. Grigori, M. Hoemmen, and J. Langou, “Communication-optimal parallel and sequential QR and LU factorizations,” SIAM Journal on Scientific Computing, vol. 34, no. 1, pp. A206–A239, 2012.
[15] E. Solomonik and J. Demmel, “Communication-optimal parallel 2.5D matrix multiplication and LU factorization algorithms,” in Proceedings of Euro-Par, 2011.
[16] A. Krishnan, K. Nagarajan, and R. K. Gupta, “Optimizing matrix multiplication using tiled parallel execution on heterogeneous architectures,” Journal of Parallel and Distributed Computing, vol. 169, pp. 1–14, 2023.
[17] M. K. Nandwana, S. Agarwal, and P. R. Panda, “Adaptive accelerator design for data-intensive computing: Opportunities and challenges,” IEEE Design & Test, vol. 40, no. 4, pp. 72–81, 2023.
[18] N. P. Jouppi et al., “In-datacenter performance analysis of a tensor processing unit,” in Proceedings of the 44th Annual International Symposium on Computer Architecture, 2017, pp. 1–12.
[19] Y. Chen, T. Krishna, J. Emer, and V. Sze, “Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks,” IEEE Journal of Solid-State Circuits, vol. 52, no. 1, pp. 127–138, 2017.
[20] M. Drumond, A. Daglis, N. Mirzadeh, D. Ustiugov, and B. Falsafi, “Coarse-grained reconfigurable arrays for dynamic acceleration,” IEEE Micro, vol. 40, no. 6, pp. 45–54, 2020.
[21] S. Mittal, “A survey of FPGA-based accelerators for convolutional neural networks,” Neural Computing and Applications, vol. 32, pp. 1109–1139, 2020.
[22] T. Hoefler and R. Belli, “Scientific benchmarking of parallel computing systems,” Proceedings of the IEEE, vol. 109, no. 7, pp. 1128–1148, 2021.
[23] J. Dongarra et al., “The international exascale software project roadmap,” International Journal of High Performance Computing Applications, vol. 25, no. 1, pp. 3–60, 2011.
[24] S. Williams, A. Waterman, and D. Patterson, “Roofline: An insightful visual performance model for multicore architectures,” Communications of the ACM, vol. 52, no. 4, pp. 65–76, 2009.
[25] J. L. Hennessy and D. A. Patterson, Computer Architecture: A Quantitative Approach, 6th ed. Cambridge, MA, USA: Morgan Kaufmann, 2019.
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 International Journal of Computer Engineering and Embedded Technologies

This work is licensed under a Creative Commons Attribution 4.0 International License.
Licensing Terms
© 2026
International Journal of Computer Engineering and Embedded Technologies.
This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.