IP Library › Granted Patent US 12,334,183
Granted Patent B2
US 12,334,183 · App. 18/306,531 · Granted Jun 17, 2025

Memory devices including processing-in-memory architecture configured to provide accumulation dispatching and hybrid partitioning

Inventors: Kevin Skadron (Charlottesville, VA); Marzieh Lenjani (Champaign, IL)
Assignee: UNIVERSITY OF VIRGINIA PATENT FOUNDATION
G11C7/1039G06F17/16G11C7/1057
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,334,183
App. No.
18/306,531
Granted
Jun 17, 2025
Kind
B2
Abstract

An integrated circuit memory device can include a plurality of banks of memory, each of the banks of memory including a first pair of sub-arrays comprising first and second sub-arrays, the first pair of sub-arrays configured to store data in memory cells of the first pair of sub-arrays, a first row buffer memory circuit located in the integrated circuit memory device adjacent to the first pair of sub-arrays and configured to store first row data received from the first pair of sub-arrays and configured to transfer the row data into and/or out of the first row buffer memory circuit, and a first sub-array level processor circuit in the integrated circuit memory device adjacent to the first pair of sub-arrays and operatively coupled to the first row data, wherein the first sub-array level processor circuit is configured to perform column oriented processing a sparse matrix kernel stored, at least in-part, in the first pair of sub-arrays, with input vector values stored, at least in part, in the first pair of sub-arrays to provide output vector values representing products of values stored in columns of the sparse matrix kernel with the input vector values.

Claims (23)

1. An integrated circuit memory device comprising:

a plurality of banks of memory, each of the banks of memory including:

a first pair of sub-arrays comprising first and second sub-arrays, the first pair of sub-arrays configured to store data in memory cells of the first pair of sub-arrays;

a first row buffer memory circuit located in the integrated circuit memory device adjacent to the first pair of sub-arrays and configured to store first row data received from the first pair of sub-arrays and configured to transfer the first row data into and/or out of the first row buffer memory circuit; and

a first sub-array level processor circuit in the integrated circuit memory device adjacent to the first pair of sub-arrays and operatively coupled to the first row data, wherein the first sub-array level processor circuit is configured to perform column oriented processing a sparse matrix kernel stored, at least in-part, in the first pair of sub-arrays, with input vector values stored, at least in part, in the first pair of sub-arrays to provide output vector values representing products of values stored in columns of the sparse matrix kernel with the input vector values.

2. The integrated circuit memory device of claim 1 wherein the values in long columns of the sparse matrix kernel and the output vector values representing accumulated values of the values in the long columns activated by the input vector values, are partitioned within the memory device so that both are stored in the same sub-array of the memory.

3. The integrated circuit memory device of claim 2 wherein short columns of the sparse matrix kernel are not partitioned within the memory device so that both are stored in the same sub-array of the memory.

4. The integrated circuit memory device of claim 2 wherein accumulation operations for the values in the long columns activated by the input vector values partitioned within the same sub-array comprise local accumulations and wherein the accumulation operations for the values in the long columns activated by the input vector values not partitioned within the same sub-array comprise local accumulations comprise remote accumulations.

5. The integrated circuit memory device of claim 4 further comprising:

a pair of dispatch sub-arrays configured to store remote accumulations to be dispatched to a remote level processor circuit including a remote sub-array level processor circuit for accumulation;

a dispatch sub-array level processor circuit operatively coupled to the pair of dispatch sub-arrays and configured to dispatch the remote accumulations from the pair of dispatch sub-arrays to the remote sub-array level processor circuit for accumulation responsive to an indication that the remote sub-array level processor circuit is idle.

6. The integrated circuit memory device of claim 5 further comprising:

indexing registers configured to represent an index value for the sparse matrix kernel beyond which accumulation operations are determined to be remote accumulations.

7. The integrated circuit memory device of claim 6 wherein the plurality of banks comprise a plurality of layers in a 3D stack.

8. The integrated circuit memory device of claim 7 wherein at least one layer in the plurality of layers comprises a logic layer configured to store the remote accumulations.

9. The integrated circuit memory device of claim 2 wherein the input vector values that activated the values in the long columns of the sparse matrix kernel are stored in the first sub-array of the memory.

10. The integrated circuit memory device of claim 7 further comprising:

a through-via extending through the plurality of layers to electrically couple data to/from the plurality of banks located among the plurality of layers.

11. The integrated circuit memory device of claim 1 wherein the first sub-array level processor circuit includes:

a control circuit operatively coupled to the first pair of sub-arrays, to the first row buffer memory circuit, and to the first sub-array level processor circuit and configured to select inputs to first and second inputs to the first sub-array level processor circuit and to provide instruction an ALU circuit included in the control circuit to operate on data in the first row buffer memory circuit.

12. The integrated circuit memory device of claim 1 wherein the first sub-array level processor circuit is configured to shift row data for the first pair of sub-arrays from or to the first row buffer memory circuit using a hot-one encoded value to randomly access a column of data stored in the first row buffer memory circuit.

13. An integrated circuit memory device comprising:

a sub-array level processor circuit in the integrated circuit memory device located adjacent to a pair of sub-arrays of the memory device and configured to perform column oriented processing on a sparse matrix kernel stored, at least in-part, in the pair of sub-arrays, with input vector values stored, at least in part, in the pair of sub-arrays to provide output vector values representing products of values stored in columns of the sparse matrix kernel with the input vector values.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2024
From: UNIVERSITY OF VIRGINIA
To: UNIVERSITY OF VIRGINIA PATENT FOUNDATION
Reel/Frame 066747/0963 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: SKADRON, KEVIN; LENJANI, MARZIEH
To: UNIVERSITY OF VIRGINIA
Reel/Frame 064915/0812 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2023
From: SKADRON, KEVIN; LENJANI, MARZIEH
To: UNIVERSITY OF VIRGINIA
Reel/Frame 063906/0283 →
Continuity (2)
Provisional Application 63334844 · Apr 26, 2022
Related Publication 20230343373A1 · Oct 26, 2023
References Cited (21)
US 20230385562A1 · Lee · 2023 [cited by examiner]
US 20240143199A1 · Poremba · 2024 [cited by examiner]
Ahn, Junwhan, et al. “A scalable processing-in-memory accelerator for parallel graph processing.” Proceedings of the 42nd Annual International Symposium on Computer Architecture. 2015. [cited by applicant]
Challapalle, Nagadastagiri, et al. “GaaS-X: Graph analytics accelerator supporting sparse data representation using crossbar architectures.” 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (IS… [cited by applicant]
Dai, Guohao, et al. “ForeGraph: Exploring large-scale graph processing on multi-FPGA architecture.” Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays. 2017. [cited by applicant]
Dai, Guohao, et al. “Graphh: A processing-in-memory architecture for large-scale graph processing.” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 38.4 (2018): 640-653. [cited by applicant]
Gui, Chuang-Yi, et al. “A survey on graph processing accelerators: Challenges and opportunities.” Journal of Computer Science and Technology 34 (2019): 339-371. [cited by applicant]
Hajinazar, Nastaran, et al. “SIMDRAM: a framework for bit-serial SIMD processing using DRAM.” Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems… [cited by applicant]
Ham, Tae Jun, et al. “Graphicionado: A high-performance and energy-efficient accelerator for graph analytics.” 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 2016. [cited by applicant]
He, Mingxuan, et al. “Newton: A DRAM-maker's accelerator-in-memory (AiM) architecture for machine learning.” 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (Micro). IEEE, 2020. [cited by applicant]
Kim, Ytong-Bin, and Tom W. Chen. “Assessing merged DRAM/logic technology.” Integration 27.2 (1999): 179-194. [cited by applicant]
Leng, Jingwen, et al. “GPUWattch: Enabling energy optimizations in GPGPUs.” ACM SIGARCH Computer Architecture News 41.3 (2013): 487-498. [cited by applicant]
Lenjani, Marzieh, et al. “Fulcrum: A simplified control and access mechanism toward flexible and practical in-situ accelerators.” 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE,… [cited by applicant]
Nai, Lifeng, et al. “Graphpim: Enabling instruction-level pim offloading in graph computing frameworks.” 2017 IEEE International symposium on high performance computer architecture (HPCA). IEEE, 2017. [cited by applicant]
Nurvitadhi, Eriko, et al. “Hardware accelerator for analytics of sparse data.” 2016 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 2016. [cited by applicant]
Ozdal, Muhammet Mustafa, et al. “Energy efficient architecture for graph analytics accelerators.” ACM SIGARCH Computer Architecture News 44.3 (2016): 166-177. [cited by applicant]
Pal, Subhankar, et al. “Outerspace: An outer product based sparse matrix multiplication accelerator.” 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 2018. [cited by applicant]
Sadredini, Elaheh, et al. “Sunder: Enabling low-overhead and scalable near-data pattern matching acceleration.” MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture. 2021. [cited by applicant]
Xie, Xinfeng, et al. “SpaceA: Sparse matrix vector multiplication on processing-in-memory accelerator.” 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 2021. [cited by applicant]
Yang, Carl, Yangzihao Wang, and John D. Owens. “Fast sparse matrix and sparse vector multiplication algorithm on the GPU.” 2015 IEEE International Parallel and Distributed Processing Symposium Workshop. IEEE, 2015. [cited by applicant]
Zhou, Minxuan, et al. “Gram: graph processing in a reram-based computational memory.” IEEE Asia and South Pacific Design Automation Conference. 2019. [cited by applicant]