IP Library › Granted Patent US 12,688,123
Granted Patent B2
US 12,688,123 · App. 18/622,245 · Granted Jul 21, 2026

Arithmetic logic unit (ALU) in a base die of a processing-in-memory component with cross-ALU data communication capability

Inventors: Vignesh Adhinarayanan (Austin, TX); Hyung-Dong Lee (Austin, TX)
Assignee: Advanced Micro Devices, Inc.
G06F12/0802G06F2212/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,123
App. No.
18/622,245
Granted
Jul 21, 2026
Kind
B2
Abstract

A system includes memory hardware including a memory and a processing-in-memory (PIM) component. A system includes a host including at least one core. The PIM component includes a memory die (e.g., a dynamic random access memory (DRAM) die) and a base die, e.g., a logic die. The base die includes one or more PIM arithmetic logic units (ALU) and one or more sense amplifiers. In at least some implementations the base die includes a shared static random-access memory (SRAM) cache that is shared between different ALU that reside in the base die.

Claims (33)

1 . A system comprising:

a host including at least one core; and

a processing-in-memory component communicatively coupled to the host, the processing-in-memory component comprising:

a memory; and

a base die operatively coupled to the memory, the base die comprising:

one or more arithmetic logic units; and

a static random-access memory cache configured to be directly accessed by the one or more arithmetic logic units, wherein the one or more arithmetic logic units and the static random-access memory cache are positioned within the base die of the processing-in-memory component, wherein the base die further comprises a multiplexer configured to route a first type of data operation to the one or more arithmetic logic units and a second type of data operation to the static random-access memory cache.

2 . The system of claim 1 , wherein the memory comprises a dynamic random-access memory die.

3 . The system of claim 1 , wherein the base die comprises a logic die of the processing-in-memory component.

4 . The system of claim 1 , wherein the base die further comprises one or more sense amplifiers communicatively coupled to the one or more arithmetic logic units.

5 . The system of claim 1 , wherein the base die further comprises multiple arithmetic logic units and wherein the multiple arithmetic logic units are configured to share the static random-access memory cache.

6 . The system of claim 5 , wherein the multiple arithmetic logic units are configured to share the static random-access memory cache to access data from a remote memory bank.

7 . The system of claim 1 , wherein the base die further comprises one or more sense amplifiers communicatively coupled to the one or more arithmetic logic units.

8 . A processing-in-memory component comprising:

a memory die comprising multiple memory banks; and

a base die communicatively coupled to the memory die and comprising:

one or more arithmetic logic units;

a static random-access memory cache; and

a multiplexer configured to route a first type of data operation to the one or more arithmetic logic units and a second type of data operation to the static random-access memory cache.

9 . The processing-in-memory component of claim 8 , wherein the memory die comprises a dynamic random-access memory die.

10 . The processing-in-memory component of claim 8 , wherein the base die comprises a logic die of the processing-in-memory component.

11 . The processing-in-memory component of claim 8 , wherein the base die further comprises one or more sense amplifiers communicatively coupled to the one or more arithmetic logic units.

12 . The processing-in-memory component of claim 8 , wherein the base die further comprises multiple arithmetic logic units, and wherein the multiple arithmetic logic units are configured to share the static random-access memory cache.

13 . The processing-in-memory component of claim 12 , wherein the multiple arithmetic logic units are configured to share the static random-access memory cache to access data from a remote memory bank.

14 . The processing-in-memory component of claim 8 , wherein the base die further comprises one or more sense amplifiers communicatively coupled to the one or more arithmetic logic units.

15 . A method comprising:

prefetching remote data from a remote memory bank into a static random-access memory cache located in a base die of a processing-in-memory component; and

initiating, using at least some of the remote data, a sequence of processing-in-memory operations via one or more arithmetic logic units located in the base die of the processing-in-memory component, wherein the base die further comprises a multiplexer configured to route a first type of data operation to the one or more arithmetic logic units and a second type of data operation to the static random-access memory cache.

16 . The method of claim 15 , further comprising locking a cache line of the static random-access memory cache prior to initiating the sequence of processing-in-memory operations.

17 . The method of claim 16 , further comprising explicitly releasing the locked cache line after initiating the sequence of processing-in-memory operations.

18 . The method of claim 15 , wherein the prefetching the remote data from the remote memory bank into the static random-access memory cache is based at least in part on one or more prefetch indicators accompanying at least one of a read command or a write command to the processing-in-memory component.

19 . The method of claim 15 , wherein the base die further comprises one or more sense amplifiers communicatively coupled to the one or more arithmetic logic units.

20 . The method of claim 15 , wherein the base die comprises a logic die of the processing-in-memory component.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2024
From: ADHINARAYANAN, VIGNESH; LEE, HYUNG-DONG
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 067004/0845 →
Continuity (1)
Related Publication 20250307142A1 · Oct 2, 2025
References Cited (87)
US 5940850A · Harish et al. · 1999 [cited by applicant]
US 6118723A · Agata et al. · 2000 [cited by applicant]
US 8161245B2 · Fields, Jr. et al. · 2012 [cited by applicant]
US 8706969B2 · Anderson et al. · 2014 [cited by applicant]
US 8947931B1 · D'Abreu · 2015 [cited by applicant]
US 9450023B1 · Konevecki et al. · 2016 [cited by applicant]
US 9582282B2 · Hayenga et al. · 2017 [cited by applicant]
US 11152056B1 · Seo et al. · 2021 [cited by applicant]
US 11251155B2 · Lee et al. · 2022 [cited by applicant]
US 11487447B2 · Islam et al. · 2022 [cited by applicant]
US 11853220B2 · Bondarenko et al. · 2023 [cited by applicant]
US 20030009632A1 · Arimilli · 2003 [cited by examiner]
US 20050198439A1 · Lange et al. · 2005 [cited by applicant]
US 20080054489A1 · Farrar et al. · 2008 [cited by applicant]
US 20080232185A1 · Bartley et al. · 2008 [cited by applicant]
US 20120087183A1 · Chang · 2012 [cited by applicant]
US 20120221785A1 · Chung · 2012 [cited by examiner]
US 20130229846A1 · Chien et al. · 2013 [cited by applicant]
US 20140019689A1 · Cain, III et al. · 2014 [cited by applicant]
US 20150063039A1 · Chen et al. · 2015 [cited by applicant]
US 20150325290A1 · Lasser et al. · 2015 [cited by applicant]
US 20160188476A1 · Yu et al. · 2016 [cited by applicant]
US 20170154667A1 · Dally · 2017 [cited by applicant]
US 20170160955A1 · Jayasena et al. · 2017 [cited by applicant]
US 20170263306A1 · Murphy · 2017 [cited by examiner]
US 20170353576A1 · Guim Bernat et al. · 2017 [cited by applicant]
US 20170358328A1 · Barkley et al. · 2017 [cited by applicant]
US 20180024948A1 · Tsai et al. · 2018 [cited by applicant]
US 20180307595A1 · Leidel et al. · 2018 [cited by applicant]
US 20190043560A1 · Sumbul et al. · 2019 [cited by applicant]
US 20190213130A1 · Madugula et al. · 2019 [cited by applicant]
US 20200042197A1 · Kotra et al. · 2020 [cited by applicant]
US 20200185370A1 · Juengling · 2020 [cited by applicant]
US 20200278923A1 · Jacob et al. · 2020 [cited by applicant]
US 20210294741A1 · Yudanov · 2021 [cited by examiner]
US 20210303307A1 · Olson · 2021 [cited by examiner]
US 20210303355A1 · Nag et al. · 2021 [cited by applicant]
US 20210406183A1 · Mashimo et al. · 2021 [cited by applicant]
US 20220019442A1 · Yudanov · 2022 [cited by examiner]
US 20220164297A1 · Sity · 2022 [cited by examiner]
US 20220223196A1 · He et al. · 2022 [cited by applicant]
US 20230091205A1 · Moga et al. · 2023 [cited by applicant]
US 20230195314A1 · Yoshihara et al. · 2023 [cited by applicant]
US 20230205693A1 · Kotra · 2023 [cited by examiner]
US 20230222064A1 · Reed · 2023 [cited by applicant]
US 20230229596A1 · Shulyak et al. · 2023 [cited by applicant]
US 20230284459A1 · Yoo · 2023 [cited by applicant]
US 20230326530A1 · Chen et al. · 2023 [cited by applicant]
US 20230376234A1 · Delacruz et al. · 2023 [cited by applicant]
US 20240004786A1 · Adhinarayanan et al. · 2024 [cited by applicant]
US 20240008259A1 · Sharma et al. · 2024 [cited by applicant]
US 20240028516A1 · Maroncelli et al. · 2024 [cited by applicant]
US 20240088099A1 · Prasad et al. · 2024 [cited by applicant]
US 20240395289A1 · Adhinarayanan et al. · 2024 [cited by applicant]
US 20240404598A1 · Hsu · 2024 [cited by applicant]
US 20250245160A1 · Scrbak et al. · 2025 [cited by applicant]
Panda, et al. “Prefetching Techniques for Near-memory Throughput Processors” Jun. 3, 2016, ICS '16: Proceedings of the 2016 International Conference on Supercomputing, https://dl.acm.org/doi/10.1145/2925426.2926282 (Yea… [cited by examiner]
Xu et al. “PIMCH: Cooperative Memory Prefetching in Processing-In-Memory Architecture” Feb. 22, 2018, 2018 23rd Asia and South Pacific Design Automation Conference (ASP-DAC), https://ieeexplore.ieee.org/abstract/documen… [cited by examiner]
Hughes et al. “Memory-side prefetching for linked data structures for processor-in-memory systems” Feb. 1, 2005, Journal of Parallel and Distributed Computing, https://www.sciencedirect.com/science/article/pii/S07437315… [cited by examiner]
U.S. Appl. No. 17/855,157 , “Final Office Action”, U.S. Appl. No. 17/855,157, Aug. 15, 2024, 16 pages. [cited by applicant]
U.S. Appl. No. 17/855,157 , “Final Office Action”, U.S. Appl. No. 17/855,157, Jan. 24, 2024, 14 pages. [cited by applicant]
U.S. Appl. No. 17/855,157 , “Non-Final Office Action”, U.S. Appl. No. 17/855,157, May 10, 2024, 15 pages. [cited by applicant]
U.S. Appl. No. 17/855,157 , “Non-Final Office Action”, U.S. Appl. No. 17/855,157, Aug. 4, 2023, 13 pages. [cited by applicant]
U.S. Appl. No. 18/215,681, “Pursuant to MPEP § 2001.06(b) the applicant brings the following co-pending application to the Examiner's attention:”, U.S. Appl. No. 18/215,681, Jun. 28, 2023, 30 pages. [cited by applicant]
U.S. Appl. No. 63/468,767, filed May 24, 2023 , “Pursuant to MPEP § 2001.06(b) the applicant brings the following co-pending application to the Examiner's attention:”, U.S. Appl. No. 63/468,767, May 24, 2023, 22 pages. [cited by applicant]
Adhinarayanan, Vignesh , et al., “Pursuant to MPEP § 2001.06(b) the applicant brings the following co-pending application to the Examiner's attention:”, U.S. Appl. No. 17/855,157, Jun. 30, 2022, 37 pages. [cited by applicant]
Bakhshalipour, Mohammad , et al., “A Survey on Recent Hardware Data Prefetching Approaches with An Emphasis on Servers”, Cornell University; Arxiv; retrieved from https://arxiv.org/abs/2009.00715 on Aug. 15, 2023, Sep. … [cited by applicant]
IARPA , “Agile—Advanced Graphic Intelligence Logical Computing Environment”, Office of the Director of National Intelligence; IARPA (Intelligence Advanced Research Projects Activity; retrieved from https://www.iarpa.gov… [cited by applicant]
Jacobs, Bryan , “Hierarchical Identify Verify Exploit (HIVE)”, DARPA [retrieved Jul. 6, 2023]. Retrieved from the Internet <https://www.darpa.mil/program/hierarchical-identify-verify-exploit>., 2 Pages. [cited by applicant]
Kepler Computing Inc. , “Fe-RAM Based Memory Technology”, retrieved from https://www.kearney.com/service/operations-performance/article/-/insights/semiconductor-production-with-kepler-computing-supply-chain-shocks-podca… [cited by applicant]
Lloyd, Scott , et al., “In-Memory Data Rearrangement for Irregular, Data-Intensive Computing”, Computer, vol. 48, No. 8 [retrieved Dec. 13, 2023]. Retrieved from the Internet <https://doi.org/10.1109/MC.2015.230>, Aug. … [cited by applicant]
O'Connor, Mike , et al., “Fine-Grained DRAM: Energy-Efficient DRAM for Extreme Bandwidth Systems”, MICRO-50, Cambridge, MA, US; Nviidia; The University of Texas at Austin; Stanford University. Retrieved from the Interne… [cited by applicant]
O'Connor, Mike , et al., “Fine-grained DRAM: energy-efficient DRAM for extreme bandwidth systems”, MICRO-50 '17: Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture [retrieved Jan. 23, 2… [cited by applicant]
Scrbak, Marko , et al., “Pursuant to MPEP § 2001.06(b) the applicant brings the following co-pending application to the Examiner's attention:”, U.S. Appl. No. 18/428,905, filed Jan. 31, 2024, 41 pages. [cited by applicant]
Seshadri, Vivek , et al., “Gather-Scatter DRAM: In-DRAM Address Translation to Improve the Spatial Locality of Non-unit Strided Accesses”, MICRO-48: Proceedings of the 48th International Symposium on Microarchitecture [… [cited by applicant]
Talati, Nishil , et al., “Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-Design”, IEEE International Symposium on High Performance Computer Architecture (HPCA), 202… [cited by applicant]
Yu, Xiangyao , et al., “IMP: Indirect Memory Prefetcher”, Massachusetts Institute of Technology; Parallel Computing Lab, Intel Labs; Micro-48, 2015, 13 pages. [cited by applicant]
U.S. Appl. No. 17/855,157 , “Non-Final Office Action”, U.S. Appl. No. 17/855,157, Nov. 5, 2024, 17 pages. [cited by applicant]
U.S. Appl. No. 18/215,681, “Non-Final Office Action”, U.S. Appl. No. 18/215,681, Jan. 22, 2025, 8 Pages. [cited by applicant]
Akin et al., “Dynamic Fine-Grained Sparse Memory Accesses,” MEMSYS '18: Proceedings of the International Symposium on Memory Systems, ACM, Oct. 1, 2018, pp. 85-97. [cited by applicant]
Choi, et al., “Adaptive Granularity Based Last-Level Cache Prefetching Method with eDRAM Prefetch Buffer for Graph Processing Applications,” Applied Sciences, vol. 11, No. 3, 2021, 23 pages. [cited by applicant]
Heirman et al., “Automatic Sublining for Efficient Sparse Memory Accesses,” ACM Transactions on Architecture and Code Optimization (TACO), vol. 18, No. 3, Article 33, Apr. 2021, 23 pages. [cited by applicant]
Non-Final Office Action issued in U.S. Appl. No. 18/428,905, mailed Dec. 2, 2025, 24 pages. [cited by applicant]
Solihin et al., “Prefetching in an Intelligent Memory Architecture Using a Helper Thread,” 5th Workshop on Multithreaded Execution, Architecture, and Compilation (MTEAC-5), MICRO-34, Dec. 2001, 8 pages. [cited by applicant]
Yoon et al., “The Dynamic Granularity Memory System,” ACM SIGARCH Computer Architecture News, vol. 40, No. 3, Jun. 9, 2012, pp. 548-559. [cited by applicant]
Zhang et al., “Fine-Grained Address Segmentation for Attention-Based Variable-Degree Prefetching,” CF '22: Proceedings of the 19th ACM International Conference on Computing Frontiers, May 17, 2022, pp. 103-112. [cited by applicant]
Final Office Action issued in U.S. Appl. No. 18/428,905, mailed Apr. 7, 2026, 23 pages. [cited by applicant]