IP Library Granted Patent US 12,730,636
Granted Patent B2
US 12,730,636 · App. 19/038,676 · Granted Sep 8, 2026

Systems and methods of programming for in-memory processing

Inventors: Marie Mai Nguyen (Pittsburgh, PA); Tong Zhang (Mountain View, CA); Yangwook Kang (San Jose, CA); Rekha Pitchumani (Oak Hill, VA); Yang Seok Ki (Palo Alto, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F9/30043G06F9/30189G06F9/544
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,730,636
App. No.
19/038,676
Granted
Sep 8, 2026
Kind
B2
Abstract

Provided are systems, methods, and apparatuses for systems and methods of programming processes in processing in memory (PIM) systems. In one or more examples, the systems, devices, and methods include enabling, via a processor, a PIM manager for processing of PIM instructions, the PIM manager being located on a base die of a stacked memory module; enabling, via the PIM manager, a plurality of processing units of one or more memory dies of the stacked memory module for processing; compiling source code comprising the PIM instructions into machine code; loading, via PIM manager, the machine code into an allocation of memory of the stacked memory module; and executing, via the plurality of processing units, the PIM instructions based on loading the machine code into the memory of the stacked memory module.

Claims (61)

1 . A method of programming processing-in-memory (PIM) processes, the method comprising:

enabling, via a processor, a PIM manager for processing of PIM instructions, the PIM manager being located on a base die of a stacked memory module;

enabling, via the PIM manager, a plurality of processing units of one or more memory dies of the stacked memory module for processing;

compiling, via a compiler of the PIM manager, source code comprising the PIM instructions into machine code, the machine code being stored in a memory of a host of the stacked memory module;

loading, via the compiler and a driver of the PIM manager, the machine code into an allocation of memory of the stacked memory module; and

executing, via the plurality of processing units, the PIM instructions based on loading the machine code into the memory of the stacked memory module.

2 . The method of claim 1 , wherein a first portion of the allocation of memory comprises a first portion of the machine code, the first portion of the machine code comprising:

a first dataset used by the plurality of processing units of the one or more memory dies in executing at least a first portion of the PIM instructions, and

a second dataset used by processing units on the base die in executing at least a second portion of the PIM instructions.

3 . The method of claim 1 , wherein a second portion of the allocation of memory comprises a second portion of the machine code, the second portion of the machine code comprising:

a first portion of the PIM instructions used by the plurality of processing units of the one or more memory dies in executing the first portion of the PIM instructions, and

a second portion of the PIM instructions used by processing units on the base die.

4 . The method of claim 3 , wherein:

a third portion of the allocation of memory comprises a third portion of the machine code, the third portion of the machine code comprising one or more load store instructions managed by the PIM manager, and

the PIM manager issues an instruction from the one or more load store instructions to trigger execution of the first portion of the PIM instructions.

5 . The method of claim 1 , wherein the PIM manager comprises a base die processor configured to direct processing on the plurality of processing units of one or more memory dies and to orchestrate, via a shared buffer of the PIM manager, data movement between the plurality of processing units of the one or more memory dies and processing units on the base die.

6 . The method of claim 1 , wherein:

the PIM manager is enabled for processing based on the processor writing an activation value to an enable register included in the allocation of memory, and

the PIM manager enables the plurality of processing units for processing based on the PIM manager being enabled.

7 . The method of claim 1 , further comprising:

determining execution of the PIM instructions is complete based on polling a done register included in the allocation of memory and determining a completion value of the done register indicates execution of the PIM instructions is complete,

wherein the PIM manager writes the completion value to the done register based on the PIM manager determining a plurality of registers respectively associated with the plurality of processing units indicate that execution of the PIM instructions by the plurality of processing units is complete.

8 . The method of claim 1 , further comprising writing a deactivation value to a stop register included in the allocation of memory to deactivate the PIM manager, wherein deactivating the PIM manager triggers the PIM manager to deactivate the plurality of processing units.

9 . The method of claim 1 , further comprising writing an exception value to an exception register included in the allocation of memory to trigger an exception based on execution of the PIM instructions.

10 . The method of claim 1 , wherein the one or more memory dies comprise one or more layers of memory dies stacked on top of the base die.

11 . The method of claim 1 , wherein:

the processor comprises at least one of a graphical processing unit (GPU) communicatively coupled to the PIM manager or a central processing unit (CPU) of the host of the stacked memory module that is communicatively coupled to the PIM manager, and

the PIM manager comprises a processor and memory for PIM management.

12 . A device comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the device to:

enable a PIM manager for processing of PIM instructions, the PIM manager being located on a base die of a stacked memory module;

enable, via the PIM manager, a plurality of processing units of one or more memory dies of the stacked memory module for processing;

compile, via a compiler of the PIM manager, source code comprising the PIM instructions into machine code, the machine code being stored in a memory of a host of the stacked memory module;

load, via the compiler and a driver of the PIM manager, the machine code into an allocation of memory of the stacked memory module; and

execute, via the plurality of processing units, the PIM instructions based on loading the machine code into the memory of the stacked memory module.

13 . The device of claim 12 , wherein a first portion of the allocation of memory comprises a first portion of the machine code, the first portion of the machine code comprising:

a first dataset used by the plurality of processing units of the one or more memory dies in executing at least a first portion of the PIM instructions, and

a second dataset used by processing units on the base die in executing at least a second portion of the PIM instructions.

14 . The device of claim 12 , wherein a second portion of the allocation of memory comprises a second portion of the machine code, the second portion of the machine code comprising:

a first portion of the PIM instructions used by the plurality of processing units of the one or more memory dies in executing the first portion of the PIM instructions, and

a second portion of the PIM instructions used by processing units on the base die.

15 . The device of claim 14 , wherein:

a third portion of the allocation of memory comprises a third portion of the machine code, the third portion of the machine code comprising one or more load store instructions managed by the PIM manager, and

the PIM manager issues an instruction from the one or more load store instructions to trigger execution of the first portion of the PIM instructions.

16 . The device of claim 12 , wherein the PIM manager is configured to orchestrate, via a shared buffer of the PIM manager, data movement between the plurality of processing units of the one or more memory dies and processing units on the base die.

17 . The device of claim 12 , wherein:

the instructions, when executed by the one or more processors, further cause the device to determine execution of the PIM instructions is complete based on polling a done register included in the allocation of memory and determining a completion value of the done register indicates execution of the PIM instructions is complete, and

the PIM manager writes the completion value to the done register based on the PIM manager determining a plurality of registers respectively associated with the plurality of processing units indicate that execution of the PIM instructions by the plurality of processing units is complete.

18 . A non-transitory computer-readable medium storing code that comprises instructions executable by a processor to:

enable a PIM manager for processing of PIM instructions, the PIM manager being located on a base die of a stacked memory module;

enable, via the PIM manager, a plurality of processing units of one or more memory dies of the stacked memory module for processing;

compile, via a compiler of the PIM manager, source code comprising the PIM instructions into machine code, the machine code being stored in a memory of a host of the stacked memory module;

load, via the compiler and a driver of the PIM manager, the machine code into an allocation of memory of the stacked memory module; and

execute, via the plurality of processing units, the PIM instructions based on loading the machine code into the memory of the stacked memory module.

19 . The non-transitory computer-readable medium of claim 18 , wherein a first portion of the allocation of memory comprises a first portion of the machine code, the first portion of the machine code comprising:

a first dataset used by the plurality of processing units of the one or more memory dies in executing at least a first portion of the PIM instructions, and

a second dataset used by processing units on the base die in executing at least a second portion of the PIM instructions.

20 . The non-transitory computer-readable medium of claim 18 , wherein a second portion of the allocation of memory comprises a second portion of the machine code, the second portion of the machine code comprising:

a first portion of the PIM instructions used by the plurality of processing units of the one or more memory dies in executing the first portion of the PIM instructions, and

a second portion of the PIM instructions used by processing units on the base die.

Continuity (3)
Provisional Application 63688814 · Aug 29, 2024
Provisional Application 63666105 · Jun 28, 2024
Related Publication 20260003617A1 · Jan 1, 2026
References Cited (144)
US 8234460B2 · Walker · 2012 [cited by applicant]
US 8335908B2 · Waugh · 2012 [cited by applicant]
US 9870339B2 · Kang et al. · 2018 [cited by applicant]
US 10467159B2 · Beard et al. · 2019 [cited by applicant]
US 10474600B2 · Malladi et al. · 2019 [cited by applicant]
US 10592121B2 · Malladi et al. · 2020 [cited by applicant]
US 10866900B2 · Chang et al. · 2020 [cited by applicant]
US 11036660B2 · Ooi et al. · 2021 [cited by applicant]
US 11093277B2 · Sankaran et al. · 2021 [cited by applicant]
US 11158357B2 · Yu et al. · 2021 [cited by applicant]
US 11276459B2 · O · 2022 [cited by applicant]
US 11468001B1 · Hassaan et al. · 2022 [cited by applicant]
US 11625245B2 · Nurvitadhi et al. · 2023 [cited by applicant]
US 11645197B2 · Seo · 2023 [cited by applicant]
US 11748077B2 · Wang et al. · 2023 [cited by applicant]
US 11860800B2 · Wang et al. · 2024 [cited by applicant]
US 11861758B2 · Agostini · 2024 [cited by applicant]
US 11880590B2 · Jeong et al. · 2024 [cited by applicant]
US 11921634B2 · Kotra et al. · 2024 [cited by applicant]
US 12038819B2 · Patle et al. · 2024 [cited by applicant]
US 12058874B1 · Farjadrad · 2024 [cited by applicant]
US 12087388B2 · Yu et al. · 2024 [cited by applicant]
US 12099455B2 · Yu et al. · 2024 [cited by applicant]
US 12182409B2 · Kang et al. · 2024 [cited by applicant]
US 12182532B1 · Kim · 2024 [cited by applicant]
US 12190038B1 · Farjadrad · 2025 [cited by applicant]
US 20060248280A1 · Al-Sukhni et al. · 2006 [cited by applicant]
US 20090300292A1 · Fang et al. · 2009 [cited by applicant]
US 20100078790A1 · Ito et al. · 2010 [cited by applicant]
US 20140043333A1 · Narayanan · 2014 [cited by examiner]
US 20140181427A1 · Jayasena et al. · 2014 [cited by applicant]
US 20140195744A1 · Fleischer · 2014 [cited by examiner]
US 20160026912A1 · Falcon et al. · 2016 [cited by applicant]
US 20170060588A1 · Choi · 2017 [cited by examiner]
US 20170255390A1 · Chang et al. · 2017 [cited by applicant]
US 20170344301A1 · Ryu et al. · 2017 [cited by applicant]
US 20180096735A1 · Pappu · 2018 [cited by applicant]
US 20180364944A1 · Grossman et al. · 2018 [cited by applicant]
US 20190042240A1 · Pappu et al. · 2019 [cited by applicant]
US 20190050325A1 · Malladi et al. · 2019 [cited by applicant]
US 20190096453A1 · Shin et al. · 2019 [cited by applicant]
US 20190138569A1 · Korthikanti et al. · 2019 [cited by applicant]
US 20190303033A1 · Noguera Serra et al. · 2019 [cited by applicant]
US 20200004690A1 · Mathew et al. · 2020 [cited by applicant]
US 20200081651A1 · Aga · 2020 [cited by examiner]
US 20200081744A1 · Siegl et al. · 2020 [cited by applicant]
US 20200097417A1 · Malladi et al. · 2020 [cited by applicant]
US 20200119736A1 · Weber et al. · 2020 [cited by applicant]
US 20200184001A1 · Gu et al. · 2020 [cited by applicant]
US 20200294558A1 · Yu et al. · 2020 [cited by applicant]
US 20210019076A1 · Kim et al. · 2021 [cited by applicant]
US 20210110876A1 · Seo et al. · 2021 [cited by applicant]
US 20210173893A1 · Luo · 2021 [cited by applicant]
US 20210200696A1 · Kang et al. · 2021 [cited by applicant]
US 20210303346A1 · Shahim et al. · 2021 [cited by applicant]
US 20210354729A1 · Ng et al. · 2021 [cited by applicant]
US 20210374607A1 · Kazakov · 2021 [cited by examiner]
US 20220036929A1 · Yu et al. · 2022 [cited by applicant]
US 20220068366A1 · Kwon et al. · 2022 [cited by applicant]
US 20220075564A1 · Moon et al. · 2022 [cited by applicant]
US 20220188606A1 · Roberts · 2022 [cited by applicant]
US 20220197656A1 · Heirman et al. · 2022 [cited by applicant]
US 20220206685A1 · Kalamatianos et al. · 2022 [cited by applicant]
US 20220206817A1 · Kotra et al. · 2022 [cited by applicant]
US 20220206869A1 · Ramachandran et al. · 2022 [cited by applicant]
US 20220222198A1 · Lanka et al. · 2022 [cited by applicant]
US 20220269436A1 · Kellam et al. · 2022 [cited by applicant]
US 20220342595A1 · Chatterjee et al. · 2022 [cited by applicant]
US 20220414013A1 · Alsop et al. · 2022 [cited by applicant]
US 20230026505A1 · Lee et al. · 2023 [cited by applicant]
US 20230051544A1 · Vanesko et al. · 2023 [cited by applicant]
US 20230053439A1 · Hornung et al. · 2023 [cited by applicant]
US 20230094148A1 · Lee et al. · 2023 [cited by applicant]
US 20230109990A1 · Pappu et al. · 2023 [cited by applicant]
US 20230128183A1 · Kang et al. · 2023 [cited by applicant]
US 20230138048A1 · Kwon et al. · 2023 [cited by applicant]
US 20230195375A1 · Puthoor et al. · 2023 [cited by applicant]
US 20230195459A1 · Puthoor et al. · 2023 [cited by applicant]
US 20230195645A1 · Puthoor et al. · 2023 [cited by applicant]
US 20230205693A1 · Kotra et al. · 2023 [cited by applicant]
US 20230260075A1 · Maiyuran et al. · 2023 [cited by applicant]
US 20230289294A1 · Swaine · 2023 [cited by applicant]
US 20230305978A1 · Davis et al. · 2023 [cited by applicant]
US 20230315651A1 · Dally et al. · 2023 [cited by applicant]
US 20230316043A1 · Darvish Rouhani et al. · 2023 [cited by applicant]
US 20230334613A1 · Daxer et al. · 2023 [cited by applicant]
US 20230343718A1 · Zhou et al. · 2023 [cited by applicant]
US 20230393849A1 · Kotra et al. · 2023 [cited by applicant]
US 20240004585A1 · Ibrahim et al. · 2024 [cited by applicant]
US 20240004653A1 · Alsop et al. · 2024 [cited by applicant]
US 20240013338A1 · Matam et al. · 2024 [cited by applicant]
US 20240036951A1 · Long et al. · 2024 [cited by applicant]
US 20240047364A1 · Blair · 2024 [cited by applicant]
US 20240063200A1 · Niu et al. · 2024 [cited by applicant]
US 20240078195A1 · Madan et al. · 2024 [cited by applicant]
US 20240103745A1 · Madan et al. · 2024 [cited by applicant]
US 20240161227A1 · Appu et al. · 2024 [cited by applicant]
US 20240168639A1 · Aga et al. · 2024 [cited by applicant]
US 20240205805A1 · Boccuzzi et al. · 2024 [cited by applicant]
US 20240403254A1 · Ma et al. · 2024 [cited by applicant]
US 20250021378A1 · Liu et al. · 2025 [cited by applicant]
US 20250031385A1 · Agarwal et al. · 2025 [cited by applicant]
US 20250036579A1 · Chang et al. · 2025 [cited by applicant]
US 20250045102A1 · Liu et al. · 2025 [cited by applicant]
US 20250053322A1 · Lee et al. · 2025 [cited by applicant]
US 20250077457A1 · Jin et al. · 2025 [cited by applicant]
US 20250085875A1 · Litt et al. · 2025 [cited by applicant]
US 20250190391A1 · Jeong · 2025 [cited by applicant]
US 20260003617A1 · Nguyen et al. · 2026 [cited by applicant]
IN 202244051173 · 2023 [cited by applicant]
JP 2019531546A · 2019 [cited by applicant]
KR 20200108768A · 2020 [cited by applicant]
KR 102430982A · 2022 [cited by applicant]
KR 20240104054 · 2024 [cited by applicant]
WO 2021028723A2 · 2021 [cited by applicant]
WO 2024137416A1 · 2024 [cited by applicant]
Samsung, “HBM-PIM: Cutting-Edge Memory Technology to Accelerate Next-Generation AI,” Tech Blog, Mar. 2023, 3 pages. [cited by applicant]
Thiruvengadam, Vijay, “Memory Processing Units (MPU),” The University of Wisconsin, Madison (Reference in WebArchive—Wayback Machine), Apr. 2024, 60 pages. [cited by applicant]
Xie, Xinfeng et al., “MPU: Memory-Centric SIMT Processor via In-DRAM Near-Bank Computing,” ACM Transactions on Architecture and Code Optimization, vol. 20, No. 3, Jul. 2023, 26 pages. [cited by applicant]
Xie, Xinfeng et al., “MPU: Towards Bandwidth-Abundant SIMT Processor via Near-Bank Computing,” arXiv, Mar. 2021, pp. 1-13. [cited by applicant]
Aspen Systems Inc., “Intel® Gaudi® 3 Accelerator”, https://www.aspsys.com/intel-gaudi-3-accelerators, 2025, 3 pages. [cited by applicant]
Aspen Systems Inc., “Processors”, 2025, https://www.aspsys.com/category/hpc-processors/, 2025, 2 pages. [cited by applicant]
Cai, Weilin et al., “A Survey on Mixture of Experts,” arXiv:2407.06204, Jun. 2024, 41 pages. [cited by applicant]
Choi, Jungwoo et al., “A Lightweight and Efficient GPU for NDP Utilizing Data Access Pattern of Image Processing,” IEEE Transactions on Computers, vol. 70, No. 1, Jan. 2022, pp. 13-26. [cited by applicant]
Cui, Weihao et al., “Optimizing Dynamic Neural Networks with Brainstorm,” 17th USENIX Symposium on Operating Systems Design and Implementation, Jul. 2023, 21 pages. [cited by applicant]
Eliseev, Artyom et al., “Fast Inference of Mixture-of-Experts Language Models with Offloading,” arXiv:2312.17238, Dec. 2023, 12 pages. [cited by applicant]
European Extended Search Report for Application No. 25184706.7, mailed Oct. 27, 2025. [cited by applicant]
European Extended Search Report for Application No. 25184833.9, mailed Nov. 3, 2025. [cited by applicant]
European Extended Search Report for Application No. 25185257.0, mailed Dec. 4, 2025. [cited by applicant]
Intel, “Intel Gaudi 3 AI Accelerator,” Technical Paper, Intel, 2024, 24 pages. [cited by applicant]
Xue, Leyang et al., “MoE-Infinity: Activation-Aware Expert Offloading for Efficient MoE Serving,” arXiv:2401.14361, Jan. 2024, 15 pages. [cited by applicant]
European Extended Search Report for Application No. 25164046.2, mailed Jul. 15, 2025. [cited by applicant]
European Extended Search Report for Application No. 25177870.0, mailed Oct. 6, 2025. [cited by applicant]
European Extended Search Report for Application No. 25183244.0, mailed Oct. 23, 2025. [cited by applicant]
Hadidi, Ramyad et al., “Performance Implications of NoCs on 3D-Stacked Memories: Insights from the Hybrid Memory Cube,” 2018 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), Jul. 20… [cited by applicant]
Horowitz, Mark, “Computing's Energy Problem (and what we can do about it),” 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC), Feb. 2014, 46 pages. [cited by applicant]
Hhsieh, Kevin et al. “Transparent Offloading and Mapping (TOM): Enabling Programmer-Transparent Near-Data Processing in GPU Systems,” 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA)], J… [cited by applicant]
Mutlu, Onur et al., “A Modern Primer on Processing in Memory,” arxiv.org, Cornell University Library, XP081831211, Dec. 2020, 42 pages. [cited by applicant]
Roy, Rishav et al., “Co-Design of Thermal Management with System Architecture and Power Management for 3D ICs,” 2022 IEEE 72ND Electronic Components and Technology Conference (ECTC), IEEE, May 2022, pp. 211-220. [cited by applicant]
Zhu, Yuxiong et al., “Integrated Thermal Analysis for Processing In Die-Stacking Memory,” MEMSYS '16: Proceedings of the Second International Symposium on Memory Systems, Oct. 2016, 13 pages. [cited by applicant]
European Office Action for Application No. 25164046.2, mailed Mar. 5, 2026. [cited by applicant]
Pattnaik, Ashutosh et al., “Scheduling Techniques for GPU Architectures with Processing-in-Memory Capabilities,” 2016 International Conference on Parallel Architecture and Compilation Techniques (PACT), Sep. 2016, pp. 3… [cited by applicant]
Office Action for U.S. Appl. No. 19/038,672, mailed Mar. 26, 2026. [cited by applicant]
Office Action for U.S. Appl. No. 19/038,683, mailed Jul. 13, 2026. [cited by applicant]