IP Library › Granted Patent US 12,632,171
Granted Patent B2
US 12,632,171 · App. 18/614,368 · Granted May 19, 2026

Memory device using multistage acceleration, operating method of memory device, and electronic device including the same

Inventor: Jinhyun Kim (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F3/0604G06F3/0629G06F3/0673G06F3/0601G06F3/0614G06F3/0656G06F3/0658G06F3/0659G06F12/0207G06F12/0284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,171
App. No.
18/614,368
Granted
May 19, 2026
Kind
B2
Abstract

An electronic device is provided. The electronic host includes: a host; a memory package including a plurality of memory devices and a first accelerator circuit configured to receive first data from the plurality of memory devices and perform a coarse acceleration operation based on the first data to obtain second data; and a memory controller including a second accelerator circuit configured to receive the second data from the first accelerator circuit and perform a fine acceleration operation based on a neural network and the second data to obtain an inference result.

Claims (52)

1 . An electronic device comprising:

a memory package comprising a plurality of memory devices and a first accelerator circuit configured to receive first data from the plurality of memory devices and perform a coarse acceleration operation based on the first data to obtain second data, wherein the coarse acceleration operation comprises a filtering operation performed based on the first data; and

a memory controller comprising a second accelerator circuit configured to receive the second data from the first accelerator circuit and perform a fine acceleration operation based on a neural network and the second data to obtain an inference result.

2 . The electronic device of claim 1 , wherein the plurality of memory devices correspond to dynamic random access memories (DRAMs) which are three-dimensionally stacked,

wherein the first accelerator circuit is provided in a master DRAM, among the DRAMs, which is in contact with a substrate, and

wherein, other than the master DRAM, the DRAMs are connected to the master DRAM through wire bonding.

3 . The electronic device of claim 1 , wherein the plurality of memory devices correspond to DRAMs which are stacked on a first die,

wherein the first accelerator circuit is provided on a second die different from the first die, and

wherein the DRAMs are connected to the first accelerator circuit through wire bonding.

4 . The electronic device of claim 1 , wherein the filtering operation comprises a zeroing operation in which, among the first data, a portion having a size smaller than a threshold value is changed into zero to obtain filtered data, and

wherein the coarse acceleration operation comprises a pruning operation in which zero data and NULL data are removed from the filtered data to obtain the second data.

5 . The electronic device of claim 1 , wherein the fine acceleration operation comprises a general matrix vector multiplication (GEMV) operation of matrix-to-vector multiplication and a general matrix-matrix multiplication (GEMM) operation of matrix-to-matrix multiplication.

6 . The electronic device of claim 1 , wherein the memory controller is configured to control one of the coarse acceleration operation and the fine acceleration operation to be selectively performed.

7 . The electronic device of claim 1 , wherein the memory controller further comprises:

a decoder configured to decode instructions received from a host to obtain decoded instructions; and

a command generator configured to generate a command for controlling the plurality of memory devices based on the decoded instructions.

8 . The electronic device of claim 1 , wherein a time required for the coarse acceleration operation and the fine acceleration operation is less than a predefined threshold value.

9 . The electronic device of claim 1 , wherein the memory controller corresponds to a compute express link (CXL) device,

wherein the electronic device further comprises:

a first interface circuit configured to interface with a host; and

a second interface circuit configured to interface with the memory package, and

wherein the first interface circuit is configured to communicate with the host based on a peripheral component interconnect express (PCIe) protocol.

10 . A memory controller comprising:

a register configured to receive instructions from a host;

a decoder configured to decode the instructions;

a command generator configured to generate a command to be provided to a memory package based on the decoded instructions, wherein the memory package comprises a plurality of memory devices; and

a first accelerator circuit configured to receive second data from the memory package and perform first operations for an inference operation based on a neural network,

wherein the second data is data obtained as a result of performing second operations on first data stored in the plurality of memory devices through a second accelerator circuit provided in the memory package, wherein the second operations comprise a filtering operation performed based on the first data.

11 . The memory controller of claim 10 , wherein the plurality of memory devices correspond to dynamic random access memories (DRAMs) which are three-dimensionally stacked inside the memory package,

wherein the second accelerator circuit is provided in a master DRAM contacting a substrate of the memory package, and

wherein, other than the master DRAM, the DRAMs are connected to the master DRAM through wire bonding.

12 . The memory controller of claim 10 , wherein the plurality of memory devices correspond to DRAMs which are three-dimensionally stacked on a first die inside the memory package,

wherein the second accelerator circuit is provided on a second die different from the first die, and

wherein the DRAMs are connected to the second accelerator circuit through wire bonding.

13 . The memory controller of claim 10 , wherein the filtering operation comprises a zeroing operation in which, among the first data, a portion having a size smaller than a threshold value is changed into zero to obtain filtered data, and

wherein the second operations comprise a pruning operation in which zero data and NULL data are removed from the filtered data to obtain the second data.

14 . The memory controller of claim 10 , wherein the first operations comprise a general matrix vector multiplication (GEMV) operation of matrix-to-vector multiplication and a general matrix-matrix multiplication (GEMM) operation of matrix-to-matrix multiplication.

15 . The memory controller of claim 10 , wherein the memory controller is configured to selectively perform one of the first operations based on the first accelerator circuit or the second operations based on the second accelerator circuit.

16 . The memory controller of claim 10 , wherein a time required for the first operations and the second operations is less than a predefined threshold value.

17 . An operating method of a memory controller connected to a memory package comprising a plurality of memory devices, the operating method comprising:

receiving output data from a first accelerator circuit which has performed a coarse acceleration operation on sparse data stored in the plurality of memory devices to obtain the output data, wherein the coarse acceleration operation comprises a filtering operation performed based on the sparse data;

obtaining an inference result by performing a fine acceleration operation on the output data by using a second accelerator circuit inside the memory controller; and

providing the inference result to a host.

18 . The operating method of claim 17 , wherein the filtering operation comprises a zeroing operation in which, among the sparse data, a portion having a size smaller than a threshold value is changed into zero to obtain filtered data,

wherein the coarse acceleration operation comprises a pruning operation in which zero data and NULL data are removed from the filtered data to obtain the output data, and

wherein the fine acceleration operation comprises a general matrix vector multiplication (GEMV) operation of matrix-to-vector multiplication and a general matrix-matrix multiplication (GEMM) operation of matrix-to-matrix multiplication.

19 . The operating method of claim 17 , wherein the plurality of memory devices correspond to dynamic random access memories (DRAMs) which are three-dimensionally stacked,

wherein the first accelerator circuit is provided in a master DRAM contacting a substrate among the DRAMs, and

wherein, other than the master DRAM, the DRAMs are connected to the master DRAM through wire bonding.

20 . The operating method of claim 17 , wherein the plurality of memory devices correspond to DRAMs which are three-dimensionally stacked on a first die,

wherein the first accelerator circuit is provided on a second die different from the first die, and

wherein the DRAMs are connected to the first accelerator circuit through wire bonding.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2024
From: KIM, JINHYUN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 066890/0643 →
Priority Claims (1)
KR 10-2023-0039280 · Mar 24, 2023 · national
Continuity (1)
Related Publication 20240319871A1 · Sep 26, 2024
References Cited (20)
US 7107412B2 · Klein et al. · 2006 [cited by applicant]
US 9965187B2 · Gokhale et al. · 2018 [cited by applicant]
US 10185499B1 · Wang et al. · 2019 [cited by applicant]
US 10762012B2 · Lee · 2020 [cited by applicant]
US 20210034957A1 · Rhu et al. · 2021 [cited by applicant]
US 20210150770A1 · Appu et al. · 2021 [cited by applicant]
US 20210203704A1 · Pohl · 2021 [cited by applicant]
US 20210240797A1 · Rashid et al. · 2021 [cited by applicant]
US 20210271680A1 · Lee et al. · 2021 [cited by applicant]
US 20220215871A1 · Kim et al. · 2022 [cited by applicant]
US 20250139434A1 · Rhu · 2025 [cited by examiner]
WO 2021202160A1 · 2021 [cited by applicant]
Extended European Search report dated Jul. 2, 2024, issued by the European Patent Office in European Application No. 24165245.2. [cited by applicant]
Hwang et al., “Centaur: A Chiplet-based, Hybrid Sparse-Dense Accelerator for Personalized Recommendations”, 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), pp. 968-981 (14 pages total). [cited by applicant]
Lee et al., “Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology”, 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA), pp. 43-56 (14 pages total). [cited by applicant]
Iskandar et al., “Near-memory Computing on FPGAs with 3D-stacked Memories: Applications, Architectures, and Optimizations”, ACM Transactions on Reconfigurable Technology and Systems, Dec. 2022, vol. 16, No. 1, Article 1… [cited by applicant]
Yu et al., “Data Stream Oriented Fine-grained Sparse CNN Accelerator with Efficient Unstructured Pruning Strategy” Session 3B: VLSI for Machine Learning and Artificial Intelligence 1, GLSVLSI '22, Jun. 6-8, 2022, pp. 24… [cited by applicant]
Yang et al., “Procrustes: a Dataflow and Accelerator for Sparse Deep Neural Network Training”, 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), pp. 711-724 (14 pages total). [cited by applicant]
Jin Hyun Kim et al., “Aquabolt-XL HBM2-PIM, LPDDR5-PIM With In-Memory Processing, and AXDIMM With Acceleration Buffer,” Hot Chips 33, IEEE, Digital Object Identifier 10.1109/MM.2022.3164651, Apr. 5, 2022, pp. 20-30. [cited by applicant]
Liu Ke et al., “Near-Memory Processing in Action: Accelerating Personalized Recommendation With AxDIMM,” In-Memory Computing, 2021 IEEE, Digital Object Identifier 10.1109/MM.2021.3097700, Jul. 16, 2021, pp. 116-127. [cited by applicant]