Memory device using multistage acceleration, operating method of memory device, and electronic device including the same
An electronic device is provided. The electronic host includes: a host; a memory package including a plurality of memory devices and a first accelerator circuit configured to receive first data from the plurality of memory devices and perform a coarse acceleration operation based on the first data to obtain second data; and a memory controller including a second accelerator circuit configured to receive the second data from the first accelerator circuit and perform a fine acceleration operation based on a neural network and the second data to obtain an inference result.
1 . An electronic device comprising:
a memory package comprising a plurality of memory devices and a first accelerator circuit configured to receive first data from the plurality of memory devices and perform a coarse acceleration operation based on the first data to obtain second data, wherein the coarse acceleration operation comprises a filtering operation performed based on the first data; and
a memory controller comprising a second accelerator circuit configured to receive the second data from the first accelerator circuit and perform a fine acceleration operation based on a neural network and the second data to obtain an inference result.
2 . The electronic device of claim 1 , wherein the plurality of memory devices correspond to dynamic random access memories (DRAMs) which are three-dimensionally stacked,
wherein the first accelerator circuit is provided in a master DRAM, among the DRAMs, which is in contact with a substrate, and
wherein, other than the master DRAM, the DRAMs are connected to the master DRAM through wire bonding.
3 . The electronic device of claim 1 , wherein the plurality of memory devices correspond to DRAMs which are stacked on a first die,
wherein the first accelerator circuit is provided on a second die different from the first die, and
wherein the DRAMs are connected to the first accelerator circuit through wire bonding.
4 . The electronic device of claim 1 , wherein the filtering operation comprises a zeroing operation in which, among the first data, a portion having a size smaller than a threshold value is changed into zero to obtain filtered data, and
wherein the coarse acceleration operation comprises a pruning operation in which zero data and NULL data are removed from the filtered data to obtain the second data.
5 . The electronic device of claim 1 , wherein the fine acceleration operation comprises a general matrix vector multiplication (GEMV) operation of matrix-to-vector multiplication and a general matrix-matrix multiplication (GEMM) operation of matrix-to-matrix multiplication.
6 . The electronic device of claim 1 , wherein the memory controller is configured to control one of the coarse acceleration operation and the fine acceleration operation to be selectively performed.
7 . The electronic device of claim 1 , wherein the memory controller further comprises:
a decoder configured to decode instructions received from a host to obtain decoded instructions; and
a command generator configured to generate a command for controlling the plurality of memory devices based on the decoded instructions.
8 . The electronic device of claim 1 , wherein a time required for the coarse acceleration operation and the fine acceleration operation is less than a predefined threshold value.
9 . The electronic device of claim 1 , wherein the memory controller corresponds to a compute express link (CXL) device,
wherein the electronic device further comprises:
a first interface circuit configured to interface with a host; and
a second interface circuit configured to interface with the memory package, and
wherein the first interface circuit is configured to communicate with the host based on a peripheral component interconnect express (PCIe) protocol.
10 . A memory controller comprising:
a register configured to receive instructions from a host;
a decoder configured to decode the instructions;
a command generator configured to generate a command to be provided to a memory package based on the decoded instructions, wherein the memory package comprises a plurality of memory devices; and
a first accelerator circuit configured to receive second data from the memory package and perform first operations for an inference operation based on a neural network,
wherein the second data is data obtained as a result of performing second operations on first data stored in the plurality of memory devices through a second accelerator circuit provided in the memory package, wherein the second operations comprise a filtering operation performed based on the first data.
11 . The memory controller of claim 10 , wherein the plurality of memory devices correspond to dynamic random access memories (DRAMs) which are three-dimensionally stacked inside the memory package,
wherein the second accelerator circuit is provided in a master DRAM contacting a substrate of the memory package, and
wherein, other than the master DRAM, the DRAMs are connected to the master DRAM through wire bonding.
12 . The memory controller of claim 10 , wherein the plurality of memory devices correspond to DRAMs which are three-dimensionally stacked on a first die inside the memory package,
wherein the second accelerator circuit is provided on a second die different from the first die, and
wherein the DRAMs are connected to the second accelerator circuit through wire bonding.
13 . The memory controller of claim 10 , wherein the filtering operation comprises a zeroing operation in which, among the first data, a portion having a size smaller than a threshold value is changed into zero to obtain filtered data, and
wherein the second operations comprise a pruning operation in which zero data and NULL data are removed from the filtered data to obtain the second data.
14 . The memory controller of claim 10 , wherein the first operations comprise a general matrix vector multiplication (GEMV) operation of matrix-to-vector multiplication and a general matrix-matrix multiplication (GEMM) operation of matrix-to-matrix multiplication.
15 . The memory controller of claim 10 , wherein the memory controller is configured to selectively perform one of the first operations based on the first accelerator circuit or the second operations based on the second accelerator circuit.
16 . The memory controller of claim 10 , wherein a time required for the first operations and the second operations is less than a predefined threshold value.
17 . An operating method of a memory controller connected to a memory package comprising a plurality of memory devices, the operating method comprising:
receiving output data from a first accelerator circuit which has performed a coarse acceleration operation on sparse data stored in the plurality of memory devices to obtain the output data, wherein the coarse acceleration operation comprises a filtering operation performed based on the sparse data;
obtaining an inference result by performing a fine acceleration operation on the output data by using a second accelerator circuit inside the memory controller; and
providing the inference result to a host.
18 . The operating method of claim 17 , wherein the filtering operation comprises a zeroing operation in which, among the sparse data, a portion having a size smaller than a threshold value is changed into zero to obtain filtered data,
wherein the coarse acceleration operation comprises a pruning operation in which zero data and NULL data are removed from the filtered data to obtain the output data, and
wherein the fine acceleration operation comprises a general matrix vector multiplication (GEMV) operation of matrix-to-vector multiplication and a general matrix-matrix multiplication (GEMM) operation of matrix-to-matrix multiplication.
19 . The operating method of claim 17 , wherein the plurality of memory devices correspond to dynamic random access memories (DRAMs) which are three-dimensionally stacked,
wherein the first accelerator circuit is provided in a master DRAM contacting a substrate among the DRAMs, and
wherein, other than the master DRAM, the DRAMs are connected to the master DRAM through wire bonding.
20 . The operating method of claim 17 , wherein the plurality of memory devices correspond to DRAMs which are three-dimensionally stacked on a first die,
wherein the first accelerator circuit is provided on a second die different from the first die, and
wherein the DRAMs are connected to the first accelerator circuit through wire bonding.