IP Library › Granted Patent US 12,566,951
Granted Patent B2
US 12,566,951 · App. 17/307,072 · Granted Mar 3, 2026

Method and apparatus for performing deep learning operations

Inventors: Dongyoung Kim (Suwon-si, KR); Sehwan Lee (Seongnam-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06N3/08G06F9/30036G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,951
App. No.
17/307,072
Granted
Mar 3, 2026
Kind
B2
Abstract

A method and apparatus for performing deep learning operations. A computation apparatus includes an adder tree-based tensor core configured to perform a tensor operation, and a multiplier and accumulator (MAC)-based vector core configured to perform a vector operation using an output of the tensor core as an input.

Claims (53)

1 . An apparatus for performing a deep learning operation, the apparatus comprising:

a tensor core configured to perform a tensor operation, the tensor core being based on an adder tree; and

a vector core configured to perform a vector operation using an output of the tensor core received from the tensor core and weight data as an input, the vector core being based on a multiplier and accumulator (MAC).

2 . The apparatus of claim 1 , wherein

the tensor operation comprises a matrix-matrix multiplication operation, and

the vector operation comprises any one or any combination of a vector-matrix multiplication operation, a vector-vector multiplication operation, and an element-wise operation.

3 . The apparatus of claim 1 , wherein the vector core comprises a MAC-based arithmetic logic unit (ALU) configured to perform the vector operation.

4 . The apparatus of claim 1 , wherein the vector core comprises a functional unit configured to perform one or both of a pooling operation and a non-linear function operation.

5 . The apparatus of claim 4 , wherein the functional unit comprises a look-up table to perform one or both of the pooling operation and the non-linear function operation.

6 . The apparatus of claim 1 , wherein the vector core comprises a weight buffer configured to store the weight data.

7 . The apparatus of claim 1 , further comprising:

a local buffer configured to store data to enable the vector core to reuse the output of the tensor core to perform the vector operation.

8 . The apparatus of claim 1 , wherein the vector core is configured to perform a first vector operation using an output of a first tensor operation as an input while the tensor core is performing a second tensor operation.

9 . The apparatus of claim 1 , wherein the vector core comprises:

a functional unit configured to perform one or both of a pooling operation and a non-linear function operation;

a MAC-based arithmetic logic unit (ALU) configured to perform the vector operation;

a weight buffer configured to store the weight data; and

a first multiplexer configured to select at least one of the output of the tensor core and an output of the ALU as an input of the functional unit.

10 . The apparatus of claim 9 , wherein the vector core comprises a second multiplexer configured to select at least one of an output of the functional unit and an output of the first multiplexer as an input of the ALU.

11 . The apparatus of claim 10 , wherein the vector core comprises a third multiplexer configured to select at least one of an output of the functional unit and the output of the ALU as an output of the vector core.

12 . The apparatus of claim 1 , wherein the tensor core is configured to perform a traversal for the tensor operation in units of blocks for performing the vector operation.

13 . The apparatus of claim 1 , wherein

the tensor core is configured to perform a convolution operation,

the vector core comprises:

a functional unit configured to perform one or both of a pooling operation and a non-linear function operation;

a MAC-based arithmetic logic unit (ALU) configured to perform the vector operation; and

a weight buffer configured to store the weight data,

the functional unit is configured to receive an output of the convolution operation as an input and to perform a first activation function operation,

the ALU is configured to perform a depth-wise convolution operation between the weight data and a result of the first activation function operation, and

the functional unit is configured to receive a result of the depth-wise convolution operation as an input and to perform a second activation function operation.

14 . A method of performing a deep learning operation, the method comprising:

performing a tensor operation using a tensor core that is based on an adder tree; and

performing a vector operation using an output of the tensor core received from the tensor core and weight data as an input, by a vector core that is based on a multiplier and accumulator (MAC).

15 . The method of claim 14 , wherein

the tensor operation comprises a matrix-matrix multiplication operation, and

the vector operation comprises any one or any combination of a vector-matrix multiplication operation, a vector-vector multiplication operation, and an element-wise operation.

16 . The method of claim 14 , wherein

performing the tensor operation comprises performing a convolution operation, and

performing the vector operation comprises:

receiving an output of the convolution operation as an input and performing a first activation function operation;

performing a depth-wise convolution operation between the weight data and a result of the first activation function operation; and

receiving a result of the depth-wise convolution operation as an input and performing a second activation function operation.

17 . The method of claim 14 , wherein performing the tensor operation comprises performing a traversal for the tensor operation in units of blocks for performing the vector operation.

18 . A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 14 .

19 . An apparatus comprising:

one or more processors configured to:

perform a first tensor operation and perform a second tensor operation subsequent to the first tensor operation, and

perform a vector operation using an output of the first tensor operation and weight data as an input while simultaneously performing the second tensor operation.

20 . The apparatus of claim 19 , wherein the one or more processors comprise:

an adder tree-based artificial neural network (ANN) accelerator configured to perform the first tensor operation and the second tensor operation; and

a multiplier and accumulator (MAC)-based processor configured to perform the vector operation.

21 . The apparatus of claim 19 , wherein the first tensor operation is a first convolution operation, the second tensor operation is a second convolution operation, and the vector operation is a depth-wise convolution operation.

22 . The apparatus of claim 21 , wherein the one or more processors are configured to perform an activation function operation on a result of the first convolution operation to generate the output of the first tensor operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 4, 2021
From: KIM, DONGYOUNG; LEE, SEHWAN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 056124/0942 →
Priority Claims (1)
KR 10-2020-0167970 · Dec 4, 2020 · national
Continuity (1)
Related Publication 20220180187A1 · Jun 9, 2022
References Cited (18)
US 20160342889A1 · Thorson et al. · 2016 [cited by applicant]
US 20180129936A1 · Young · 2018 [cited by examiner]
US 20190042251A1 · Nurvitadhi et al. · 2019 [cited by applicant]
US 20190138567A1 · Martin et al. · 2019 [cited by applicant]
US 20190392287A1 · Ovsiannikov et al. · 2019 [cited by applicant]
US 20200028444A1 · Tochikawa · 2020 [cited by examiner]
US 20200134417A1 · Mohapatra · 2020 [cited by examiner]
US 20200293867A1 · Shao et al. · 2020 [cited by applicant]
US 20210397933A1 · Boesch · 2021 [cited by examiner]
CN 210924662U · 2020 [cited by applicant]
KR 1020190084088A · 2019 [cited by applicant]
KR 102034659B1 · 2019 [cited by applicant]
KR 1020200070089A · 2020 [cited by applicant]
KR 1020200128360A · 2020 [cited by applicant]
Chen, Chixiao, et al., “iFPNA: A Flexible and Efficient Deep Learning Processor in 28-nm CMOS Using a Domain-Specific Instruction Set and Reconfigurable Fabric”, [cited by applicant]
Extended European search report issued on Feb. 8, 2022, in counterpart European Patent Application No. 21185561.4 (7 pages in English). [cited by applicant]
Zhu, Maohua, et al., “Sparse Tensor Core: Algorithm and Hardware Co-Design for Vector-wise Sparse Neural Networks on Modern GPUs”, Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, 20… [cited by applicant]
Korean Office Action issued on Jan. 23, 2025, in corresponding Korean Patent Application No. 10-2020-0167970. (2pages in English, 6 pages in Korean). [cited by applicant]