IP Library › Granted Patent US 12,265,912
Granted Patent B2
US 12,265,912 · App. 17/151,966 · Granted Apr 1, 2025

Deep neural network training accelerator and operation method thereof

Inventors: Jong Sun Park (Seoul, KR); Dong Yeob Shin (Seoul, KR); Geon Ho Kim (Gunpo-si, KR); Joong Ho Jo (Seoul, KR)
Assignee: Korea University Research and Business Foundation
G06N3/084G06F9/5027G06F18/2148G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,912
App. No.
17/151,966
Granted
Apr 1, 2025
Kind
B2
Abstract

A deep neural network training accelerator includes an operational unit sequentially performing first and second operations on a plurality of input data of a sub-set according to a mini-batch gradient descent, a determination unit determining each of the input data as one of skip data and training data based on a confidence matrix obtained by the first operation, and a control unit controlling the operational unit to skip the second operation with respect to the skip data.

Claims (38)

1. A deep neural network training accelerator implemented in hardware and configured to increase learning speed and reduce learning energy, the deep neural network training accelerator comprising:

an operational unit, executed by the deep neural network training accelerator, to sequentially perform first and second operations on a plurality of input data of a sub-set according to a mini-batch gradient descent;

a determination unit, executed by the deep neural network training accelerator, to determine each of the plurality of input data as one of skip data and training data based on a confidence matrix obtained by the first operation; and

a control unit, executed by the deep neural network training accelerator, to control the operational unit to skip the second operation with respect to the skip data,

wherein the determination unit comprises a comparator, the comparator executed by the deep neural network training accelerator to:

compare a largest element among elements of the confidence matrix with a predetermined threshold,

output a low signal corresponding to the skip data to the control unit when a value of the largest element is equal to or greater than the predetermined threshold, and output a high signal corresponding to the training data to the control unit when a value of the largest element is smaller than the predetermined threshold, and

wherein the control unit, executed by the deep neural network training accelerator, to output a parallelization control signal to the operation unit in response to the low signal being received to control the operation unit to skip the second operation on the skip data and perform the second operation on the training data based on the parallelization control signal,

wherein the performing of the second operation comprises:

initializing, by the operational unit, any one operational device corresponding to the skip data among operational devices in response to the parallelization control signal;

reassigning, by the operational unit, a portion of each of the training data assigned to the other operational devices to the any one operational device; and

processing, by the operational unit, the second operation with respect to the training data in parallel using the operational devices after a predetermined time has elapsed from a time at which the first operation is performed.

2. The deep neural network training accelerator of claim 1 , wherein the operational unit performs the second operation with respect to the training data after a predetermined time has elapsed from a time at which the first operation is performed.

3. The deep neural network training accelerator of claim 1 , wherein the first operation is a first training stage of the mini-batch gradient descent, which uses a forward propagation algorithm.

4. The deep neural network training accelerator of claim 1 , wherein the second operation is a second training stage of the mini-batch gradient descent, which sequentially uses a backward propagation algorithm and a weight update algorithm.

5. The deep neural network training accelerator of claim 1 , wherein the control unit parallelizes the second operation with respect to the training data in response to the low signal.

6. The deep neural network training accelerator of claim 1 , wherein a number of the low signals is inversely proportional to an operation time of the second operation.

7. The deep neural network training accelerator of claim 1 , further comprising: an input unit assigning each of the input data arbitrarily selected from total input data to the operational unit; and an output unit summing each variation in weight output through the operational unit to output a variation in output weight corresponding to a gradient of the sub-set.

8. The deep neural network training accelerator of claim 1 , wherein the operational unit has a systolic array structure and comprises operational devices that sequentially perform the first and second operations.

9. The deep neural network training accelerator of claim 8 , wherein the operational unit initializes any one operational device corresponding to the skip data among the operational devices in response to the parallelization control signal applied thereto from the control unit.

10. The deep neural network training accelerator of claim 9 , wherein the operational unit reassigns a portion of the training data assigned to the other operational devices among the operational devices to the any one operational device.

11. The deep neural network training accelerator of claim 8 , wherein the control unit reassigns a plurality of sub-data divided from each of the training data to the operational devices according to a data flow.

12. The deep neural network training accelerator of claim 11 , wherein the data flow refers to a data movement path for reading and storing data.

13. A method of operating a deep neural network training accelerator implemented in hardware and configured to increase learning speed and reduce learning energy, the method comprising:

performing, by an operational unit, first and second operations on a plurality of input data of a sub-set according to a mini-batch gradient descent;

determining, by a determination unit, each of the plurality of input data as one of skip data and training data based on a confidence matrix obtained by the first operation;

controlling, by a control unit, the operational unit to skip the second operation with respect to the skip data;

wherein the determination unit comprises a comparator, configured to the comparator executed by the deep neural network training accelerator to:

compare a largest element among elements of the confidence matrix with a predetermined threshold,

output a low signal corresponding to the skip data to the control unit when a value of the largest element is equal to or greater than the predetermined threshold, and

output a high signal corresponding to the training data to the control unit when a value of the largest element is smaller than the predetermined threshold; and

output, by the control unit, a parallelization control signal to the operation unit in response to the low signal being received to control the operation unit to skip the second operation on the skip data and perform the second operation on the training data based on the parallelization control signal,

wherein the performing of the second operation comprises:

initializing, by the operational unit, any one operational device corresponding to the skip data among operational devices in response to the parallelization control signal;

reassigning, by the operational unit, a portion of each of the training data assigned to the other operational devices to the any one operational device; and

processing, by the operational unit, the second operation with respect to the training data in parallel using the operational devices after a predetermined time has elapsed from a time at which the first operation is performed.

14. The method of claim 13 , wherein the first operation is a first training stage of the mini-batch gradient descent, which uses a forward propagation algorithm.

15. The method of claim 13 , wherein the second operation is a second training stage of the mini-batch gradient descent, which sequentially uses a backward propagation algorithm and a weight update algorithm.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2021
From: PARK, JONG SUN; SHIN, DONG YEOB; KIM, GEON HO; JO, JOONG HO
To: KOREA UNIVERSITY RESEARCH AND BUSINESS FOUNDATION
Reel/Frame 054951/0917 →
Priority Claims (1)
KR 10-2020-0089210 · Jul 17, 2020 · national
Continuity (1)
Related Publication 20220019897A1 · Jan 20, 2022
References Cited (22)
US 8706798B1 · Suchter et al. · 2014 [cited by applicant]
US 9547818B2 · Osogami · 2017 [cited by examiner]
US 10423861B2 · Gao · 2019 [cited by examiner]
US 10699189B2 · Lie · 2020 [cited by examiner]
US 11062202B2 · James · 2021 [cited by examiner]
US 11544545B2 · Baum · 2023 [cited by examiner]
US 20170357896A1 · Tsatsin · 2017 [cited by examiner]
US 20180314941A1 · Lie · 2018 [cited by examiner]
US 20190049231A1 · Choi · 2019 [cited by examiner]
US 20190205740A1 · Judd et al. · 2019 [cited by applicant]
US 20190244083A1 · Franca-Neto · 2019 [cited by examiner]
US 20210142155A1 · James · 2021 [cited by examiner]
US 20210142167A1 · Lie · 2021 [cited by examiner]
KR 101987475B1 · 2019 [cited by applicant]
KR 102034659B1 · 2019 [cited by applicant]
KR 1020200071865A · 2020 [cited by applicant]
KR 1020200076800A · 2020 [cited by applicant]
Master et al. “Revisting Small Batch Training For Deeep Neural Networks,” Graphcore Research Bistrol, UK, pp. 1-18, 2018. [cited by examiner]
Shrestha et al. “Review of Deep Learning Algorithms and Architectures,” in IEEE Access, vol. 7, pp. 53040-53065, 2019. [cited by examiner]
Rojas, Raul, and Raúl Rojas “The backpropagation algorithm”, Neural networks: a systematic introduction, 1996, pp. 149-182. [cited by examiner]
Bendelac, Shiri. “Enhanced Neural Network Training Using Selective Backpropagation and Forward Propagation” Diss. Virginia Tech, 2018, (97 pages in English). [cited by applicant]
Jiang, Angela H., et al. “Accelerating Deep Learning by Focusing on the Biggest Losers.” arXiv preprint arXiv:1910.00762, 2019, (14 pages in English). [cited by applicant]