IP Library › Granted Patent US 12,205,036
Granted Patent B2
US 12,205,036 · App. 16/174,050 · Granted Jan 21, 2025

Apparatus and methods for training in fully connected layers of convolutional networks

Inventors: Qi Guo (Beijing, CN); Shijin Zhang (Beijing, CN); Yunji Chen (Beijing, CN); Tianshi Chen (Beijing, CN)
Assignee: CAMBRICON TECHNOLOGIES CORPORATION LIMITED
G06N3/084G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,036
App. No.
16/174,050
Granted
Jan 21, 2025
Kind
B2
Abstract

Aspects for backpropagation in a fully connect layer of a convolutional neural network are described herein. The aspects may include a direct memory access unit configured to receive input data and one or more first data gradients from a storage device. The aspects may further include a master computation module configured to transmit the input data and the one or more first data gradients to one or more slave computation modules. The slave computation modules are respectively configured to multiply one of the one or more first data gradients with the input data to generate a default weight gradient vector.

Claims (35)

1. An integrated circuit (IC) chip for backpropagation in a fully connected layer of a neural network, comprising:

a controller circuit configured to receive an instruction; and

one or more computation circuits that include:

a master computation circuit,

one or more slave computation circuits, and

an interconnection circuit communicatively connected to the master computation circuit and the one or more slave computation circuits,

wherein the master computation circuit configured to

receive input data and one or more first data gradients in response to the instruction, and

transmit the input data and the one or more first data gradients to the one or more slave computation circuits, and

wherein the one or more slave computation circuits are respectively configured to multiply one of the one or more first data gradients with the input data to generate a default weight gradient vector,

wherein the master computation circuit is further configured to update one or more weight values based on the default weight gradient vector,

wherein the master computation circuit is further configured to apply a derivative of an activation function to the one or more first data gradients to generate one or more input gradients,

wherein the one or more slave computation circuits are respectively configured to multiply one of the one or more input gradients with one or more weight vectors in a weight matrix to generate one or more multiplication results, and

wherein the interconnection circuit is configured to combine the one or more multiplication results of a lower dimension into an output gradient vector of a higher dimension.

2. The IC chip of claim 1 , wherein the master computation circuit is further configured to

calculate a scaled weight gradient vector based on the default weight gradient vector and a predetermined threshold value; and

update one or more weight values based on the scaled weight gradient vector.

3. The IC chip of claim 1 , wherein the interconnection circuit is configured to channel data between the master computation circuit and the one or more slave computation circuits.

4. The IC chip of claim 1 , wherein each of the one or more slave computation circuits includes a slave neuron caching circuit, wherein the slave neuron caching circuit is configured to store the one or more first data gradients with the input data.

5. The IC chip of claim 1 , wherein each of the one or more slave computation circuits includes a weight value caching circuit, wherein the weight value caching circuit is configured to store the weight matrix that includes the one or more weight vectors.

6. A method for backpropagation in a fully connected layer of a neural network, comprising:

receiving, by a controller circuit, an instruction;

receiving, by a master computation circuit, input data and one or more first data gradients in response to the instruction;

transmitting, by the master computation circuit, the input data and the one or more first data gradients to one or more slave computation circuits, wherein the master computation circuit is communicatively connected to the one or more slave computation circuits via an interconnection circuit;

respectively multiplying, by the one or more slave computation circuits, one of the one or more first data gradients with the input data to generate a default weight gradient vector;

updating, by the master computation circuit, one or more weight values based on the default weight gradient vector;

applying, by the master computation circuit, a derivative of an activation function to the one or more first data gradients to generate one or more input gradients;

respectively multiplying, by the one or more slave computation circuits, one of the one or more input gradients with one or more weight vectors in a weight matrix to generate one or more multiplication results; and

combining, by the interconnection circuit, the one or more multiplication results of a lower dimension into an output gradient vector of a higher dimension.

7. The method of claim 6 , further comprising

calculating, by the master computation circuit, a scaled weight gradient vector based on the default weight gradient vector and a predetermined threshold value; and

updating, by the master computation circuit, one or more weight values based on the scaled weight gradient vector.

8. The method of claim 6 , further comprising channeling, by the interconnection circuit, data between the master computation circuit and the one or more slave computation circuits.

9. The method of claim 6 , further comprising storing, by a slave neuron caching circuit of each of the one or more slave computation circuits, the one or more first data gradients with the input data.

10. The method of claim 6 , further comprising storing, by a weight value caching circuit of each of the one or more slave computation circuits, the weight matrix that includes the one or more weight vectors.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2018
From: GUO, QI; ZHANG, SHIJIN; CHEN, YUNJI; CHEN, TIANSHI
To: CAMBRICON TECHNOLOGIES CORPORATION LIMITED
Reel/Frame 047871/0682 →
Priority Claims (1)
CN 201610285062.0 · Apr 29, 2016 · national
Continuity (2)
Continuation In Part PCTCN2016081114 · May 5, 2016
Related Publication 20190065958A1 · Feb 28, 2019
References Cited (30)
CN 105488565A · 2016 [cited by applicant]
CN 106991478A · 2017 [cited by applicant]
WO WO2017185394A1 · 2017 [cited by applicant]
WO WO2017185248A1 · 2017 [cited by examiner]
Touretzky, Backpropagation Learning, Lecture 15-486/782: Artificial Neural Networks, Computer Science Carnegie Mellon University, 2006 (Year: 2006). [cited by examiner]
Lee Performance Analysis of Bit-Width Reduced FPU in FPGAs, Journal of Embedded Systems vol. 2009 (Year: 2009). [cited by examiner]
Zhang Optimizing FPGA-based Accelerator Design for Deep Convolutional Neural Networks, FPGA' 15 ACM, (Year: 2015). [cited by examiner]
Li Arithmetic formats for implementing artificial neural network, Can. J. Elect. Comput. Eng., vol. 31, No. 1, Winter, 2006 (Year: 2006). [cited by examiner]
Hamalainen, TUTNC: a general purpose parallel computer for neural network computations, Microprocessors and Microsystems vol. 19, 1995 (Year: 1995). [cited by examiner]
Fowers, A High Memory Bandwidth FPGA Accelerator for Sparse Matrix-Vector Multiplication, 2014 IEEE 22nd Annual International Symposium on Field-Programmable Custom Computing Machines, IEEE, 2014 (Year: 2014). [cited by examiner]
T. Chen, et al., “A Small-Footprint Accelerator for Large-Scale Neural Networks”, ACM Transactions on Computer Systems, vol. 33, No. 2, Article 6, May 2015, 27 pages. [cited by applicant]
Z. Du, et al., “An Accelerator for High Efficient Vision Processing”, IEEE Transactions on Computer-aided Design of Integrated Circuits and System, vol. 36, No. 2, Feb. 2017, pp. 227-240. [cited by applicant]
S. Liu, et al., “Cambricon: An Instruction Set Architecture for Neural Networks”, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Oct. 12, 2016, pp. 393-405. [cited by applicant]
S. Zhang, et al., “Cambricon-X″ An Accelerator for Sparse Neural Networks”, The 49th Annual IEEE/ACM International Symposium on Microarchitecture Article No. 20, Oct. 15, 2016, 12 pages. [cited by applicant]
Y. Chen, et al., “DaDianNao: A Machine-Learning Supercomputer”, 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, Dec. 13, 2014, pp. 609-622. [cited by applicant]
T. Luo, et al., “DaDianNao: A Neural Network Supercomputer”, IEEE Transaction on Computers, vol. 66, No. 1, Jan. 2017, pp. 73-88. [cited by applicant]
T. Chen, et al., “DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning”, ASPLOS '14, Proceedings of the 19th international conference on Architectural support for programming languages … [cited by applicant]
Y. Chen, et al., “DianNao Family: Energy-Efficient Hardware Accelerators for Machine Learning”, Communications of the ACM, vol. 59, No. 11, Nov. 2016, pp. 105-112. [cited by applicant]
D. Liu, et al., “PuDianNao: A Polyvalent Machine Learning Accelerator”, ASPLOS '15 Proceedings of the Twentieth International Conference on Architectural Support for Programming Languages and Operating Systems, Mar. 14,… [cited by applicant]
Z. Du, et al., “ShiDianNao: Shifting Vision Processing Closer to the Sensor”, ISCA '15 Proceedings of the 42nd Annual International Symposium on Computer Architecture, Jun. 13, 2015, pp. 92-104. [cited by applicant]
EP 16899905.0—European Search Report, mailed Nov. 15, 2019, 4 pages. [cited by applicant]
Ramon J. Aliaga, et al., “System-on-Chip Implementation of Neural Network Training on FPGA”, International Journal of Advances in Systems and Measurement, 2009, 12 pages. [cited by applicant]
Pedro O. Domingos, et al., “An Efficient and Scalable Architecture for Neural Networks With Backpropagation Learning” IEEE, 2005, 6 pages. [cited by applicant]
EP 16899905.0—Response to the Communication under Article 94(3) EPC, filed Jun. 24, 2020, 33 pages. [cited by applicant]
EP 16899905.0—Summons to attend Oral Proceedings, mailed Sep. 9, 2020, 12 pages. [cited by applicant]
Timo Hamalainen, et al., “TUTNC: a general purpose parallel computer for neural network computations”, Elsevier Science B.V. Microprocessors and Microsystems vol. 19 No. 8 Oct. 1995, 19 pages. [cited by applicant]
Harri Klapuri, et. al., “Mapping artificial neural networks to a tree shape neurocomputer”, 1996 Elsevier Science B.V., 10 pages. [cited by applicant]
CN 201610285062.0—Office Action, mailed Jul. 10, 2020, 9 pages. (no English translation). [cited by applicant]
PCT/CN2016/081114—International Search Report, mailed Feb. 3, 2017, 13 pages. (no English translation). [cited by applicant]
KR 1020187033950—Office Action, mailed Apr. 25, 2022, 8 pages, (with English translation). [cited by applicant]