IP Library › Granted Patent US 12,554,991
Granted Patent B2
US 12,554,991 · App. 16/174,108 · Granted Feb 17, 2026

Device and method for performing self-learning operations of an artificial neural network

Inventors: Zhen Li (Beijing, CN); Qi Guo (Beijing, CN); Yunji Chen (Beijing, CN); Tianshi Chen (Beijing, CN)
Assignee: CAMBRICON TECHNOLOGIES CORPORATION LIMITED
G06N3/08G06F17/16G06N3/04G06N3/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,554,991
App. No.
16/174,108
Granted
Feb 17, 2026
Kind
B2
Abstract

Aspects for self-learning operations of an artificial neural network are described herein. The aspects may include a master computation module configured to transmit an input vector via an interconnection unit and one or more slave computation modules connected to the master computation module via the interconnection unit. Each of the one or more slave computation modules may be configured to respectively store a column weight vector of a weight matrix and multiply the input vector with the column weight vector to generate a first multiplication result. The interconnection unit may be configured to combine the one or more first multiplication results into a first multiplication vector and transmit the first multiplication vector to the master computation module.

Claims (70)

1 . An apparatus for neural network operations, the apparatus comprising:

a master computation circuit configured to transmit an input vector via an interconnection circuit; and

one or more slave computation circuits connected to the master computation circuit via the interconnection circuit,

wherein each of the one or more slave computation circuits is configured to:

respectively store a column weight vector of a weight matrix, and

multiply the input vector with the column weight vector to generate a first multiplication result, and

wherein the interconnection circuit is configured to:

combine one or more first multiplication results of the one or more slave computation circuits into a first multiplication vector, wherein each first multiplication result of the one or more first multiplication results is associated with a respective slave computation circuit of the one or more slave computation circuits, and

transmit the first multiplication vector to the master computation circuit,

wherein the master computation circuit is further configured to:

add a first bias vector to the first multiplication vector to generate a first biased vector,

activate the first biased vector by applying a first activation function to the first biased vector to generate a first activated vector,

sample the first activated vector by a Gibbs sampler to generate a first phase hidden layer vector, and

transmit the first phase hidden layer vector to the one or more slave computation circuits via the interconnection circuit,

wherein each of the one or more slave computation circuits is configured to:

multiply the column weight vector with an element of the first phase hidden layer vector to generate a second multiplication result vector, and

transmit the second multiplication result vector to the interconnection circuit, and

wherein the interconnection circuit is configured to:

add one or more second multiplication result vectors of the one or more slave computation circuits into a second multiplication vector, wherein each second multiplication result vector of the one or more second multiplication result vectors is associated with a respective slave computation circuit of the one or more slave computation circuits, and

transmit the second multiplication vector to the master computation circuit, wherein the master computation circuit is configured to:

add a second bias vector to the second multiplication vector to generate a second biased vector,

activate the second biased vector by applying a second activation function to the second biased vector to generate a second activated vector, and

sample the second activated vector by the Gibbs sampler to generate a second phase visible layer vector,

wherein the master computation circuit is configured to transmit the second phase visible layer vector to the one or more slave computation circuits via the interconnection circuit,

wherein each of the one or more slave computation circuits is configured to:

multiply the second phase visible layer vector with the column weight vector to generate a third multiplication result, and

transmit the third multiplication result to the interconnection circuit, and

wherein the interconnection circuit is configured to:

combine one or more third multiplication results of the one or more slave computation circuits into a third multiplication vector, wherein each third multiplication result of the one or more third multiplication results is associated with a respective slave computation circuit of the one or more slave computation circuits, and

transmit the third multiplication vector to the master computation circuit,

wherein the master computation circuit is configured to:

add a third bias vector to the third multiplication vector to generate a third biased vector, and

activate the third biased vector by applying a third activation function to the third biased vector to generate a third phase hidden layer vector,

wherein the master computation circuit is configured to transmit the input vector, the first phase hidden layer vector, the second phase visible layer vector, and the third phase hidden layer vector to the one or more slave computation circuits via the interconnection circuit,

wherein the one or more slave computation circuits are configured to:

calculate a first cross product between a transpose of the input vector and the first phase hidden layer vector,

calculate a second cross product between a transpose of the second phase visible layer vector and the third phase hidden layer vector, and

update the weight matrix based on a learning rate and a difference between the first cross product and the second cross product.

2 . The apparatus of claim 1 , wherein the one or more slave computation circuits are configured to update the first bias vector based on a learning rate and a difference between the first phase hidden layer vector and the third phase hidden layer vector.

3 . The apparatus of claim 1 , wherein the one or more slave computation circuits are configured to update the second bias vector based on a learning rate and a difference between the input vector and the second phase visible layer vector.

4 . A method for neural network operations, the method comprising:

transmitting, by a master computation circuit, an input vector via an interconnection circuit;

respectively storing, by each of one or more slave computation circuits connected to the master computation circuit via an interconnection circuit, a column weight vector of a weight matrix;

multiplying, by each of the one or more slave computation circuits, the input vector with the column weight vector to generate a first multiplication result;

combining, by the interconnection circuit, one or more first multiplication results of the one or more slave computation circuits into a first multiplication vector, wherein each first multiplication result of the one or more first multiplication results is associated with a respective slave computation circuit of the one or more slave computation circuits;

transmitting, by the interconnection circuit, the first multiplication vector to the master computation circuit,

adding, by the master computation circuit, a first bias vector to the first multiplication vector to generate a first biased vector,

activating, by the master computation circuit, the first biased vector by applying a first activation function to the first biased vector to generate a first activated vector, and

sampling, by the master computation circuit, the first activated vector by a Gibbs sampler to generate a first phase hidden layer vector,

transmitting, by the master computation circuit, the first phase hidden layer vector to the one or more slave computation circuits via the interconnection circuit;

multiplying, by each of the one or more slave computation circuits, the column weight vector with an element of the first phase hidden layer vector to generate a second multiplication result vector;

transmitting, by each of the one or more slave computation circuits, the second multiplication result vectors to the interconnection circuit;

adding, by the interconnection circuit, one or more second multiplication result vectors of the one or more slave computation circuits into a second multiplication vector, wherein each second multiplication result vector of the one or more second multiplication result vectors is associated with a respective slave computation circuit of the one or more slave computation circuits;

transmitting, by the interconnection circuit, the second multiplication vector to the master computation circuit,

adding, by the master computation circuit, a second bias vector to the second multiplication vector to generate a second biased vector,

activating, by the master computation circuit, the second biased vector by applying a second activation function to the second biased vector to generate a second activated vector, and

sampling, by the master computation circuit, the second activated vector by the Gibbs sampler to generate a second phase visible layer vector,

transmitting, by the master computation circuit, the second phase visible layer vector to the one or more slave computation circuits via the interconnection circuit;

multiplying, by each of the one or more slave computation circuits, the second phase visible layer vector with the column weight vector to generate a third multiplication result;

transmitting, by each of the one or more slave computation circuits, the third multiplication result to the interconnection circuit; and

combining, by the interconnection circuit, one or more third multiplication results of the one or more slave computation circuits into a third multiplication vector, wherein each third multiplication result of the one or more third multiplication results is associated with a respective slave computation circuit of the one or more slave computation circuits;

transmitting, by the interconnection circuit, the third multiplication vector to the master computation circuit,

adding, by the master computation circuit, a third bias vector to the third multiplication vector to generate a third biased vector,

activating, by the master computation circuit, the third biased vector by applying a third activation function to the third biased vector to generate a third phase hidden layer vector,

transmitting, by the master computation circuit, the input vector, the first phase hidden layer vector, the second phase visible layer vector, and the third phase hidden layer vector to the one or more slave computation circuits via the interconnection circuit,

calculating, by the one or more slave computation circuits, a first cross product between a transpose of the input vector and the first phase hidden layer vector,

calculating, by the one or more slave computation circuits, a second cross product between a transpose of the second phase visible layer vector and the third phase hidden layer vector, and

updating, by the one or more slave computation circuits, the weight matrix based on a learning rate and a difference between the first cross product and the second cross product.

5 . The method of claim 4 , further comprising updating, by the one or more slave computation circuits, the first bias vector based on a learning rate and a difference between the first phase hidden layer vector and the third phase hidden layer vector.

6 . The method of claim 4 , further comprising updating, by the one or more slave computation circuits, the second bias vector based on a learning rate and a difference between the input vector and the second phase visible layer vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2019
From: LI, ZHEN; GUO, QI; CHEN, YUNJI; CHEN, TIANSHI
To: CAMBRICON TECHNOLOGIES CORPORATION LIMITED
Reel/Frame 048388/0514 →
Continuity (2)
Continuation In Part PCTCN2016080320 · Apr 27, 2016
Related Publication 20190065953A1 · Feb 28, 2019
References Cited (28)
US 5909681A · Passera · 1999 [cited by examiner]
US 6058206A · Kortge · 2000 [cited by examiner]
US 10521715B1 · Loffe · 2019 [cited by examiner]
CN 104757992A · 2015 [cited by applicant]
CN 105488565A · 2016 [cited by applicant]
WO WO2017185248A1 · 2017 [cited by applicant]
Hämäläinen, Timo, Jukka Saarinen, and Kimmo Kaski. “TUTNC: A general purpose parallel computer for neural network computations.” Microprocessors and Microsystems 19.8: 447-465. (Year: 1995). [cited by examiner]
EP 16899762.5—European Search Report, mailed Nov. 28, 2019, 4 pages. [cited by applicant]
EP 16899762.5—Filing Receipt, mailed Oct. 12, 2018, 49 pages. [cited by applicant]
EP 16899762.5—Communication from Examining Division, mailed Jun. 22, 2020, 7 pages. [cited by applicant]
EP 16899762.5—Response to Communication pursuant to Article 94(3) EPC, mailed Apr. 27, 2020, 62 pages. [cited by applicant]
EP 16899762.5—Summons to Attend Oral Proceedings Pursuant to Rule 115(1) EPC, mailed Jun. 22, 2020, 2 pages. [cited by applicant]
NPL Copyright Statement, 1 page. [cited by applicant]
Timo Hamalainen, et al., “TUTNC: a general purpose parallel computer for neural network computations”, 1995 Elsevier Science B.V., 19 pages. [cited by applicant]
Harri Klapuri, et al., “Mapping Artificial Neural Networks to a Tree Shape Neurocomputer”, Microprocessors and Microsystems, Jan. 5, 1996, 10 pages. [cited by applicant]
CN 201610267211.0—First Office Action, mailed Sep. 2, 2020, 13 pages. [cited by applicant]
PCT/CN2016/080320—International Search Report, mailed Feb. 3, 2017, 14 pages. (no English translation). [cited by applicant]
EP16899762.5, Office Action mailed Dec. 12, 2020, 6 pages. [cited by applicant]
T. Chen, et al., “A Small-Footprint Accelerator for Large-Scale Neural Networks”, ACM Transactions on Computer Systems, vol. 33, No. 2, Article 6, May 2015, 27 pages. [cited by applicant]
Z. Du, et al., “An Accelerator for High Efficient Vision Processing”, IEEE Transactions on Computer-aided Design of Integrated Circuits and System, vol. 36, No. 2, Feb. 2017, pp. 227-240. [cited by applicant]
S. Liu, et al., “Cambricon: An Instruction Set Architecture for Neural Networks”, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Oct. 12, 2016, pp. 393-405. [cited by applicant]
S. Zhang, et al., “Cambricon-X” An Accelerator for Sparse Neural Networks, The 49th Annual IEEE/ACM International Symposium on Microarchitecture Article No. 20, Oct. 15, 2016, 12 pages. [cited by applicant]
Y. Chen, et al., “DaDianNao: A Machine-Learning Supercomputer”, 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, Dec. 13, 2014, pp. 609-622. [cited by applicant]
T. Luo, et al., “DaDianNao: A Neural Network Supercomputer”, IEEE Transaction on Computers, vol. 66, No. 1, Jan. 2017, pp. 73-88. [cited by applicant]
T. Chen, et al., “DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning”, ASPLOS '14, Proceedings of the 19th international conference on Architectural support for programming languages … [cited by applicant]
Y. Chen, et al., “DianNao Family: Energy-Efficient Hardware Accelerators for Machine Learning”, Communications of the ACM, vol. 59, No. 11, Nov. 2016, pp. 105-112. [cited by applicant]
D. Liu, et al., “PuDianNao: A Polyvalent Machine Learning Accelerator”, ASPLOS '15 Proceedings of the Twentieth International Conference on Architectural Support for Programming Languages and Operating Systems, Mar. 14,… [cited by applicant]
Z. Du, et al., “ShiDianNao: Shifting Vision Processing Closer to the Sensor”, ISCA '15 Proceedings of the 42nd Annual International Symposium on Computer Architecture, Jun. 13, 2015, pp. 92-104. [cited by applicant]