IP Library › Granted Patent US 12,271,807
Granted Patent B2
US 12,271,807 · App. 17/250,892 · Granted Apr 8, 2025

Convolutional neural network computing method and system based on weight kneading

Inventors: Xiaowei Li (Beijing, CN); Xin Wei (Beijing, CN); Hang Lu (Beijing, CN)
Assignee: Institute of Computing Technology, Chinese Academy of Sciences
G06N3/048G06F5/01G06F7/50G06F7/5443G06F17/16G06N3/04G06N3/063H03M7/40G06F2207/386
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,271,807
App. No.
17/250,892
Granted
Apr 8, 2025
Kind
B2
Abstract

Disclosed embodiments relate to a convolutional neural network computing method and system based on weight kneading, comprising: arranging original weights in a computation sequence and aligning by bit to obtain a weight matrix, removing slack bits in the weight matrix, allowing essential bits in each column of the weight matrix to fill the vacancies according to the computation sequence to obtain an intermediate matrix, removing null rows in the intermediate matrix, obtain a kneading matrix, wherein each row of the kneading matrix serves as a kneading weight; obtaining positional information of the activation corresponding to each bit of the kneading weight; divides the kneading weight by bit into multiple weight segments, processing summation of the weight segments and the corresponding activations according to the positional information, and sending a processing result to an adder tree to obtain an output feature map by means of executing shift-and-add on the processing result.

Claims (16)

1. A convolutional neural network computing system based on weight kneading, comprising:

a weight kneading module for acquiring multiple groups of activations to be operated and corresponding original weights, arranging the original weights in a computation sequence and aligning by bit to obtain a weight matrix, removing slack bits in the weight matrix to obtain a reduced matrix with vacancies, allowing essential bits in each column of the reduced matrix to fill the vacancies according to the computation sequence to obtain an intermediate matrix, removing null rows in the intermediate matrix, and placing zeros at vacancies of the intermediate matrix to obtain a kneading matrix, wherein each row of the kneading matrix serves as a kneading weight; and

a split accumulation module for obtaining, according to a correspondence relationship between the activations and the essential bits in the original weights, positional information of the activation corresponding to each bit of the kneading weight, sending the kneading weight to a split accumulator, which divides the kneading weight by bit into multiple weight segments, processing summation of the weight segments and the corresponding activations according to the positional information, and sending a processing result to an adder tree to obtain an output feature map by means of executing shift-and-add on the processing result,

wherein the original weights are 16-bit fixed-point numbers, and the split accumulator comprises splitters for dividing the kneading weight by bit.

2. The convolutional neural network computing system based on weight kneading according to claim 1 , wherein the split accumulation module comprises saving the positional information of the activation corresponding to each bit of the kneading weight with Huffman coding.

3. The convolutional neural network computing system based on weight kneading according to claim 1 , wherein the activations are pixel values of an image.

4. A convolutional neural network computing method based on weight kneading, comprising:

step 1, acquiring multiple groups of activations to be operated and corresponding original weights, arranging the original weights in a computation sequence and aligning by bit to obtain a weight matrix, removing slack bits in the weight matrix to obtain a reduced matrix with vacancies, allowing essential bits in each column of the reduced matrix to fill the vacancies according to the computation sequence to obtain an intermediate matrix, removing null rows in the intermediate matrix, and placing zeros at vacancies of the intermediate matrix to obtain a kneading matrix, wherein each row of the kneading matrix serves as a kneading weight;

step 2, obtaining, according to a correspondence relationship between the activations and the essential bits in the original weights, positional information of the activation corresponding to each bit of the kneading weight;

step 3, sending the kneading weight to a split accumulator, which divides the kneading weight by bit into multiple weight segments, processing summation of the weight segments and the corresponding activations according to the positional information, and sending a processing result to an adder tree to obtain an output feature map by means of executing shift-and-add on the processing result.

5. The convolutional neural network computing method based on weight kneading according to claim 4 , wherein the split accumulator in the step 3 comprises splitters for dividing the kneading weight by bit, and segment adders for processing summation of the weight segments and the corresponding activations.

6. The convolutional neural network computing method based on weight kneading according to claim 4 , wherein the step 2 comprises saving the positional information of the activation corresponding to each bit of the kneading weight with Huffman coding.

7. The convolutional neural network computing method based on weight kneading according to claim 4 , wherein the activations are pixel values of an image.

8. The convolutional neural network computing system based on weight kneading according to claim 2 , wherein the activations are pixel values of an image.

9. The convolutional neural network computing method based on weight kneading according to claim 5 , wherein the step 2 comprises saving the positional information of the activation corresponding to each bit of the kneading weight with Huffman coding.

10. The convolutional neural network computing method based on weight kneading according to claim 6 , wherein the activations are pixel values of an image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2021
From: LI, XIAOWEI; WEI, XIN; LU, HANG
To: INSTITUTE OF COMPUTING TECHNOLOGY, CHINESE ACADEMY OF SCIENCES
Reel/Frame 055657/0814 →
Priority Claims (1)
CN 201811100309.2 · Sep 20, 2018 · national
Continuity (1)
Related Publication 20210350214A1 · Nov 11, 2021
References Cited (10)
US 20210004668A1 · Moshovos · 2021 [cited by examiner]
CN 103679185A · 2014 [cited by applicant]
CN 107086910A · 2017 [cited by applicant]
CN 109543816S · 2019 [cited by applicant]
US Office Action, as issued in connection with U.S. Appl. No. 17/250,889, dated Feb. 29, 2024, 21 pgs. [cited by applicant]
Delmas, Alberto, et al. “Bit-tactical: Exploiting ineffectual computations in convolutional neural networks: Which, why, and how.” arXiv preprint arXiv: 1803.03688 (Mar. 2018). (Year: 2018). [cited by applicant]
Judd, Patrick, et al. “Stripes: Bit-serial deep neural network computing.” 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 2016. (Year: 2016). [cited by applicant]
International Search report mailed Jul. 30, 2019, in PCT Application No. PCT/CN2019/087767. [cited by applicant]
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, et al., “DaDianNao: A Machine-Learning Supercomputer,” presented at the Proceedings of the 47th Annual IEEE/ACM International Symposium on Microarchitecture, Cambridge,… [cited by applicant]
J. Albericio, A. Delm, P. Judd, S. Sharify, G. O'Leary, et al., “Bit-pragmatic deep neural network computing,” presented at the Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture, Cambr… [cited by applicant]