IP Library › Granted Patent US 12,423,560
Granted Patent B2
US 12,423,560 · App. 16/709,961 · Granted Sep 23, 2025

Apparatus and methods for forward propagation in convolutional neural networks

Inventors: Tianshi Chen (Beijing, CN); Dong Han (Beijing, CN); Yunji Chen (Beijing, CN); Shaoli Liu (Beijing, CN); Qi Guo (Beijing, CN)
Assignee: CAMBRICON TECHNOLOGIES CORPORATION LIMITED
G06N3/045G06F9/30G06F9/3001G06F13/362G06F17/16G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,560
App. No.
16/709,961
Granted
Sep 23, 2025
Kind
B2
Abstract

Aspects for forward propagation of a convolutional artificial neural network are described herein. The aspects may include a direct memory access unit configured to receive input data from a storage device and a master computation module configured to select one or more portions of the input data based on a predetermined convolution window. Further, the aspects may include one or more slave computation modules respectively configured to convolute a convolution kernel with one of the one or more portions of the input data to generate a slave output value. Further still, the aspects may include an interconnection unit configured to combine the one or more slave output values into one or more intermediate result vectors, wherein the master computation module is further configured to merge the one or more intermediate result vectors into a merged intermediate vector.

Claims (50)

1. An apparatus for neural network operations, comprising:

a direct memory access (DMA) circuit configured to retrieve input data and an instruction from an off-chip external memory;

a controller circuit configured to receive the instruction from the DMA circuit;

a computation circuit that includes a master computation circuit and one or more slave computation circuits,

wherein the DMA circuit is configured to write the input data to multiple on-chip caching circuits of the computation circuit,

wherein the master computation circuit is configured to:

receive the input data from the DMA circuit,

select, in response to an instruction, one or more portions of the input data from one of the multiple on-chip caching circuits, based on a predetermined convolution window, wherein the instruction includes a first address of the one or more portions of the input data, a first size of the one or more portions of the input data, a second address of a portion of a convolution kernel, and a second size of the portion of the convolution kernel, and

wherein the one or more slave computation circuits are respectively configured to convolute the portion of the convolution kernel with one of the one or more portions of the input data to generate a slave output value; and

an interconnection circuit configured to combine the one or more slave output values into one or more intermediate result vectors, and wherein the master computation circuit is further configured to merge the one or more intermediate result vectors into a merged intermediate vector.

2. The apparatus of claim 1 , wherein each of the one or more slave computation circuits includes a slave neuron caching circuit configured to store one of the one or more portions of the input data.

3. The apparatus of claim 1 , wherein each of the one or more slave computation circuits includes a weight value caching circuit configured to store the portion of the convolution kernel that corresponds to the slave computation circuit.

4. The apparatus of claim 1 , wherein each of the one or more slave computation circuits includes a vector multiplier configured to multiply the portion of the convolution kernel with each of the one or more portions of the input data.

5. The apparatus of claim 4 , wherein each of the one or more slave computation circuits includes an adder configured to sum results of a multiplication of the portion of the convolution kernel with each of the one or more portions of the input data to generate the slave output value.

6. The apparatus of claim 1 , wherein the master computation circuit includes a merging circuit configured to merge the one or more intermediate result vectors into the merged intermediate vector.

7. The apparatus of claim 6 , wherein the master computation circuit includes

a master neuron caching circuit configured to store a bias value; and,

an adder configured to add the bias value to the merged intermediate vector to generate a biased intermediate vector.

8. The apparatus of claim 1 , wherein the master computation circuit includes an activator configured to activate the biased intermediate vector by applying an activation function to the biased intermediate vector.

9. The apparatus of claim 8 , wherein the activation function is a function indicated by the instruction and selected from the group consisting of a sigmoid function, a tan h function, a relu function, and a softmax function.

10. The apparatus of claim 1 , wherein the instruction is selected from the group consisting of a convolution network sigmoid instruction, a convolution network tan h instruction, a convolution network relu instruction, and a convolution network group instruction.

11. The apparatus of claim 10 , wherein the convolution network sigmoid instruction includes an indication of a sigmoid function as the activation function.

12. The apparatus of claim 10 , wherein the convolution network tan h instruction includes an indication of a tan h function as the activation function.

13. The apparatus of claim 10 , wherein the convolution network relu instruction includes an indication of a relu function as the activation function.

14. The apparatus of claim 10 , wherein the convolution network group instruction includes an output address.

15. A method for neural network operations, comprising:

retrieving, by a direct memory access (DMA) circuit, input data and an instruction from an off-chip external memory;

receiving, by a controller circuit, the instruction from the DMA circuit;

receiving, by a computation circuit that includes a master computation circuit and one or more slave computation circuits, input data;

writing, by the DMA circuit, the input data to multiple on-chip caching circuits of the computation circuit;

selecting, by the master computation circuit of the computation circuit in response to an instruction, one or more portions of the input data from one of the multiple on-chip caching circuits, based on a predetermined convolution window, wherein the instruction includes a first address of the one or more portions of the input data, a first size of the one or more portions of the input data, a second address of a portion of a convolution kernel, and a second size of the portion of the convolution kernel;

convoluting, by the one or more slave computation circuits of the computation circuit, the portion of the convolution kernel with one of the one or more portions of the input data to generate a slave output value;

combining, by an interconnection circuit, the one or more slave output values into one or more intermediate result vectors; and

merging, by the master computation circuit of the computation circuit, the one or more intermediate result vectors into a merged intermediate vector.

16. The method of claim 15 , further comprising storing, by a slave neuron caching circuit of each of the one or more slave computation circuits, one of the one or more portions of the input data.

17. The method of claim 15 , further comprising storing, by a weight value caching circuit of each of the one or more slave computation circuits, the portion of the convolution kernel that corresponds to the slave computation circuit.

18. The method of claim 15 , further comprising multiplying, by a vector multiplier of each of the one or more slave computation circuits, the portion of the convolution kernel with each of the one or more portions of the input data.

19. The method of claim 18 , further comprising summing, by an adder of each of the one or more slave computation circuits, results of a multiplication of the portion of the convolution kernel with each of the one or more portions of the input data to generate the slave output value.

20. The method of claim 15 , further comprising merging, by a merging circuit of the master computation circuit, the one or more intermediate result vectors into the merged intermediate vector.

21. The method of claim 20 , further comprising:

storing, by a master neuron caching unit of the master computation circuit, a bias value; and

adding, by an adder of the master computation circuit, the bias value to the merged intermediate vector to generate a biased intermediate vector.

22. The method of claim 15 , further comprising activating, by an activator of the master computation circuit, the biased intermediate vector by applying an activation function to the biased intermediate vector.

23. The method of claim 22 , wherein the activation function is a function indicated by the instruction and selected from the group consisting of a sigmoid function, a tan h function, a relu function, and a softmax function.

24. The method of claim 15 , wherein the instruction is selected from the group consisting of a convolution network sigmoid instruction, a convolution network tan h instruction, a convolution network relu instruction, and a convolution network group instruction.

25. The method of claim 24 ,

wherein the convolution network sigmoid instruction includes an indication of a sigmoid function as the activation function,

wherein the convolution network tan h instruction includes an indication of a tan h function as the activation function,

wherein the convolution network relu instruction includes an indication of a relu function as the activation function, and

wherein the convolution network group instruction includes an output address.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2019
From: CHEN, TIANSHI; HAN, DONG; CHEN, YUNJI; LIU, SHAOLI; GUO, QI
To: CAMBRICON TECHNOLOGIES CORPORATION LIMITED
Reel/Frame 051240/0476 →
Priority Claims (1)
CN 201610282534.7 · Apr 29, 2016 · national
Continuity (3)
Continuation 16174155 · Oct 29, 2018
Continuation In Part PCTCN2016080967 · May 4, 2016
Related Publication 20200110983A1 · Apr 9, 2020
References Cited (59)
US 5204938A · Skapura et al. · 1993 [cited by applicant]
US 9153230B2 · Maaninen · 2015 [cited by applicant]
US 20110239032A1 · Kato et al. · 2011 [cited by applicant]
US 20140180989A1 · Krizhevsky et al. · 2014 [cited by applicant]
CN 1584824A · 2005 [cited by applicant]
CN 1700250A · 2005 [cited by applicant]
CN 101681450A · 2010 [cited by applicant]
CN 103150596A · 2013 [cited by applicant]
CN 103199806A · 2013 [cited by applicant]
CN 104063719A · 2014 [cited by applicant]
CN 104346622A · 2015 [cited by applicant]
CN 104809426A · 2015 [cited by applicant]
CN 104916322A · 2015 [cited by applicant]
CN 105184366A · 2015 [cited by applicant]
CN 105488565A · 2016 [cited by applicant]
CN 105512723A · 2016 [cited by applicant]
EP 2891946A1 · 2015 [cited by applicant]
WO 2016030230A1 · 2016 [cited by applicant]
WO 2017185386A1 · 2017 [cited by applicant]
CN 201610282534.7—First Office Action, mailed Jun. 3, 2019, 10 pages. (no English translation). [cited by applicant]
CN 201610282534.7—Second Office Action, mailed Mar. 9, 2020, 8 pages. (with brief English explanation). [cited by applicant]
PCT/CN2016/080967—International Search Report, mailed Jan. 26, 2017, 16 pages. (with brief English explanation). [cited by applicant]
CN 201710903509.0—Office Action, mailed Jun. 5, 2019, 12 pages. (with brief English explanation). [cited by applicant]
CN 201811148189.3—Office Action, mailed Sep. 19, 2019, 11 pages. (with brief English explanation). [cited by applicant]
EP 16899897.9—Article 94(3) EPC, mailed Sep. 7, 2021, 7 pages. [cited by applicant]
KR 10-2018-7033947, First Office Action, mailed Jun. 29, 2021, 12 pages. (with English translation). [cited by applicant]
KR 10-2018-7033947, Second Office Action, mailed Dec. 15, 2021, 29 pages. (with English translation). [cited by applicant]
T. Chen, et al., “A Small-Footprint Accelerator for Large-Scale Neural Networks”, ACM Transactions on Computer Systems, vol. 33, No. 2, Article 6, May 2015, 27 pages. [cited by applicant]
Z. Du, et al., “An Accelerator for High Efficient Vision Processing”, IEEE Transactions on Computer-aided Design of Integrated Circuits and System, vol. 36, No. 2, Feb. 2017, pp. 227-240. [cited by applicant]
S. Liu, et al., “Cambricon: An Instruction Set Architecture for Neural Networks”, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, Oct. 12, 2016, pp. 393-405. [cited by applicant]
S. Zhang, et al., “Cambricon-X” An Accelerator for Sparse Neural Networks, The 49th Annual IEEE/ACM International Symposium on Microarchitecture Article No. 20, Oct. 15, 2016, 12 pages. [cited by applicant]
Y. Chen, et al., “DaDianNao: A Machine-Learning Supercomputer”, 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, Dec. 13, 2014, pp. 609-622. [cited by applicant]
T. Luo, et al., “DaDianNao: A Neural Network Supercomputer”, IEEE Transaction on Computers, vol. 66, No. 1, Jan. 2017, pp. 73-88. [cited by applicant]
T. Chen, et al., “Dian Nao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning”, ASPLOS 14, Proceedings of the 19th international conference on Architectural support for programming languages … [cited by applicant]
Y. Chen, et al., “DianNao Family: Energy-Efficient Hardware Accelerators for Machine Learning”, Communications of the ACM, vol. 59, No. 11, Nov. 2016, pp. 105-112. [cited by applicant]
D. Liu, et al., “Pu Dian Nao: A Polyvalent Machine Learning Accelerator”, ASP LOS '15 Proceedings of the Twentieth International Conference on Architectural Support for Programming Languages and Operating Systems, Mar. … [cited by applicant]
Z. Du, et al., “ShiDianNao: Shifting Vision Processing Closer to the Sensor”, ISCA '15 Proceedings of the 42nd Annual International Symposium on Computer Architecture, Jun. 13, 2015, pp. 92-104. [cited by applicant]
EP 16899897.9—Communication Pursuant to Article 94(3) EPC, mailed Mar. 13, 2020, 12 pages. [cited by applicant]
Kalin Ovtcharov, et al., “Accelerating Deep Convolutional Neural Networks Using Specialized Hardware”, Microsoft Research, Feb. 22, 2015, 4 pages. [cited by applicant]
CN 201610282534.7—Second Office Action, mailed Mar. 9, 2020, 7 pages. (no English translation). [cited by applicant]
PCT/Cn2016/080967—International Search Report, mailed Jan. 26, 2017, 15 pages. (no English translation). [cited by applicant]
EP 16899897.9—Communication pursuant to Arcticle 94 (3) EPC, mailed Nov. 8, 2022, 8 pages. [cited by applicant]
Chakradhar et al., “A dynamically configurable coprocessor for convolutional neural networks”, ISCA '10: Proceedings of the 37th annual international symposium on Computer architecture, Jun. 19, 2010, pp. 247-257. [cited by applicant]
CN201610282534.7—Notification of grant of patent right for invention mailed on Jun. 22, 2020, 3 pages. [cited by applicant]
CN201710903509.0—Notification of grant of patent right for invention mailed on Apr. 2, 2020, 3 pages. [cited by applicant]
CN201710903509.0—Second Office Action malled on Nov. 1, 2019, 10 pages. [cited by applicant]
CN201811148189.3—Notification of grant of patent right for invention mailed on Mar. 26, 2020, 3 pages. [cited by applicant]
CN201811148189.3—Second Office Action mailed on Dec. 31, 2019, 8 pages. [cited by applicant]
CN202010616975.2—First Office Action mailed on May 15, 2023, 9 pages. [cited by applicant]
CN202010616975.2—Notification of grant of patent right for invention mailed on Nov. 17, 2023, 3 pages. [cited by applicant]
EP16899897.9—Communication under Rule 71(3) mailed on May 10, 2023, 9 pages. [cited by applicant]
EP16899897.9—Decision to Grant mailed on Aug. 31, 2023, 2 pages. [cited by applicant]
EP16899897.9—Supplementary European search report mailed on Dec. 12, 2019, 4 pages. [cited by applicant]
KR20187033947—Written Decision on registration mailed on Apr. 7, 2022, 4 pages. [cited by applicant]
KR20187033947—Written opinion mailed on Aug. 23, 2021, 19 pages. [cited by applicant]
KR20187033947—Written Opinion mailed on Feb. 11, 2022, 22 pages. [cited by applicant]
U.S. Appl. No. 16/174,155—Non-Final Office Action mailed on Jul. 25, 2019, 19 pages. [cited by applicant]
U.S. Appl. No. 16/174,155—Notice of Allowance mailed on Nov. 6, 2019, 9 pages. [cited by applicant]
Venkata Nanda Kishore Buddhiraju, “Parallelized Convolution”, Retrieved from internet URL—https://cse.buffalo.edu/faculty/miller/Courses/CSE633/Venkata-Nanda-Kishore-Buddhiraju-Fall-2011.pdf, Apr. 23, 2025, 20pages. [cited by applicant]