IP Library Granted Patent US 12,217,162
Granted Patent B2
US 12,217,162 · App. 18/085,273 · Granted Feb 4, 2025

Integrated circuit chip apparatus

Inventors: Shaoli Liu (Beijing, CN); Xinkai Song (Beijing, CN); Bingrui Wang (Beijing, CN); Yao Zhang (Beijing, CN); Shuai Hu (Beijing, CN)
Assignee: CAMBRICON TECHNOLOGIES CORPORATION LIMITED
G06N3/063G06F7/483G06F7/5443G06F17/153G06F17/16G06N3/04G06N3/06G06N3/08H01L25/065G06F2207/4824
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,162
App. No.
18/085,273
Granted
Feb 4, 2025
Kind
B2
Abstract

Provided are an integrated circuit chip apparatus and a related product, the integrated circuit chip apparatus being used for executing a multiplication operation, a convolution operation or a training operation of a neural network. The present technical solution has the advantages of a small amount of calculation and low power consumption.

Claims (55)

1. An integrated circuit chip apparatus comprising: a main processing circuit and a plurality of basic processing circuits, wherein

the plurality of basic processing circuits are configured to perform a first set of neural network computations in parallel on data transferred by the main processing circuit to obtain a plurality of computation results, and transfer the plurality of computation results to the main processing circuit, and

the main processing circuit is configured to perform a second set of neural network computations in series on the plurality of computation results.

2. The integrated circuit chip apparatus of claim 1 , wherein the main processing circuit or at least one of the plurality of basic processing circuits includes a data type conversion circuit configured to convert data between a floating point data type and a fixed point data type.

3. The integrated circuit chip apparatus of claim 2 , wherein the main processing circuit is further configured to:

receive a data block; and

convert the data block to a fixed point data block using the data type conversion circuit.

4. The integrated circuit chip apparatus of claim 3 , wherein the main processing circuit is further configured to:

divide the fixed point data block into a distribution data block and a broadcasting data block;

partition the distribution data block to obtain a plurality of basic data blocks;

distribute the plurality of basic data blocks to the plurality of basic processing circuits; and

broadcast the broadcasting data block to the plurality of basic processing circuits.

5. The integrated circuit chip apparatus of claim 4 , wherein the basic processing circuits are further configured to:

perform the first set of neural network computations on the basic data blocks and the broadcasting data block in the fixed point data type to obtain the plurality of computation results.

6. The integrated circuit chip apparatus of claim 5 , wherein the main processing circuit is further configured to:

convert the plurality of computation results to the floating point data type; and

perform the second set of neural network computations on the computation results in the floating point data type.

7. The integrated circuit chip apparatus of claim 1 , wherein the main processing circuit is further configured to:

receive a data block;

divide the data block into a distribution data block and a broadcasting data block;

partition the distribution data block to obtain a plurality of basic data blocks;

distribute the plurality of basic data blocks to the plurality of basic processing circuits; and

broadcast the broadcasting data block to the plurality of basic processing circuits.

8. The integrated circuit chip apparatus of claim 7 , wherein the basic processing circuits are further configured to:

convert the basic data blocks and the broadcasting data block into data blocks of a fixed point data type; and

perform the first set of neural network computations on the basic data blocks and the broadcasting data block in the fixed point data type to obtain fixed point computation results.

9. The integrated circuit chip apparatus of claim 8 , wherein the basic processing circuits are further configured to:

convert the computation results from the fixed point data type to a floating point data type; and

transfer the computation results in the floating point data type to the main processing circuit.

10. The integrated circuit chip apparatus of claim 8 ,

wherein the basic processing circuits are further configured to:

transfer the computation results in fixed point data type to the main processing circuit, wherein the main processing circuit is further configured to:

convert the plurality of computation results to a floating point data type; and

perform the second set of neural network computations on the computation results in the floating point data type.

11. The integrated circuit chip apparatus of claim 7 , wherein the main processing circuit is configured to broadcast the broadcasting data block as a whole to the plurality of basic processing circuits.

12. The integrated circuit chip apparatus of claim 7 , wherein the broadcasting data block is a multiplier data block and the distribution data block a multiplicand data block.

13. The integrated circuit chip apparatus of claim 7 , wherein the broadcasting data block is an input data block for convolution and the distribution data block is a convolution kernel.

14. The integrated circuit chip apparatus of claim 1 , wherein

the main processing circuit includes a main register or a main on-chip caching circuit, and

each of the plurality of basic processing circuits includes a basic register or a basic on-chip caching circuit.

15. The integrated circuit chip apparatus of claim 1 , wherein the main processing circuit includes one or more of a vector computing unit circuit, an arithmetic and logic unit circuit, an accumulator circuit, a matrix transposition circuit, a direct memory access circuit, a data type conversion circuit, or a data rearrangement circuit.

16. The integrated circuit chip apparatus of claim 1 , wherein the first set of neural network computations are inner product computations, wherein the second set of neural computations are accumulation computations that accumulate the computation results obtained by the inner product computations.

17. The integrated circuit chip apparatus of claim 1 , further comprising: a branch processing circuit, wherein the branch processing circuit is located between the main processing circuit and at least one basic processing circuit, wherein the branch processing circuit is configured to forward data between the main processing circuit and at least one basic processing circuit.

18. A processing system, comprising:

a neural network computing apparatus;

a general interconnection interface; and

a general-purpose processing apparatus connected to the neural network computing apparatus via the general interconnection interface,

wherein the neural network computing apparatus further comprises: a main processing circuit and a plurality of basic processing circuits, wherein

the plurality of basic processing circuits are configured to perform a first set of neural network computations in parallel on data transferred by the main processing circuit to obtain a plurality of computation results, and transfer the plurality of computation results to the main processing circuit, and

the main processing circuit is configured to perform a second set of neural network computations in series on the plurality of computation results.

19. A method for performing neural network operations using an integrated circuit chip apparatus comprising a main processing circuit, and a plurality of basic processing circuits, the method comprising:

performing, by the plurality of basic processing circuits, a first set of neural network computations in parallel on data transferred by the main processing circuit to obtain a plurality of computation results, and transfer the plurality of computation results to the main processing circuit, and

performing, by the main processing circuit, a second set of neural network computations in series on the plurality of computation results.

20. The method of claim 19 , wherein the main processing circuit or at least one of the plurality of basic processing circuits includes a data type conversion circuit, wherein the method further comprises:

converting data, by the data type conversion circuit, between a floating point data type and a fixed point data type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2022
From: LIU, SHAOLI; SONG, XINKAI; WANG, BINGRUI; ZHANG, YAO; HU, SHUAI
To: CAMBRICON TECHNOLOGIES CORPORATION LIMITED
Reel/Frame 062163/0167 →
Priority Claims (7)
CN 201711343642.1 · Dec 14, 2017 · national
CN 201711346333.X · Dec 14, 2017 · national
CN 201711347310.0 · Dec 14, 2017 · national
CN 201711347406.7 · Dec 14, 2017 · national
CN 201711347407.1 · Dec 14, 2017 · national
CN 201711347408.6 · Dec 14, 2017 · national
CN 201711347767.1 · Dec 14, 2017 · national
Continuity (3)
Continuation 16721875 · Dec 19, 2019
Continuation PCTCN2019073453 · Jan 28, 2019
Related Publication 20230120704A1 · Apr 20, 2023
References Cited (129)
US 5161117A · Waggener, Jr. · 1992 [cited by examiner]
US 5956703A · Turner et al. · 1999 [cited by applicant]
US 6094715A · Wilkinson et al. · 2000 [cited by applicant]
US 6272577B1 · Leung · 2001 [cited by examiner]
US 7236995B2 · Hinds · 2007 [cited by examiner]
US 7426501B2 · Nugent · 2008 [cited by examiner]
US 7571303B2 · Smith et al. · 2009 [cited by applicant]
US 7945607B2 · Hinds · 2011 [cited by examiner]
US 8156057B2 · Nugent · 2012 [cited by examiner]
US 8601013B2 · Dlugosch · 2013 [cited by examiner]
US 9607355B2 · Zou et al. · 2017 [cited by applicant]
US 9990687B1 · Kaufhold et al. · 2018 [cited by applicant]
US 10073816B1 · Lu et al. · 2018 [cited by applicant]
US 11295196B2 · Chen et al. · 2022 [cited by applicant]
US 20050125477A1 · Genov et al. · 2005 [cited by applicant]
US 20100122070A1 · Guevorkian et al. · 2010 [cited by applicant]
US 20110026640A1 · Milbar · 2011 [cited by applicant]
US 20110119467A1 · Cadambi et al. · 2011 [cited by applicant]
US 20110314256A1 · Callahan, II et al. · 2011 [cited by applicant]
US 20150019840A1 · Anderson et al. · 2015 [cited by applicant]
US 20150277912A1 · Gueron · 2015 [cited by applicant]
US 20160191204A1 · Kim et al. · 2016 [cited by applicant]
US 20160342888A1 · Yang et al. · 2016 [cited by applicant]
US 20160350645A1 · Brothers et al. · 2016 [cited by applicant]
US 20170061279A1 · Yang et al. · 2017 [cited by applicant]
US 20170102921A1 · Henry et al. · 2017 [cited by applicant]
US 20170103305A1 · Henry et al. · 2017 [cited by applicant]
US 20170344880A1 · Nekuii et al. · 2017 [cited by applicant]
US 20170357891A1 · Judd et al. · 2017 [cited by applicant]
US 20180046894A1 · Yao · 2018 [cited by applicant]
US 20180046903A1 · Yao et al. · 2018 [cited by applicant]
US 20180121240A1 · Cai et al. · 2018 [cited by applicant]
US 20180157969A1 · Xie et al. · 2018 [cited by applicant]
US 20180218275A1 · Arrigoni et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurviladhi et al. · 2018 [cited by applicant]
US 20190087716A1 · Du et al. · 2019 [cited by applicant]
US 20190114534A1 · Teng et al. · 2019 [cited by applicant]
US 20190339972A1 · Valentine et al. · 2019 [cited by applicant]
CN 103199806A · 2013 [cited by applicant]
CN 103631761A · 2014 [cited by applicant]
CN 104134349A · 2014 [cited by applicant]
CN 104463324A · 2015 [cited by applicant]
CN 104572011A · 2015 [cited by applicant]
CN 104992430A · 2015 [cited by applicant]
CN 105426344A · 2016 [cited by applicant]
CN 105956659A · 2016 [cited by applicant]
CN 106126481A · 2016 [cited by applicant]
CN 106570559A · 2017 [cited by applicant]
CN 106575379A · 2017 [cited by applicant]
CN 106844294A · 2017 [cited by applicant]
CN 106940815A · 2017 [cited by applicant]
CN 106991476A · 2017 [cited by applicant]
CN 106991478A · 2017 [cited by applicant]
CN 107016175A · 2017 [cited by applicant]
CN 107229967A · 2017 [cited by applicant]
CN 107239829A · 2017 [cited by applicant]
CN 107315574A · 2017 [cited by applicant]
CN 107329734A · 2017 [cited by applicant]
CN 107330515A · 2017 [cited by applicant]
CN 109726806A · 2019 [cited by applicant]
CN 11136897A · 2020 [cited by applicant]
CN 107608715B · 2020 [cited by applicant]
JP 2001188767A · 2001 [cited by applicant]
KR 1020160140394A · 2016 [cited by applicant]
WO 2017106469A1 · 2017 [cited by applicant]
WO 2017185412A1 · 2017 [cited by applicant]
WO 2017185414A1 · 2017 [cited by applicant]
First Office action issued in related Chinese Application No. 201711346333.X, dated Sep. 3, 2019, 8 pages. [cited by applicant]
International Search Report in corresponding International Application No. PCT/CN2019/073453, mailed Apr. 18, 2019, 4 pages. [cited by applicant]
Second Office action issued in related Chinese Application No. 201711347406.7, dated Nov. 27, 2019, 7 pages. [cited by applicant]
Third Office action issued in related Chinese Application No. 201711346333.X, dated Feb. 21, 2020, 9 pages. [cited by applicant]
First Office action issued in related Chinese Application No. 201711347408.6, dated Sep. 12, 2019, 7 pages. [cited by applicant]
First Office action issued in related Chinese Application No. 201711455397.3, dated Nov. 14, 2019, 8 pages. [cited by applicant]
Second Office action issued in related Chinese Application No. 201711455397.3, dated Mar. 3, 2020, 7 pages. [cited by applicant]
Office action issued in related Taiwan Application No. 107144036, dated Dec. 7, 2021, 10 pages. [cited by applicant]
Office action issued in related Taiwan Application No. 107144037, dated Dec. 7, 2021, 9 pages. [cited by applicant]
Liu, Shaoli et al., “Cambricon: An Instruction Set Architecture for Neural Networks”, IEEE Computer Society, 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, 13 pages. [cited by applicant]
Zhang, Shijin et al., “Cambricon-X: An Accelerator for Sparse Neural Networks”, 978-1-5090-3/16/$31.00, 2016 IEEE, 12 pages. [cited by applicant]
Chen, Yunji et al., “DaDianNao: A Machine-Learning Supercomputer”, IEEE Computer Society, 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, 14 pages. [cited by applicant]
Chen, Tianshi et al., “DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning”, ASPLOS '14, Mar. 1-5, 2014, Salt Lake City, Utah, USA, 15 pages. [cited by applicant]
Liu, Daofu et al., “PuDianNao: A Polyvalent Machine Learning Accelerator”, ASPLOS '15, Mar. 14-18, 2015, Istanbul, Turkey, 13 pages. [cited by applicant]
Du, Zidong et al., “ShiDianNao: Shifting Vision Processing Closer to the Sensor”, ISCA '15, Jun. 13-17, 2015, Portland, OR, USA, 13 pages. [cited by applicant]
First Office action issued in related Chinese Application No. 201811462676.7, dated Sep. 17, 2019, 9 pages. [cited by applicant]
First Office action issued in related Chinese Application No. 201811462969.5, dated Sep. 30, 2019, 9 pages. [cited by applicant]
International Search Report and Written Opinion in corresponding International Application No. PCT/CN2017/099991, mailed May 31, 2018, 8 pages. [cited by applicant]
First Office action issued in related Chinese Application No. 201780002287.3, dated Dec. 2, 2019, 12 pages. [cited by applicant]
Second Office action issued in related Chinese Application No. 201811462969.5, dated Feb. 3, 2020, 10 pages. [cited by applicant]
First Office action issued in related Chinese Application No. 201910102972.4, dated Nov. 29, 2019, 7 pages. [cited by applicant]
First Office action issued in related Chinese Application No. 201910534118.5, dated Nov. 18, 2019, 8 pages. [cited by applicant]
First Office action issued in related Chinese Application No. 201910531031.2, dated Nov. 6, 2019, 7 pages. [cited by applicant]
First Office action issued in related Chinese Application No. 201910530860.9, dated Nov. 19, 2019, 6 pages. [cited by applicant]
First Office action issued in related Chinese Application No. 201910534527.5, dated Dec. 11, 2019, 7 pages. [cited by applicant]
Extended European search report in related European Application No. 19211995.6, dated Apr. 6, 2020, 11 pages. [cited by applicant]
Jonghoon Jin et al: “Flattened Convolutional Neural Networks for Feedforward Acceleration”, Arxiv.org, Nov. 20, 2015, 11 pages. [cited by applicant]
The Tensorflow Authors: “tensorflow/conv_grad_input_ops.cc at 19881 lc64d3139d52eb074fdf20c8156c42f9d0etensorflow/tensorflow . GitHub”, GitHub TensorFlow repository, Aug. 2, 2017, 21 pages. [cited by applicant]
Vincent Dumoulin et al:“A guide to convolution arithmetic for deep learning”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Mar. 23, 2016, 28 pages. [cited by applicant]
Extended European search report in related European Application No. 19212002.0, dated Apr. 8, 2020, 11 pages. [cited by applicant]
Minsik Cho et al: “MEC: Memory-efficient Convolution for Deep Neural Network”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jun. 21, 2017, 10 pages. [cited by applicant]
Extended European search report in related European Application No. 19212010.3, dated Apr. 20, 2020, 9 pages. [cited by applicant]
Extended European search report in related European Application No. 19212365.1, dated Apr. 21, 2020, 10 pages. [cited by applicant]
Extended European search report in related European Application No. 19212368.5, dated Apr. 22, 2020, 10 pages. [cited by applicant]
Second Office action issued in related Chinese Application No. 201910534528.X, dated Feb. 25, 2020, 8 pages. [cited by applicant]
Yunji Chen, “DaDianNao: Machine-Learning Supercomputer” <<2014 47th Annual IEEE/ACM International Symposium on Microarchitecture>> , Jan. 19, 2015, 15 pages. [cited by applicant]
First Office action issued in related Japanese Application No. 2019 553977, dated Feb. 2, 2021, 5 pages. [cited by applicant]
Lili Song et al., “C-Brain:A Deep Learning Accelerator that Tames the Diversity of CNNs through Adaptive Data-level Parallelization” Proceedings of the 53rd ACM/EDAC/IEEE Design Automation Conference, US IEEE, Jun. 5, 2… [cited by applicant]
First Office action issued in related Japanese Application No. 2019 221533, dated Nov. 4, 2020, 4 pages. [cited by applicant]
Third Office action issued in related Chinese Application No. 201910534528.X, dated May 22, 2020, 9 pages. [cited by applicant]
Third Office action issued in related Chinese Application No. 201910531031.2, dated Jul. 3, 2020, 11 pages. [cited by applicant]
Yu Wang et al, “Low Power Convolutional Neural Networks on a Chip”, 2016 IEEE International Symposium on Circuits and Systems(ISCAS), IEEE, May 22, 2016, pp. 129-132, XP 32941496A. [cited by applicant]
Office Action issued in related European Application No. 19211995.6, dated Dec. 8, 2021, 11 pages. [cited by applicant]
Office Action issued in related Korean Application No. 10-2019-7029020, dated Februay 26, 2022, 11 pages. [cited by applicant]
Ngo, Kalle. “FPGA hardware acceleration on inception style parameter reduced convolution neural networks.” (2016). [cited by applicant]
Reyes IV, R., Fedyushkina, I. V., Skvortsov, V. S., & Filimonov, D. A. (2013). Prediction of progesterone receptor Inhibition by high-performance neural network algorithm. International journal of mathematical models an… [cited by applicant]
Karam, R., Paul, S., Puri, R., & Bhunia, S. (2017). Memory-centric reconfigurable accelerator for classification and machine learning applications. ACM Journal on Emerging Technologies in Computing Systems (JETC), 13(3)… [cited by applicant]
Sankaradas, M., Jakkula, V., Cadambi, S., et al. (2009, July). A massively parallel coprocessor for convolutional neural networks. In 2009 20th IEEE International Conference on Application-specific Systems, Architecture… [cited by applicant]
Chen, Yunji et al., “DianNao Family: Energy-Efficient Hardware Accelerators for Machine Learning”, DOI: 10.1145/2996864, Nov. 2016, vol. 59, No. 11, Communications of the ACM, 8 pages. [cited by applicant]
Bianconi, G., Dorogovtsev, S. N., & Mendes, J. F. (2015). Mutually connected component of networks of networks with replica nodes. Physical Review E, 91(1), 012804. (Year: 2015). [cited by applicant]
Yang et al., (2016). A Systematic approach to blocking convolutional neural networks. arXiv: 1606.04209v1, 12 pages. [cited by applicant]
Pande et al., Matrix Convolution using parallel programming, International Journal of Sciene and Researach (IJSR), India Online ISSN:2319-7064 (Jul. 2013) 6 pages. [cited by applicant]
Chellapilla et al., High performance convolutional neural networks for document processing, HAL Id: inria-00112631. https://halinria.fr/inria-00112631 (Nov. 2006) 7 pages. [cited by applicant]
Hy et al., DLPlib: A library for deep learning processor, J. Computer Science & Tech 32(2):286-296 (Mar. 2017) 11 pages. [cited by applicant]
“Brito, R., Fong, S., Cho, K., Song, W., Wong, R., Mohammed, S., & Fiaidhi, J. (2016). GPU-enabled back-propagation artificialneural network for digit recognition in parallel. The Journal of Supercomputing, 72(10), 3868… [cited by applicant]
“Han, S., Liu, X., Mao, H., Pu, J., Pedram, A., Horowitz, M.A., & Dally, W. J. (2016). EIE: efficient inference engine on compressed deep neural network. ACM Sigarch Computer Architecture News, 44(3), 243-254(Year: 2016… [cited by applicant]
Zhang, J., & Li, J. (2017, February). Improving the performance of OpenCL-based FPGA accelerator for convolutional neuralnetwork. In Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate A… [cited by applicant]
Azarkhish, et al., (2017). Neurostream: Scalable and Energy Efficient Deep Learning with Smart Memory Cues. arXiv preprint arXiv:1701.06420. (Year 2017). [cited by applicant]
Moini S., et al., A resource-limited hardware accelerator for convoultional neural networks in embeded vision applications. IEEE Transactions on Circuits and Systems II: Express Briefs. 2017, Apr. 4:64(10):1217-21 (Year… [cited by applicant]
Chinese Office Action in related Chinese Application No. 201911335145.6 dated Feb. 27, 2023 (34 pages). [cited by applicant]
Yuanyuan Li et al., New Materials Science and Technology-Metal Materials Sep. 30, 2012 (7 pages). [cited by applicant]
Second Office action issued in related Chinese Application No. 201911163257.8, dated Sep. 21, 2023, 6 pages. [cited by applicant]