IP Library › Granted Patent US 12,518,148
Granted Patent B2
US 12,518,148 · App. 17/254,372 · Granted Jan 6, 2026

Data processing method, device, computer equipment and storage medium

Inventors: Shaoli Liu (Anhui, CN); Di Huang (Anhui, CN); Xishan Zhang (Anhui, CN); Yao Zhang (Anhui, CN)
Assignee: ANHUI CAMBRICON INFORMATION TECHNOLOGY CO., LTD.
G06N3/063G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,148
App. No.
17/254,372
Granted
Jan 6, 2026
Kind
B2
Abstract

The present disclosure provides a data processing method, a board card device, computer equipment, and a storage medium for data quantization. The board card provided in the present disclosure includes a storage device, an interface apparatus, a control device, and an artificial intelligence chip of a data processing device, where the artificial intelligence chip is connected to the storage device, the control device, and the interface apparatus, respectively. The storage device is configured to store data; the interface apparatus is configured to implement data transmission between the artificial intelligence chip and an external apparatus; and the control device is configured to monitor the state of the artificial intelligence chip. The data to be quantized is quantized according to corresponding quantization parameters, which reduce the storage space of data while ensuring precision, accuracy and reliability of the operation result, and improving the operation efficiency.

Claims (80)

1 . A neural network quantization method, comprising:

determining a plurality of pieces of data to be quantized in target data of a layer in a neural network to be quantized, wherein each piece of data to be quantized is a subset of the target data, the target data is data to be operated and quantized in the layer to be quantized, and the data to be operated and quantized includes at least one of an input neuron, a weight, a bias, and a gradient;

and wherein a quantity of the plurality of pieces of data depends on one or more dimensions of the target data and the quantity of the plurality of pieces of data is inversely correlated to a processing capability of a device implementing the neural network quantization;

determining, for each piece of data, a corresponding quantization parameter based on a pre-determined correspondence, wherein the corresponding quantization parameter comprises one or more of the following: a decimal point position, a scaling factor, or an offset;

quantizing each piece of data to be quantized according to the corresponding quantization parameter to obtain a piece of quantized data corresponding to each piece of data to be quantized, wherein the plurality of pieces of data are quantized in parallel;

obtaining a quantization result of the target data according to the piece of quantized data corresponding to each piece of data to be quantized, so that an operation is performed in the layer to be quantized according to the quantization result of the target data;

determining a quantization error corresponding to the target data, wherein the determining of the quantization error comprises:

determining a difference between the plurality of pieces of data to be quantized and a plurality of pieces of de-quantized data of the plurality of quantized data; or

determining a difference as a total quantization interval over the plurality of pieces of data to be quantized;

wherein the quantization error is an average of the differences over a sum of the plurality of pieces of data to be quantized; and

according to the quantization error and an error threshold determined based on an empirical value, adjusting a data bit width of each piece of data to be quantized to obtain an adjusted bit width corresponding to each piece of data to be quantized.

2 . The neural network quantization method of claim 1 , wherein the layer to be quantized is a convolution layer, the target data is an input neuron, and the determining a plurality of pieces of data to be quantized in the target data of the layer to be quantized includes:

in the input neuron of the convolution layer, determining a plurality of pieces of data to be quantized corresponding to a convolution kernel according to a dimension and stride of the convolution kernel, where the dimension of the convolution kernel includes height, width, and a count of channels.

3 . The neural network quantization method of claim 1 , wherein the determining the plurality of pieces of data to be quantized in the target data of the layer to be quantized includes:

determining a plurality of pieces of data to be quantized in the target data of the layer to be quantized according to a dimension of the target data, where the dimension of the target data includes a count of batches, channel, height, and width, wherein the determining a plurality of pieces of data to be quantized in the target data of the layer to be quantized according to the dimension of the target data includes:

determining one or more batches of data in the target data of the layer to be quantized as one piece of data to be quantized; or

determining one or more channels of data in the target data of the layer to be quantized as one piece of data to be quantized.

4 . The neural network quantization method of claim 1 , wherein the determining a plurality of pieces of data to be quantized in the target data of the layer to be quantized includes:

according to a real-time processing capability of a device running the neural network, determining a plurality of pieces of data to be quantized in the target data of the layer to be quantized, where a size of each piece of data to be quantized is positively correlated with the real-time processing capability.

5 . The neural network quantization method of claim 1 , further comprising:

computing a corresponding quantization parameter according to each piece of data to be quantized and a corresponding data bit width, wherein the computing the corresponding quantization parameter according to each piece of data to be quantized and the corresponding data bit width includes:

when the quantization parameter does not include an offset of each piece of data to be quantized, obtaining a first type of point position of each piece of data to be quantized according to a maximum of an absolute value of each piece of data to be quantized and the corresponding data bit width; or

when the quantization parameter does not include an offset of each piece of data to be quantized, obtaining a maximum of the piece of quantized data according to each piece of data to be quantized and the corresponding data bit width, obtaining a first type of scaling factor of each piece of data to be quantized according to the maximum of the absolute value of each piece of data to be quantized and the maximum of the piece of quantized data; or

when the quantization parameter includes an offset of each piece of data to be quantized, obtaining a second type of point position of each piece of data to be quantized according to a maximum and a minimum of each piece of data to be quantized and the corresponding data bit width; or

when the quantization parameter includes an offset of each piece of data to be quantized, obtaining the maximum of the piece of quantized data according to each piece of data to be quantized and the corresponding data bit width, and obtaining a second type of scaling factor of each piece of data to be quantized according to the maximum and the minimum of each piece of data to be quantized, and the maximum of the piece of quantized data.

6 . The neural network quantization method of claim 1 , further comprising:

updating the data bit width corresponding to each piece of data to be quantized to a corresponding adjusted bit width, and computing a corresponding adjusted quantization parameter according to each piece of data to be quantized and the corresponding adjusted bit width to quantize each piece of data to be quantized according to the corresponding adjusted quantization parameter.

7 . The neural network quantization method of claim 6 , wherein the adjusting the data bit width of each piece of data to be quantized to obtain an adjusted bit width corresponding to each piece of data to be quantized according to the quantization error and an error threshold corresponding to each piece of data to be quantized includes when the quantization error is greater than a first error threshold, increasing the corresponding data bit width to obtain the corresponding adjusted bit width; and

further comprising:

computing an adjusted quantization error of each piece of data to be quantized according to the piece of data to be quantized and the corresponding adjusted bit width;

continuing to increase the corresponding adjusted bit width according to the adjusted quantization error and the first error threshold until the adjusted quantization error is less than or equal to the first error threshold;

computing the adjusted quantization error of the data to be quantized according to the adjusted bit width and the piece of data to be quantized; and

continuing to decrease the corresponding adjusted bit width according to the adjusted quantization error and the second error threshold until the adjusted quantization error obtained according to the adjusted bit width and the piece of data to be quantized is greater than or equal to the second error threshold

wherein the adjusting the data bit width of each piece of data to be quantized to obtain an adjusted bit width corresponding to each piece of data to be quantized according to the quantization error and an error threshold corresponding to each piece of data to be quantized includes when the quantization error is less than a second error threshold, increasing the corresponding data bit width to obtain the corresponding adjusted bit width, where the second error threshold is less than the first error threshold.

8 . The neural network quantization method of claim 1 , wherein, during a fine-tuning stage and/or training stage of a neural network operation, the method further comprises:

obtaining a variation range of the data to be quantized in a current iteration and historical iterations, where the historical iterations are iterations before the current iteration; and

according to the variation range of the data to be quantized, determining a target iteration interval corresponding to the data to be quantized to enable the layer to be quantized to update the quantization parameter of the data to be quantized according to the target iteration interval, where the target iteration interval includes at least one iteration.

9 . The neural network quantization method of claim 8 , further comprising:

according to the data bit width of the data to be quantized in the current iteration, determining the data bit width of the data to be quantized in the iterations within the target iteration interval to enable the neural network to determine the quantization parameter; and

according to a point position of the data to be quantized in the current iteration, determining the point position of the data to be quantized in the iterations within the target iteration interval, where the point position includes the first type of point position and/or the second type of point position.

10 . A neural network quantization device comprising one or more processing circuits, wherein for any layer to be quantized in a neural network, the device comprises:

a data determination module implemented on the one or more processing circuits and configured to determine a plurality of pieces of data to be quantized in target data of a layer to be quantized, wherein each piece of data to be quantized is a subset of the target data, the target data is any kind of data to be operated and quantized in the layer to be quantized, and the data to be operated includes at least one of an input neuron, a weight, a bias, and a gradient, and wherein a quantity of the plurality of pieces of data depends on one or more dimensions of the target data and the quantity of the plurality of pieces of data is inversely correlated to a processing capability of a device implementing the neural network quantization;

a data quantization module implemented on the one or more processing circuits and configured to determine, for each piece of data, a corresponding quantizing parameter based on a pre-determined correspondence, wherein the corresponding quantization parameter comprises one or more of the following: a decimal point position, a scaling factor, or an offset, and quantize each piece of data to be quantized according to the corresponding quantization parameter to obtain a piece of quantized data corresponding to each piece of data to be quantized, wherein the plurality of pieces of data are quantized in parallel;

a data operation module implemented on the one or more processing circuits and configured to obtain a quantization result of the target data according to the piece of quantized data corresponding to each piece of data to be quantized, so that an operation may be performed in the layer to be quantized according to the quantization result of the target data;

a first quantization error determination module implemented on one or more processing circuits and configured to determine a quantization error corresponding to the target data by

determining a difference between the plurality of pieces of data to be quantized and a plurality of pieces of de-quantized data of the plurality of quantized data; or

determining a difference as a total quantization interval over the plurality of pieces of data to be quantized;

wherein the quantization error is an average of the differences over a sum of the plurality of pieces of data to be quantized; and

an adjusted bit width determination module implemented on one or more processing circuits and configured to, according to the quantization error and an error threshold determined based on an empirical value, adjust a data bit width of each piece of data to be quantized to obtain an adjusted bit width corresponding to each piece of data to be quantized.

11 . The neural network quantization device of claim 10 , wherein the layer to be quantized is a convolution layer, the target data is an input neuron, and the data determination module includes:

a first determination sub-module configured to, in the input neuron of the convolution layer, determine a plurality of pieces of data to be quantized corresponding to a convolution kernel according to a dimension and stride of the convolution kernel, where the dimension of the convolution kernel includes height, width, and a count of channels.

12 . The neural network quantization device of claim 10 , wherein the data determination module includes:

a second determination sub-module configured to determine a plurality of pieces of data to be quantized in the target data of the layer to be quantized according to the dimension of the target data, where the dimension of the target data includes a count of batches, channel, height, and width,

wherein the second determination sub-module includes a determination sub-module based on a count of batches or channel configured to determine one or more batches of data in the target data of the layer to be quantized as one piece of data to be quantized.

13 . The neural network quantization device of claim 10 , wherein the data determination module includes:

a third determination sub-module configured to, according to a real-time processing capability of a device running the neural network, determine a plurality of pieces of data to be quantized in the target data of the layer to be quantized, where a size of each piece of data to be quantized is positively correlated with the real-time processing capability.

14 . The neural network quantization device of claim 10 , further comprising:

a parameter determination sub-module configured to compute the corresponding quantization parameter according to each piece of data to be quantized and a corresponding data bit width;

wherein the parameter determination sub-module includes:

a first point position determination sub-module configured to, when the quantization parameter does not include an offset of each piece of data to be quantized, obtain a first type of point position of each piece of data to be quantized according to a maximum of an absolute value of each piece of data to be quantized and the corresponding data bit width;

a first maximum determination sub-module configured to, when the quantization parameter does not include the offset of each piece of data to be quantized, obtain a maximum of the piece of quantized data according to each piece of data to be quantized and the corresponding data bit width;

a first scaling factor determination sub-module configured to obtain a first type of scaling factor of each piece of data to be quantized according to the maximum of the absolute value of each piece of data to be quantized and the maximum of the piece of quantized data;

a second point position determination sub-module configured to, when the quantization parameter includes the offset of each piece of data to be quantized, obtain a second type of point position of each piece of data to be quantized according to the maximum and a minimum of each piece of data to be quantized and the corresponding data bit width;

a second maximum determination sub-module configured to, when the quantization parameter includes the offset of each piece of data to be quantized, obtain the maximum of the piece of quantized data according to each piece of data to be quantized and the corresponding data bit width; and

a second scaling factor determination sub-module configured to obtain a second type of scaling factor of each piece of data to be quantized according to the maximum and the minimum of each piece of data to be quantized, and the maximum of the piece of quantized data.

15 . The neural network quantization device of 10 , further comprising:

an adjusted quantization parameter determination module configured to update the data bit width corresponding to each piece of data to be quantized to a corresponding adjusted bit width, and compute a corresponding adjusted quantization parameter according to each piece of data to be quantized and the corresponding adjusted bit width to quantize each piece of data to be quantized according to the corresponding adjusted quantization parameter;

a first adjusted quantization error module configured to compute an adjusted quantization error of each piece of data to be quantized according to the data to be quantized and the corresponding adjusted bit width; and

a first adjusting bit width cycle determination module configured to continue to increase the corresponding adjusted bit width according to the adjusted quantization error and the first error threshold until the adjusted quantization error is less than or equal to the first error threshold;

wherein the adjusted bit width determination module includes:

a first adjusted bit width determination sub-module configured to, when the quantization error is greater than a first error threshold, increase the corresponding data bit width to obtain the corresponding adjusted bit width; and

a second adjusted bit width determination sub-module configured to, when the quantization error is less than a second error threshold, decrease the corresponding data bit width to obtain the corresponding adjusted bit width, where the second error threshold is less than the first error threshold.

16 . The neural network quantization device of claim 10 , wherein, during a fine-tuning stage and/or training stage of a neural network operation, the device further includes:

a first data variation range determination module configured to obtain a variation range of the data to be quantized in a current iteration and historical iterations, where the historical iterations are iterations before the current iteration; and

a target iteration interval determination module configured to, according to the variation range of the data to be quantized, determine a target iteration interval corresponding to the data to be quantized to enable the layer to be quantized to update the quantization parameter of the data to be quantized according to the target iteration interval, where the target iteration interval includes at least one iteration.

17 . The neural network quantization device of claim 16 , further comprising:

a first target iteration interval application module configured to, according to the data bit width of the data to be quantized in the current iteration, determine the data bit width of the data to be quantized in the iterations within the target iteration interval to enable the neural network to determine the quantization parameter according to the data bit width of the data to be quantized in the iterations within the target iteration interval; and

a second target iteration interval application module configured to, according to a point position of the data to be quantized in the current iteration, determine the point position of the data to be quantized in the iterations within the target iteration interval, where the point position includes the first type of point position and/or the second type of point position.

18 . An artificial intelligence chip, comprising the neural network quantization device of claim 10 .

19 . A non-transitory computer readable storage medium, on which a computer program instruction is stored, wherein when the computer program instruction is executed by a processor, the neural network quantization method of claim 1 is realized.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 21, 2020
From: HUANG, DI; LIU, SHAOLI; ZHANG, XISHAN; ZHANG, YAO
To: ANHUI CAMBRICON INFORMATION TECHNOLOGY CO., LTD.
Reel/Frame 054705/0815 →
Priority Claims (2)
CN 201910784982.0 · Aug 23, 2019 · national
CN 201910888599.X · Sep 19, 2019 · national
Continuity (1)
Related Publication 20210374510A1 · Dec 2, 2021
References Cited (219)
US 5052043A · Gaborski · 1991 [cited by applicant]
US 6704757B1 · Ohmi et al. · 2004 [cited by applicant]
US 6715065B1 · Ebata et al. · 2004 [cited by applicant]
US 6931639B1 · Eickemeyer · 2005 [cited by applicant]
US 7242414B1 · Thekkath et al. · 2007 [cited by applicant]
US 7406451B2 · Mrziglod et al. · 2008 [cited by applicant]
US 7721128B2 · Johns et al. · 2010 [cited by applicant]
US 7945607B2 · Hinds · 2011 [cited by applicant]
US 8694572B2 · Samy et al. · 2014 [cited by applicant]
US 8924455B1 · Barman et al. · 2014 [cited by applicant]
US 9412366B2 · Wilensky et al. · 2016 [cited by applicant]
US 9916531B1 · Zivkovic et al. · 2018 [cited by applicant]
US 10187568B1 · Tran et al. · 2019 [cited by applicant]
US 10224954B1 · Madduri et al. · 2019 [cited by applicant]
US 10360304B1 · Alvarez et al. · 2019 [cited by applicant]
US 10427306B1 · Quinlan et al. · 2019 [cited by applicant]
US 20020138714A1 · Leibholz et al. · 2002 [cited by applicant]
US 20030167460A1 · Desai et al. · 2003 [cited by applicant]
US 20050138327A1 · Tabei · 2005 [cited by applicant]
US 20060161375A1 · Duberstein et al. · 2006 [cited by applicant]
US 20090113186A1 · Kato et al. · 2009 [cited by applicant]
US 20090125293A1 · Lefurgy et al. · 2009 [cited by applicant]
US 20100073068A1 · Cho et al. · 2010 [cited by applicant]
US 20110060587A1 · Phillips et al. · 2011 [cited by applicant]
US 20110301777A1 · Cox et al. · 2011 [cited by applicant]
US 20120316845A1 · Grey et al. · 2012 [cited by applicant]
US 20130054110A1 · Sata · 2013 [cited by applicant]
US 20130332610A1 · Beveridge · 2013 [cited by applicant]
US 20140081625A1 · Wilensky et al. · 2014 [cited by applicant]
US 20140164737A1 · Collange et al. · 2014 [cited by applicant]
US 20140249814A1 · Nakano et al. · 2014 [cited by applicant]
US 20150134581A1 · Doeding et al. · 2015 [cited by applicant]
US 20150370303A1 · Krishnaswamy et al. · 2015 [cited by applicant]
US 20160026231A1 · Ignowski et al. · 2016 [cited by applicant]
US 20160054922A1 · Awasthi et al. · 2016 [cited by applicant]
US 20160124710A1 · Lutz et al. · 2016 [cited by applicant]
US 20160170866A1 · Ioualalen et al. · 2016 [cited by applicant]
US 20160328645A1 · Lin et al. · 2016 [cited by applicant]
US 20160328647A1 · Lin et al. · 2016 [cited by applicant]
US 20170090956A1 · Linsky · 2017 [cited by applicant]
US 20170103022A1 · Kreinin et al. · 2017 [cited by applicant]
US 20170142327A1 · Bayani · 2017 [cited by applicant]
US 20170161604A1 · Craddock et al. · 2017 [cited by applicant]
US 20170221176A1 · Munteanu et al. · 2017 [cited by applicant]
US 20170257079A1 · Jain et al. · 2017 [cited by applicant]
US 20170262959A1 · Lee et al. · 2017 [cited by applicant]
US 20170316307A1 · Koster et al. · 2017 [cited by applicant]
US 20170316312A1 · Goyal et al. · 2017 [cited by applicant]
US 20170344882A1 · Ambrose et al. · 2017 [cited by applicant]
US 20170353163A1 · Gazneli et al. · 2017 [cited by applicant]
US 20170357530A1 · Shih et al. · 2017 [cited by applicant]
US 20170357910A1 · Sommer et al. · 2017 [cited by applicant]
US 20180046903A1 · Yao et al. · 2018 [cited by applicant]
US 20180088996A1 · Rossi et al. · 2018 [cited by applicant]
US 20180096243A1 · Patil et al. · 2018 [cited by applicant]
US 20180107925A1 · Choi et al. · 2018 [cited by applicant]
US 20180157464A1 · Lutz et al. · 2018 [cited by applicant]
US 20180288440A1 · Chao · 2018 [cited by applicant]
US 20180293517A1 · Browne et al. · 2018 [cited by applicant]
US 20180300931A1 · Vembu et al. · 2018 [cited by applicant]
US 20180322391A1 · Wu et al. · 2018 [cited by applicant]
US 20180357541A1 · Chen et al. · 2018 [cited by applicant]
US 20180367729A1 · Parasnis et al. · 2018 [cited by applicant]
US 20180373976A1 · Woo · 2018 [cited by applicant]
US 20190042925A1 · Choe et al. · 2019 [cited by applicant]
US 20190050710A1 · Wang et al. · 2019 [cited by applicant]
US 20190057696A1 · Ogawa · 2019 [cited by applicant]
US 20190114142A1 · Yoda et al. · 2019 [cited by applicant]
US 20190122094A1 · Chen et al. · 2019 [cited by applicant]
US 20190122119A1 · Husain · 2019 [cited by applicant]
US 20190138372A1 · Tee · 2019 [cited by applicant]
US 20190147322A1 · Kim et al. · 2019 [cited by applicant]
US 20190164285A1 · Nye et al. · 2019 [cited by applicant]
US 20190180170A1 · Huang et al. · 2019 [cited by applicant]
US 20190199370A1 · Madduri et al. · 2019 [cited by applicant]
US 20190220734A1 · Ferdman et al. · 2019 [cited by applicant]
US 20190228762A1 · Wang et al. · 2019 [cited by applicant]
US 20190251429A1 · Du et al. · 2019 [cited by applicant]
US 20190265949A1 · Ito · 2019 [cited by applicant]
US 20190278677A1 · Terechko et al. · 2019 [cited by applicant]
US 20190294968A1 · Vantrease et al. · 2019 [cited by applicant]
US 20190303765A1 · Gou et al. · 2019 [cited by applicant]
US 20190339937A1 · Lo et al. · 2019 [cited by applicant]
US 20200005424A1 · Appu et al. · 2020 [cited by applicant]
US 20200097799A1 · Divakar et al. · 2020 [cited by applicant]
US 20200117453A1 · Zhang et al. · 2020 [cited by applicant]
US 20200117614A1 · Zhang et al. · 2020 [cited by applicant]
US 20200125508A1 · Liu et al. · 2020 [cited by applicant]
US 20200126554A1 · Chen et al. · 2020 [cited by applicant]
US 20200126555A1 · Chen et al. · 2020 [cited by applicant]
US 20200142748A1 · Liu et al. · 2020 [cited by applicant]
US 20200159527A1 · Zhang et al. · 2020 [cited by applicant]
US 20200159530A1 · Zhang et al. · 2020 [cited by applicant]
US 20200159532A1 · Zhang et al. · 2020 [cited by applicant]
US 20200159533A1 · Zhang et al. · 2020 [cited by applicant]
US 20200160162A1 · Zhang et al. · 2020 [cited by applicant]
US 20200160163A1 · Liu et al. · 2020 [cited by applicant]
US 20200160219A1 · Zhang et al. · 2020 [cited by applicant]
US 20200160220A1 · Zhang et al. · 2020 [cited by applicant]
US 20200160221A1 · Zhang et al. · 2020 [cited by applicant]
US 20200160222A1 · Zhang et al. · 2020 [cited by applicant]
US 20200168227A1 · Chen et al. · 2020 [cited by applicant]
US 20200174547A1 · Fang et al. · 2020 [cited by applicant]
US 20200183752A1 · Liu et al. · 2020 [cited by applicant]
US 20200241874A1 · Chen et al. · 2020 [cited by applicant]
US 20200257972A1 · Miniskar et al. · 2020 [cited by applicant]
US 20200302283A1 · Zhu et al. · 2020 [cited by applicant]
US 20200334041A1 · Zhang et al. · 2020 [cited by applicant]
US 20200334522A1 · Zhang et al. · 2020 [cited by applicant]
US 20200334525A1 · Tsai et al. · 2020 [cited by applicant]
US 20200334572A1 · Zhang et al. · 2020 [cited by applicant]
US 20200394522A1 · Liu et al. · 2020 [cited by applicant]
US 20200394523A1 · Liu et al. · 2020 [cited by applicant]
US 20210042889A1 · Pei · 2021 [cited by applicant]
US 20210061028A1 · Da Deppo et al. · 2021 [cited by applicant]
US 20210117768A1 · Liu et al. · 2021 [cited by applicant]
US 20210117810A1 · Liu · 2021 [cited by applicant]
US 20210182177A1 · Su et al. · 2021 [cited by applicant]
US 20210192349A1 · Lian et al. · 2021 [cited by applicant]
US 20210264270A1 · Liu et al. · 2021 [cited by applicant]
US 20210286688A1 · Liu et al. · 2021 [cited by applicant]
US 20210334007A1 · Liu et al. · 2021 [cited by applicant]
US 20210334137A1 · Zhang et al. · 2021 [cited by applicant]
US 20210341989A1 · Chen et al. · 2021 [cited by applicant]
US 20210374510A1 · Liu et al. · 2021 [cited by applicant]
US 20210374511A1 · Liu et al. · 2021 [cited by applicant]
US 20210374540A1 · Yuan et al. · 2021 [cited by applicant]
CN 1503858A · 2004 [cited by applicant]
CN 1503958A · 2004 [cited by applicant]
CN 1851668A · 2006 [cited by applicant]
CN 101572829A · 2009 [cited by applicant]
CN 102270042A · 2011 [cited by applicant]
CN 102789413A · 2012 [cited by applicant]
CN 102903089A · 2013 [cited by applicant]
CN 104914977A · 2015 [cited by applicant]
CN 105389158A · 2016 [cited by applicant]
CN 107665364A · 2016 [cited by applicant]
CN 103534664A · 2016 [cited by applicant]
CN 105893419A · 2016 [cited by applicant]
CN 106156310A · 2016 [cited by applicant]
CN 106354568A · 2017 [cited by applicant]
CN 106406812A · 2017 [cited by applicant]
CN 106469291A · 2017 [cited by applicant]
CN 106650922A · 2017 [cited by applicant]
CN 106814639A · 2017 [cited by applicant]
CN 107197297A · 2017 [cited by applicant]
CN 106951587A · 2017 [cited by applicant]
CN 106951962A · 2017 [cited by applicant]
CN 106997236A · 2017 [cited by applicant]
CN 107003988A · 2017 [cited by applicant]
CN 107025629A · 2017 [cited by applicant]
CN 107368174A · 2017 [cited by applicant]
CN 107451654A · 2017 [cited by applicant]
CN 107644254A · 2018 [cited by applicant]
CN 107797913A · 2018 [cited by applicant]
CN 104899641A · 2018 [cited by applicant]
CN 108337000A · 2018 [cited by applicant]
CN 108717570A · 2018 [cited by applicant]
CN 109062540A · 2018 [cited by applicant]
CN 109063820A · 2018 [cited by applicant]
CN 109146057A · 2019 [cited by applicant]
CN 109190754A · 2019 [cited by applicant]
CN 109214509A · 2019 [cited by applicant]
CN 109389219A · 2019 [cited by applicant]
CN 109472353A · 2019 [cited by applicant]
CN 110008952A · 2019 [cited by applicant]
CN 109740739A · 2019 [cited by applicant]
CN 109800877A · 2019 [cited by applicant]
CN 109902745A · 2019 [cited by applicant]
CN 110020616A · 2019 [cited by applicant]
CN 109993296A · 2019 [cited by applicant]
CN 112446460A · 2021 [cited by examiner]
EP 0789296A1 · 1997 [cited by applicant]
EP 2703945A2 · 2014 [cited by applicant]
EP 3106997A2 · 2016 [cited by applicant]
EP 3407268A1 · 2018 [cited by applicant]
JP H03075860A · 1989 [cited by applicant]
JP H09265379A · 1997 [cited by applicant]
JP 2009134433A · 2012 [cited by applicant]
JP 2013514570A · 2013 [cited by applicant]
JP 2015509183A · 2015 [cited by applicant]
JP 1996087475B2 · 2015 [cited by applicant]
JP 2015176158A · 2015 [cited by applicant]
JP 2014199464A · 2017 [cited by applicant]
JP 201810618A · 2018 [cited by applicant]
JP 201826114A · 2018 [cited by applicant]
JP 2018514872A · 2018 [cited by applicant]
JP 2019519852A · 2019 [cited by applicant]
KR 20100087845A · 2009 [cited by applicant]
KR 20190034985A · 2019 [cited by applicant]
KR 20190054454A · 2019 [cited by applicant]
WO 2008153194A1 · 2008 [cited by applicant]
WO 2016186823A1 · 2016 [cited by applicant]
WO 2018103736A1 · 2018 [cited by applicant]
WO 2018140294A1 · 2018 [cited by applicant]
WO 2018192500A1 · 2018 [cited by applicant]
WO WO2021036362A1 · 2021 [cited by examiner]
Choukroun et al, Low-bit Quantization of Neural Networks for Efficient Inference; Dated Mar. 25, 2019; 10 pages. [cited by applicant]
Prado et al, QUENN: Quantization Engine for low-power Neural Networks; Dated Nov. 17, 2018; 9 pages. [cited by applicant]
Yang et al, Deploy Large-Scale Deep Neural Networks in Resource Constrained IoT Devices with Local Quantization Region; Dated May 28, 2018; 8 pages. [cited by applicant]
Li et al., “Using Artificial Neural Network for Predicting Thread Partitioning in Speculative Multithreading”, IEEE, 2015, pp. 823-826. [cited by applicant]
Kalathingal Sajith et al., “Dynamic Inter-Thread Vectorization Architecture: Extracting OLP from TLP”, 2016 28th International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD), IEEE, Oct. 26,… [cited by applicant]
Na et al., “Speeding up Convolutional Neural Network Training with Dynamic Precision Scaling and Flexible MultiplierAccumulator”, Section 2 Proposed Approach: Concept, ACM, Aug. 8-10, 2016, 6 pages. [cited by applicant]
Hanlon, Jamie, “Why is so much memory needed for deep neural networks?”, URL: https://www.graphcore.ai/posts/why-is-so-much-memory-needed-for-deep-neural-networks, Jan. 31, 2017, 6 pages. [cited by applicant]
Extended European Search Report for Application No. 19215861.6 mailed May 15, 2020. [cited by applicant]
Extended European Search Report for Application No. 19215862.4 mailed May 15, 2020. [cited by applicant]
Sumina Yamashita, et al., “A Method to create illustrate images using DCGAN,” JISJ SIG Technical Report, vol. 2017-MPS-112 No. 16, Feb. 27, 2017; translation of abstract included. [cited by applicant]
Gysel Philipp et al., “Ristretto: A Framework for Empirical Study of Resource-Efficient Inference in Convolutional Neural Networks”, IEEE Transactions on Neural Networks and Learning Systems, IEEE, Piscataway, NJ, USA, … [cited by applicant]
Yi Yang et al., “Deploy Large-Scale Deep Neural Networks in Resource Constrained Io T Devices with Local Quantization Region”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,… [cited by applicant]
European Patent Office, Extended European Search Report for European Application No. 19218382.0 dated Apr. 24, 2020. [cited by applicant]
Olariu Cristian et al., “A Cloud-Based AI Framework for Machine Learning Orchestration: A “Driving or Not-Driving” Case-Study for Self-Driving Cars”, 2019 IEEE Intelligent Vehicles Symposium (IV). IEEE, Jun. 9, 2019 (Ju… [cited by applicant]
Kallam Suresh et al., “Evaluating the Performance of Deep Learning Techniques on Classification Using Tensor Flow Application”, 2018 International Conference on Advances in Computing and Communication Engineering (ICACC… [cited by applicant]
Song Mingcong et al., “In-Situ AI: Towards Autonomous and Incremental Deep Leaming for IoT Systems”, 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA), IEEE, Feb. 24, 2018 (Feb. 24, 2018… [cited by applicant]
Hsu Jeremy, “For sale: deep learning [News]”, IEEE Spectrum, IEEE Inc. New York, US, vol. 53, No. 8, Aug. 1, 2016 (Aug. 1, 2016), pp. 12-13, XP011620787, ISSN: 0018-9235, DOI: 10.1109/MSPEC.2016.7524158 [retrieved on Ju… [cited by applicant]
European Patent Office, extended European search report for Application No. 19216754.2 mailed May 8, 2020. [cited by applicant]
Extended European Search Report for EP Application No. 19214324.6 mailed Oct. 1, 2020. [cited by applicant]
Yan JL, Zhang Y, Tu FB, et al; Research on low-power neural network computing chip technology; Sci Sin Inform, 2019, 49: 314-333, doi: 10.1360/N112018-00282, 20 pages. [cited by applicant]
Wang, Peisong, et al. “Two-step quantization for low-bit neural networks.” Proceedings of the IEEE Conference on computer vision and pattern recognition. 2018. (Year: 2018) 9 pages. [cited by applicant]
Final Office Action received in U.S. Appl. No. 17/622,647, mailed on Nov. 6, 2025 (47 pages). [cited by applicant]