IP Library › Granted Patent US 12,585,926
Granted Patent B2
US 12,585,926 · App. 16/237,308 · Granted Mar 24, 2026

Adjusting precision and topology parameters for neural network training based on a performance metric

Inventors: Bita Darvish Rouhani (Bellevue, WA); Eric S. Chung (Woodinville, WA); Daniel Lo (Bothell, WA); Douglas C. Burger (Bellevue, WA)
Assignee: Microsoft Technology Licensing, LLC
G06N3/063G06F9/30025G06F18/217G06N3/047G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,585,926
App. No.
16/237,308
Granted
Mar 24, 2026
Kind
B2
Abstract

Apparatus and methods for training neural networks based on a performance metric, including adjusting numerical precision and topology as training progresses are disclosed. In some examples, block floating-point formats having relatively lower accuracy are used during early stages of training. Accuracy of the floating-point format can be increased as training progresses based on a determined performance metric. In some examples, values for the neural network are transformed to normal precision floating-point formats. The performance metric can be determined based on entropy of values for the neural network, accuracy of the neural network, or by other suitable techniques. Accelerator hardware can be used to implement certain implementations, including hardware having direct support for block floating-point formats.

Claims (68)

1 . A computing system comprising:

one or more processors;

at least one memory in communication with the one or more processors; and

the computing system being configured to dynamically update a parameter of a neural network during training of the neural network, the training of the neural network comprising:

for each of multiple iterations in an epoch of training of the neural network, in a given iteration of the multiple iterations:

for a first set of one or more training values:

with at least one of the processors, perform forward propagation for at least one layer of a neural network to propagate activation values for edges of the at least one layer using a first parameter value for the parameter of the neural network, wherein the parameter is a precision parameter or a topology parameter;

with at least one of the processors, perform backward propagation for at least one layer of the neural network to propagate gradients for nodes of the at least one layer using the first parameter value;

for the given iteration, with at least one of the processors, determine a performance metric for the neural network, the performance metric being selected from accuracy of the neural network, entropy of the neural network, mean square error of the neural network, perplexity of the neural network, or gradient signal to noise ratio of the neural network;

based on the performance metric determined during the given iteration, dynamically adjust the parameter of the neural network to produce an adjusted parameter, the adjusted parameter being selected to improve the performance metric in a neural network having the adjusted parameter, the performance; and

for a second set of one or more training values in another iteration of the multiple iterations, perform one or more additional propagation operations for the at least one layer of the neural network having the adjusted parameter, the one or more additional propagation operations comprising a propagation operation that updates activation values and/or weights for the neural network using the adjusted parameter.

2 . The computing system of claim 1 , further configured to, after adjusting the parameter of the neural network:

perform forward propagation for the at least one layer of the neural network having the adjusted parameter to propagate activation values for edges of the at least one layer, producing updated activation values;

perform backward propagation for at least one layer of the neural network having the adjusted parameter to propagate gradients for nodes of the at least one layer, producing updated weights; and

store the updated activation values and/or the updated weights in the at least one memory.

3 . The computing system of claim 1 , wherein:

the activation values and the weights are stored in a block floating-point format; and

the adjusted parameter comprises at least increasing a number of mantissa bits of the block floating-point format, increasing a number of exponent bits of the block floating-point format, or changing a sharing parameter of a shared exponent.

4 . The computing system of claim 1 , wherein the computing system is configured to:

update the activation values, the gradients, and/or node weights by increasing a number of bits used to store mantissa values in the activation values the gradients, and/or node weights, respectively, producing updated activation values, updated gradients, and/or updated node weights; and

store the updated activation values, updated gradients, and/or updated node weights in the at least one memory.

5 . The computing system of claim 1 , wherein the computing system determines the performance metric by:

determining differences between an output of a layer of the neural network from an expected output; and

based on the determining differences, adjusting the parameter by increasing a number of mantissa bits used to store updated activation values, updated gradients, and/or updated node weights in the neural network having the adjusted parameter.

6 . The computing system of claim 1 , wherein the computing system is further configured to store values for the neural network having the adjusted parameter in a computer-readable storage device or media, the values including at least one of: updated activation values, updated gradients, or updated node weights generated after the adjusting the parameter.

7 . The computing system of claim 1 , wherein the performance metric is based on at least one of the following for a layer of the neural network: number of true positives, number of true negatives, number of false positives, or number of false negatives.

8 . A method of operating a computing system implementing a neural network, the method comprising:

with the computing system, dynamically updating a parameter of the neural network during training of the neural network, the training of the neural network comprising:

for each of multiple iterations in an epoch of training of the neural network, in a given iteration of the multiple iterations:

for a first set of one or more training values, training at least one layer of the neural network by forward propagating and backward propagating activation or gradient values, respectively, for a number of training epochs using a first parameter value for a parameter of the neural network;

determining a performance metric for the neural network during the training of the neural network, the performance metric being selected from accuracy of the neural network, entropy of the neural network, mean square error of the neural network, perplexity of the neural network, or gradient signal to noise ratio of the neural network;

for the given iteration, using the performance metric, dynamically adjusting the parameter of the neural network during the training of the neural network to provide an adjusted parameter of the neural network and updating the neural network using the adjusted parameter to produce an updated neural network; and

with the updated neural network and during the training of the neural network, performing additional training for at least one layer of the neural network by forward propagating and backward propagating activation or gradient values, respectively, for at least one additional training epoch using a second set of one or more training values in another iteration of the multiple iterations.

9 . The method of claim 8 , wherein the parameter is adjusted according to a predetermined schedule.

10 . The method of claim 8 , wherein the parameter is a precision parameter or a topology parameter.

11 . The method of claim 8 , wherein the performance metric comprises at least one of the following:

accuracy of the neural network;

change in accuracy of the neural network over two or more of the training epochs;

accuracy of at least one layer of the neural network;

change in accuracy of at least one layer of the neural network over two or more training epochs;

entropy of at least one layer of the neural network; or

change in entropy of at least one layer of the neural network.

12 . The method of claim 8 , wherein the performance metric is based on accuracy or change in accuracy of at least one layer of the neural network, and wherein the accuracy or change in accuracy is measured based at least in part on one or more of the following for the at least one layer of the neural network: a true positive rate, a true negative rate, a positive predictive rate, a negative predictive value, a false negative rate, a false positive rate, a false discovery rate, a false omission rate, or an accuracy rate.

13 . The method of claim 8 , wherein the performance metric is based on one or more of the following: mean square error of at least one layer of the neural network, perplexity of at least one layer of the neural network, gradient signal to noise ratio of at least one layer of the neural network, or entropy of the neural network.

14 . The method of claim 8 , wherein the parameter is for a network topology parameter, and wherein the adjusting the precision parameter comprises at least one of the following: adjusting a number of layers of the neural network, adjusting a number of nodes of a layer of the neural network, adjusting sparsity of edges of a layer of the neural network or adjusting a number of non-zero edges of a layer of the neural network.

15 . One or more non-transitory computer-readable storage devices or media comprising:

computer-executable instructions that, when executed by a computing system comprising at least one processor and at least one memory coupled to the at least one processor, cause the computer system to implement a first layer of the neural network using first node weights and/or first activation values expressed in a first floating-point format;

computer-executable instructions that, when executed by the computing system, cause the computer system to, during training of the neural network, for a first set of one or more training values, forward propagate values from the first layer of the neural network to a second layer of the neural network using a first value for a parameter of the neural network, wherein the parameter is a precision parameter or a topology parameter, and to perform backward propagation for at least one layer of the neural network using the first parameter value;

computer-executable instructions that, when executed by the computing system, cause the computer system to, for the given iteration, a performance metric for the neural network, the performance metric being selected from accuracy of the neural network, entropy of the neural network, mean square error of the neural network, perplexity of the neural network, or gradient signal to noise ratio of the neural network, determine a training performance metric for the neural network;

computer-executable instructions that, when executed by the computing system, cause the computer system to dynamically adjust the parameter for the neural network based on the performance metric determined for the given iteration to provide an adjusted parameter, the adjusted parameter being selected to improve the performance metric in a neural network having the adjusted parameter; and

computer-executable instructions that, when executed by the computing system, cause the computer system to modify the neural network during the training of the neural network based on the adjusted precision parameter;

computer-executable instructions that, when executed by the computing system, cause the computer system to, for a second set of one or more training values in another iteration of the training, perform additional forward and backward propagation operations for at least one layer of the neural network using the adjusted parameter.

16 . The one or more non-transitory computer-readable storage devices or media of claim 15 , wherein the performance metric comprises at least one of the following:

accuracy of the neural network;

change in accuracy of the neural network over two or more of training epochs of the neural network;

accuracy of at least one layer of the neural network;

change in accuracy of at least one layer of the neural network over two or more training epochs;

entropy of at least one layer of the neural network; or

change in entropy of at least one layer of the neural network.

17 . The one or more non-transitory computer-readable storage devices or media of claim 15 , wherein:

the first floating point format is a block floating-point format; and

the adjusted precision parameter causes the computer system to increase a number of mantissa bits of the first floating point format so that at least some of the activation values and/or weights of the modified neural network are stored in a second floating point format having an increased number of mantissa bits.

18 . The one or more non-transitory computer-readable storage devices or media of claim 17 , wherein:

the first floating point format stores at least one value with a mantissa having one, two, three, four, five, or size bits; and

the second floating point format stores the at least one value with a mantissa having at least one more bits than the value in the first floating point format.

19 . The one or more non-transitory computer-readable storage devices or media of claim 17 , wherein the increased number of bits of mantissas in the second floating point format is selected using a rate distortion function.

20 . The one or more non-transitory computer-readable storage devices or media of claim 15 , wherein the first floating point format is a block floating-point format, and wherein the instructions further comprise:

computer-executable instructions that, when executed by the computing system, cause the computer system to, based on the adjusted precision parameter, modify the neural network to convert values from the first block floating-point format to a normal precision floating point format.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2019
From: DARVISH ROUHANI, BITA; CHUNG, ERIC S.; LO, DANIEL; BURGER, DOUGLAS C.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 047921/0728 →
Continuity (1)
Related Publication 20200210840A1 · Jul 2, 2020
References Cited (133)
US 5283559A · Kalendra et al. · 1994 [cited by applicant]
US 6144977A · Giangarra et al. · 2000 [cited by applicant]
US 6708068B1 · Sakaue · 2004 [cited by applicant]
US 8319738B2 · Taylor · 2012 [cited by applicant]
US 10167800B1 · Chung et al. · 2019 [cited by applicant]
US 10366217B2 · Finzi et al. · 2019 [cited by applicant]
US 11537859B2 · Cassidy · 2022 [cited by applicant]
US 11604647B2 · Sun · 2023 [cited by applicant]
US 12165038B2 · Lo · 2024 [cited by applicant]
US 20040041842A1 · Lippincott · 2004 [cited by applicant]
US 20060209041A1 · Studt et al. · 2006 [cited by applicant]
US 20070258641A1 · Srinivasan · 2007 [cited by examiner]
US 20080282046A1 · Yuuki · 2008 [cited by applicant]
US 20110154006A1 · Natu et al. · 2011 [cited by applicant]
US 20120165038A1 · Soma · 2012 [cited by applicant]
US 20130222320A1 · Huang et al. · 2013 [cited by applicant]
US 20140289445A1 · Savich · 2014 [cited by applicant]
US 20150084900A1 · Hodges · 2015 [cited by applicant]
US 20160070414A1 · Shukla et al. · 2016 [cited by applicant]
US 20160098149A1 · Baumgartner · 2016 [cited by applicant]
US 20160328646A1 · Lin et al. · 2016 [cited by applicant]
US 20180157465A1 · Bittner et al. · 2018 [cited by applicant]
US 20180157899A1 · Xu · 2018 [cited by examiner]
US 20180322607A1 · Mellempudi et al. · 2018 [cited by applicant]
US 20180341857A1 · Lee · 2018 [cited by examiner]
US 20190075301A1 · Chou · 2019 [cited by examiner]
US 20190386717A1 · Shattil · 2019 [cited by examiner]
US 20200042287A1 · Chalamalasetti · 2020 [cited by examiner]
US 20200202201A1 · Shirahata · 2020 [cited by examiner]
US 20200264876A1 · Lo et al. · 2020 [cited by applicant]
US 20200272213A1 · Sridharan et al. · 2020 [cited by applicant]
US 20250061320A1 · Lo · 2025 [cited by applicant]
CN 107636697A · 2018 [cited by applicant]
Drumond et al., “Training DNNs with Hybrid Block Floating Point”, Dec. 2, 2018, arXiv:1804.01526v4, pp. 1-11 (Year: 2018). [cited by examiner]
Na et al., “Speeding up Convolutional Neural Network Training with Dynamic Precision Scaling and Flexible Multiplier-Accumulator”, Aug. 8, 2016, ISLPED '16: Proceedings of the 2016 International Symposium on Low Power E… [cited by examiner]
Chakrabarti, et al., “Backprop with Approximate Activations for Memory-efficient Network Training”, In Journal of Computing Research Repository, Jan. 23, 2019, 09 Pages. [cited by applicant]
Krishnamoorthi, Raghuraman, “Quantizing Deep Convolutional Networks for Efficient Inference: A whitepaper”, In Journal of Computing Research Repository, Jun. 21, 2018, 36 Pages. [cited by applicant]
Park, et al., “Value-Aware Quantization for Training and Inference of Neural Networks”, In Proceedings of the European Conference on Computer Vision, Sep. 8, 2018, pp. 608-624. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US2020/015765”, Mailed Date: May 15, 2020, 14 Pages. [cited by applicant]
Wang, et al., “Training Deep Neural Networks with 8-bit Floating Point Numbers”, In Journal of Computing Research Repository, Dec. 19, 2018, 11 Pages. [cited by applicant]
Anonymous, Artificial Intelligence Index 2017 Annual Report, Nov. 2017, 101 pages. [cited by applicant]
Baydin et al., “Automatic Differentiation in Machine Learning: a Survey,” Journal of Machine Learning Research 18 (2018), Feb. 5, 2018, 43 pages (also published as arXiv:1502.05767V4 [cs.SC]Feb. 5, 2018). [cited by applicant]
Bulcò et al., “In-Place Activated BatchNorm for Memory-Optimized Training of DNNs,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 5639-5647 (also published as arXiv:1712.02616 [cs.C… [cited by applicant]
Burger, “Accelerating Persistent Neural Networks at Datacenter Scale,” Microsoft Corporation, 52 pp. accessed Apr. 18, 2018, available at: https://www.microsoft.com/en-us/research/blog/rnicrosoft-unveils-project-brainwa… [cited by applicant]
Burger, “Microsoft Unveils Project Brainwave for Real-Time AI,” Microsoft Corporation, 3 pp (Aug. 18, 2018). [cited by applicant]
Chen et al., “Compressing Neural Networks with the Hashing Trick,” In International Conference on Machine Learning, pages 2285-2294, 2015 (also cited as arXiv:1504.04788Vl [cs.LG] Apr. 19, 2015). [cited by applicant]
Chiu et al., State-of-the-art Speech Recognition with Sequence-to-Sequence Models. CoRR, abs/171201769, 2017 (also cited as arXiv:1712.01769V6 [cs.CL]Feb. 23, 2018). [cited by applicant]
Chung et al., “Serving DNNs in Real Time at Datacenter Scale with Project Brainwave,” IEEE Micro Pre-Print, 11 pages accessed Apr. 4, 2018, available at https://www.microsoft.com/en-us/research/uploads/prod/2018/03/rni0… [cited by applicant]
Colah, “Understanding LSTM Networks,” posted on Aug. 27, 2015, 13 pages. [cited by applicant]
Courbariaux et al., “Low precision arithmetic for deep learning,” also available as arXiv:1412.7024V1, Dec. 2014. [cited by applicant]
Courbariaux et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or −1,” arXiv preprint arXiv:1602.02830V3, Mar. 2016, 11 pages. [cited by applicant]
Courbariaux et al., “Binaryconnect: Training Deep Neural Networks with Binary Weights During Propagations,” In Proceedings of the 28th International Conference on Neural Information Processing Systems, vol. 2, Dec. 2015… [cited by applicant]
Courbariaux et al., “Training Deep Neural Networks with Low Precision Multiglications,” Sep. 23, 2015, 10 pages. [cited by applicant]
CS231n Convoiutional Neural Netwofks for Visual Recognition, downloaded from cs231n.github.io/ogtimization-2, Dec. 20, 2018, 9 pages. [cited by applicant]
Denil et al., Predicting Parameters in Deep Learning, In Advances in Neural Information Processing Sxstems, Dec. 2013, pp. 2148-2156. [cited by applicant]
Elam et al., “A Block Floating Point Implementation for an N-Point FFT on the TMS320C55X DSP,” Texas Instruments Application Report SPRA948, Sep. 2003, 13 pages. [cited by applicant]
“FFT/IFFT Block Floating Point Scaling,” Altera Corporation Application Note 404, Oct. 2005, ver. 1.0, 7 pages. [cited by applicant]
Goodfellow et al., “Deep Learning,” downloaded from http://www.deeplearningbook.org/ on May 2, 2018, (document dated 2016), 766 pages. [cited by applicant]
Gomez, “Backpropogating an LSTM: A Numerical Example,” Apr. 18, 2016, downloaded from medium.com/@aidangomez/let-s-do-this-f9b699de31d9, Dec. 20, 2018, 8 pages. [cited by applicant]
Gupta et al., “Deep Learning with Limited Numerical Precision,” Feb. 9, 2015, 10 pages. [cited by applicant]
Han et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding,” arXiV preprint arXiv:1510.00149V5 [cszCV], Feb. 15 2016, 14 pages. [cited by applicant]
Hassan et al., “Achieving Human Parity on Automatic Chinese to English News Translation,” CoRR, abs/180305567, 2018 (also published as arXiv:1803.05567V2 [cs.CL] Jun. 29, 2018). [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” arXiv preprint arXiv:1512.03385v1 [cs.CV] Dec. 10, 2015. [cited by applicant]
Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” arXiv:1502.03167v3 [cs.LG], Mar. 2015, 11 pages. [cited by applicant]
Jain et al., “Gist: Efficient Data Encoding for Deep Neural Network Training,” 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture, Jun. 2018, 14 pages. [cited by applicant]
Karl N's Blog., “Batch Normalization—What the hey?,” Posted on Jun. 7, 2016, downloaded from gab41.1ab41.org/batch-normalization-what-the-hey-d480039a9e3b, Jan. 9, 2019, 7 pages. [cited by applicant]
Kevin's Blog, “Deriving the Gradient for the Backward Pass of Batch Normalization,” Posted on Sep. 14, 2016, downloaded from kevinzakka.github.io/2016/09/14/batch_normalization/, Jan. 9, 2019, 7 pages. [cited by applicant]
Köster et al., “Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks,” In Advances in Neural Information Processing Systems, pp. 1742-1752, 2017 (also published as arXiv:1711.02213v2 [c… [cited by applicant]
Kratzert's Blog, “Understanding the backward pass through Batch Normalization Layer,” Posted on Feb. 12, 2016, downloaded from kratzert.github.io/2016/02/12/understanding-the-gradient-flow-through-the-bathchnor. . . on … [cited by applicant]
Langhammer et al., “Floating-Point DSP Block Architecture for FPGAs,” Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, Feb. 2015, pp. 117-125. [cited by applicant]
Le et al., “Neural Architecture Search with Reinforcement Learning,” PowerPoint presentation, 37 pages. [cited by applicant]
Le, “A Tutorial on Deep Learning, Part 1: Nonlinear Classifiers and the Backgrogagation Algorithm,” Dec. 2015, 18 pages. [cited by applicant]
Le, “A Tutorial on Deep Learning, Part 2: Autoencoders, Convolutional Neural Networks and Recurrent Neural Networks,” Oct. 2015, 20 pages. [cited by applicant]
Lecun et al., “Optimal Brain Damage,” In Advances in Neural Information Processing Systems, Nov. 1989, pp. 598-605. [cited by applicant]
Li et al., “Stochastic Modified Equations and Adaptive Stochastic Gradient Algorithms,” Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 2017, 10 pages. [cited by applicant]
Li et al., “Ternary eight networks,” arXiV preprint arXiv:1605.04711v2 [cszCV] Nov. 19, 2016. [cited by applicant]
Lin et al., “Fixed Point Quantization of Deep Convolutional Networks,” In International Conference on Machine Learning, pp. 2849-2858, 2016 (also ublished as arXiv:1511.06393v3 [cs:LG] Jun. 2, 2016). [cited by applicant]
Liu, “DARTS: Differentiable Architecture Search,” arXiv:1806.09055v1 [cs.LG], Jun. 24, 2018, 12 pages. [cited by applicant]
Mellempudi et al., “Ternary Neural Networks with Fine-Grained Quantization,” May 2017, 11 pages. [cited by applicant]
Mendis et al., “Helium: Lifting High-Performance Stencil Kernals from Stripped x86 Binaries to Halide DSL Code,” Proceedings of the 36th ACM SIGPLAN Conference on Programming Languate Design and Implementation, Jun. 201… [cited by applicant]
Mishra et al., “Apprentice: Using Knowledge Distillation Techniques to Improve Low-Precision Network Accuracy,” arXiV preprint arXiv:1711.05852v1 [cs:LG] Nov. 15, 2017. [cited by applicant]
Muller et al., “Handbook of Floating-Point Arithmetic,” Birkhäuser Boston (New York 2010), 78 pages including pp. 269-320. [cited by applicant]
Nielsen, “Neural Networks and Deep Learning,” downloaded from http://neuralnetworksanddeeplearning.com/index.html on May 2, 2018, document dated Dec. 2017, 314 pages. [cited by applicant]
Nvidia. Nvidia tensorrt optimizer, https://developer.nvidia.com/tensorrt downloaded on Mar. 4, 2019, 9 pages. [cited by applicant]
Page, “Neural Networks and Deep Learning,” www.cs.wise.edu/˜dpage/cs760/, 73 pp. [cited by applicant]
Park et al., “Energy-efficient Neural Network Accelerator Based on Outlier-aware Low-precision Computation,” 2018 ACM/IEEE 4th Annual International Symposium on Computer Architecture, Jun. 2018, pp. 688-698. [cited by applicant]
Rajagopal et al., “Synthesizing a Protocol Converter from Executable Protocol Traces,” IEEE Transactions on Computers, vol. 40, No. 4, Apr. 1991, pp. 487-499. [cited by applicant]
Rajpurkar et al., “SQuAD: 100,000+ Questions for Machine Comprehension of Text,” Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Nov. 2016, pp. 2383-2392. [cited by applicant]
Rastegari et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” In Proceedings of 14th Annual European Conference on Computer Vision, pp. 525-542. Oct. 2016. [cited by applicant]
“Russakovsky et al., ““ImageNet Large Scale Visual Recognition Challenge,”” International Journal of Computer Vision (IJCV), vol. 115, Issue 3, Dec. 2015, pp. 211-252 (also published as zrXiv:1409.0575v3 [cs.CV]Jan. 30,… [cited by applicant]
Russinovich, “Inside the Microsoft FPGA-based Configurable Cloud,” Microsoft Corporation, https://channel9.msdn.com/Events/Build/2017/B8063, 8 pp. (May 8,2017). [cited by applicant]
Russinovich, “Inside the Microsoft FPGA-based Configurable Cloud,” Microsoft Corporation, Powerpoint Presentation; 41 pp. (May 8, 2017). [cited by applicant]
Smith et al., “A Bayesian Perspective on Generalization and Stochastic Gradient Descent,” 6th International Conference on Learning Representations, Apr.-May 2018, 13 pages. [cited by applicant]
Szegedy et al., “Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning,” arXiv:1602.07261v2, Aug. 23, 2016, 12 pages. [cited by applicant]
Szegedy et al., “Rethinking the Inception Architecture for Computer Vision,” arXiv:1512.00567v3 [cs.CV] Dec. 11, 2015, 10 pages. [cited by applicant]
Tensorflow-slim image classification model library. https://github.com/tensorflow/models/tree/master/research/slim, downloaded on Mar. 4, 2019, 8 pages. [cited by applicant]
TITU1994 Blog, “Neural Architecture Search with Controller RNN,” downloaded from github.com/titu1994/neural-architecture-search on Jan. 9, 2019, 3 pages. [cited by applicant]
Vanhoucke et al., “Improving the speed of neural networks on CPUs,” In Deep Learning and Unsupervised Feature Learning Workshop, Dec. 2011, 8 pages. [cited by applicant]
Vucha et al., “Design and FPGA Implementation of Systolic Array Architecture for Matrix Multiplication,” International Journal of Computer Applications, vol. 26, No. 3, Jul. 2011, 5 pages. [cited by applicant]
Weinberger et al., “Feature Hashing for Large Scale Multitask Learning,” In Proceedings of the 26th Annual International Conference on Machine Learning, Jun. 2009, 8 pages. [cited by applicant]
Wen et al., “Learning Structured Sparsity in Deep Neural Networks,” In Advances in Neural Information Processing Systems, Dec. 2016, pp. 2074-2082 (also published as arXiv:1608.036654v4 [cs.NE] Oct. 18, 2016). [cited by applicant]
Wilkinson, “Rounding Errors in Algebraic Processes,” Notes on Applied Science No. 32, Department of Scientific and Industrial Research, National Physical Laboratory (United Kingdom) (London 1963), 50 pages including pp.… [cited by applicant]
Wired, “Microsoft's Internet Business Gets a New Kind of Processor,” 11 pp. Apr. 19, 2018, available at: https://www.wired.com/2016/09/microsoft-bets-future-chip-reprogram-fly/. [cited by applicant]
Xiong et al., “Achieving Human Parity in Conversational Speech Recognition,” arXiv:1610.05256v2 [cs:CL] Feb. 17, 2017, 13 pages. [cited by applicant]
Yeh, “Deriving Batch-Norm Backprop Equations,”downloaded from chrisyeh96.github.io/2017/08/28/deriving-batchnorm-backprop on Dec. 20, 2018, 5 pages. [cited by applicant]
Zhou et al., “DoReFa-net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients,” arXiv:1606.06160v3 [cs.NE] Feb. 2, 2018, 13 pages. [cited by applicant]
Zoph et al., “Learning Transferable Architectures for Scalable Image Recognition,” arXiv.1707.07012v1, Jul. 2017, 14 pages. [cited by applicant]
Zoph et al., “Neural Architecture Search with Reinforcement Learning,” 5th International Conference on Learning Representations, Apr. 2017, 16 pages. [cited by applicant]
Drumond, et al., “End-to-End DNN Training with Block Floating Point Arithmetic”, In Repository of arXiv:1804.01526v2, Apr. 9, 2018, 9 Pages. [cited by applicant]
Drumond, et al., “Training DNNs with Hybrid Block Floating Point”, In Repository of arXiv:1804.01526v4, Dec. 2, 2018, 11 Pages. [cited by applicant]
Han, et al., “Learning Both Weights and Connections for Efficient Neural Networks”, In Repository of arXiv:1506.02626v3, Oct. 30, 2015, 9 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US2019/066676”, Mailed Date: Apr. 15, 2020, 15 Pages. [cited by applicant]
Raghavan, et al., “Bit-Regularized Optimization of Neural Nets”, In Repository of arXiv:1708.04788, Aug. 16, 2017, 11 Pages. [cited by applicant]
Song, et al., “Computation Error Analysis of Block Floating Point Arithmetic Oriented Convolution Neural Network Accelerator Design”, In Repository of arXiv:1709.07776v2, Nov. 24, 2017, 8 Pages. [cited by applicant]
Wiedemann, et al., “Entropy-Constrained Training of Deep Neural Network”, In Repository of arXiv:1812.07520v2, Dec. 19, 2019, 8 Pages. [cited by applicant]
Park, et al., “Cell division: weight bit-width reduction technique for convolutional neural network hardware accelerators”, In Proceedings of the 24th Asia and South Pacific Design Automation Conference, Jan. 21, 2019, … [cited by applicant]
Zhao,, et al., “Improving Neural Network Quantization without Retraining using Outlier Channel Splitting”, In Proceedings of the 36th International Conference on Machine Learning, May 24, 2019, 10 pages. [cited by applicant]
“Office Action Issued in Indian Patent Application No. 202147035445”, Mailed Date: Feb. 16, 2023, 6 Pages. [cited by applicant]
Non-Final Office Action mailed on Mar. 14, 2024, in U.S. Appl. No. 16/276,395, 13 pages. [cited by applicant]
Communication Pursuant to Article 94 (3) EPC Received for European Patent Application No. 19839023.9, mailed on Jan. 2, 2024, 09 pages. [cited by applicant]
Controller Calibration on Resistive Touch Screens, DMC is a touch screen manufacturer, DMC Co. LTD., Jan. 22, 2019, 2 pages. [cited by applicant]
First Examination Report Received for Indian Patent Application No. 202147036372, mailed on Feb. 13, 2023, 6 pages. [cited by applicant]
First Office Action Received for Chinese Application No. 202080014556.X, mailed on Apr. 15, 2024, 24 pages (English Translation Provided). [cited by applicant]
International Search Report and written Opinion Received for PCT Application No. PCT/US2021/017406, mailed on Apr. 28, 2020, 10 pages. [cited by applicant]
Intimation of Grant received for Indian patent Application No. 202147035445, mailed on Oct. 10, 2024, 1 Page. [cited by applicant]
Non-Final Office Action mailed on Apr. 6, 2023, in U.S. Appl. No. 16/276,395, 18 pages. [cited by applicant]
Non-Final Office Action mailed on Nov. 2, 2020, in U.S. Appl. No. 16/286,098, 10 Pages. [cited by applicant]
Notice of Allowance mailed on Feb. 24, 2021, in U.S. Appl. No. 16/286,098, 5 pages. [cited by applicant]
Notice of Allowance mailed on Jul. 30, 2024, in U.S. Appl. No. 16/276,395, 5 pages. [cited by applicant]
Notification on Grant of Patent Right for Invention Received for Chinese Application No. 202080014556.X, mailed on Oct. 25, 2024, 4 pages (English Translation Provided). [cited by applicant]
Zoph, et al., “Neural Architecture Search with Reinforcement Learning,” In Proceedings of 5th International conference on Learning Representations, Apr. 2017, 16 Pages. [cited by applicant]
Communication pursuant to Article 94(3) EPC, Received for European Application No. 207083718, mailed on Mar. 20, 2024, 08 pages. [cited by applicant]
Communication pursuant to Article 94(3) EPC Received for European Application No. 20708371.8, mailed on Jan. 15, 2026, 07 pages. [cited by applicant]