IP Library › Granted Patent US 12,566,962
Granted Patent B2
US 12,566,962 · App. 16/240,514 · Granted Mar 3, 2026

Dithered quantization of parameters during training with a machine learning tool

Inventors: Thomas M. Annau (San Carlos, CA); Haishan Zhu (Redmond, WA); Daniel Lo (Bothell, WA); Eric S. Chung (Woodinville, WA)
Assignee: Microsoft Technology Licensing, LLC
G06N3/084G06F7/49963G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,962
App. No.
16/240,514
Granted
Mar 3, 2026
Kind
B2
Abstract

A machine learning tool uses dithered quantization of parameters during training of a machine learning model such as a neural network. The machine learning tool receives training data and initializes certain parameters of the machine learning model (e.g., weights for connections between nodes of a neural network, biases for nodes). The machine learning tool trains the parameters in one or more iterations based on the training data. In particular, in a given iteration, the machine learning tool applies the machine learning model to at least some of the training data and, based at least in part on the results, determines parameter updates to the parameters. The machine learning tool updates the parameters using the parameter updates and a dithered quantizer function, which can add random values before a rounding or truncation operation.

Claims (60)

1 . A method, implemented in a computer system comprising at least one hardware processor and at least one memory coupled to the at least one hardware processor, the method comprising:

receiving training data;

initializing parameters of a machine learning model;

training at least some of the parameters in multiple iterations based on the training data, including, in each of the multiple iterations during the training:

applying the machine learning model to at least some of the training data;

based at least in part on results of the applying the machine learning model, determining a measure of loss as a difference between an output of the machine learning model for the at least some of the training data and an expected output of the machine learning model for the at least some of the training data;

based at least in part on the measure of loss for the at least some of the training data, determining a set of parameter updates to the at least some of the parameters;

combining predetermined respective dither values from a random value generator function with respective combinations of respective parameter updates of the set of parameter updates and respective starting parameter values of respective parameters of the at least some of the parameters used during the applying the machine learning model to at least some of the training data, wherein at least a portion of the respective predetermined dither values differ from one another, to provide respective quantizer function operands; and

after the combining the predetermined respective dither values, quantizing the respective quantizer function operands using a quantizer function to provide one or more updated parameter values for at least a portion of the at least some of the parameters, wherein the one or more updated parameter values are used in at least one subsequent iteration of the multiple iterations.

2 . The method of claim 1 , wherein the machine learning model is a neural network having multiple layers, each of the multiple layers including one or more nodes, and wherein the parameters include one or more of:

weights for connections between the one or more nodes;

biases for at least some of the one or more nodes;

a count of the multiple layers;

for each of the multiple layers, a count of nodes in the layer; and

a control parameter for batch normalization.

3 . The method of claim 1 , wherein the combining and the quantizing are implemented as, where w t is a starting parameter value, Δw t is a parameter update, and d is a dither value:

round(w t +Δw t +d), wherein round( ) is a rounding function;

round((w t +Δw t +d)/p)×p, wherein p is a level of precision;

floor(w t +Δw t +d), wherein floor( ) is a floor function;

floor((w t +Δw t +d)/p)×p, wherein p is a level of precision; or

quant(w t +Δw t +d), wherein quant( ) is a quantizer function.

4 . The method of claim 1 , wherein the at least some of the parameters are in a floating-point format having a given level of precision for mantissa values, and wherein the combining applies dithering values d associated with the at least some of the parameters, respectively, by adding mantissa values at a higher level of precision before rounding or truncating to a nearest mantissa value for the given level of precision.

5 . The method of claim 1 , wherein the at least some of the parameters are in an integer format or fixed-point format, and wherein the combining applies dithering values d associated with the at least some of the parameters, respectively, by adding values at a higher level of precision before rounding or truncating to a nearest integer value.

6 . The method of claim 1 , wherein the quantizing is implemented using a rounding function, wherein the at least some of the parameters are in a format having a given level of precision after quantization is applied, and wherein dithering values d for the combining are selected in a range of −0.5 to 0.5 of an increment of a value at the given level of precision after quantization.

7 . The method of claim 1 , wherein the quantizing is implemented using a floor function, wherein the at least some of the parameters are in a format having a given level of precision after quantization is applied, and wherein dithering values d for the combining are selected in a range of 0.0 to 1.0 of an increment of a value at the given level of precision after quantization.

8 . The method of claim 1 , wherein the determining the set of parameter updates, Δw t , includes multiplying unscaled parameter updates Δw′ t by a learning rate η, as Δw t =η×Δw′ t .

9 . The method of claim 1 , wherein the training the at least some of the parameters further includes:

converting values, including the at least some of the parameters, from a first format to a second format, the second format having a lower precision than the first format, wherein the first format is a first floating-point format having m 1 bits of precision for mantissa values and e1 bits of precision for exponent values, wherein the second format is a second floating-point format:

having m 2 bits of precision for mantissa values and e 2 bits of precision for exponent values, wherein m 1 >m 2 , and wherein e 1 >e 2 ; or

having a shared exponent value.

10 . The method of claim 1 , wherein the machine learning model is a deep neural network, a support vector machine, a Bayesian network, a decision tree, or a linear classifier.

11 . The method of claim 1 , wherein at least a portion of the respective dither values are based at least in part on output of a random number generator, and wherein the respective dither values are random values having a power spectrum of white noise or blue noise.

12 . The method of claim 1 , wherein the quantizer function is the same in each of the multiple iterations.

13 . The method of claim 1 , wherein the training data includes multiple examples, each of the multiple examples having one or more attributes and a label, and wherein each of the multiple iterations uses a single example or mini-batch of examples randomly selected from among any remaining examples, for an epoch, of the multiple examples.

14 . The method of claim 1 , wherein the initializing the parameters includes:

for an initial iteration of an initial epoch, setting the parameters to random values.

15 . A computer system comprising:

a buffer, in memory of the computer system, configured to receive training data; and

a machine learning tool, implemented with one or more processors of the computer system, configured to perform operations comprising:

initializing parameters of a machine learning model;

training at least some of the parameters in multiple iterations based on the training data, including, in each of the multiple iterations during the training:

applying the machine learning model to at least some of the training data;

based at least in part on results of the applying the machine learning model, determining a measure of loss as a difference between an output of the machine learning model for the at least some of the training data and an expected output of the machine learning model for the at least some of the training data;

based at least in part on the measure of loss for the at least some of the training data, determining a set of parameter updates to the at least some of the parameters;

combining predetermined respective dither values from a random value generator function with respective combinations of respective parameter updates of the set of parameter updates and respective starting parameter values of respective parameters of the at least some of the parameters used during the applying the machine learning model to at least some of the training data, wherein at least a portion of the respective predetermined dither values differ from one another, to provide respective quantizer function operands;

after the combining the predetermined respective dither values, quantizing the respective quantizer function operands using a quantizer function to provide one or more updated parameter values for at least a portion of the at least some of the parameters, wherein the one or more updated parameter values are used in at least one subsequent iteration of training of the machine learning model; and

a buffer, in memory of the computer system, configured to store the parameters.

16 . One or more computer-readable media having stored thereon computer-executable instructions for causing one or more processing units, when programmed thereby, to perform operations comprising:

receiving training data;

initializing parameters of a machine learning model;

training at least some of the parameters in multiple iterations based on the training data, including, in each of the multiple iterations during the training:

applying the machine learning model to at least some of the training data;

based at least in part on results of the applying the machine learning model, determining a measure of loss as a difference between an output of the machine learning model for the at least some of the training data and an expected output of the machine learning model for the at least some of the training data;

based at least in part on the measure of loss for the at least some of the training data, determining a set of parameter updates to the at least some of the parameters;

combining predetermined respective dither values from a random value generator function with respective combinations of respective parameter updates of the set of parameter updates and respective starting parameter values of respective parameters of the at least some of the parameters used during the applying the machine learning model to at least some of the training data, wherein at least a portion of the respective predetermined dither values differ from one another, to provide respective quantizer function operands; and

after the combining the predetermined respective dither values, quantizing the respective quantizer function operands using a quantizer function to provide one or more updated parameter values for at least a portion of the at least some of the parameters, wherein the one or more updated parameter values are used in at least one subsequent iteration of training of the machine learning model.

17 . The method of claim 1 , wherein (1) the quantizer function is different between at least a portion of the multiple iterations, or (2) the quantizer function is not applied during one or more iterations of the multiple iterations.

18 . The computer system of claim 15 , wherein (1) the dithered quantizer function is different between at least a portion of the multiple iterations, or (2) the dithered quantization function is not applied during one or more iterations of the multiple iterations.

19 . The one or more computer-readable storage media of claim 16 , wherein (1) the dithered quantizer function is different between at least a portion of the multiple iterations, or (2) the dithered quantization function is not applied during one or more iterations of the multiple iterations.

20 . The method of claim 1 , wherein the at least a portion of the respective dither values are independently selected.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2019
From: ANNAU, THOMAS M.; ZHU, HAISHAN; LO, DANIEL; CHUNG, ERIC S.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 047928/0305 →
Continuity (1)
Related Publication 20200218982A1 · Jul 9, 2020
References Cited (100)
US 6144977A · Giangarra et al. · 2000 [cited by applicant]
US 10167800B1 · Chung et al. · 2019 [cited by applicant]
US 20140289445A1 · Savich · 2014 [cited by applicant]
US 20160328646A1 · Lin et al. · 2016 [cited by applicant]
US 20180107926A1 · Choi et al. · 2018 [cited by applicant]
US 20180157465A1 · Bittner et al. · 2018 [cited by applicant]
US 20190347554A1 · Choi · 2019 [cited by examiner]
US 20200242466A1 · Mohassel · 2020 [cited by examiner]
Yoojin Choi and Mostafa El-Khamy and Jungwon Lee (2018). Compression of Deep Convolutional Neural Networks under Joint Sparsity Constraints. CoRR, abs/1805.08303. (Year: 2018). [cited by examiner]
K. Ando, K. Ueyoshi, Y. Oba, K. Hirose, R. Uematsu, T. Kudo, M. Ikebe, T. Asai, S. Takamaeda-Yamazaki, & M. Motomura (2018). Dither NN: An Accurate Neural Network with Dithering for Low Bit-Precision Hardware. In 2018 I… [cited by examiner]
Andrew J. R. Simpson (2015). Taming the ReLU with Parallel Dither in a Deep Neural Network. CoRR, abs/1509.05173. (Year: 2015). [cited by examiner]
I. Boybat et al., “Improved Deep Neural Network Hardware-Accelerators Based on Non-Volatile-Memory: The Local Gains Technique,” 2017 IEEE International Conference on Rebooting Computing (ICRC), Washington, DC, 2017, pp.… [cited by examiner]
Yi-Te Hsu, Yu-Chen Lin, Szu-Wei Fu, Yu Tsao, & Tei-Wei Kuo. (2018). A study on speech enhancement using exponent-only floating point quantized neural network (EOFP-QNN). (Year: 2018). [cited by examiner]
Julian Faraone and Nicholas J. Fraser and Michaela Blott and Philip Heng Wai Leong (2018). SYQ: Learning Symmetric Quantization For Efficient Deep Neural Networks. CoRR, abs/1807.00301. p. 1-10 (Year: 2018). [cited by examiner]
R. M. Gray and T. G. Stockham, “Dithered quantizers,” in IEEE Transactions on Information Theory, vol. 39, No. 3, May 1993, doi: 10.1109/18.256489. p. 1-8 (Year: 1993). [cited by examiner]
Aldrich, N. (2005). Exploring Dither in Floating-Point Systems. Researchgate.net (Year: 2005). [cited by examiner]
Alistarh, D., Grubic, D., Li, J., Tomioka, R., & Vojnovic, M. (2017). QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding. In Advances in Neural Information Processing Systems. Curran Associates, In… [cited by examiner]
Zhang, D., Yang, J., Ye, D., & Hua, G.. (2018). LQ-Nets: Learned Quantization for Highly Accurate and Compact Deep Neural Networks. (Year: 2018). [cited by examiner]
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., & Bengio, Y.. (2016). Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations. (Year: 2016). [cited by examiner]
Katarzyna Janocha, & Wojciech Marian Czarnecki. (2017). On Loss Functions for Deep Neural Networks in Classification. (Year: 2017). [cited by examiner]
Hao Li, Soham De, Zheng Xu, Christoph Studer, Hanan Samet, & Tom Goldstein. (2017). Training Quantized Nets: A Deeper Understanding. (Year: 2017). [cited by examiner]
Kevin Kepp “Quantized Back Propagation for Efficient Deep Learning” Oct. 19, 2018 Master's Thesis, The Fraunhofer Institute for Telecommunications, URL: https://www.kevinkepp.de/assets/docs/masters-thesis.pdf (Year: 201… [cited by examiner]
Andrew Ng “CS294A Lecture notes : Sparse autoencoder” CS294A/CS294W Deep Learning and Unsupervised Feature Learning, Winter 2011, Undergraduate class (Year: 2011). [cited by examiner]
Maclin, Richard & Shavlik, Jude. (1999). Combining the Predictions of Multiple Classifiers: Using Competitive Learning to Initialize Neural Networks. (Year: 1999). [cited by examiner]
Ando et al., “Dither NN: An Accurate Neural Network with Dithering for Low Bit-Precision Hardware,” [cited by applicant]
Choi et al., “Compression of Deep Convolutional Neural Networks under Joint Sparsity Constraints,” arXiv:1805.08303v2, 16 pp. (May 2018). [cited by applicant]
International Search Report and Written Opinion dated Apr. 20, 2020, from International Patent Application No. PCT/US2019/067303, 15 pp. [cited by applicant]
Anonymous, Artificial Intelligence Index 2017 Annual Report, Nov. 2017, 101 pages. [cited by applicant]
Baydin et al., “Automatic Differentiation in Machine Learning: a Survey,” Journal of Machine Learning Research 18 (2018), Feb. 5, 2018, 43 pages (also published as arXiv:1502.05767v4 [cs.SC] Feb. 5, 2018). [cited by applicant]
Bulò et al., “In-Place Activated BatchNorm for Memory-Optimized Training of DNNs,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 5639-5647 (also published as arXiv:1712.02616 [cs.CV… [cited by applicant]
Burger, “Accelerating Persistent Neural Networks at Datacenter Scale,” Microsoft Corporation, 52 pp. accessed Apr. 18, 2018, available at: https://www.microsoft.com/en-us/research/blog/microsoft-unveils-project-brainwav… [cited by applicant]
Burger, “Microsoft Unveils Project Brainwave for Real-Time AI,” Microsoft Corporation, 3 pp (Aug. 18, 2018). [cited by applicant]
Chen et al., “Compressing Neural Networks with the Hashing Trick,” In International Conference on Machine Learning, pp. 2285-2294, 2015 (also cited as arXiv:1504.04788v1 [cs.LG] Apr. 19, 2015). [cited by applicant]
Chiu et al., State-of-the-art Speech Recognition with Sequence-to-Sequence Models. CoRR, abs/1712.01769, 2017 (also cited as arXiv:1712.01769v6 [cs.CL] Feb. 23, 2018). [cited by applicant]
Chung et al., “Serving DNNs in Real Time at Datacenter Scale with Project Brainwave,” IEEE Micro Pre-Print, 11 pages accessed Apr. 4, 2018, available at https://www.microsoft.com/en-us/research/uploads/prod/2018/03/mi02… [cited by applicant]
Colah, “Understanding LSTM Networks,” posted on Aug. 27, 2015, 13 pages. [cited by applicant]
Courbariaux et al., “Low precision arithmetic for deep learning,” also available as arXiv:1412.7024v1, Dec. 2014. [cited by applicant]
Courbariaux et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or −1,” arXiv preprint arXiv:1602.02830v3, Mar. 2016, 11 pages. [cited by applicant]
Courbariaux et al., “Binaryconnect: Training Deep Neural Networks with Binary Weights During Propagations,” In Proceedings of the 28th International Conference on Neural Information Processing Systems, vol. 2, Dec. 2015… [cited by applicant]
Courbariaux et al., “Training Deep Neural Networks with Low Precision Multiplications,” Sep. 23, 2015, 10 pages. [cited by applicant]
CS231n Convolutional Neural Networks for Visual Recognition, downloaded from cs231n.github.io/optimization-2, Dec. 20, 2018, 9 pages. [cited by applicant]
Denil et al., Predicting Parameters in Deep Learning, In Advances in Neural Information Processing Systems, Dec. 2013, pp. 2148-2156. [cited by applicant]
Dunay et al., “Dithering for Floating-Point Number Representation,” 1st International On-Line Workshop on Dithering in Measurement, 12 pp. (1998). [cited by applicant]
Elam et al., “A Block Floating Point Implementation for an N-Point FFT on the TMS320C55x DSP,” Texas Instruments Application Report SPRA948, Sep. 2003, 13 pages. [cited by applicant]
“FFT/IFFT Block Floating Point Scaling,” Altera Corporation Application Note 404, Oct. 2005, ver. 1.0, 7 pages. [cited by applicant]
Goodfellow et al., “Deep Learning,” downloaded from http://www.deeplearningbook.org/ on May 2, 2018, (document dated 2016), 766 pages. [cited by applicant]
Gomez, “Backpropogating an LSTM: A Numerical Example,” Apr. 18, 2016, downloaded from medium.com/@aidangomez/let-s-do-this-f9b699de31d9, Dec. 20, 2018, 8 pages. [cited by applicant]
Gupta et al., “Deep Learning with Limited Numerical Precision,” Feb. 9, 2015, 10 pages. [cited by applicant]
Han et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding,” arXiv preprint arXiv:1510.00149v5 [cs:CV], Feb. 15, 2016, 14 pages. [cited by applicant]
Hassan et al., “Achieving Human Parity on Automatic Chinese to English News Translation,” CoRR, abs/1803.05567, 2018 (also published as arXiv:1803.05567v2 [cs.CL] Jun. 29, 2018). [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” arXiv preprint arXiv:1512.03385v1 [cs.CV] Dec. 10, 2015. [cited by applicant]
Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” arXiv:1502.03167v3 [cs.LG], Mar. 2015, 11 pages. [cited by applicant]
Jain et al., “Gist: Efficient Data Encoding for Deep Neural Network Training,” 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture, Jun. 2018, 14 pages. [cited by applicant]
Karl N's Blog., “Batch Normalization—What the hey?,” Posted on Jun. 7, 2016, downloaded from gab41.lab41.org/batch-normalization-what-the-hey-d480039a9e3b, Jan. 9, 2019, 7 pages. [cited by applicant]
Kevin's Blog, “Deriving the Gradient for the Backward Pass of Batch Normalization,” Posted on Sep. 14, 2016, downloaded from kevinzakka.github.io/2016/09/14/batch_normalization/, Jan. 9, 2019, 7 pages. [cited by applicant]
Köster et al., “Flexpoint: An Adaptive Numerical Format for Efficient Training of Deep Neural Networks,” In Advances in Neural Information Processing Systems, pp. 1742-1752, 2017 (also published as arXiv:1711.02213v2 [c… [cited by applicant]
Kratzert's Blog, “Understanding the backward pass through Batch Normalization Layer,” Posted on Feb. 12, 2016, downloaded from kratzert.github.io/2016/02/12/understanding-the-gradient-flow-through-the-bathch-nor . . . o… [cited by applicant]
Langhammer et al., “Floating-Point DSP Block Architecture for FPGAs,” Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, Feb. 2015, pp. 117-125. [cited by applicant]
Le et al., “Neural Architecture Search with Reinforcement Learning,” PowerPoint presentation, 37 pages. [cited by applicant]
Le, “A Tutorial on Deep Learning, Part 1: Nonlinear Classifiers and the Backpropagation Algorithm,” Dec. 2015, 18 pages. [cited by applicant]
Le, “A Tutorial on Deep Learning, Part 2: Autoencoders, Convolutional Neural Networks and Recurrent Neural Networks,” Oct. 2015, 20 pages. [cited by applicant]
Lecun et al., “Optimal Brain Damage,” In Advances in Neural Information Processing Systems, Nov. 1989, pp. 598-605. [cited by applicant]
Li et al., “Stochastic Modified Equations and Adaptive Stochastic Gradient Algorithms,” Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 2017, 10 pages. [cited by applicant]
Li et al., “Ternary eight networks,” arXiv preprint arXiv:1605.04711v2 [cs:CV] Nov. 19, 2016. [cited by applicant]
Lin et al., “Fixed Point Quantization of Deep Convolutional Networks,” In International Conference on Machine Learning, pp. 2849-2858, 2016 (also published as arXiv:1511.06393v3 [cs:LG] Jun. 2, 2016). [cited by applicant]
Liu, “DARTS: Differentiable Architecture Search,” arXiv:1806.09055v1 [cs.LG], Jun. 24, 2018, 12 pages. [cited by applicant]
Mellempudi et al., “Ternary Neural Networks with Fine-Grained Quantization,” May 2017, 11 pages. [cited by applicant]
Mendis et al., “Helium: Lifting High-Performance Stencil Kernals from Stripped x86 Binaries to Halide DSL Code,” Proceedings of the 36th ACM SIGPLAN Conference on Programming Languate Design and Implementation, Jun. 201… [cited by applicant]
Mishra et al., “Apprentice: Using Knowledge Distillation Techniques to Improve Low-Precision Network Accuracy,” arXiv preprint arXiv:1711.05852v1 [cs:LG] Nov. 15, 2017. [cited by applicant]
Muller et al., “Handbook of Floating-Point Arithmetic,” Birkhäuser Boston (New York 2010), 78 pages including pp. 269-320. [cited by applicant]
Nielsen, “Neural Networks and Deep Learning,” downloaded from http://neuralnetworksanddeeplearning.com/index.html on May 2, 2018, document dated Dec. 2017, 314 pages. [cited by applicant]
Nvidia. Nvidia tensorrt optimizer, https://developer.nvidia.com/tensorrt downloaded on Mar. 4, 2019, 9 pages. [cited by applicant]
Page, “Neural Networks and Deep Learning,” www.cs.wise.edu/˜dpage/cs760/, 73 pp. [cited by applicant]
Park et al., “Energy-efficient Neural Network Accelerator Based on Outlier-aware Low-precision Computation,” 2018 ACM/IEEE 4th Annual International Symposium on Computer Architecture, Jun. 2018, pp. 688-698. [cited by applicant]
Rajagopal et al., “Synthesizing a Protocol Converter from Executable Protocol Traces,” IEEE Transactions on Computers, vol. 40, No. 4, Apr. 1991, pp. 487-499. [cited by applicant]
Rajpurkar et al., “SQuAD: 100,000+ Questions for Machine Comprehension of Text,” Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Nov. 2016, pp. 2383-2392. [cited by applicant]
Rastegari et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” In Proceedings of 14th Annual European Conference on Computer Vision, pp. 525-542. Oct. 2016. [cited by applicant]
“Russakovsky et al., ““ImageNet Large Scale Visual Recognition Challenge,”” International Journal of Computer Vision (IJCV), vol. 115, Issue 3, Dec. 2015, pp. 211-252 (also published as zrXiv:1409.0575v3 [cs.CV] Jan. 30… [cited by applicant]
Russinovich, “Inside the Microsoft FPGA-based Configurable Cloud,” Microsoft Corporation, https://channel9.msdn.com/Events/Build/2017/B8063, 8 pp. (May 8, 2017). [cited by applicant]
Russinovich, “Inside the Microsoft FPGA-based Configurable Cloud,” Microsoft Corporation, Powerpoint Presentation; 41 pp. (May 8, 2017). [cited by applicant]
Smith et al., “A Bayesian Perspective on Generalization and Stochastic Gradient Descent,” 6th International Conference on Learning Representations, Apr.-May 2018, 13 pages. [cited by applicant]
Szegedy et al., “Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning,” arXiv:1602.07261v2, Aug. 23, 2016, 12 pages. [cited by applicant]
Szegedy et al., “Rethinking the Inception Architecture for Computer Vision,” arXiv:1512.00567v3 [cs.CV] Dec. 11, 2015, 10 pages. [cited by applicant]
Tensorflow-slim image classification model library. https://github.com/tensorflow/models/tree/master/research/slim, downloaded on Mar. 4, 2019, 8 pages. [cited by applicant]
TITU1994 Blog, “Neural Architecture Search with Controller RNN,” downloaded from github.com/titu1994/neural-architecture-search on Jan. 9, 2019, 3 pages. [cited by applicant]
Vanhoucke et al., “Improving the speed of neural networks on CPUs,” In Deep Learning and Unsupervised Feature Learning Workshop, Dec. 2011, 8 pages. [cited by applicant]
Vucha et al., “Design and FPGA Implementation of Systolic Array Architecture for Matrix Multiplication,” International Journal of Computer Applications, vol. 26, No. 3, Jul. 2011, 5 pages. [cited by applicant]
Weinberger et al., “Feature Hashing for Large Scale Multitask Learning,” In Proceedings of the 26th Annual International Conference on Machine Learning, Jun. 2009, 8 pages. [cited by applicant]
Wen et al., “Learning Structured Sparsity in Deep Neural Networks,” In Advances in Neural Information Processing Systems, Dec. 2016, pp. 2074-2082 (also published as arXiv:1608.036654v4 [cs.NE] Oct. 18, 2016). [cited by applicant]
Wikipedia, “Dither,” 9 pages (document dated Jan. 2, 2019). [cited by applicant]
Wilkinson, “Rounding Errors in Algebraic Processes,” Notes on Applied Science No. 32, Department of Scientific and Industrial Research, National Physical Laboratory (United Kingdom) (London 1963), 50 p. including pp. 26… [cited by applicant]
Wired, “Microsoft's Internet Business Gets a New Kind of Processor,” 11 pp. Apr. 19, 2018, available at: https://www.wired.com/2016/09/microsoft-bets-future-chip-reprogram-fly/. [cited by applicant]
Xiong et al., “Achieving Human Parity in Conversational Speech Recognition,” arXiv:1610.05256v2 [cs:CL] Feb. 17, 2017, 13 pages. [cited by applicant]
Yeh, “Deriving Batch-Norm Backprop Equations,” downloaded from chrisyeh96.github.io/2017/08/28/deriving-batchnorm-backprop on Dec. 20, 2018, 5 pages. [cited by applicant]
Zhou et al., “DoReFa-net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients,” arXiv:1606.06160v3 [cs.NE] Feb. 2, 2018, 13 pages. [cited by applicant]
Zoph et al., “Learning Transferable Architectures for Scalable Image Recognition,” arXiv.1707.07012v1, Jul. 2017, 14 pages. [cited by applicant]
Zoph et al., “Neural Architecture Search with Reinforcement Learning,” 5th International Conference on Learning Representations, Apr. 2017, 16 pages. [cited by applicant]
Communication pursuant to Rules 161(1) and 162 EPC dated Aug. 11, 2021, from European Patent Application No. 19839308.4, 3 pp. [cited by applicant]
Communication pursuant to Article 94(3) Received in European Patent Application No. 19839308.4 (MS# 405473-EP01-PCT), mailed on Jul. 23, 2024, 7 pages. [cited by applicant]
Decision to refuse Received for European Application No. 19839308.4, (MS#405473-EP01-PCT), mailed on Jul. 28, 2025, 16 pages. [cited by applicant]