IP Library Granted Patent US 12,299,555
Granted Patent B2
US 12,299,555 · App. 17/894,798 · Granted May 13, 2025

Training network with discrete weight values

Inventors: Steven L. Teig (Menlo Park, CA); Eric A. Sather (Palo Alto, CA)
Assignee: Amazon Technologies, Inc.
G06N3/04G06F17/12G06F17/18G06F18/211G06N3/08G06N3/084G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,555
App. No.
17/894,798
Granted
May 13, 2025
Kind
B2
Abstract

Some embodiments provide an electronic device that includes a set of processing units and a set of machine-readable media. The set of machine-readable media stores sets of instructions for applying a network of computation nodes to an input received by the device. The set of machine-readable media stores at least two sets of machine-trained parameters for configuring the network for different types of inputs. A first of the sets of parameters is used for applying the network to a first type of input and a second of the sets of parameters is used for applying the network to a second type of input.

Claims (28)

1. A method for training a plurality of weights of a neural network comprising a plurality of layers, the method comprising:

propagating a plurality of inputs through the plurality of layers of the neural network to generate a plurality of outputs, each input having a corresponding expected output;

calculating a value for a loss function that comprises (i) a loss term that measures, for each input, a difference between the generated output and the expected output for the input and (ii) a constraint term that penalizes weights for differentiating from a ternary set of allowed values for the weights, wherein the ternary set of allowed values for each weight comprises zero, a positive value for a layer to which the weight belongs, and a negation of the positive value for the layer to which the weight belongs; and

using the calculated value for the loss function comprising the loss term and the constraint term to train the weights of the neural network, wherein constraining the weights to the ternary set causes weight data for the network to be represented using two or fewer bits for each weight and stored on a processor without the use of a separate memory, for a device in which the neural network is embedded, wherein the processor for the device (i) stores each weight as a ternary value of the set {0, 1, −1} and (ii) computes dot products involving the weights using the ternary value and multiplies the dot products by the positive value for the layer to which the weight belongs when executing the neural network.

2. The method of claim 1 , wherein the constraint term is a continuously-differentiable function of each of the weights that is equal to zero for a particular weight when the particular weight is equal to any one of the ternary set of allowed values for the particular weight.

3. The method of claim 1 , wherein the ternary set of allowed values for a first layer is different than the ternary set of allowed values for a second layer.

4. The method of claim 1 , wherein propagating the plurality of inputs through the plurality of layers of the neural network comprises rounding each respective weight to one of the respective allowed weight values for the respective weight value in order to propagate the plurality of inputs.

5. The method of claim 1 further comprising iteratively propagating inputs through the neural network, calculating a new value for the loss function, and using the new calculated value to train the weights of the neural network.

6. The method of claim 5 , wherein a size of the constraint term is increased relative to the loss term during later iterations such that the constraint term has a larger effect on the training during the later iterations.

7. The method of claim 1 , wherein the neural network is trained to perform a specific function when embedded in the device.

8. The method of claim 1 further comprising imposing an additional constraint on the weights of the network that at least a particular percentage of the weights with ternary values be equal to zero.

9. The method of claim 1 , wherein using the calculated value for the loss function to train the weights of the neural network comprises:

backpropagating the calculated value through the neural network to determine, for each weight, a rate of change in the calculated value relative to a rate of change in the weight; and

modifying each particular weight according to the determined rate of change for the particular weight.

10. A non-transitory machine-readable medium storing a program which when executed by at least one processing unit trains a plurality of weights of a neural network comprising a plurality of layers, the program comprising sets of instructions for:

propagating a plurality of inputs through the plurality of layers of the neural network to generate a plurality of outputs, each input having a corresponding expected output;

calculating a value for a loss function that comprises (i) a loss term that measures, for each input, a difference between the generated output and the expected output for the input and (ii) a constraint term that penalizes weights for differentiating from a ternary set of allowed values for the weights, wherein the ternary set of allowed values for each weight comprises zero, a positive value for a layer to which the weight belongs, and a negation of the positive value for the layer to which the weight belongs; and

using the calculated value for the loss function comprising the loss term and the constraint term to train the weights of the neural network, wherein constraining the weights to the ternary set causes weight data for the network to be represented using two or fewer bits for each weight and stored on a processor without the use of a separate memory, for a device in which the neural network is embedded, wherein the processor for the device (i) stores each weight as a ternary value of the set {0, 1, −1} and (ii) computes dot products involving the weights using the ternary value and multiplies the dot products by the positive value for the layer to which the weight belongs when executing the neural network.

11. The non-transitory machine-readable medium of claim 10 , wherein the constraint term is a continuously-differentiable function of each of the weights that is equal to zero for a particular weight when the particular weight is equal to any one of the ternary set of allowed values for the particular weight.

12. The non-transitory machine-readable medium of claim 10 , wherein the ternary set of allowed values for a first layer is different than the ternary set of allowed values for a second layer.

13. The non-transitory machine-readable medium of claim 10 , wherein the set of instructions for propagating the plurality of inputs through the plurality of layers of the neural network comprises a set of instructions for rounding each respective weight to one of the respective allowed weight values for the respective weight value in order to propagate the plurality of inputs.

14. The non-transitory machine-readable medium of claim 10 , wherein the program further comprises sets of instructions for iteratively propagating inputs through the neural network, calculating a new value for the loss function, and using the new calculated value to train the weights of the neural network.

15. The non-transitory machine-readable medium of claim 14 , wherein a size of the constraint term is increased relative to the loss term during later iterations such that the constraint term has a larger effect on the training during the later iterations.

16. The non-transitory machine-readable medium of claim 10 , wherein the neural network is trained to perform a specific function when embedded in the device.

17. The non-transitory machine-readable medium of claim 10 , wherein the program further comprises a set of instructions for imposing an additional constraint on the weights of the network that at least a particular percentage of the weights with ternary values be equal to zero.

18. The non-transitory machine-readable medium of claim 10 , wherein the set of instructions for using the calculated value for the loss function to train the weights of the neural network comprises sets of instructions for:

backpropagating the calculated value through the neural network to determine, for each weight, a rate of change in the calculated value relative to a rate of change in the weight; and

modifying each particular weight according to the determined rate of change for the particular weight.

Assignments (2)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
Continuity (3)
Continuation 15815235 · Nov 16, 2017
Provisional Application 62492940 · May 1, 2017
Related Publication 20220405591A1 · Dec 22, 2022
References Cited (98)
US 4918618A · Tomlinson, Jr. · 1990 [cited by applicant]
US 5255347A · Matsuba et al. · 1993 [cited by applicant]
US 5477225A · Young et al. · 1995 [cited by applicant]
US 5956703A · Turner et al. · 1999 [cited by applicant]
US 6571225B1 · Oles et al. · 2003 [cited by applicant]
US 9633282B2 · Sharma et al. · 2017 [cited by applicant]
US 9904874B2 · Shoaib et al. · 2018 [cited by applicant]
US 10387531B1 · Vanhoucke · 2019 [cited by examiner]
US 10515303B2 · Lie et al. · 2019 [cited by applicant]
US 10657438B2 · Lie et al. · 2020 [cited by applicant]
US 10740434B1 · Duong et al. · 2020 [cited by applicant]
US 10796198B2 · Franca-Neto · 2020 [cited by applicant]
US 10949736B2 · Deisher et al. · 2021 [cited by applicant]
US 11017295B1 · Teig et al. · 2021 [cited by applicant]
US 11113603B1 · Teig et al. · 2021 [cited by applicant]
US 11138292B1 · Nair et al. · 2021 [cited by applicant]
US 11429861B1 · Teig et al. · 2022 [cited by applicant]
US 20040078403A1 · Scheuermann et al. · 2004 [cited by applicant]
US 20150161995A1 · Sainath et al. · 2015 [cited by applicant]
US 20160086078A1 · Ji et al. · 2016 [cited by applicant]
US 20160174902A1 · Georgescu et al. · 2016 [cited by applicant]
US 20160328643A1 · Liu et al. · 2016 [cited by applicant]
US 20170011288A1 · Brothers et al. · 2017 [cited by applicant]
US 20170161640A1 · Shamir · 2017 [cited by applicant]
US 20170243110A1 · Esquivel et al. · 2017 [cited by applicant]
US 20180025268A1 · Teig et al. · 2018 [cited by applicant]
US 20180046900A1 · Dally et al. · 2018 [cited by applicant]
US 20180107925A1 · Choi et al. · 2018 [cited by applicant]
US 20180114113A1 · Ghahramani et al. · 2018 [cited by applicant]
US 20180121796A1 · Deisher et al. · 2018 [cited by applicant]
US 20180189229A1 · Desoli et al. · 2018 [cited by applicant]
US 20180189638A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180218518A1 · Yan et al. · 2018 [cited by applicant]
US 20180246855A1 · Redfern et al. · 2018 [cited by applicant]
US 20180285719A1 · Baum et al. · 2018 [cited by applicant]
US 20180285727A1 · Baum et al. · 2018 [cited by applicant]
US 20180285736A1 · Baum et al. · 2018 [cited by applicant]
US 20180293493A1 · Kalamkar et al. · 2018 [cited by applicant]
US 20180293691A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180300600A1 · Ma et al. · 2018 [cited by applicant]
US 20180307980A1 · Barik et al. · 2018 [cited by applicant]
US 20180307985A1 · Appu et al. · 2018 [cited by applicant]
US 20180315158A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180373975A1 · Yu et al. · 2018 [cited by applicant]
US 20190026078A1 · Bannon et al. · 2019 [cited by applicant]
US 20190114499A1 · Delaye et al. · 2019 [cited by applicant]
US 20190138896A1 · Deng · 2019 [cited by applicant]
US 20190147338A1 · Pau et al. · 2019 [cited by applicant]
US 20190187983A1 · Ovsiannikov et al. · 2019 [cited by applicant]
US 20190205358A1 · Diril et al. · 2019 [cited by applicant]
US 20190303750A1 · Kumar et al. · 2019 [cited by applicant]
US 20190354868A1 · Wierstra et al. · 2019 [cited by applicant]
US 20200380344A1 · Lie et al. · 2020 [cited by applicant]
US 20210110236A1 · Shibata · 2021 [cited by applicant]
CN 108876698A · 2018 [cited by applicant]
CN 108280514B · 2020 [cited by applicant]
WO 2012024329A1 · 2012 [cited by applicant]
WO 2020044527A1 · 2020 [cited by applicant]
Zhu et al. “Trained Ternary Quantization”, ICLR, 2017, pp. 10. [cited by examiner]
Achterhold, Jan, et al., “Variational Network Quantization,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 18 pages, ICLR, Vancouver, BC, Canada. [cited by applicant]
Andri, Renzo, et al., “YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, Mar. 14, 2017, 14 pages, IEEE, New York,… [cited by applicant]
Ardakani, Arash, et al., “Sparsely-Connected Neural Networks: Towards Efficient VLSI Implementation of Deep Neural Networks,” Proceedings of the 5th International Conference on Learning Representations (ICLR 2017), Apr.… [cited by applicant]
Bang, Suyoung, et al., “A 288 μW Programmable Deep-Learning Processor with 270KB On-Chip Weight Storage Using Non-Uniform Memory Hierarchy for Mobile Intelligence,” Proceedings of 2017 IEEE International Solid-State Cir… [cited by applicant]
Bruns, Erich, et al., “Mobile Phone-Enabled Museum Guidance with Adaptive Classification,” IEEE Computer Graphics and Applications, Jul. 9, 2008, 5 pages, vol. 28, Issue 4, IEEE. [cited by applicant]
Castelli, Ilaria, et al., “Combination of Supervised and Unsupervised Learning for Training the Activation Functions of Neural Networks,” Pattern Recognition Letters, Jun. 26, 2013, 14 pages, vol. 37, Elsevier B.V. [cited by applicant]
Chakraborty, Bishwajit, et al., “Acoustic Seafloor Sediment Classification Using Self-Organizing Feature Maps,” IEEE Transactions on Geoscience and Remote Sensing, Dec. 2001, 4 pages, vol. 39, Issue 12, IEEE. [cited by applicant]
Courbariaux, Matthieu, et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or −1,” Mar. 17, 2016, 11 pages, arXiv:1602.02830v3, Computing Research Repository (CoRR… [cited by applicant]
Courbariaux, Matthieu, et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations,” Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS 15),… [cited by applicant]
Duda, Jarek, “Asymmetric Numeral Systems: Entropy Coding Combining Speed of Huffman Coding with Compression Rate of Arithmetic Coding,” Jan. 6, 2014, 24 pages, arXiv:1311.2540v2, Computer Research Repository (CoRR)—Corn… [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Forssell, Mats, “Hardware Implementation of Artificial Neural Networks,” Information Flow in Networks, Month Unknown 2013, 4 pages. [cited by applicant]
Gao, Mingyu, et al., “Tetris: Scalable and Efficient Neural Network Acceleration with 3D Memory,” Proceedings of the 22nd International Conference on Architectural Support for Programming Languages and Operating Systems… [cited by applicant]
Guo, Yiwen, et al., “Network Sketching: Exploring Binary Structure in Deep CNNs,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 9 pages, IEEE, Honolulu, HI. [cited by applicant]
Harmon, Mark, et al., “Activation Ensembles for Deep Neural Networks,” Feb. 24, 2017, 9 pages, arXiv:1702.07790v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
He, Zhezhi, et al., “Optimize Deep Convolutional Neural Network with Ternarized Weights and High Accuracy,” Jul. 20, 2018, 8 pages, arXiv:1807.07948v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY… [cited by applicant]
Hegde, Kartik, et al., “UCNN: Exploiting Computational Reuse in Deep Neural Networks via Weight Repetition,” Proceedings of the 45th Annual International Symposium on Computer Architecture (ISCA '18), Jun. 2-6, 2018, 14… [cited by applicant]
Jain, Anil K., et al., “Artificial Neural Networks: A Tutorial,” Computer, Mar. 1996, 14 pages, vol. 29, Issue 3, IEEE. [cited by applicant]
Karan, Oguz, et al., “Diagnosing Diabetes using Neural Networks on Small Mobile Devices,” Expert Systems with Applications, Jan. 2012, 7 pages, vol. 39, Issue 1, Elsevier, Ltd. [cited by applicant]
Leng, Cong, et al., “Extremely Low Bit Neural Network: Squeeze the Last Bit Out with ADMM,” Proceedings of 32nd AAAI Conference on Artificial Intelligence (AAAI-18), Feb. 2-7, 2018, 16 pages, Association for the Advance… [cited by applicant]
Li, Fengfu, et al., “Ternary Weight Networks,” May 16, 2016, 9 pages, arXiv:1605.04711v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Li, Hong-Xing, et al., “Interpolation Functions of Feedforward Neural Networks,” Computers & Mathematics with Applications, Dec. 2003, 14 pages, vol. 46, Issue 12, Elsevier Ltd. [cited by applicant]
Louizos, Christos, et al., “Bayesian Compression for Deep Learning,” Proceedings of Advances in Neural Information Processing Systems 30 (NIPS 2017), Dec. 4-9, 2017, 17 pages, Neural Information Processing Systems Found… [cited by applicant]
Marcheret, Etienne, et al., “Detecting Audio-Visual Synchrony Using Deep Neural Networks,” Interspeech 2015, Sep. 6-10, 2015, 5 pages, ISCA, Dresden, Germany. [cited by applicant]
Marchesi, M., et al., “Multi-layer Perceptrons with Discrete Weights”, 1990 International Joint Conference on Neural Networks, Jun. 17-21, 1990, 8 pages, IEEE, San Diego, CA, USA. [cited by applicant]
Merolla, Paul, et al., “Deep Neural Networks are Robust to Weight Binarization and Other Non-linear Distortions,” Jun. 7, 2016, 10 pages, arXiv:1606.01981v1, Computing Research Repository (CoRR)—Cornell University, Itha… [cited by applicant]
Oskouei, Seyyed Salar Latifi, et al., “CNNdroid: GPU-Accelerated Execution of Trained Deep Convolutional Neural Networks on Android,” MM '16, Oct. 15-19, 2016, 5 pages, ACM, Amsterdam, Netherlands. [cited by applicant]
Özkan, Coskun, et al., “The Comparison of Activation Functions for Multispectral Landsat TM Image Classification,” Photogrammetric Engineering & Remote Sensing, Nov. 2003, 10 pages, vol. 69, No. 11, American Society for… [cited by applicant]
Rastegari, Mohammad, et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” Proceedings of 2016 European Conference on Computer Vision (ECCV '16), Oct. 8-16, 2016, 17 pages, Lecture Note… [cited by applicant]
Ravanbakhsh, Siamak, et al., “Stochastic Neural Networks with Monotonic Activation Functions,” Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, May 9-11, 2016, 10 pages, Cadiz,… [cited by applicant]
Rojas, Raul, “The Backpropagation Algorithm,” Neural Networks: A Systematic Introduction, Jul. 12, 1996, 36 pages, Springer-Verlag Berlin Heidelberg. [cited by applicant]
Shamsuddin, Siti Mariyam, et al., “Weight Changes for Learning Mechanisms in Two-Term Back Propagation Network,” Artificial Neural Networks—Architectures and Applications, Jan. 2013, 31 pages, InTech. [cited by applicant]
Stelmack, Marc A., et al., “Neural Network Approximation of Mixed Continuous/Discrete Systems in Multidisciplinary Design,” 36th Aerospace Sciences Meeting and Exhibit, Jan. 12-15, 1998, 16 pages, AIAA, Reno, NV, USA. [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Uffelman, Malcolm Rucj, “Information Theory, Learning Systems, and Self-Organizing Control Systems: Final Report,” Contract NONR-4467(00), Amendment 1, Mar. 31, 1967, 51 pages, Scope Incorporated, Falls Church, VA, USA. [cited by applicant]
Vaswani, Sharan, “Exploiting Sparsity in Supervised Learning,” Month Unknown 2014, 9 pages, retrieved from https://vaswanis.github.io > optimization_report. [cited by applicant]
Xiong, Chao, et al., “Conditional Neural Network for Modality-aware Face Recognition,” 2015 IEEE Conference on Computer Vision (ICCV), Dec. 11-18, 2015, 9 pages, IEEE, Los Condes, Chile. [cited by applicant]
Yan, Shi, “L1 Norm Regularization and Sparsity Explained for Dummies,” Aug. 27, 2016, 13 pages, retrieved from https://blog.mlreview.com/l1-norm-regularization-and-sparsity-explained-for-dummies-5b0e4be3938a. [cited by applicant]
Han, Song, “Efficient Methods and Hardware for Deep Learning,” Sep. 2017, 125 pages, Stanford University, Palo Alto, CA, USA. [cited by applicant]