IP Library Granted Patent US 12,367,661
Granted Patent B1
US 12,367,661 · App. 18/088,726 · Granted Jul 22, 2025

Weighted selection of inputs for training machine-trained network

Inventors: Steven L. Teig (Menlo Park, CA); Eric A. Sather (Palo Alto, CA); Andrew F. Siegel (Shoreline, WA); Evgeny Sorkin (Vancouver, CA)
Assignee: AMAZON TECHNOLOGIES, INC.
G06V10/774G06V10/776
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,661
App. No.
18/088,726
Granted
Jul 22, 2025
Kind
B1
Abstract

Some embodiments provide a method for training a machine-trained network that includes multiple parameters. The method propagates a batch of input training items through the network to generate output values and compute values of a loss function for each of the input training items. The method computes a weight for each input training item based on the computed loss function values for each of the input training items. The method selects input training items with larger weights more often than input training items with smaller weights for subsequent batches of input training items.

Claims (38)

1. A method for training a machine-trained network comprising a plurality of parameters, the method comprising:

propagating a batch of input training items through the network to generate output values and compute values of a loss function for each of the input training items;

computing a weight for each input training item based on the computed loss function values for each of the input training items; and

selecting input training items with larger weights more often than input training items with smaller weights for subsequent batches of input training items.

2. The method of claim 1 , wherein:

each input training item has a corresponding expected output value; and

computing a value of a loss function for a particular input training item comprises comparing the corresponding expected output value to the generated output value for the input training item.

3. The method of claim 2 , wherein the loss function values for the particular input training item increases as a distance between the corresponding expected output value and the generated output value for the input training item increases.

4. The method of claim 2 , wherein the loss function is a measure of unhappiness.

5. The method of claim 1 , wherein the weights for the input training items are proportional to the computed loss function values for the input training items.

6. The method of claim 1 , wherein:

the input training items are selected from a plurality of available input training items; and

a number of available input training items is larger than a number of input training items in each batch of input training items.

7. The method of claim 6 , wherein each input training item is selected at most once per batch of input training items.

8. The method of claim 6 , wherein each input training item is selected at least once in the subsequent batches of input training items.

9. The method of claim 1 , wherein:

the network is trained for classifying items into a predefined set of classes; and

the generated output value for a particular input training item comprises, for each class, a probability that the particular input training item belongs to the class.

10. The method of claim 1 , wherein selecting input training items with larger weights more often enables the parameters of the machine-trained network to converge more quickly to optimal values.

11. A non-transitory machine-readable medium storing a program which when executed by at least one processing unit trains a machine-trained network comprising a plurality of parameters, the program comprising sets of instructions for:

propagating a batch of input training items through the network to generate output values and compute values of a loss function for each of the input training items;

computing a weight for each input training item based on the computed loss function values for each of the input training items; and

selecting input training items with larger weights more often than input training items with smaller weights for subsequent batches of input training items.

12. The non-transitory machine-readable medium of claim 11 , wherein:

each input training item has a corresponding expected output value; and

the set of instructions for computing a value of a loss function for a particular input training item comprises a set of instructions for comparing the corresponding expected output value to the generated output value for the input training item.

13. The non-transitory machine-readable medium of claim 12 , wherein the loss function values for the particular input training item increases as a distance between the corresponding expected output value and the generated output value for the input training item increases.

14. The non-transitory machine-readable medium of claim 12 , wherein the loss function is a measure of unhappiness.

15. The non-transitory machine-readable medium of claim 11 , wherein the weights for the input training items are proportional to the computed loss function values for the input training items.

16. The non-transitory machine-readable medium of claim 11 , wherein:

the input training items are selected from a plurality of available input training items; and

a number of available input training items is larger than a number of input training items in each batch of input training items.

17. The non-transitory machine-readable medium of claim 16 , wherein each input training item is selected at most once per batch of input training items.

18. The non-transitory machine-readable medium of claim 16 , wherein each input training item is selected at least once in the subsequent batches of input training items.

19. The non-transitory machine-readable medium of claim 11 , wherein:

the network is trained for classifying items into a predefined set of classes; and

the generated output value for a particular input training item comprises, for each class, a probability that the particular input training item belongs to the class.

20. The non-transitory machine-readable medium of claim 11 , wherein the selection of input training items with larger weights more often enables the parameters of the machine-trained network to converge more quickly to optimal values.

Assignments (3)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2023
From: TEIG, STEVEN L.; SATHER, ERIC A.; SIEGEL, ANDREW F.; SORKIN, EVGENY
To: PERCEIVE CORPORATION
Reel/Frame 063899/0429 →
Continuity (1)
Provisional Application 63294571 · Dec 29, 2021
References Cited (122)
US 8934722B2 · Peleg · 2015 [cited by examiner]
US 10019654B1 · Pisoni · 2018 [cited by applicant]
US 10572979B2 · Vogels et al. · 2020 [cited by applicant]
US 10579907B1 · Kim · 2020 [cited by examiner]
US 10970441B1 · Zhang et al. · 2021 [cited by applicant]
US 10984560B1 · Appalaraju et al. · 2021 [cited by applicant]
US 11227047B1 · Vashisht et al. · 2022 [cited by applicant]
US 11270188B2 · Baker · 2022 [cited by applicant]
US 11531879B1 · Teig et al. · 2022 [cited by applicant]
US 11568179B2 · Posner · 2023 [cited by examiner]
US 11568367B2 · Hajian · 2023 [cited by examiner]
US 11610154B1 · Teig et al. · 2023 [cited by applicant]
US 11783007B2 · Edridge · 2023 [cited by examiner]
US 12045723B2 · Cho · 2024 [cited by examiner]
US 12061966B2 · Lapuschkin · 2024 [cited by examiner]
US 12165017B1 · Grossman · 2024 [cited by examiner]
US 20150340032A1 · Gruenstein · 2015 [cited by applicant]
US 20170061326A1 · Talathi et al. · 2017 [cited by applicant]
US 20170061625A1 · Estrada et al. · 2017 [cited by applicant]
US 20170091615A1 · Liu et al. · 2017 [cited by applicant]
US 20170206464A1 · Clayton et al. · 2017 [cited by applicant]
US 20180068221A1 · Brennan et al. · 2018 [cited by applicant]
US 20180101783A1 · Savkli · 2018 [cited by applicant]
US 20180114113A1 · Ghahramani et al. · 2018 [cited by applicant]
US 20180293713A1 · Vogels et al. · 2018 [cited by applicant]
US 20190005358A1 · Pisoni · 2019 [cited by applicant]
US 20190114544A1 · Sundaram et al. · 2019 [cited by applicant]
US 20190138882A1 · Choi et al. · 2019 [cited by applicant]
US 20190220741A1 · Garcia et al. · 2019 [cited by applicant]
US 20190258935A1 · Umeda · 2019 [cited by examiner]
US 20190286970A1 · Karaletsos et al. · 2019 [cited by applicant]
US 20190340492A1 · Burger et al. · 2019 [cited by applicant]
US 20200034751A1 · Kobayashi et al. · 2020 [cited by applicant]
US 20200051550A1 · Baker · 2020 [cited by applicant]
US 20200072610A1 · Hofmann et al. · 2020 [cited by applicant]
US 20200134461A1 · Chai et al. · 2020 [cited by applicant]
US 20200202213A1 · Rouhani et al. · 2020 [cited by applicant]
US 20200234144A1 · Such et al. · 2020 [cited by applicant]
US 20200285898A1 · Dong · 2020 [cited by examiner]
US 20200285939A1 · Baker · 2020 [cited by applicant]
US 20200311186A1 · Wang et al. · 2020 [cited by applicant]
US 20200311207A1 · Kim et al. · 2020 [cited by applicant]
US 20200334569A1 · Moghadam et al. · 2020 [cited by applicant]
US 20200401929A1 · Duerig · 2020 [cited by examiner]
US 20210056444A1 · Shimazu · 2021 [cited by examiner]
US 20210142170A1 · Ozcan et al. · 2021 [cited by applicant]
US 20210295205A1 · Ranco · 2021 [cited by examiner]
US 20210342642A1 · Shabtay · 2021 [cited by examiner]
US 20210342688A1 · Wang · 2021 [cited by examiner]
US 20220004921A1 · Balaraman · 2022 [cited by examiner]
US 20220027536A1 · Dutta · 2022 [cited by examiner]
US 20220092359A1 · Zhang · 2022 [cited by examiner]
US 20220129760A1 · Ravikumar · 2022 [cited by examiner]
US 20220138564A1 · Da Costa · 2022 [cited by examiner]
US 20220156577A1 · Jha · 2022 [cited by examiner]
US 20220208377A1 · Plummer · 2022 [cited by examiner]
US 20220237449A1 · Chung · 2022 [cited by examiner]
US 20220253647A1 · Perkins · 2022 [cited by examiner]
US 20220292805A1 · Feng · 2022 [cited by examiner]
US 20220301288A1 · Higuchi · 2022 [cited by examiner]
US 20220366040A1 · Marbouti · 2022 [cited by examiner]
US 20230017505A1 · Menon · 2023 [cited by examiner]
US 20230054130A1 · Wang · 2023 [cited by examiner]
US 20230085127A1 · Byun · 2023 [cited by examiner]
US 20230098994A1 · Labatie · 2023 [cited by examiner]
US 20230116417A1 · Taccari · 2023 [cited by examiner]
US 20230120256A1 · Terjek · 2023 [cited by examiner]
US 20230162005A1 · Cheng · 2023 [cited by examiner]
US 20230194640A1 · Christodoulou · 2023 [cited by examiner]
US 20230237333A1 · Shao · 2023 [cited by examiner]
US 20230385687A1 · Mahmood · 2023 [cited by examiner]
US 20230394807A1 · Sawada · 2023 [cited by examiner]
US 20240241286A1 · Cha · 2024 [cited by examiner]
US 20240370698A1 · Kosuge · 2024 [cited by examiner]
Abolfazli, Mojtaba, et al., “Differential Description Length for Hyperparameter Selection in Machine Learning,” May 22, 2019, 19 pages, arXiv:1902.04699, arXiv.org. [cited by applicant]
Achterhold, Jan, et al., “Variational Network Quantization,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 18 pages, ICLR, Vancouver, BC, Canada. [cited by applicant]
Agostinelli, Forest, et al., “Learning Activation Functions to Improve Deep Neural Networks,” Apr. 21, 2015, 9 pages, retrieved from https://arxiv.org/abs/1412.6830. [cited by applicant]
Author Unknown, “Renderman Support: Multi-Camera Rendering,” Feb. 2007, 8 pages, Pixar Animation Studios. [cited by applicant]
Author Unknown, “Renderman Support: REYES Options,” Month Unknown 2015, 42 pages, Pixar Animation Studios. [cited by applicant]
Author Unknown,“ 3.1. Cross-validation: evaluating estimator performance,” SciKit-learn 0.20.3 documentation, Apr. 16, 2019, 15 pages, retrieved from https://web.archive.org/web/20190416122230/https://scikit-learn.org/s… [cited by applicant]
Babaeizadeh, Mohammad, et al., “NoiseOut: A Simple Way to Prune Neural Networks,” Proceedings of the 29th Conference on Neural Information Processing Systems (NIPS 2016), Nov. 18, 2016, 5 pages, ACM, Barcelona, Spain. [cited by applicant]
Bagherinezhad, Hessam, et al., “LCNN: Look-up Based Convolutional Neural Network,” Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Jul. 21-26, 2017, 10 pages, IEEE, Honolulu, … [cited by applicant]
Castelli, Ilaria, et al., “Combination of Supervised and Unsupervised Learning for Training the Activation Functions of Neural Networks,” Pattern Recognition Letters, Jun. 26, 2013, 14 pages, vol. 37, Elsevier B.V. [cited by applicant]
Chandra, Pravin, et al., “An Activation Function Adapting Training Algorithm for Sigmoidal Feedforward Networks,” Neurocomputing, Jun. 25, 2004, 9 pages, vol. 61, Elsevier. [cited by applicant]
Cochrane, Courtney, “Time Series Nested Cross-Validation,” Towards Data Science, May 18, 2018, 6 pages, retrieved from https://towardsdatascience.com/time-series-nested-cross-validation-76adba623eb9. [cited by applicant]
Courbariaux, Matthieu, et al., “BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations,” Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS 15),… [cited by applicant]
Duda, Jarek, “Asymmetric Numeral Systems: Entropy Coding Combining Speed of Huffman Coding with Compression Rate of Arithmetic Coding,” Jan. 6, 2014, 24 pages, arXiv:1311.2540v2, Computer Research Repository (CoRR)—Corn… [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Falkner, Stefan, et al., “BOHB: Robust and Efficient Hyperparameter Optimization at Scale,” Proceedings of the 35th International Conference on Machine Learning, Jul. 10-15, 2018, 19 pages, Stockholm, Sweden. [cited by applicant]
Hansen, Katja, et al., “Assessment and Validation of Machine Learning Methods for Predicting Molecular Atomization Energies,” Journal of Chemical Theory and Computation, Jul. 11, 2013, 16 pages, vol. 9, American Chemica… [cited by applicant]
Harmon, Mark, et al., “Activation Ensembles for Deep Neural Networks,” Feb. 24, 2017, 9 pages, arXiv:1702.07790v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
He, Kaiming, et al., “Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification,” Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Dec. 7-13, 2015, 9 pag… [cited by applicant]
Huynh, Thuan Q., et al., “Effective Neural Network Pruning Using Cross-Validation,” Proceedings of International Joint Conference on Natural Networks, Jul. 31-Aug. 4, 2005, 6 pages, IEEE, Montreal, Canada. [cited by applicant]
Jain, Anil K., et al., “Artificial Neural Networks: A Tutorial,” Computer, Mar. 1996, 14 pages, vol. 29, Issue 3, IEEE. [cited by applicant]
Kingma, Diederik P., et al., “Auto-Encoding Variational Bayes,” May 1, 2014, 14 pages, arXiv:1312.6114v10, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Kingma, Diederik P., et al., “Variational Dropout and the Local Reparameterization Trick,” Proceedings of the 28th International Conference on Neural Information Processing Systems (NIPS '15), Dec. 7-12, 2015, 14 pages,… [cited by applicant]
Kumar, Ashish, et al., “Resource-efficient Machine Learning in 2 KB RAM for the Internet of Things,” Proceedings of the 34th International Conference on Machine Learning, Aug. 6-11, 2017, 10 pages, vol. 70, PMLR, Sydney… [cited by applicant]
Leyton-Brown, Kevin, et al., “Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algorithms,” Aug. 9, 2012, 12 pages, retrieved from https://arxiv.org/abs/1208.3719. [cited by applicant]
Li, Hong-Xing, et al., “Interpolation Functions of Feedforward Neural Networks,” Computers & Mathematics with Applications, Dec. 2003, 14 pages, vol. 46, Issue 12, Elsevier Ltd. [cited by applicant]
Louizos, Christos, et al., “Bayesian Compression for Deep Learning,” Proceedings of Advances in Neural Information Processing Systems 30 (NIPS 2017), Dec. 4-9, 2017, 17 pages, Neural Information Processing Systems Found… [cited by applicant]
M., Sanjay, “Why and how to Cross Validate a Model?,” Towards Data Science, Nov. 12, 2018, 4 pages, retrieved from https://towardsdatascience.com/why-and-how-to-cross-validate-a-model-d6424b45261f. [cited by applicant]
Mackay, Matthew, et al., “Self-Tuning Networks: Bilevel Optimization of Hyperparameters Using Structured Best-Response Functions,” Proceedings of Seventh International Conference on Learning Representations (ICLR '19), … [cited by applicant]
Maclaurin, Dougal, et al., “Gradient-based Hyperparameter Optimization through Reversible Learning,” Proceedings of the 32nd International Conference on Machine Learning, Jul. 7-19, 2015, 10 pages, JMLR, Lille, France. [cited by applicant]
Mitchell, Tom, “Artificial Neural Networks,” Machine Learning—Chapter 4, Month Unknown 1997, 47 pages, McGraw Hill. [cited by applicant]
Molchanov, Dmitry, et al., “Variational Dropout Sparsifies Deep Neural Networks,” Feb. 27, 2017, 10 pages, arXiv:1701.05369v2, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Nair, Vinod, et al., “Rectified Linear Units Improve Restricted Boltzmann Machines,” Proceedings of the 27th International Conference on Machine Learning, Jun. 21-24, 2010, 8 pages, Omnipress, Haifa, Israel. [cited by applicant]
Neklyudov, Kirill, et al., “Structured Bayesian Pruning via Log-Normal Multiplicative Noise, ” Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Dec. 4-9, 2017, 10 pages, ACM, Long… [cited by applicant]
Non-Published Commonly Owned Related U.S. Appl. No. 18/088,725 with similar specification, filed Dec. 26, 2022, 61 pages, Perceive Corporation. [cited by applicant]
Non-Published Commonly Owned Related U.S. Appl. No. 18/088,727 with similar specification, filed Dec. 26, 2022, 61 pages, Perceive Corporation. [cited by applicant]
Park, Jongsoo, et al., “Faster CNNs with Direct Sparse Convolutions and Guided Pruning,” Jul. 28, 2017, 12 pages, arXiv:1608.01409v5, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Russell, Stuart, et al., “Artificial Intelligence: A Modern Approach,” Second Edition, Month Unknown 2003, 145 pages, Prentice Hall, Inc. [cited by applicant]
Shayer, Oran, et al., “Learning Discrete Weights Using the Local Reparameterization Trick,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 12 pages, ICLR, Vanco… [cited by applicant]
Srivastava, Nitish, et al., “Dropout: A Simple Way to Prevent Neural Networks from Overfitting,” Journal of Machine Learning Research, Jun. 2014, 30 pages, vol. 15, JMLR.org. [cited by applicant]
Srivastava, Rupesh Kumar, et al., “Highway Networks,” Nov. 3, 2015, 6 pages, arXiv:1505.00387v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Wen, Wei, et al., “Learning Structured Sparsity in Deep Neural Networks,” Oct. 18, 2016, 10 pages, arXiv:1608.03665v4, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Yang, Tien-Ju, et al., “Designing Energy-Efficient Convolutional Neural Networks using Energy-Aware Pruning,” Apr. 18, 2017, 9 pages, arXiv:1611.05128v4, Computer Research Repository (CoRR)—Cornell University, Ithaca, N… [cited by applicant]
Zhang, Dongqing, et al., “LQ-Nets: Learned Quantization for Highly Accurate and Compact Deep Neural Networks,” Jul. 26, 2018, 21 pages, arXiv:1807.10029v1, Computer Research Repository (CoRR)—Cornell University, Ithaca,… [cited by applicant]
Zhang, Hongbo, “Artificial Neuron Network Hyperparameter Tuning by Evolutionary Algorithm and Pruning Technique,” Month Unknown 2018, 8 pages. [cited by applicant]
Zhen, Hui-Ling, et al., “Nonlinear Collaborative Scheme for Deep Neural Networks,” Nov. 4, 2018, 11 pages, retrieved from https://arxiv.org/abs/1811.01316. [cited by applicant]
Zhu, Chenzhuo, et al., “Trained Ternary Quantization,” Dec. 4, 2016, 9 pages, arXiv:1612.01064v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Zilly, Julian Georg, et al., “Recurrent Highway Networks,” Jul. 4, 2017, 12 pages, arXiv:1607.03474v5, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]