IP Library Granted Patent US 12,387,092
Granted Patent B1
US 12,387,092 · App. 18/115,553 · Granted Aug 12, 2025

Neural network loss function that incorporates incorrect category probabilities

Inventors: Philip Sharos (Limerick, IE); Steven L. Teig (Menlo Park, CA)
Assignee: Amazon Technologies, Inc.
G06N3/063G06V10/774
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,092
App. No.
18/115,553
Granted
Aug 12, 2025
Kind
B1
Abstract

Some embodiments provide a method for training a machine-trained (MT) network to classify inputs into multiple categories. The method propagates a set of input training items through the MT network to generate a set of corresponding outputs. Each input training item belongs to a category and the output for each input training item includes, for each category, a computed probability of the input belonging to the category. The method computes a value for a loss function based on the generated outputs. The loss function includes a first term based on the computed probabilities of each input belonging to its category and not based on individual computed probabilities of the inputs belonging to other categories and a second term based on the individual computed probabilities of each input belonging to each of the categories. The method uses the computed value for the loss function to train the MT network.

Claims (30)

1. A method for training a machine-trained (MT) network to classify inputs into a plurality of categories, the method comprising:

propagating a set of input training items through the MT network to generate a set of corresponding outputs, wherein (i) each respective input training item belongs to a respective category and (ii) the respective output for each respective input training item comprises, for each category, a computed probability of the input belonging to the category;

computing a value for a loss function based on the generated outputs, the loss function comprising (i) a first term based on the computed probabilities of each respective input belonging to its respective category and not based on individual computed probabilities of the inputs belonging to other categories and (ii) a second term based on the individual computed probabilities of each respective input belonging to each of the categories; and

using the computed value for the loss function to train the MT network.

2. The method of claim 1 , wherein the first term is a cross-entropy term that, for each respective input, is based on a logarithm of the computed probability of the input belonging to its respective category.

3. The method of claim 2 , wherein the loss function further comprises a complementary cross-entropy term that, for each respective input, is based on a logarithm of the computed probability of the input not belonging to its respective category.

4. The method of claim 1 , wherein each of the loss function terms is continuously differentiable.

5. The method of claim 1 , wherein the second term approximates, for each respective input, a maximum of a function applied to each of the individual computed probabilities of the respective input belonging to categories other than its respective category, wherein the function increases as probability increases.

6. The method of claim 1 , wherein the second term approximates, for each respective input, a minimum of a function applied to each of the individual computed probabilities of the respective input belonging to categories other than its respective category, wherein the function increases as probability increases.

7. The method of claim 1 , wherein the second term is a log-sum-exponent term that emphasizes, for each respective input, a largest of the individual computed probabilities of the respective input belonging to categories other than its respective category.

8. The method of claim 1 , wherein the second term is a log-sum-exponent term that emphasizes, for each respective input, a smallest of the individual computed probabilities of the respective input belonging to categories other than its respective category.

9. The method of claim 1 , wherein the second term is a regularizing term that, for each respective input, is smallest when the individual computed probabilities of the respective input belonging to categories other than its respective category are evenly distributed.

10. The method of claim 1 , wherein the input training items are images and the plurality of different categories are types of objects depicted in the images.

11. The method of claim 1 , wherein the input training items are audio recordings and the plurality of different categories are different acoustic scenes in which the audio recordings are captured.

12. The method of claim 1 , wherein the loss function for a particular input is minimized for a given computed probability of the input belonging to its category when the computed probabilities of the input belonging to categories other than its respective category are evenly distributed.

13. The method of claim 1 , wherein using the computed value for the loss function to train the MT network comprises:

backpropagating the computed loss function through the MT network to compute gradients of the loss function with respect to each of a plurality of parameters of the network; and

using the computed gradients to modify the plurality of parameters of the network.

14. A non-transitory machine-readable medium storing a program which when executed by at least one processor trains a machine-trained (MT) network to classify inputs into a plurality of categories, the program comprising sets of instructions for:

propagating a set of input training items through the MT network to generate a set of corresponding outputs, wherein (i) each respective input training item belongs to a respective category and (ii) the respective output for each respective input training item comprises, for each category, a computed probability of the input belonging to the category;

computing a value for a loss function based on the generated outputs, the loss function comprising (i) a first term based on the computed probabilities of each respective input belonging to its respective category and not based on individual computed probabilities of the inputs belonging to other categories and (ii) a second term based on the individual computed probabilities of each respective input belonging to each of the categories; and

using the computed value for the loss function to train the MT network.

15. The non-transitory machine-readable medium of claim 14 , wherein:

the first term is a cross-entropy term that, for each respective input, is based on a logarithm of the computed probability of the input belonging to its respective category; and

the loss function further comprises a complementary cross-entropy term that, for each respective input, is based on a logarithm of the computed probability of the input not belonging to its respective category.

16. The non-transitory machine-readable medium of claim 14 , wherein the second term approximates, for each respective input, a maximum of a function applied to each of the individual computed probabilities of the respective input belonging to categories other than its respective category, wherein the function increases as probability increases.

17. The non-transitory machine-readable medium of claim 14 , wherein the second term approximates, for each respective input, a minimum of a function applied to each of the individual computed probabilities of the respective input belonging to categories other than its respective category, wherein the function increases as probability increases.

18. The non-transitory machine-readable medium of claim 14 , wherein the second term is a log-sum-exponent term that emphasizes, for each respective input, a largest of the individual computed probabilities of the respective input belonging to categories other than its respective category.

19. The non-transitory machine-readable medium of claim 14 , wherein the second term is a log-sum-exponent term that emphasizes, for each respective input, a smallest of the individual computed probabilities of the respective input belonging to categories other than its respective category.

20. The non-transitory machine-readable medium of claim 14 , wherein the second term is a regularizing term that, for each respective input, is smallest when the individual computed probabilities of the respective input belonging to categories other than its respective category are evenly distributed.

Assignments (3)
BILL OF SALE Recorded Oct 31, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069288/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069288/0731 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2023
From: TEIG, STEVEN L.; SHAROS, PHILIP
To: PERCEIVE CORPORATION
Reel/Frame 063494/0248 →
Continuity (1)
Provisional Application 63388044 · Jul 11, 2022
References Cited (67)
US 4866634A · Reboh et al. · 1989 [cited by applicant]
US 5255347A · Matsuba et al. · 1993 [cited by applicant]
US 5579436A · Chou et al. · 1996 [cited by applicant]
US 6601052B1 · Lee et al. · 2003 [cited by applicant]
US 6985781B2 · Keeler et al. · 2006 [cited by applicant]
US 7333923B1 · Yamanishi et al. · 2008 [cited by applicant]
US 8000538B2 · Sarkar · 2011 [cited by applicant]
US 8145662B2 · Chen et al. · 2012 [cited by applicant]
US 9373087B2 · Nowozin · 2016 [cited by applicant]
US 10019654B1 · Pisoni · 2018 [cited by applicant]
US 10586151B1 · Teig · 2020 [cited by applicant]
US 10671888B1 · Sather · 2020 [cited by examiner]
US 11475310B1 · Teig · 2022 [cited by examiner]
US 20030033263A1 · Cleary · 2003 [cited by applicant]
US 20040243954A1 · Devgan et al. · 2004 [cited by applicant]
US 20060010093A1 · Fan et al. · 2006 [cited by applicant]
US 20070078630A1 · Filatov et al. · 2007 [cited by applicant]
US 20070239642A1 · Sindhwani et al. · 2007 [cited by applicant]
US 20070271287A1 · Acharya et al. · 2007 [cited by applicant]
US 20080256011A1 · Rice · 2008 [cited by applicant]
US 20090106173A1 · Andrew et al. · 2009 [cited by applicant]
US 20100082639A1 · Li et al. · 2010 [cited by applicant]
US 20110105346A1 · Beattie et al. · 2011 [cited by applicant]
US 20110120718A1 · Craig · 2011 [cited by applicant]
US 20110182345A1 · Lei et al. · 2011 [cited by applicant]
US 20120271791A1 · Laan et al. · 2012 [cited by applicant]
US 20120283992A1 · Ji et al. · 2012 [cited by applicant]
US 20120331025A1 · Gemulla et al. · 2012 [cited by applicant]
US 20130257873A1 · Isozaki · 2013 [cited by applicant]
US 20140124265A1 · Al-Yami et al. · 2014 [cited by applicant]
US 20140207837A1 · Taniguchi et al. · 2014 [cited by applicant]
US 20140222747A1 · Zhou et al. · 2014 [cited by applicant]
US 20150100530A1 · Winih et al. · 2015 [cited by applicant]
US 20150161995A1 · Sainath et al. · 2015 [cited by applicant]
US 20150262083A1 · Xu et al. · 2015 [cited by applicant]
US 20160132786A1 · Balan et al. · 2016 [cited by applicant]
US 20160260222A1 · Paglieroni et al. · 2016 [cited by applicant]
US 20170061326A1 · Talathi et al. · 2017 [cited by applicant]
US 20170068844A1 · Friedland · 2017 [cited by applicant]
US 20170091615A1 · Liu et al. · 2017 [cited by applicant]
US 20170140298A1 · Wabnig et al. · 2017 [cited by applicant]
US 20170154425A1 · Pierce et al. · 2017 [cited by applicant]
US 20170161640A1 · Shamir · 2017 [cited by applicant]
US 20170183836A1 · Ahmed et al. · 2017 [cited by applicant]
US 20170262735A1 · Sanchez et al. · 2017 [cited by applicant]
US 20180033024A1 · Latapie et al. · 2018 [cited by applicant]
US 20180068221A1 · Brennan et al. · 2018 [cited by applicant]
US 20180101783A1 · Savkli · 2018 [cited by applicant]
US 20180174041A1 · Imam et al. · 2018 [cited by applicant]
US 20180219888A1 · Apostolopoulos · 2018 [cited by applicant]
US 20190005358A1 · Pisoni · 2019 [cited by applicant]
US 20190230913A1 · Chen et al. · 2019 [cited by applicant]
US 20190286970A1 · Karaletsos et al. · 2019 [cited by applicant]
US 20200007934A1 · Ortiz et al. · 2020 [cited by applicant]
US 20200249996A1 · Addepalli et al. · 2020 [cited by applicant]
US 20230040889A1 · Teig et al. · 2023 [cited by applicant]
Chen, Minghua, et al., “Markov Approximation for Combinatorial Network Optimization,” IEEE Transactions on Information Theory, Oct. 2013, 27 pages, vol. 59, No. 10, IEEE. [cited by applicant]
Duda, Jarek, “Asymmetric Numeral Systems: Entropy Coding Combining Speed of Huffman Coding with Compression Rate of Arithmetic Coding,” Jan. 6, 2014, 24 pages, arXiv:1311.2540v2, Computer Research Repository (CoRR)—Corn… [cited by applicant]
Eisele, Robert, “The log-sum-exp trick in Machine Learning,” Computer Science & Machine Learning, Jun. 22, 2016, 3 pages. [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
Huang, Gao, et al., “Multi-Scale Dense Networks for Resource Efficient Image Classification,” Proceedings of the 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 14 pages, ICLR,… [cited by applicant]
Jain, Anil K., et al., “Artificial Neural Networks: A Tutorial,” Computer, Mar. 1996, 14 pages, vol. 29, Issue 3, IEEE. [cited by applicant]
Li, Hong-Xing, et al., “Interpolation Functions of Feedforward Neural Networks,” Computers & Mathematics with Applications, Dec. 2003, 14 pages, vol. 46, Issue 12, Elsevier Ltd. [cited by applicant]
Mandelbaum, Amit et al., “Distance-based Confidence Score for Neural Network Classifiers,” Sep. 28, 2017, 10 pages, arXiv:1709.09844v1, Computer Research Repository (CoRR) , Cornell University, Ithaca, NY, USA. [cited by applicant]
Shalev-Shwartz, Shai, et al., “Minimizing the Maximal Loss: How and Why,” Proceedings of the 33rd International Conference on Machine Learning, Jun. 19-24, 2016, 13 pages, JMLR, New York, NY, USA. [cited by applicant]
Srivastava, Nitish, et al., “Dropout: A Simple Way to Prevent Neural Networks from Overfitting,” Journal of Machine Learning Research, Jun. 2014, 30 pages, vol. 15, JMLR.org. [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]