IP Library Granted Patent US 12,248,880
Granted Patent B2
US 12,248,880 · App. 18/238,507 · Granted Mar 11, 2025

Using batches of training items for training a network

Inventors: Eric A. Sather (Palo Alto, CA); Steven L. Teig (Menlo Park, CA); Andrew C. Mihal (San Jose, CA)
Assignee: Amazon Technologies, Inc.
G06N3/084G06F18/214G06F18/217G06N3/04G06N3/08G06T7/97G06V10/454G06V10/764G06V10/82G06V40/167G06V40/172G06T2207/20081G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,880
App. No.
18/238,507
Granted
Mar 11, 2025
Kind
B2
Abstract

Some embodiments provide a method for training a machine-trained (MT) network that processes inputs using network parameters. The method propagates a set of input training items through the MT network to generate a set of output values. The set of input training items comprises multiple training items for each of multiple categories. The method identifies multiple training item groupings in the set of input training items. Each grouping includes at least two training items in a first category and at least one training item in a second category. The method calculates a value of a loss function as a summation of individual loss functions for each of the identified training item groupings. The individual loss function for each particular training item grouping is based on the output values for the training items of the grouping. The method trains the network parameters using the calculated loss function value.

Claims (54)

1. A method for training a machine-trained (MT) network that classifies inputs into a plurality of categories, the method comprising:

propagating a plurality of input training items through the MT network to generate a respective output value for each respective input training item, the plurality of input training items comprising input training items for each of the categories;

identifying a plurality of triplets from the plurality of input training items, wherein each respective triplet comprises two input training items in a respective first category and one input training item in a respective second category, wherein a plurality of the input training items belong to at least two different triplets;

calculating a value of a loss function as a summation of respective individual loss functions for each of the respective identified triplets, the respective individual loss function for each respective triplet based on the output values generated for the input training items of the triplet; and

training the MT network using the calculated loss function value.

2. The method of claim 1 , wherein the input items are images and the categories comprise different types of objects found in the images.

3. The method of claim 1 , wherein:

the input items are images and the categories comprise a plurality of different people, wherein each input training item for a particular person comprises a different image of the particular person's face; and

the machine-trained network is for embedding into a device to perform facial recognition.

4. The method of claim 1 , wherein:

the set of input images comprises N i images for each of N p categories; and

a number of identified triplets is of the order (N i ){circumflex over ( )}3*(N p ){circumflex over ( )}2.

5. The method of claim 1 , wherein:

each output value is a vector; and

the respective individual loss function for each respective triplet is a function of the proximity of the vectors for one of the input training items in the first respective category and the input training item in the respective second category to the vector for the other input training item in the first respective category.

6. The method of claim 1 , wherein:

each output value is a vector; and

the respective individual loss function for each respective triplet is a function of the probability of a misclassification of one of the input training items in the respective first category based on the vector of the three input training items of the triplet.

7. The method of claim 1 , wherein:

the MT network comprises input nodes, output nodes, and interior nodes between the input nodes and output nodes;

each node produces a node output value and each interior node and output node receives as node input values a set of node output values of other nodes and applies weights to each received input value; and

training the MT network comprises training the weights.

8. The method of claim 1 , wherein the propagating, identifying, calculating, and training are performed iteratively.

9. The method of claim 1 , wherein training the MT network comprises:

backpropagating the calculated loss function value through the MT network to determine, for each of a set of network parameters, a rate of change in the calculated loss function value relative to a rate of change in the network parameter; and

modifying each network parameter in the set according to the determined rate of change for the network parameter.

10. The method of claim 1 , wherein identifying the plurality of triplets comprises identifying each possible triplet formed by the plurality of input training items.

11. A non-transitory machine-readable medium storing a program which when executed by at least one processing unit trains a machine-trained (MT) network that classifies inputs into a plurality of categories, the program comprising sets of instructions for:

propagating a plurality of input training items through the MT network to generate a respective output value for each respective input training item, the plurality of input training items comprising input training items for each of the categories;

identifying a plurality of triplets from the plurality of input training items, wherein each respective triplet comprises two input training items in a respective first category and one input training item in a respective second category, wherein a plurality of the input training items belong to at least two different triplets;

calculating a value of a loss function as a summation of respective individual loss functions for each of the respective identified triplets, the respective individual loss function for each respective triplet based on the output values generated for the input training items of the triplet; and

training the MT network using the calculated loss function value.

12. The non-transitory machine-readable medium of claim 11 , wherein the input items are images and the categories comprise different types of objects found in the images.

13. The non-transitory machine-readable medium of claim 11 , wherein:

the input items are images and the categories comprise a plurality of different people, wherein each input training item for a particular person comprises a different image of the particular person's face; and

the machine-trained network is for embedding into a device to perform facial recognition.

14. The non-transitory machine-readable medium of claim 11 , wherein:

the set of input images comprises N i images for each of N p categories; and

a number of identified triplets is of the order (N i ){circumflex over ( )}3*(N p ){circumflex over ( )}3.

15. The non-transitory machine-readable medium of claim 11 , wherein:

each output value is a vector; and

the respective individual loss function for each respective triplet is a function of the proximity of the vectors for one of the input training items in the first respective category and the input training item in the respective second category to the vector for the other input training item in the first respective category.

16. The non-transitory machine-readable medium of claim 11 , wherein:

each output value is a vector; and

the respective individual loss function for each respective triplet is a function of the probability of a misclassification of one of the input training items in the respective first category based on the vector of the three input training items of the triplet.

17. The non-transitory machine-readable medium of claim 11 , wherein:

the MT network comprises input nodes, output nodes, and interior nodes between the input nodes and output nodes;

each node produces a node output value and each interior node and output node receives as node input values a set of node output values of other nodes and applies weights to each received input value; and

the set of instructions for training the MT network comprises a set of instructions for training the weights.

18. The non-transitory machine-readable medium of claim 11 , wherein the propagating, identifying, calculating, and training are performed iteratively.

19. The non-transitory machine-readable medium of claim 11 , wherein the set of instructions for training the MT network comprises sets of instructions for:

backpropagating the calculated loss function value through the MT network to determine, for each of a set of network parameters, a rate of change in the calculated loss function value relative to a rate of change in the network parameter; and

modifying each network parameter in the set according to the determined rate of change for the network parameter.

20. The non-transitory machine-readable medium of claim 11 , wherein the set of instructions for identifying the plurality of triplets comprises a set of instructions for identifying each possible triplet formed by the plurality of input training items.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2024
From: PERCEIVE CORPORATION
To: AMAZON.COM SERVICES LLC
Reel/Frame 069026/0067 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2024
From: AMAZON.COM SERVICES LLC
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 069026/0252 →
Continuity (5)
Continuation 17514701 · Oct 29, 2021
Continuation 16852329 · Apr 17, 2020
Continuation 15901456 · Feb 21, 2018
Provisional Application 62599013 · Dec 14, 2017
Related Publication 20230409918A1 · Dec 21, 2023
References Cited (64)
US 5255347A · Matsuba et al. · 1993 [cited by applicant]
US 5461698A · Schwanke et al. · 1995 [cited by applicant]
US 8429106B2 · Downs · 2013 [cited by examiner]
US 9928448B1 · Merler · 2018 [cited by examiner]
US 9928449B2 · Chavez et al. · 2018 [cited by applicant]
US 10019654B1 · Pisoni · 2018 [cited by applicant]
US 10592732B1 · Sather et al. · 2020 [cited by applicant]
US 10671888B1 · Sather et al. · 2020 [cited by applicant]
US 11163986B2 · Sather et al. · 2021 [cited by applicant]
US 11741369B2 · Sather et al. · 2023 [cited by applicant]
US 20030033263A1 · Cleary · 2003 [cited by applicant]
US 20110282897A1 · Li et al. · 2011 [cited by applicant]
US 20140079297A1 · Tadayon et al. · 2014 [cited by applicant]
US 20160132786A1 · Balan et al. · 2016 [cited by applicant]
US 20160379352A1 · Zhang et al. · 2016 [cited by applicant]
US 20170124385A1 · Ganong et al. · 2017 [cited by applicant]
US 20170161640A1 · Shamir · 2017 [cited by applicant]
US 20170278289A1 · Marino et al. · 2017 [cited by applicant]
US 20170357896A1 · Tsatsin et al. · 2017 [cited by applicant]
US 20180075849A1 · Khoury et al. · 2018 [cited by applicant]
US 20180165554A1 · Zhang et al. · 2018 [cited by applicant]
US 20180232566A1 · Griffin et al. · 2018 [cited by applicant]
US 20190005358A1 · Pisoni · 2019 [cited by applicant]
US 20190065957A1 · Movshovitz-Attias et al. · 2019 [cited by applicant]
US 20190130231A1 · Liu et al. · 2019 [cited by applicant]
US 20190138896A1 · Deng · 2019 [cited by applicant]
US 20190180176A1 · Yudanov · 2019 [cited by examiner]
US 20190258925A1 · Li · 2019 [cited by examiner]
US 20190279046A1 · Han · 2019 [cited by examiner]
US 20190362233A1 · Aizawa et al. · 2019 [cited by applicant]
US 20200250476A1 · Sather et al. · 2020 [cited by applicant]
US 20220051002A1 · Sather et al. · 2022 [cited by applicant]
WO 2016118402A1 · 2016 [cited by applicant]
Bhavsar, Hetal, et al., “Support Vector Machine Classification using Mahalanobis Distance Function,” International Journal of Scientific & Engineering Research, Jan. 2015, 12 pages, vol. 6, No. 1, IJSER. [cited by applicant]
Chung, Eric, et al., “Accelerating Persistent Neural Networks at Datacenter Scale,” HC29: Hot Chips: A Symposium on High Performance Chips 2017, Aug. 20-22, 2017, 52 pages, IEEE, Cupertino, CA, USA. [cited by applicant]
Emer, Joel, et al., “Hardware Architectures for Deep Neural Networks,” CICS/MTL Tutorial, Mar. 27, 2017, 258 pages, Massachusetts Institute of Technology, Cambridge, MA, USA, retrieved from http://www.rle.mit.edu/eems/w… [cited by applicant]
He, Qin, “Neural Network and Its Application in IR,” Month Unknown 1999, 31 pages. [cited by applicant]
Hermans, Alexander, et al., “In Defense of the Triplet Loss for Person Re-Identification,” Nov. 21, 2017, 17 pages, retrieved from https://arxiv.org/abs/1703.07737. [cited by applicant]
Huang, Gao, et al., “Multi-Scale Dense Networks for Resource Efficient Image Classification,” Proceedings of the 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 14 pages, ICLR,… [cited by applicant]
Jain, Anil K., et al., “Artificial Neural Networks: A Tutorial,” Computer, Mar. 1996, 14 pages, vol. 29, Issue 3, IEEE. [cited by applicant]
Jouppi, Norman, P., et al., “In-Datacenter Performance Analysis of a Tensor Processing Unit,” Proceedings of the 44th Annual International Symposium on Computer Architecture (ISCA '17), Jun. 24-28, 2017, 17 pages, ACM, … [cited by applicant]
Karresand, Martin, et al., “File Type Identification of Data Fragments by Their Binary Structure,” Proceedings of the 2006 IEEE Workshop on Information Assurance, Jun. 21-23, 2006, 8 pages, IEEE, West Point, NY, USA. [cited by applicant]
Keller, Michel, et al., “Learning Deep Descriptors with Scale-Aware Triplet Networks,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18-23, 2018, 9 pages, IEEE, Salt Lake City, UT, USA. [cited by applicant]
Kohrs, Arnd, et al., “Improving Collaborative Filtering with Multimedia Indexing Techniques to Create User-Adapting Web Sites,” Proceedings of the Seventh ACM International Conference on Multimedia (Part 1), Oct. 1999, … [cited by applicant]
Li, Hong-Xing, et al., “Interpolation Functions of Feedforward Neural Networks,” Computers & Mathematics with Applications, Dec. 2003, 14 pages, vol. 46, Issue 12, Elsevier Ltd. [cited by applicant]
Liu, Yishu, et al., “Scene Classification via Triplet Networks,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, Jan. 2018, 18 pages, vol. 11, No. 1, IEEE. [cited by applicant]
Mandelbaum, Amit et al., “Distance-based Confidence Score for Neural Network Classifiers,” Sep. 28, 2017, 10 pages, arXiv:1709.09844v1, Computer Research Repository (CoRR) , Cornell University, Ithaca, NY, USA. [cited by applicant]
Nguyen, Bac, et al., “Supervised Distance Metric Learning Through Maximization of the Jeffrey Divergence,” Pattern Recognition, Nov. 16, 2016, 11 pages, vol. 64, Elsevier Ltd. [cited by applicant]
Qian, Qi, et al., Efficient Distance Metric Learning by Adaptive Sampling and Mini-Batch Stochastic Gradient Descent (SGD), Machine Learning, Month Unknown 2015, 20 pages, Springer. [cited by applicant]
Rastegari, Mohammad, et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” Proceedings of 2016 European Conference on Computer Vision (ECCV '16), Oct. 8-16, 2016, 17 pages, Lecture Note… [cited by applicant]
Schroff, Florian, et al., “FaceNet: A Unified Embedding for Face Recognition and Clustering,” Proceedings of 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2015), Jun. 7-12, 2015, 9 pages, IEEE, B… [cited by applicant]
Shayer, Oran, et al., “Learning Discrete Weights Using the Local Reparameterization Trick,” Proceedings of 6th International Conference on Learning Representations (ICLR 2018), Apr. 30-May 3, 2018, 12 pages, ICLR, Vanco… [cited by applicant]
Sim, Jaehyeong, et al., “A 1.42TOPS/W Deep Convolutional Neural Network Recognition Processor for Intelligent IoE Systems,” Proceedings of 2016 IEEE International Solid-State Circuits Conference (ISSCC 2016), Jan. 31-Fe… [cited by applicant]
Sze, Vivienne, et al., “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Aug. 13, 2017, 32 pages, arXiv:1703.09039v2, Computer Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]
Urtasun, R., “Lecture 9: Support Vector Machines,” CSC2515 Fall 2015 Introduction to Machine Learning, Month Unknown 2015, 25 pages. [cited by applicant]
Wang, Zilei, et al., “Linear Distance Coding for Image Classification,” IEEE Transactions on Image Processing, Feb. 2013, 12 pages, vol. 22, No. 2, IEEE. [cited by applicant]
Worfolk, Patrick, et al., “ULP DNN Processing Unit Market Opportunity,” Month Unknown 2017, 40 pages, Synaptics Incorporated, San Jose, CA, USA. [cited by applicant]
Xiang, Shiming, et al., “Learning a Mahalanobis distance metric for data clustering and classification,” Pattern Recognition, Dec. 1, 2008, 13 pages, vol. 41, No. 12, Elsevier Ltd. [cited by applicant]
Yang, Liu, “Distance Metric Learning: A Comprehensive Survey,” May 19, 2006, 51 pages. [cited by applicant]
Zhang, Dongqing, et al., “LQ-Nets: Learned Quantization for Highly Accurate and Compact Deep Neural Networks,” Jul. 26, 2018, 21 pages, arXiv: 1807.10029v1, Computer Research Repository (CoRR)—Cornell University, Ithaca… [cited by applicant]
Zhao, Bin, et al., “Maximum Margin Clustering with Multivariate Loss Function,” 2009 Ninth IEEE International Conference on Data Mining, Dec. 6-9, 2009, 10 pages, IEEE, Miami Beach, FL. [cited by applicant]
Zheng, Wei-Shi, et al., “Person Re-identification by Probabilistic Relative Distance Comparison,” 2011 Proceedings of CVPR—IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Jun. 2011, 8 pages,… [cited by applicant]
Zhou, Shuchang, et al., “DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients,” Jul. 17, 2016, 14 pages, arXiv:1606.06160v2, Computer Research Repository (CoRR)—Cornell University,… [cited by applicant]
Zhu, Chenzhuo, et al., “Trained Ternary Quantization,” Dec. 4, 2016, 9 pages, arXiv: 1612.01064v1, Computing Research Repository (CoRR)—Cornell University, Ithaca, NY, USA. [cited by applicant]