IP Library › Granted Patent US 12,530,576
Granted Patent B2
US 12,530,576 · App. 17/375,960 · Granted Jan 20, 2026

Accounting for long-tail training data through logit adjustment

Inventors: Aditya Krishna Menon (New York, NY); Sanjiv Kumar (Jericho, NY); Himanshu Jain (Jersey City, NJ); Andreas Veit (New York, NY); Ankit Singh Rawat (New York, NY); Gayan Sadeep Jayasumana Hirimbura Matara Kankanamge (Houston, TX)
Assignee: Google LLC
G06N3/08G06F18/2113G06F18/2415G06F18/2431
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,576
App. No.
17/375,960
Granted
Jan 20, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for accounting for long-tail training data.

Claims (36)

1 . A method performed by one or more computers, the method comprising:

obtaining training data comprising a plurality of training examples, each training example including a training input and a label for the training input that identifies a ground truth category for the training input from a plurality of categories; and

training a classifier neural network on the training data, the classifier neural network having a plurality of network parameters and the training comprising repeatedly performing operations comprising:

obtaining a batch of one or more training examples from the training data;

for each training example in the batch, processing the training input in the training example using the classifier neural network and in accordance with current values of the network parameters to generate a set of scores for the training input that includes a respective score for each of the plurality of categories;

for each training example in the batch, generating a respective adjusted score for each of the plurality of categories, wherein the respective adjusted scores are calculated based on respective prior probability estimates for each of the plurality of categories and the set of scores generated for the training input in the training example, and wherein the respective prior probability estimates for the categories are determined based on a distribution of the training data;

determining gradients of a logit adjusted loss function with respect to the network parameters, wherein the logit adjusted loss function measures, for each training example in the batch, a cross entropy loss between (i) a first probability distribution that assigns a probability of one to the ground truth category for the training input in the training example and a probability of zero to all other categories in the plurality of categories and (ii) a second probability distribution, wherein the second probability distribution is generated from the respective adjusted scores for each of the plurality of categories; and

updating the current values of the network parameters using the gradients.

2 . The method of claim 1 , wherein the respective prior probabilities for each of the categories are determined based on how many training examples in the training data are in each of the categories.

3 . The method of claim 2 , wherein the respective prior probability estimate is higher for categories that have larger numbers of training examples than for categories that have smaller numbers of training examples.

4 . The method of claim 2 , wherein the respective prior probability estimate for a given category is equal to a ratio of (i) a number of training examples in the training data that are in the given category to (ii) a total number of training examples in the training data.

5 . The method of claim 1 , wherein the adjusted score for a given category y is equal to:

f y (x)+τlog(π y ),

where f y (x) is the respective score for the given category y generated by the classifier neural network f by processing the training input x, τ is a positive value, and π y is the prior probability estimate.

6 . A method performed by one or more computers for classifying an image input using an inference system, the inference system comprising a classifier neural network that has been trained on training data comprising a plurality of image training examples, each image training example including an image training input and a label for the image training input that identifies a ground truth category for the image training input from a plurality of categories, and the method comprising:

receiving, by the inference system, an image input;

processing the image input using the classifier neural network to generate a set of scores for the image input that includes a respective score for each of the plurality of categories;

for each of the plurality of categories, generating, by the inference system, a respective adjusted score for the category from (i) the respective score for the category generated by the classifier neural network and (ii) a prior probability estimate for the category, wherein the prior probability estimate for the category is determined based on a distribution of the training data; and

identifying, by the inference system, one or more categories having the highest adjusted scores, wherein each adjusted score was calculated based on the respective score generated by the classifier neural network, and assigning the one or more categories having the highest adjusted scores as a classification for the image input.

7 . The method of claim 6 , wherein the respective prior probability estimates for each of the categories are determined based on how many image training examples in the training data are in each of the categories.

8 . The method of claim 7 , wherein the respective prior probability estimate is higher for categories that have larger numbers of training examples than for categories that have smaller numbers of image training examples.

9 . The method of claim 7 , wherein the respective prior probability estimate for a given category is equal to a ratio of (i) a number of image training examples in the training data that are in the given category to (ii) a total number of image training examples in the training data.

10 . The method of claim 6 , wherein the adjusted score for a given category y is equal to:

f y (x)−τlog(π y ),

where f y (x) is the respective score for the given category y generated by the classifier neural network f by processing the image input x, τ is a positive value, and π y is the prior probability estimate.

11 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations for classifying an image input using an inference system, the inference system comprising a classifier neural network that has been trained on training data comprising a plurality of image training examples, each image training example including an image training input and a label for the image training input that identifies a ground truth category for the image training input from a plurality of categories, and the operations comprising:

receiving, by the inference system, an image input;

processing the image input using the classifier neural network to generate a set of scores for the image input that includes a respective score for each of the plurality of categories;

, for each of the plurality of categories, generating, by the inference system, a respective adjusted score for the category from (i) the respective score for the category generated by the classifier neural network and (ii) a prior probability estimate for the category, wherein the prior probability estimate for the category is determined based on a distribution of the training data; and

identifying, by the inference system, one or more categories having the highest adjusted scores, wherein each adjusted score was calculated based on the respective score generated by the classifier neural network, and assigning the one or more categories having the highest adjusted scores as a classification for the image input.

12 . The system of claim 11 , wherein the respective prior probability estimates for each of the categories are determined based on how many image training examples in the training data are in each of the categories.

13 . The system of claim 12 , wherein the respective prior probability estimate is higher for categories that have larger numbers of image training examples than for categories that have smaller numbers of image training examples.

14 . The system of claim 12 , wherein the respective prior probability estimate for a given category is equal to a ratio of (i) a number of image training examples in the training data that are in the given category to (ii) a total number of image training examples in the training data.

15 . The system of claim 11 , wherein the adjusted score for a given category y is equal to:

f y (x)−τlog(π y ),

where f y (x) is the respective score for the given category y generated by the classifier neural network f by processing the image input x, τ is a positive value, and π y is the prior probability estimate.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2021
From: MENON, ADITYA KRISHNA; KUMAR, SANJIV; JAIN, HIMANSHU; VEIT, ANDREAS; RAWAT, ANKIT SINGH; HIRIMBURA MATARA KANKANAMGE, GAYAN SADEEP JAYASUMANA
To: GOOGLE LLC
Reel/Frame 057119/0568 →
Continuity (1)
Related Publication 20230017505A1 · Jan 19, 2023
References Cited (66)
US 6728690B1 · Meek · 2004 [cited by examiner]
Nguyen-Trang, T., & Vo-Van, T. (May 2016). A new approach for determining the prior probabilities in the classification problem by Bayesian method. Advances in Data Analysis and Classification, 11, 629-643. (Year: 2016). [cited by examiner]
Zhou, B., Cui, Q., Wei, X. S., & Chen, Z. M. (Jun. 2020). Bbn: Bilateral-branch network with cumulative learning for long-tailed visual recognition. In Proceedings of the IEEE/CVF conference on computer vision and patte… [cited by examiner]
Kim, Y., & Zhang, O. (Jun. 2014). Credibility adjusted term frequency: A supervised term weighting scheme for sentiment analysis and text classification. arXiv preprint arXiv:1405.3518. (Year: 2014). [cited by examiner]
Kull, M., & Flach, P. (Jan. 2015). Novel decompositions of proper scoring rules for classification: Score adjustment as precursor to calibration. In Machine Learning and Knowledge Discovery in Databases, ECML PKDD 2015,… [cited by examiner]
Bartlett et al., “Convexity, Classification, and Risk Bounds,” Journal of the American Statistical Association, Jun. 16, 2005, 63 pages. [cited by applicant]
Bengio, et al., “Adaptive Importance Sampling to Accelerate Training of a Neural Probabilistic Language Model,” Trans. Neur. Netw., Apr. 2008, p. 713-722. [cited by applicant]
Brodersen et al., “The Balanced Accuracy and it's Posterior Distribution,” Proceedings of the International Conference on Pattern Recognition, Aug. 2010, p. 3121-3124. [cited by applicant]
Buda et al., “A Systematic Study of the Class Imbalance Problem in Convolutional Neural Networks,” Neural Networks, 2018, p. 249-259. [cited by applicant]
Byrd et al., “What is the Effect of Importance Weighting in Deep Learning,” Proceedings of the 36th International Conference on Machine Learning, Jun. 2019, p. 872-881. [cited by applicant]
Cao et al, “Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss,” 33rd Conference on Neural Information Processing Systems, Oct. 27, 2019, 18 pages. [cited by applicant]
Cardie et al., “Improving Minority Class Prediction Using Case-Specific Feature Weights,” Proceedings of the International Conference on Machine Learning, 1997, 10 pages. [cited by applicant]
Chan et al., “Learning with Non-Uniform Class and Cost Distributions: Effects and a Distributed Multi-Classifier Approach,” KDD—98 Workshop on Distributed Data Mining, 1998, 27 pages. [cited by applicant]
Chawla et al., “Synthetic Minority Over-Sampling Technique,” Journal of Artificial Intelligence Research (JAIR), Jun. 2002, pp. 321-357. [cited by applicant]
Cui et al., “Class-Balanced Loss Based on Effective Number of Samples,” CVPR, 2019, 10 pages. [cited by applicant]
Dmochowski et al., “Maximum Likelihood in Cost-Sensitive Learning: Model Specification, Approximations and Upper Bounds,” Journal of Machine Learning Research, 2010, p. 3313-3332. [cited by applicant]
Elkan, “The Foudnations of Cost-Sensitive Learning,” Proceedings of the International Joint Conference on Artifical Intelligence, 2001, 6 pages. [cited by applicant]
Fan et al., “Learning with Average Top-K Loss,” Advances in Neural Information Processing Systems, 2017, p. 497-505. [cited by applicant]
Fawcett et al., “Combining Data Mining and Machine Learning for Effective User Profiling,” Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1996, p. 8-13. [cited by applicant]
Gneiting et al., “Probabilistic Forecasts, Calibration and Sharpness,” Journal of the Royal Statistical Society: Series B Statistical Methodology, 2007, p. 243-268. [cited by applicant]
Gneiting et al., “Strictly Proper Scoring Rules, Prediction, and Estimation,” Journal of the American Statistical Association, 2007, p. 359-378. [cited by applicant]
Guo et al., “On Calibration of Modern Neural Networks,” Proceedings of the 34th International Conference on Machine Learning, 2017, p. 1321-1330. [cited by applicant]
Hazan et al., “Approximated Structed Prediction for Learning Large Scale Graphical Models,” arXiv prints 1006.2899v2, Jul. 9, 2012, 5 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition, 2016, 9 pages. [cited by applicant]
Hinton et al., “Distilling the Knowledge in a Neural Network,” arXiv prints 1503.0253lvl, Mar. 9, 2015, 9 pages. [cited by applicant]
Iranmehr et al., “Cost-Sensitive Support Vector Machines,” Neurocomputing, 2019, vol. 343, p. 50-64. [cited by applicant]
Jamal et al., “Rethinking Class-Balanced Methods for Long-Tailed Visual Recognition from a Domain Adaptation Perspective,” Google Research, 2020, 10 pages. [cited by applicant]
Kang et al., “Decoupling Representation and Classifier for Long-Tailed Recognition,” Eighth Conference on Learning Representations (ICLR), Feb. 19, 2020, 16 pages. [cited by applicant]
Khan et al., “Stiriking the Right Balance with Uncertainty,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, p. 103-112. [cited by applicant]
Kim et al., “Adjusting Decision Boundary for Class Imbalanced Learning,” arXiv prints, Mar. 11, 2020, 11 pages. [cited by applicant]
King et al., “Logistic Regression in Rare Events Data,” Political Analysis, 2001, vol. 9(2), p. 137-163. [cited by applicant]
Koltchinskii et al., “Some New Bounds on the Generalization Error of Combined Classifiers,” Advances in Neural Information Processing Systems, MIT Press, 1997, p. 245-251. [cited by applicant]
Kubat et al., “Addressing the Curse of Imbalanced Training Sets: One Sided Selection,” Proceedings of the International Conference on Machine Learning, 1997, 8 pages. [cited by applicant]
Kuleshov et al., “Accurate Uncertainties for Deep Learning Using Calibrated Regression,” Proceedings of the 35th International Conference on Machine Learning, vol. 80 of Proceedings of Machine Learning Research, 2018, p… [cited by applicant]
Lin, “A Note on Margin Based Loss Functions in Classification,” Statistics & Probability Letters, vol. 68(1), p. 73-82. [cited by applicant]
Liu et al., “Large Scale Long-Tailed Recognition in an Open World,” IEEE Conference on Computer Vision and Pattern Recognition, 2019, p. 2537-2546. [cited by applicant]
Liu et al., “Large-Margin Softmax Loss for Convolutional Neural Networks,” Proceedings of the 33rd International Conference on Machine Learning, vol. 48, 2016, p. 507-516. [cited by applicant]
Liu et al., “Sphereface: Deep Hypersphere Embedding for Face Recognition,” 2017 IEEE Conference on Computer Vision and Pattern Recognition, 2017, p. 6738-6746. [cited by applicant]
Mahajan et al., “Exploring the Limits of Weakly Supervised Pretraining,” Computer Vision—ECCV 2018, 2018, p. 185-201. [cited by applicant]
Maloof et al., “Learning when Data Sets are Imbalanced and when costs are unequal and unknown,” ICML 2003 Workshop on Learning from Imbalanced Datasets, 2003, 8 pages. [cited by applicant]
Menon et al., “On the Statistical Consistency of Algorithms for Binary Classification under Class Imbalance,” Proceedings of the 30th International Conference on Machine Learning, 2013, p. 603-611. [cited by applicant]
Mikolov et al., “Distributed Representations of Words and Phrases and their Compositionality,” Proceedings of the 26th International Conference on Neural Information Processing Systems, 2013, p. 3111-3119. [cited by applicant]
Morik et al., “Combining Statistical Learning with a knowledge-based approach—a case study in intensive care monitoring,” Proceedings of the Sixteenth International Conference on Machine Learning, 1999, p. 268-277. [cited by applicant]
Muller et al., “When does label smoothing help,” Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems, 2019, p. 4696-4705. [cited by applicant]
Murphy et al., “A general framework for forecast verification,” Monthly Weather Review, 1987, vol. 115(7), p. 1330-1338. [cited by applicant]
Pletscher et al., “Entropy and Margin Maximization for Structured Output Learning,” Machine Learning and Knowledge Discovery in Databases, 2010, p. 83-98. [cited by applicant]
Provost, “Machine Learning from Imbalanced Data Sets 101,” Proceedings of the AAAI—2000 Workshop on Imbalanced Data Sets, 2000, 3 pages. [cited by applicant]
Qiao et al., “Adaptive weighted lerarning for unbalanced multicategory classification,” Biometrics, 2009, vol. 65(1), p. 159-168. [cited by applicant]
Reid et al., “Composite Binary Losses,” Journal of Machine Learning Research, 2010, vol. 11, p. 2387-2422. [cited by applicant]
Shirazi et al., “Risk Minimization, Probability Elicitation and Cost Sensitive SVM's,” Proceedings of the 27th International Conference on Machine Learning, 2010, p. 759-766. [cited by applicant]
Soudry et al., “The Implicit bias of gradient descent on separable data,” Journal of Machine Learning Res., Jan. 2018, vol. 19(1), p. 2822-2878. [cited by applicant]
Szegedy et al., “Rethinking the Inception Architecture for Computer Vision,” 2016 IEEE Conference on Computer Vision and Pattern Recognition, 2016, p. 2818-2826. [cited by applicant]
Tan et al., “Equalization Loss for Long-Tailed Object Recognition,” Computer Vision Foundation, 2020, 10 pages. [cited by applicant]
Tatsumi et al., “Performance Evaluation of Multiobjective Multiclass Support Vector Machines Maximizing Geometric Margins,” Numerical Algebra, Control and Optimization, 2011, vol. 1:151, 19 pages. [cited by applicant]
Van Horn et al., “The Devil is in the Tails: Fine Grained Classification in the Wild,” arXiv print 1709.01450, 2017, 22 pages. [cited by applicant]
Wallace et al., “Class Imbalance,” IEEE, 2011, 10 pages. [cited by applicant]
Wang et al., “Additive Margin Softmax for Face Verification,” IEEE Signal Processing Letters, 2018, vol. 25(7), p. 926-930. [cited by applicant]
Wu et al., “Asymmetric Support Vector Machines: Low false-positive learning under the user tolerance,” Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2008, p. 749-757. [cited by applicant]
Xie et al., “The Logit Model and Response-based samples,” Sociolofical Methods & Research, 1989, vol. 17(3), p. 283-302. [cited by applicant]
Ye et al., “Identifying and Compensating for Feature Deviation in Imbalanced Deep Learning,” arXiv prints 2001.01385v3, Nov. 8, 2020, 18 pages. [cited by applicant]
Yi et al., “Sampling bias corrected neural modeling for large corpus item recommendations,” Proceedings of the 13th ACM Conference on Recommender Systems, 2019, p. 269-277. [cited by applicant]
Zadrozny et al., “Learning and Making Decisions when costs and probabilities are both unknown,” Proceedings of the Seventh ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2001, p. 204-213. [cited by applicant]
Zhang et al., “Range loss for deep face recognition with long tailed training data,” 2017 IEEE International Conference on Computer Vision, 2017, p. 5419-5428. [cited by applicant]
Zhang et al., “To balance or not to balance: A simple yet effective approach for learning with long-tailed distributions,” arXiv prints 1912.04486v2, Mar. 10, 2020, 17 pages. [cited by applicant]
Zhang, “Class Size independent generalization analysis of some discriminative multicategory classification methods,” Proceedings of the 17th International Conference on Neural Information Processing Systems, 2004, p. 16… [cited by applicant]
Zhou et al., “Training Cost-Sensitive Neural Networks with Methods Addressing the Class Imbalance Problem,” IEEE Transactions on Knowledge and Data Engineering (TKDE), 2006, vol. 18(1), 14 pages. [cited by applicant]