IP Library › Granted Patent US 12,333,775
Granted Patent B2
US 12,333,775 · App. 17/526,886 · Granted Jun 17, 2025

Efficient neural networks via ensembles and cascades

Inventors: Yair Alon (Mountain View, CA); Elad Eban (Mountain View, CA); Xiaofeng Wang (Mountain View, CA)
Assignee: Google LLC
G06V10/255G06N3/045G06N20/20G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,775
App. No.
17/526,886
Granted
Jun 17, 2025
Kind
B2
Abstract

A combination of two or more trained machine learning models can exhibit a combined accuracy greater than the accuracy of any one of the constituent models. However, this increase accuracy comes at additional computational cost. Cascades of machine learning models are provided herein that result in increased model accuracy and/or reduced model compute cost. These benefits are obtained by conditionally executing one or more of the models of the cascade based on the estimated correctness of already-executed models. The estimated correctness can be obtained as an additional output of the already-executed model(s) or could be determined as an entropy, maximum class probability, maximum class logit, or other function of the output(s) of the already-executed model(s). The expected computational cost of executing the model cascade is reduced by only executing the downstream model(s) when the upstream model(s) has resulted in an output whose accuracy is suspect.

Claims (49)

1. A method comprising:

applying a first machine learning model to an input to generate a first model output, wherein the first machine learning model has been trained on a first set of training examples;

determining, based on the first model output, a correctness metric for the first model output;

determining that the correctness metric exceeds a threshold, wherein the threshold has a value that has been determined for the first machine learning model and a second machine learning model based on a set of training inputs; and

responsive to determining that the correctness metric exceeds the threshold:

applying the second machine learning model to the input to generate a second model output, wherein the second machine learning model has been trained by (i) selecting, from the first set of training examples, a second set of training examples that, when applied to the first machine learning model, result in the generation of outputs corresponding to sub-threshold correctness metric values, and (ii) training the second machine learning model using the second set of training examples; and

combining the first model output and the second model output to generate a combined output, wherein the combining comprises at least one of: summing of the first model output and the second model output, taking the mean of the first model output and the second model output, taking a mode of the first model output and the second model output, taking an average of the first model output and the second model output, taking a weighted average of the first model output and the second model output, performing an elementwise operation on the first model output and the second model output, voting based on the first model output and the second model output, or summing logit values of the first model output and logit values of the second model output.

2. The method of claim 1 , wherein the input is an input image.

3. The method of claim 2 , wherein the first machine learning model comprises a convolutional neural network.

4. The method of claim 2 , wherein the first model output includes a version of the input image modified by the first machine learning model.

5. The method of claim 1 , wherein the first model output is indicative of membership of the input in one or more classes from an enumerated set of classes.

6. The method of claim 1 , wherein the first model output includes the correctness metric.

7. The method of claim 1 , wherein determining the correctness metric for the first model output comprises determining the correctness metric based on at least one of: a probability of a highest-confidence class represented by the first model output, an entropy of a probability distribution represented by the first model output, a magnitude of a difference between a highest probability represented by the first model output and a second-highest probability represented by the first model output, a magnitude of a difference between a logit of a highest probability represented by the first model output and a logit of a second-highest probability represented by the first model output, a conditional entropy bottleneck value of the first model output, or a cross entropy of the first model output.

8. The method of claim 1 , wherein combining the first model output and the second model output to generate a combined output comprises summing logit values of the first model output and logit values of the second model output.

9. The method of claim 1 , wherein the first machine learning model and the second machine learning model have different model structures.

10. The method of claim 1 , wherein the first machine learning model and the second machine learning model are the same convolutional neural network evaluated with different input sizes.

11. The method of claim 1 , wherein the first machine learning model and the second machine learning model are associated with respective first and second computational costs, and wherein the threshold has a value that has been determined based on the first and second computational costs such that the expected accuracy of an output of the method is increased while maintaining the expected computational cost of executing the method less than a specified computational cost.

12. The method of claim 1 , wherein the first machine learning model and the second machine learning model are associated with respective first and second computational costs, and wherein the threshold has a value that has been determined based on the first and second computational costs such that the expected computational cost of executing the method is reduced while maintaining the expected accuracy of an output of the method greater than a specified accuracy.

13. The method of claim 1 , further comprising:

applying an additional input to the first machine learning model to generate a third model output;

determining, based on the third model output, a correctness metric for the third model output, wherein the correctness metric for the third model output is indicative of a degree of confidence in the accuracy of the third model output;

determining that the correctness metric for the third model exceeds the threshold; and

responsive to determining that the correctness metric exceeds the threshold:

applying the additional input to the second machine learning model to generate a fourth model output;

determining, based on the fourth model output, a correctness metric for the fourth model output, wherein the correctness metric for the fourth model output is indicative of a degree of confidence in the accuracy of the fourth model output;

determining that the correctness metric for the fourth model exceeds an additional threshold; and

responsive to determining that the correctness metric for the fourth model output exceeds the additional threshold:

applying the additional input to a third machine learning model to generate a fifth model output; and

combining the third model output, the fourth model output, and the fifth model output to generate an additional combined output for the additional input.

14. An article of manufacture including a non-transitory computer-readable medium, having stored therein instructions executable by a computing device to cause the computing device to perform a method comprising:

applying a first machine learning model to an input to generate a first model output, wherein the first machine learning model has been trained on a first set of training examples;

determining, based on the first model output, a correctness metric for the first model output;

determining that the correctness metric exceeds a threshold, wherein the threshold has a value that has been determined for the first machine learning model and a second machine learning model based on a set of training inputs; and

responsive to determining that the correctness metric exceeds the threshold:

applying the second machine learning model to the input to generate a second model output, wherein the second machine learning model has been trained by (i) selecting, from the first set of training examples, a second set of training examples that, when applied to the first machine learning model, result in the generation of outputs corresponding to sub-threshold correctness metric values, and (ii) training the second machine learning model using the second set of training examples; and

combining the first model output and the second model output to generate a combined output, wherein the combining comprises at least one of: summing of the first model output and the second model output, taking the mean of the first model output and the second model output, taking a mode of the first model output and the second model output, taking an average of the first model output and the second model output, taking a weighted average of the first model output and the second model output, performing an elementwise operation on the first model output and the second model output, voting based on the first model output and the second model output, or summing logit values of the first model output and logit values of the second model output.

15. The article of manufacture of claim 14 , wherein the first machine learning model and the second machine learning model are the same convolutional neural network evaluated with different input sizes.

16. The article of manufacture of claim 14 , wherein the first machine learning model and the second machine learning model are associated with respective first and second computational costs, and wherein the threshold has a value that has been determined based on the first and second computational costs such that one of: (i) the expected accuracy of an output of the method is increased while maintaining the expected computational cost of executing the method less than a specified computational cost, or (ii) the expected computational cost of executing the method is reduced while maintaining the expected accuracy of an output of the method greater than a specified accuracy.

17. A system comprising:

one or more processors; and

a non-transitory computer-readable medium, having stored therein instructions executable by the one or more processors to cause the system to perform a method comprising:

applying a first machine learning model to an input to generate a first model output, wherein the first machine learning model has been trained on a first set of training examples;

determining, based on the first model output, a correctness metric for the first model output;

determining that the correctness metric exceeds a threshold, wherein the threshold has a value that has been determined for the first machine learning model and a second machine learning model based on a set of training inputs; and

responsive to determining that the correctness metric exceeds the threshold:

applying the second machine learning model to the input to generate a second model output, wherein the second machine learning model has been trained by (i) selecting, from the first set of training examples, a second set of training examples that, when applied to the first machine learning model, result in the generation of outputs corresponding to sub-threshold correctness metric values, and (ii) training the second machine learning model using the second set of training examples; and

combining the first model output and the second model output to generate a combined output, wherein the combining comprises at least one of: summing of the first model output and the second model output, taking the mean of the first model output and the second model output, taking a mode of the first model output and the second model output, taking an average of the first model output and the second model output, taking a weighted average of the first model output and the second model output, performing an elementwise operation on the first model output and the second model output, voting based on the first model output and the second model output, or summing logit values of the first model output and logit values of the second model output.

18. The system of claim 17 , wherein the first machine learning model and the second machine learning model are the same convolutional neural network evaluated with different input sizes.

19. The system of claim 17 , wherein the first machine learning model and the second machine learning model are associated with respective first and second computational costs, and wherein the threshold has a value that has been determined based on the first and second computational costs such that one of: (i) the expected accuracy of an output of the method is increased while maintaining the expected computational cost of executing the method less than a specified computational cost, or (ii) the expected computational cost of executing the method is reduced while maintaining the expected accuracy of an output of the method greater than a specified accuracy.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2021
From: ALON, YAIR; EBAN, ELAD; WANG, XIAOFENG
To: GOOGLE LLC
Reel/Frame 058310/0617 →
Continuity (2)
Provisional Application 63114205 · Nov 16, 2020
Related Publication 20220156524A1 · May 19, 2022
References Cited (56)
US 20160148079A1 · Shen · 2016 [cited by examiner]
US 20210073686A1 · Ding · 2021 [cited by examiner]
US 20210216831A1 · Ben-Itzhak · 2021 [cited by examiner]
US 20230112076A1 · Enomoto · 2023 [cited by examiner]
Inoue, Hiroshi. “Adaptive ensemble prediction for deep neural networks based on confidence level.” The 22nd International Conference on Artificial Intelligence and Statistics. PMLR, 2019. (Year: 2019). [cited by examiner]
Li, Xiaoxiao, et al. “Not All Pixels Are Equal: Difficulty-Aware Semantic Segmentation via Deep Layer Cascade.” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2017. (Year: 2017). [cited by examiner]
Wang, Xin, et al. “Idk cascades: Fast deep learning by learning not to overthink.” arXiv preprint arXiv:1706.00885v4 (2018). (Year: 2018). [cited by examiner]
Teplyakov, L. M., et al. “Training of neural network-based cascade classifiers.” Journal of Communications Technology and Electronics 64 (2019): 846-853. (Year: 2019). [cited by examiner]
Nguyen, Tien Thanh, et al. “Ensemble selection based on classifier prediction confidence.” Pattern Recognition 100 (2020): 107104. (Year: 2019). [cited by examiner]
Mohandes, Mohamed, Mohamed Deriche, and Salihu O. Aliyu. “Classifiers combination techniques: A comprehensive review.” IEEE Access 6 (2018): 19626-19639. (Year: 2018). [cited by examiner]
Leo Breiman, “Bagging predictors”, Machine learning, 24(2): 123-140, 1996. [cited by applicant]
Cao et al., “Learnable embedding space for efficient neural architecture compression”, 2019. [cited by applicant]
Carreira et al., “A short note about kinetics-600”, 2018. [cited by applicant]
Chaudhuri et al., “Fine-grained stochastic architecture search”, 2020. [cited by applicant]
Chen et al., “Rethinking atrous convolution for semantic image segmentation”, 2017. [cited by applicant]
Cordts et al., “The cityscapes dataset for semantic urban scene understanding”, 2016. [cited by applicant]
Cubuk et al., “Autoaugment: Learning augmentation policies from data”, 2019. [cited by applicant]
Christoph Feichtenhofer, “X3d: Expanding architectures for efficient video recognition”, 2020. [cited by applicant]
Fort et al., “Deep ensembles: A loss landscape perspective”, 2019. [cited by applicant]
Freund et al., “A decision-theoretic generalization of on-line learning and an application to boosting”, Journal of computer and system sciences, 55(1):119-139, 1997. [cited by applicant]
Guan et al., “Energy-efficient amortized inference with cascaded deep classifiers”, 2018. [cited by applicant]
He et al., “Deep residual learning for image recognition”, 2016. [cited by applicant]
Howard et al., “Mobilenets: Efficient convolutional neural networks for mobile vision applications”, 2017. [cited by applicant]
Hu et al., “Squeeze-and-excitation networks”, 2018. [cited by applicant]
Huang et al., “Multi-scale dense networks for resource efficient image classification”, 2018. [cited by applicant]
Huang et al., “Snapshot ensembles: Train 1, get M for free”, 2017. [cited by applicant]
Huang et al., “Densely connected convolutional networks”, 2017. [cited by applicant]
Kondratyuk et al., “When ensembling smaller models is more efficient than single large models”, 2020. [cited by applicant]
Lakshminarayanan et al., “Simple and scalable predictive uncertainty estimation using deep ensembles”, 2017. [cited by applicant]
Liu et al., “Progressive neural architecture search”, 2018. [cited by applicant]
Liu et al., “DARTS: Differentiable architecture search”, 2019. [cited by applicant]
Lobacheva et al. “On power laws in deep ensembles”, 2020. [cited by applicant]
Qiu et al., “Trimmed action recognition, dense-captioning events in videos, and spatio-temporal action localization with focus on activitynet challenge”, 2019. [cited by applicant]
Real et al., “Regularized evolution for image classifier architecture search”, 2019. [cited by applicant]
Russakovsky et al., “Imagenet large scale visual recognition challenge”, 2015. [cited by applicant]
Sandler et al., “Mobilenetv2: Inverted residuals and linear bottlenecks”, 2018. [cited by applicant]
Schapire et al., “The strength of weak learnability”, Machine learning, 5(2):197-227, 1990. [cited by applicant]
Shazeer et al., “Outrageously large neural networks: The sparsely-grated mixture-of-experts layer”, 2017. [cited by applicant]
Streeter et al., “Approximation algorithms for cascading prediction models”, 2018. [cited by applicant]
Szegedy et al., “Inception-v4, inception-resnet and the impact of residual connections on learning”, 2017. [cited by applicant]
Szegedy et al., “Going deeper with convolutions”, 2015. [cited by applicant]
Szegedy et al., “Rethinking the inception architecture for computer vision”, 2016. [cited by applicant]
Tan et al., “MnasNet: Platform-aware neural architecture search for mobile”, 2019. [cited by applicant]
Tan et al., “Efficientnet: Rethinking model scaling for convolutional neural networks”, 2019. [cited by applicant]
Touvron et al., “Fixing the train-test resolution discrepancy”, 2019. [cited by applicant]
Veit et al., “Convolutional networks with adaptive inference graphs”, 2018. [cited by applicant]
Viola et al., “Rapid object detection using a boosted cascade of simple features”, 2001. [cited by applicant]
Wang et al., “Skipnet: Learning dynamic routing in convolutional networks”, 2018. [cited by applicant]
Wen et al., “Batchensemble: an alternative approach to efficient ensemble and lifelong learning”, 2020. [cited by applicant]
Wenzel et al., “Hyperparameter ensembles for robustness and uncertainty quantification”, 2020. [cited by applicant]
Wu et al., “Blockdrop: Dynamic inference paths in residual networks”, 2018. [cited by applicant]
Xie et al., “Aggregated residual transformations for deep neural networks”, 2017. [cited by applicant]
Zhang et al., “Shufflenet: An extremely efficient convolutional neural network for mobile devices”, 2018. [cited by applicant]
Zoph et al., “Learning transferable architectures for scalable image recognition”, 2018. [cited by applicant]
Howard et al., “Searching for mobilenetv3”, 2019. [cited by applicant]
Bolukbasi et al., “Adaptive neural networks for efficient inference”, 2017. [cited by applicant]
Cited By (1)
US 12,579,479