IP Library › Granted Patent US 12,561,557
Granted Patent B2
US 12,561,557 · App. 17/551,065 · Granted Feb 24, 2026

Meta pseudo-labels

Inventors: Hieu Hy Pham (Redwood City, CA); Zihang Dai (Pittsburgh, PA); Qizhe Xie (Pittsburgh, PA); Quoc V. Le (Sunnyvale, CA)
Assignee: Google LLC
G06N3/08G06F18/2155G06F18/217G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,557
App. No.
17/551,065
Granted
Feb 24, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network using meta pseudo-labels. One of the methods includes training a student neural network using pseudo-labels generated by a teacher neural network that is being trained jointly with the student neural network.

Claims (63)

1 . A method performed by one or more computers and for training a student neural network having a plurality of student parameters to perform a machine learning task, the method comprising:

training the student neural network jointly with a teacher neural network, wherein the teacher neural network has a plurality of teacher parameters, the joint training comprising repeatedly performing the following:

obtaining a first plurality of unlabeled training inputs;

processing each of the unlabeled training inputs in the first plurality of unlabeled training inputs using the teacher neural network and in accordance with current values of the teacher parameters to generate a respective teacher output for the machine learning task for each of the unlabeled training inputs;

generating by the teacher neural network, for each of the unlabeled training inputs, a respective pseudo-label for the unlabeled training input from the respective teacher output for the unlabeled training input;

training the student neural network to determine updated values of the student parameters from current values of the student parameters by optimizing a student objective function that measures, for each of the unlabeled training inputs in the first plurality of unlabeled training inputs, an error between (i) a respective student output for the unlabeled training input generated by processing the unlabeled training input using the student neural network in accordance with the current values of the student parameters and (ii) the respective pseudo-label for the unlabeled training input, wherein each teacher and student output for the machine learning task specifies a respective probability distribution over a plurality of classes and wherein generating, for each of the unlabeled training inputs, the respective pseudo-label for the unlabeled training input comprises:

selecting one of the classes using the probability distribution specified by the teacher output; and

generating a pseudo-label that identifies the selected class as the ground-truth output for the unlabeled training input;

obtaining a first plurality of labeled training inputs and, for each labeled training input in the first plurality of labeled training inputs, a respective ground truth output for the machine learning task; and

training the teacher neural network to determine updated values of the teacher parameters using student outputs generated by the student neural network for the labeled training inputs, wherein training the teacher neural network comprises optimizing a teacher objective function that includes a first term that measures, for each of the labeled training inputs in the first plurality of labeled training inputs, an error between (i) a respective student output for the labeled training input generated by processing the labeled training input using the student neural network in accordance with the updated values of the student parameters and (ii) the respective ground truth output for the labeled training input, and wherein optimizing the teacher objective comprises an approximate gradient of the first term of the teacher objective function with respect to the teacher parameters.

2 . The method of claim 1 , wherein the total number of teacher parameters of the teacher neural network is greater than the total number of student parameters of the student neural network.

3 . The method of claim 1 , wherein the teacher objective function also includes a supervised learning term that measures, for each labeled training input in a second plurality of labeled training inputs, an error between (i) a respective teacher output for the labeled training input generated by processing the labeled training input using the teacher neural network in accordance with the current values of the teacher parameters and (ii) a respective ground truth output for the labeled training input.

4 . The method of claim 3 , wherein the first plurality of labeled training inputs are the same as the second plurality of labeled training inputs.

5 . The method of claim 1 , wherein the teacher objective function also includes a semi-supervised learning term that measures, for a second plurality of unlabeled training inputs, a performance of the teacher neural network in accordance with the current values of the teacher parameters on a semi-supervised learning task as measured on the second plurality of unlabeled training inputs.

6 . The method of claim 5 , wherein the first plurality of unlabeled training inputs are the same as the second plurality of unlabeled training inputs.

7 . The method of claim 1 , wherein, for each of the unlabeled training inputs, the respective pseudo-label for the unlabeled training input is the same as the respective teacher output for the unlabeled training input.

8 . The method of claim 1 , wherein selecting one of the classes using the probability distribution specified by the teacher output comprises:

sampling one of the classes from the probability distribution specified by the teacher output.

9 . The method of claim 1 , wherein computing an approximate gradient of the first term of the teacher objective function comprises:

computing a first student gradient with respect to the student parameters of the student objective function evaluated at the current values of the student parameters and for the first plurality of unlabeled training inputs;

computing a second student gradient with respect to the student parameters of the first term of the teacher objective function evaluated at the updated values of the student parameters and for the first plurality of labeled training inputs;

computing a teacher gradient with respect to the teacher parameters of a second objective that measures, for each of the first plurality of unlabeled training inputs, an error between (i) the respective pseudo-label for the unlabeled training input and (ii) the respective teacher output for the unlabeled training input generated by the teacher neural network in accordance with the current values of the teacher parameters; and

computing the approximation from the first student gradient, the second student gradient, and the teacher gradient.

10 . The method of claim 9 , wherein computing the approximation from the first student gradient, the second student gradient, and the teacher gradient comprises:

determining a feedback coefficient from the first and second student gradients; and

multiplying the teacher gradient by the feedback coefficient.

11 . The method of claim 1 , further comprising:

after the joint training, further training the student neural network on a third plurality of labeled training inputs through supervised learning.

12 . The method of claim 1 , further comprising:

before the joint training, training the teacher neural network on a fourth plurality of labeled training inputs through supervised learning.

13 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for training a student neural network having a plurality of student parameters to perform a machine learning task, the operations comprising:

training the student neural network jointly with a teacher neural network, wherein the teacher neural network has a plurality of teacher parameters, the joint training comprising repeatedly performing the following:

obtaining a first plurality of unlabeled training inputs;

processing each of the unlabeled training inputs in the first plurality of unlabeled training inputs using the teacher neural network and in accordance with current values of the teacher parameters to generate a respective teacher output for the machine learning task for each of the unlabeled training inputs;

generating by the teacher neural network, for each of the unlabeled training inputs, a respective pseudo-label for the unlabeled training input from the respective teacher output for the unlabeled training input;

training the student neural network to determine updated values of the student parameters from current values of the student parameters by optimizing a student objective function that measures, for each of the unlabeled training inputs in the first plurality of unlabeled training inputs, an error between (i) a respective student output for the unlabeled training input generated by processing the unlabeled training input using the student neural network in accordance with the current values of the student parameters and (ii) the respective pseudo-label for the unlabeled training input, wherein each teacher and student output for the machine learning task specifies a respective probability distribution over a plurality of classes and wherein generating, for each of the unlabeled training inputs, the respective pseudo-label for the unlabeled training input comprises:

selecting one of the classes using the probability distribution specified by the teacher output; and

generating a pseudo-label that identifies the selected class as the ground-truth output for the unlabeled training input;

obtaining a first plurality of labeled training inputs and, for each labeled training input in the first plurality of labeled training inputs, a respective ground truth output for the machine learning task; and

training the teacher neural network to determine updated values of the teacher parameters using student outputs generated by the student neural network for the labeled training inputs, wherein training the teacher neural network comprises optimizing a teacher objective function that includes a first term that measures, for each of the labeled training inputs in the first plurality of labeled training inputs, an error between (i) a respective student output for the labeled training input generated by processing the labeled training input using the student neural network in accordance with the updated values of the student parameters and (ii) the respective ground truth output for the labeled training input, and wherein optimizing the teacher objective comprises an approximate gradient of the first term of the teacher objective function with respect to the teacher parameters.

14 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations for training a student neural network having a plurality of student parameters to perform a machine learning task, the operations comprising:

training the student neural network jointly with a teacher neural network, wherein the teacher neural network has a plurality of teacher parameters, the joint training comprising repeatedly performing the following:

obtaining a first plurality of unlabeled training inputs;

processing each of the unlabeled training inputs in the first plurality of unlabeled training inputs using the teacher neural network and in accordance with current values of the teacher parameters to generate a respective teacher output for the machine learning task for each of the unlabeled training inputs;

generating by the teacher neural network, for each of the unlabeled training inputs, a respective pseudo-label for the unlabeled training input from the respective teacher output for the unlabeled training input;

training the student neural network to determine updated values of the student parameters from current values of the student parameters by optimizing a student objective function that measures, for each of the unlabeled training inputs in the first plurality of unlabeled training inputs, an error between (i) a respective student output for the unlabeled training input generated by processing the unlabeled training input using the student neural network in accordance with the current values of the student parameters and (ii) the respective pseudo-label for the unlabeled training input, wherein each teacher and student output for the machine learning task specifies a respective probability distribution over a plurality of classes and wherein generating, for each of the unlabeled training inputs, the respective pseudo-label for the unlabeled training input comprises:

selecting one of the classes using the probability distribution specified by the teacher output; and

generating a pseudo-label that identifies the selected class as the ground-truth output for the unlabeled training input;

obtaining a first plurality of labeled training inputs and, for each labeled training input in the first plurality of labeled training inputs, a respective ground truth output for the machine learning task; and

training the teacher neural network to determine updated values of the teacher parameters using student outputs generated by the student neural network for the labeled training inputs, wherein training the teacher neural network comprises optimizing a teacher objective function that includes a first term that measures, for each of the labeled training inputs in the first plurality of labeled training inputs, an error between (i) a respective student output for the labeled training input generated by processing the labeled training input using the student neural network in accordance with the updated values of the student parameters and (ii) the respective ground truth output for the labeled training input, and wherein optimizing the teacher objective comprises an approximate gradient of the first term of the teacher objective function with respect to the teacher parameters.

15 . The system of claim 14 , wherein the teacher objective function also includes a supervised learning term that measures, for each labeled training input in a second plurality of labeled training inputs, an error between (i) a respective teacher output for the labeled training input generated by processing the labeled training input using the teacher neural network in accordance with the current values of the teacher parameters and (ii) a respective ground truth output for the labeled training input.

16 . The system of claim 14 , wherein the teacher objective function also includes a semi-supervised learning term that measures, for a second plurality of unlabeled training inputs, a performance of the teacher neural network in accordance with the current values of the teacher parameters on a semi-supervised learning task as measured on the second plurality of unlabeled training inputs.

17 . The system of claim 14 , wherein, for each of the unlabeled training inputs, the respective pseudo-label for the unlabeled training input is the same as the respective teacher output for the unlabeled training input.

18 . The system of claim 14 , wherein selecting one of the classes using the probability distribution specified by the teacher output comprises:

sampling one of the classes from the probability distribution specified by the teacher output.

19 . The system of claim 14 , wherein computing an approximate gradient of the first term of the teacher objective function comprises:

computing a first student gradient with respect to the student parameters of the student objective function evaluated at the current values of the student parameters and for the first plurality of unlabeled training inputs;

computing a second student gradient with respect to the student parameters of the first term of the teacher objective function evaluated at the updated values of the student parameters and for the first plurality of labeled training inputs;

computing a teacher gradient with respect to the teacher parameters of a second objective that measures, for each of the first plurality of unlabeled training inputs, an error between (i) the respective pseudo-label for the unlabeled training input and (ii) the respective teacher output for the unlabeled training input generated by the teacher neural network in accordance with the current values of the teacher parameters; and

computing the approximation from the first student gradient, the second student gradient, and the teacher gradient.

20 . The system of claim 19 , wherein computing the approximation from the first student gradient, the second student gradient, and the teacher gradient comprises:

determining a feedback coefficient from the first and second student gradients; and

multiplying the teacher gradient by the feedback coefficient.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2022
From: PHAM, HIEU HY; DAI, ZIHANG; XIE, QIZHE; LE, QUOC V.
To: GOOGLE LLC
Reel/Frame 058525/0199 →
Continuity (2)
Provisional Application 63125363 · Dec 14, 2020
Related Publication 20220188636A1 · Jun 16, 2022
References Cited (96)
US 20160078339A1 · Li · 2016 [cited by examiner]
US 20190205748A1 · Fukuda · 2019 [cited by examiner]
US 20200134427A1 · Oh · 2020 [cited by examiner]
Asit Mishra et al, “Apprentice: Using Knowledge Distillation Techniques to Improve Low-Precision Network Accuracy”, Nov. 15, 2017, arXiv:1711.05852v1 (Year: 2017). [cited by examiner]
Qizhe Xie et al., “Self-training with Noisy Student improves ImageNet classification”, Nov. 11, 2019, arXiv:1911.04252v1 (Year: 2019). [cited by examiner]
Geoffrey Hinton et al., “Distilling the Knowledge in a Neural Network”, Mar. 9, 2015, arxiv:1503.02531v1 (Year: 2015). [cited by examiner]
Aakash Nain, “Self-training with Noisy Student”, Nov. 16, 2019, https://medium.com/@nainaakash012/self-training-with-noisy-student-f33640edbab2#:˜:text=Annotating%20data%20is%20mundane.,called%20it%20a%20noisy%20student… [cited by examiner]
Khaled Boudaoud, “The motivation behind Residual Neural Networks”, Nov. 5, 2019, https://medium.com/analytics-vidhya/introduction-to-residual-neural-networks-8af5b7c4afd4 (Year: 2019). [cited by examiner]
Lin Wang et al., “Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks”, Apr. 13, 2020, arXiv:2004.05937v1 (Year: 2020). [cited by examiner]
Abadi et al., “TensorFlow: A System for Large-Scale Machine Learning,” Presented at 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI '16), Savannah, GA, USA, Nov. 2-4, 2016, pp. 265-283. [cited by applicant]
Arazo et al., “Pseudo-Labeling and Confirmation Bias in Deep Semi-Supervised Learning,” arXiv, Sep. 25, 2019, 12 pages. [cited by applicant]
Baydin et al., “Online Learning Rate Adaptation with Hypergradient Descent,” Presented at Sixth International Conference on Learning Representations, Vancouver, Canada, April 30-May 3, 2018, 11 pages. [cited by applicant]
Berthelot et al., “MixMatch: A Holistic Approach to Semi-Supervised Learning,” Advances in Neural Information Processing Systems, 2019, 11 pages. [cited by applicant]
Berthelot et al., “ReMixMatch: Semi-supervised learning with distribution alignment and augmentation anchoring,” arXiv, Feb. 13, 2020, 13 pages. [cited by applicant]
Beyer et al., “Are we done with ImageNet?”, arXiv, Jun. 12, 2020, 15 pages. [cited by applicant]
Chapelle et al., “Semi-Supervised Learning,” The MIT press, 2010, 524 pages. [cited by applicant]
Chen et al., “A Simple Framework for Contrastive Learning of Visual Representations,” Proceedings of the 37th International Conference on Machine Learning, 2020, 11 pages. [cited by applicant]
Chen et al., “Big Self-Supervised Models are Strong Semi-Supervised Learners,” Advances in Neural Information Processing Systems, 2020, 13 pages. [cited by applicant]
Chen et al., “Improved Baselines with Momentum Contrastive Learning,” arXiv, Mar. 9, 2020, 3 pages. [cited by applicant]
Chollet et al. “Xception: Deep Learning With Depthwise Separable Convolutions,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jul. 2017, pp. 1251-1258. [cited by applicant]
Cubuk et al., “AutoAugment: Learning Augmentation Strategies From Data,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 113-123. [cited by applicant]
Cubuk et al., “RandAugment: Practical Automated Data Augmentation with a Reduced Search Space,” Advances in Neural Information Processing Systems, 2020, 12 pages. [cited by applicant]
Dosovitskiy et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale,” arXiv, Oct. 22, 2020, 21 pages. [cited by applicant]
Finn et al., “Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks,” Proceedings of the 34th International Conference on Machine Learning, 2017, 10 pages. [cited by applicant]
Furlanello et al., “Born Again Neural Networks,” Proceedings of the 35th International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Ghiasi et al., “DropBlock: A regularization method for convolutional networks,” Advances in Neural Information Processing Systems, 2018, 11 pages. [cited by applicant]
Gidaris et al., “Unsupervised representation learning by predicting image rotations,” arXiv, Mar. 21, 2018, 6 pages. [cited by applicant]
Grandvalet et al., “Semi-supervised Learning by Entropy Minimization,” Advances in Neural Information Processing Systems, 2004, 8 pages. [cited by applicant]
Grill et al., “Bootstrap Your Own Latent—A New Approach to Self-Supervised Learning,” Advances in Neural Information Processing Systems, 2020, 14 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2016, pp. 770-778. [cited by applicant]
He et al., “Momentum Contrast for Unsupervised Visual Representation Learning,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2020, pp. 9729-9738. [cited by applicant]
He et al., “Revisiting Self-Training for Neural Sequence Generation,” Presented at Eighth International Conference on Learning Representations, virtual conference, Apr. 26-May 1, 2020, 15 pages. [cited by applicant]
Henaff et al., “Data-Efficient Image Recognition with Contrastive Predictive Coding,” arXiv, Jul. 1, 2020, 13 pages. [cited by applicant]
Hinton et al., “Distilling the Knowledge in a Neural Network,” arXiv, Mar. 9, 2015, 9 pages. [cited by applicant]
Hu et al., “Squeeze- and-Excitation Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 7132-7141. [cited by applicant]
Huang et al., “Densely Connected Convolutional Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jul. 2017, pp. 4700-4708. [cited by applicant]
Huang et al., “GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism,” Advances in Neural Information Processing Systems, 2019, 10 pages. [cited by applicant]
Jackson et al., “Semi-Supervised Learning by Label Gradient Alignment,” arXiv, Feb. 6, 2019, 12 pages. [cited by applicant]
Kahn et al., “Self-Training for End-to-End Speech Recognition,” Presented at IEEE International Conference on Acoustics, Speech and Signal Processing, Barcelona, Spain, May 4-8, 2020, pp. 7084-7088. [cited by applicant]
Ke et al., “Dual Student: Breaking the Limits of the Teacher in Semi-Supervised Learning,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2019, pp. 6728-6736. [cited by applicant]
Kolesnikov et al., “Big Transfer (BiT): General Visual Representation Learning,” Proceedings of the European Conference on Computer Vision, 2020, 10 pages. [cited by applicant]
Krizhevsky, “Learning multiple layers of features from tiny images,” Apr. 8, 2009, 60 pages. [cited by applicant]
Laine et al., “Temporal ensembling for semi-supervised learning,” Presented at 5th International Conference on Learning Representations, Toulon, France, Apr. 24-26, 2017, 13 pages. [cited by applicant]
Lee et al., “Pseudo-Label: The simple and efficient semi-supervised learning method for deep neural networks,” Presented at International Conference on Machine Learning, Atlanta, Georgia, Jun. 16-21, 2013, 6 pages. [cited by applicant]
Lepikhin et al., “GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding,” arXiv, Jun. 30, 2020, 35 pages. [cited by applicant]
Li et al., “Prototypical Contrastive Learning of Unsupervised Representations,” arXiv, Jul. 28, 2020, 16 pages. [cited by applicant]
Liu et al., “DARTS: Differentiable Architecture Search,” Presented at Seventh International Conference on Learning Representations, New Orleans, LA, USA, May 6-9, 2019, 13 pages. [cited by applicant]
Liu et al., “Progressive neural architecture search,” Proceedings of the European Conference on Computer Vision, Sep. 2018, 16 pages. [cited by applicant]
Loshchilov et al., “SGDR: Stochastic Gradient Descent with Warm Restarts,” Presented at 5th International Conference on Learning Representations, Toulon, France, Apr. 24-26, 2017, 16 pages. [cited by applicant]
Mahajan et al., “Exploring the Limits of Weakly Supervised Pretraining,” Proceedings of the European Conference on Computer Vision, Sep. 2018, pp. 181-196. [cited by applicant]
Miller et al., “When does label smoothing help?,” Advances in Neural Information Processing Systems, 2019, 10 pages. [cited by applicant]
Misra et al., “Self-Supervised Learning of Pretext-Invariant Representations,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2020, pp. 6707-6717. [cited by applicant]
Miyato et al., “Virtual Adversarial Training: A Regularization Method for Supervised and Semi-Supervised Learning,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Jul. 23, 2018, 14 pages. [cited by applicant]
Netzer et al., “Reading Digits in Natural Images with Unsupervised Feature Learning,” Advances in Neural Information Processing Systems Workshop on Deep Learning and Unsupervised Feature Learning, 2011, 9 pages. [cited by applicant]
Noroozi et al., “Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles,” arXiv, Jun. 26, 2016, 18 pages. [cited by applicant]
Oliver et al., “Realistic Evaluation of Deep Semi-Supervised Learning Algorithms,” Advances in Neural Information Processing Systems, 2018. [cited by applicant]
Park et al., “Improved noisy student training for automatic speech recognition,” Interspeech, 2020, 5 pages. [cited by applicant]
Pathak et al., “Context Encoders: Feature Learning by Inpainting,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2016, pp. 2536-2544. [cited by applicant]
Radosavovic et al., “Data Distillation: Towards Omni-Supervised Learning,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 4119-4128. [cited by applicant]
Real et al., “Regularized Evolution for Image Classifier Architecture Search,” AAAI Technical Track: Machine Learning, Jul. 17, 2019, pp. 4780-4789. [cited by applicant]
Ren et al., “Learning to Reweight Examples for Robust Deep Learning,” Proceedings of the 35th International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Ren et al., “Not All Unlabeled Data are Equal: Learning to Weight Data in Semi-supervised Learning,” Advances in Neural Information Processing Systems, 2020, 12 pages. [cited by applicant]
Riloff et al., “Automatically generating extraction patterns from untagged text,” Proceedings of the national conference on artificial intelligence, 1996, pp. 1044-1049. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision, Apr. 11, 2015, pp. 211-252. [cited by applicant]
Scudder, “Probability of error of some adaptive pattern-recognition machines,” IEEE Transactions on Information Theory, Jul. 1965, pp. 363-371. [cited by applicant]
Sohn et al., “FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence,” Advances in Neural Information Processing Systems, 2020, 13 pages. [cited by applicant]
Such et al., “Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training Data,” Proceedings of the 37th International Conference on Machine Learning, 2020, 11 pages. [cited by applicant]
Sun et al., “Revisiting Unreasonable Effectiveness of Data in Deep Learning Era,” Proceedings of the IEEE International Conference on Computer Vision, Oct. 2017, pp. 843-852. [cited by applicant]
Szegedy et al., “Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning,” Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, 2017, pp. 4278-4284. [cited by applicant]
Szegedy et al., “Rethinking the Inception Architecture for Computer Vision,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2016, pp. 2818-2826. [cited by applicant]
Tan et al., “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” Proceedings of the 36th International Conference on Machine Learning, 2019, 10 pages. [cited by applicant]
Tarvainen et al., “Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results,” Advances in Neural Information Processing Systems, 2017, 10 pages. [cited by applicant]
Thomee et al., “YFCC100M: the new data in multimedia research,” Communications of the ACM, Feb. 2016, 59(2):64-73. [cited by applicant]
Torralba et al., “80 Million Tiny Images: A Large Data Set for Nonparametric Object and Scene Recognition,” IEEE Transactions on Pattern Analysis and Machine Learning, Nov. 2008, pp. 1958-1970. [cited by applicant]
Touvron et al., “Fixing the train-test resolution discrepancy,” Advances in Neural Information Processing Systems, 2019, 11 pages. [cited by applicant]
Touvron et al., “Fixing the train-test resolution discrepancy: FixEfficientNet,” arXiv, Nov. 18, 2020, 5 pages. [cited by applicant]
Verma et al., “Interpolation consistency training for semi-supervised learning,” Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, 2019, pp. 3635-3641. [cited by applicant]
Wang et al., “EnAET: Self-Trained Ensemble AutoEncoding Transformations for Semi-Supervised Learning,” arXiv, Nov. 21, 2019, 10 pages. [cited by applicant]
Wang et al., “Meta-Semi: A Meta-learning Approach for Semi-supervised Learning,” arXiv, Jul. 5, 2020, 18 pages. [cited by applicant]
Wang et al., “Optimizing Data Usage via Differentiable Rewards,” Proceedings of the 37th International Conference on Machine Learning, 2020, 13 pages. [cited by applicant]
Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning, May 1992, pp. 229-256. [cited by applicant]
Xie et al., “Aggregated Residual Transformations for Deep Neural Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jul. 2017, pp. 1492-1500. [cited by applicant]
Xie et al., “Self-Training With Noisy Student Improves ImageNet Classification,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2020, pp. 10687-10698. [cited by applicant]
Xie et al., “Unsupervised Data Augmentation for Consistency Training,” Advances in Neural Information Processing Systems, 2020, 13 pages. [cited by applicant]
Yalniz et al., “Billion-scale semi-supervised learning for image classification,” arXiv, May 2, 2019. [cited by applicant]
Yamada et al., “Shakedrop Regularization for Deep Residual Learning,” IEEE Access, Dec. 18, 2019, 7:186126-186136. [cited by applicant]
Yarowsky, “Unsupervised Word Sense Disambiguation Rivaling Supervised Methods,” 33rd Annual Meeting of the Association for Computational Linguistics, Jun. 1995, pp. 189-196. [cited by applicant]
You et al., “Large Batch Training of Convolutional Networks,” arXiv, Sep. 13, 2017, 8 pages. [cited by applicant]
Yun et al., “CutMix: Regularization Strategy to Train Strong Classifiers With Localizable Features,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2019, pp. 6023-6032. [cited by applicant]
Zagoruyko et al., “Wide Residual Networks,” Proceedings of the British Machine Vision Conference, Sep. 2016, 12 pages. [cited by applicant]
Zhang et al., “mixup: Beyond Empirical Risk Minimization,” International Conference on Learning Representations, 2018, 13 pages. [cited by applicant]
Zhang et al., “Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2019, pp. 3713-3722. [cited by applicant]
Zhang et al., “PolyNet: A Pursuit of Structural Diversity in Very Deep Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jul. 2017, pp. 718-726. [cited by applicant]
Zheng et al., “Meta Label Correction for Learning with Weak Supervision,” arXiv, Nov. 10, 2019, 5 pages. [cited by applicant]
Zoph et al., “Rethinking Pre-training and Self-training,” Advances in Neural Information Processing Systems, 2020, 13 pages. [cited by applicant]
Zoph et al., “Learning Transferable Architectures for Scalable Image Recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 8697-8710. [cited by applicant]