IP Library › Granted Patent US 12,488,237
Granted Patent B2
US 12,488,237 · App. 17/488,166 · Granted Dec 2, 2025

Training neural networks using transfer learning

Inventors: Joan Puigcerver i Perez (Zurich, CH); Basil Mustafa (Zurich, CH); André Susano Pinto (Zurich, CH); Carlos Riquelme Ruiz (Zurich, CH); Neil Matthew Tinmouth Houlsby (Zurich, CH); Daniel M. Keysers (Stallikon, CH)
Assignee: Google LLC
G06N3/08G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,237
App. No.
17/488,166
Granted
Dec 2, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training neural networks using transfer learning. One of the methods includes training a neural network to perform a first prediction task, including: obtaining trained model parameters for each of a plurality of candidate neural networks, wherein each candidate neural network has been pre-trained to perform a respective second prediction task that is different from the first prediction task; obtaining a plurality of training examples corresponding to the first prediction task; selecting a proper subset of the plurality of candidate neural networks using the plurality of training examples; generating, for each candidate neural network, one or more fine-tuned neural networks, wherein each fine-tuned neural network is generated by updating the model parameters of the candidate neural network using the plurality of training examples; and determining model parameters for the neural network using the respective fine-tuned neural networks.

Claims (82)

1 . A method for training a neural network to perform a first prediction task, the method comprising:

obtaining trained model parameters for each of a plurality of candidate neural networks, wherein each candidate neural network has been pre-trained to perform a respective second prediction task that is different from the first prediction task;

obtaining a plurality of training examples corresponding to the first prediction task;

prior to fine-tuning any of the plurality of candidate neural networks for the first prediction task:

predicting, for each plurality of candidate neural networks and using the plurality of training examples, a respective performance of the candidate neural network on the first prediction task, and

selecting a proper subset of the plurality of candidate neural networks using the respective predicted performance on the first prediction task for each of the candidate neural networks;

after selecting the proper subset, fine-tuning only the candidate neural networks in the proper subset for the first prediction task by generating, for each candidate neural network in the proper subset, one or more fine-tuned neural networks, wherein each of the one or more fine-tuned neural networks is generated by updating the model parameters of the candidate neural network using the plurality of training examples; and

determining model parameters for the neural network using the one or more fine-tuned neural networks.

2 . The method of claim 1 , wherein the neural network is an ensemble neural network comprising each fine-tuned neural network.

3 . The method of claim 1 , wherein:

the neural network is an ensemble neural network comprising a plurality of member neural networks; and

determining model parameters for the neural network comprises:

selecting, from the fine-tuned neural networks, the plurality of member neural networks of the ensemble neural network according to a performance of the fine-tuned neural networks on the first prediction task.

4 . The method of claim 3 , wherein selecting, from the fine-tuned neural networks, the plurality of member neural networks of the ensemble neural network according to a performance of the fine-tuned neural networks on the first prediction task comprises:

at a first time step, selecting the fine-tuned neural network that has a highest performance on the first prediction task; and

at each of one or more subsequence time steps, determining, from the remaining fine-tuned neural networks, a particular fine-tuned neural network that, when added to the ensemble neural network, causes a highest increase in the performance of the ensemble neural network on the first prediction task.

5 . The method of claim 1 , wherein the plurality of training examples comprises a plurality of training inputs and corresponding ground-truth outputs, and wherein selecting the proper subset of the plurality of candidate neural networks using the respective predicted performance on the first prediction task for each of the candidate neural networks comprises:

for each of the plurality of candidate neural networks:

processing at least some of the plurality of training inputs using the candidate neural network to generate a respective representation of the training input;

configuring a machine learning model to process representations of training inputs generated by the candidate neural network and to generate predictions of the corresponding ground-truth outputs; and

determining a measure of performance of the machine learning model, wherein the measure of performance is used to determine a predicted performance of the candidate neural network on the first prediction task; and

selecting the proper subset of the plurality of candidate neural networks using the measures of performance of the machine learning models corresponding to the plurality of candidate neural networks.

6 . The method of claim 5 , wherein determining the measure of performance for a candidate neural network comprises determining the measure of performance using leave-one-out cross validation.

7 . The method of claim 5 , wherein the machine learning model is a nearest-neighbor classifier.

8 . The method of claim 1 , wherein:

each of the plurality of candidate neural networks was trained using a respective plurality of second training examples, wherein at least some of the respective pluralities of second training examples are different; and

selecting the proper subset of the plurality of candidate neural networks using the respective predicted performance on the first prediction task for each of the candidate neural networks comprises, for each of the plurality of candidate neural networks, determining a predicted similarity between i) the plurality of training examples and ii) the plurality of second training examples that was used to train the candidate neural network, wherein the predicted similarity is used to determine a predicted performance of the candidate neural network on the first prediction task.

9 . The method of claim 8 , wherein:

the plurality of training examples comprises a plurality of training inputs and corresponding ground-truth outputs;

the respective plurality of second training examples corresponding to each candidate neural network comprises a plurality of second training inputs and corresponding second ground-truth outputs; and

for each candidate neural network, determining a predicted similarity between i) the plurality of training examples and ii) the plurality of second training examples that were used to train the candidate neural network comprises one or more of:

determining a similarity between a distribution of the training inputs and a distribution of the second training inputs, or

determining a similarity between a distribution of the ground-truth outputs and a distribution of the second ground-truth outputs.

10 . The method of claim 8 , wherein:

the plurality of training examples comprises a plurality of training inputs;

the respective plurality of second training examples corresponding to each candidate neural network comprises a respective plurality of second training inputs; and

determining a respective predicted similarity between i) the plurality of training examples and ii) the plurality of second training examples that were used to train each candidate neural network comprises:

training, using each of the pluralities of second training inputs, a machine learning model to process a particular second training input and to generate a model output that comprises, for each candidate neural network, a respective likelihood value represented a predicted likelihood that the candidate neural network was trained using the particular second training input;

processing at least some of the plurality of training inputs using the machine learning model to generate respective model outputs; and

combining the respective model outputs to generate, for each candidate neural network, a respective similarity value representing the predicted similarity between i) the plurality of training examples and ii) the plurality of second training examples that were used to train the candidate neural network.

11 . The method of claim 8 , wherein:

each candidate neural network has been pre-trained to perform a same second prediction task;

the plurality of training examples comprises a plurality of training inputs;

the respective plurality of second training examples corresponding to each candidate neural network comprises a respective plurality of second ground-truth outputs; and

determining a respective predicted similarity between i) the plurality of training examples and ii) the plurality of second training examples that were used to train each candidate neural network comprises:

training a machine learning model to perform the same second prediction task using a training data set comprising the respective plurality of second training examples corresponding to each candidate neural network;

processing at least some of the plurality of training inputs using the machine learning model to generate respective model outputs; and

for each candidate neural network, determining a similarity between (i) the model outputs and (ii) the plurality of second ground-truth outputs corresponding to the candidate neural network.

12 . The method of claim 1 , wherein each of a plurality of particular candidate neural network has been pre-trained to perform a same second prediction task,

the pre-training comprising:

training one or more base neural networks to perform the same second prediction task using a training data set;

determining a plurality of strict subsets of the training data set; and

for each of the one or more base neural networks and for each of one or more respective strict subsets, fine-tuning the base neural network using the strict subset to generate a respective expert neural network.

13 . The method of claim 12 , wherein at least some of the plurality of strict subsets of the training data set correspond to respective categories of ground-truth labels in the training data set.

14 . The method of claim 1 , wherein, for at least a subset of the plurality of candidate neural networks, one or more of:

initial values for the model parameters of each candidate neural network in the subset were different;

each candidate neural network in the subset was trained using a respective different subset of a same training data set;

each candidate neural network in the subset has a different network architecture; or

each candidate neural network in the subset was trained to perform a respective different second prediction task.

15 . The method of claim 1 , wherein generating, for each candidate neural network in the proper subset, one or more fine-tuned neural networks comprises:

generating, for each candidate neural network in the proper subset, multiple fine-tuned neural networks by updating the model parameters of the candidate neural network according to each of a plurality of different sets of hyperparameter values.

16 . The method of claim 1 , wherein a number of candidate neural networks in the proper subset is determined using respective measures of predicted performance of the candidate neural networks on the first prediction task.

17 . The method of claim 16 , wherein a candidate neural network is selected to be in the subset if a difference between i) the measure of predicted performance of the candidate neural network and ii) a highest measure of predicted performance of all the candidate neural networks is less than a threshold.

18 . The method of claim 16 , wherein:

a maximum number of candidate neural networks in the proper subset is predetermined; and

if the number of candidate neural networks selected to be in the proper subset is less than the maximum number, then one or more additional fine-tuned neural networks are generated corresponding to respective candidate neural networks in the proper subset and according to respective different sets of hyperparameter values.

19 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one more computers to perform operations for training a neural network to perform a first prediction task, the operations comprising:

obtaining trained model parameters for each of a plurality of candidate neural networks, wherein each candidate neural network has been pre-trained to perform a respective second prediction task that is different from the first prediction task;

obtaining a plurality of training examples corresponding to the first prediction task;

prior to fine-tuning any of the plurality of candidate neural networks for the first prediction task:

predicting, for each plurality of candidate neural networks and using the plurality of training examples, a respective performance of the candidate neural network on the first prediction task, and

selecting a proper subset of the plurality of candidate neural networks using the respective predicted performance on the first prediction task for each of the candidate neural networks;

after selecting the proper subset, fine-tuning only the candidate neural networks in the proper subset for the first prediction task by generating, for each candidate neural network in the proper subset, one or more fine-tuned neural networks, wherein each of the one or more fine-tuned neural networks is generated by updating the model parameters of the candidate neural network using the plurality of training examples; and

determining model parameters for the neural network using the one or more respective fine-tuned neural networks.

20 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one more computers to perform operations for training a neural network to perform a first prediction task, the operations comprising:

obtaining trained model parameters for each of a plurality of candidate neural networks, wherein each candidate neural network has been pre-trained to perform a respective second prediction task that is different from the first prediction task;

obtaining a plurality of training examples corresponding to the first prediction task;

prior to fine-tuning any of the plurality of candidate neural networks for the first prediction task:

predicting, for each plurality of candidate neural networks and using the plurality of training examples, a respective performance of the candidate neural network on the first prediction task, and

selecting a proper subset of the plurality of candidate neural networks using the respective predicted performance on the first prediction task for each of the candidate neural networks;

after selecting the proper subset, fine-tuning only the candidate neural networks in the proper subset for the first prediction task by generating, for each candidate neural network in the proper subset, one or more fine-tuned neural networks, wherein each of the one or more fine-tuned neural networks is generated by updating the model parameters of the candidate neural network using the plurality of training examples; and

determining model parameters for the neural network using the one or more fine-tuned neural networks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2021
From: PEREZ, JOAN PUIGCERVER I; MUSTAFA, BASIL; SUSANO PINTO, ANDRE; RIQUELME RUIZ, CARLOS; HOULSBY, NEIL MATTHEW TINMOUTH; KEYSERS, DANIEL M.
To: GOOGLE LLC
Reel/Frame 057794/0406 →
Continuity (2)
Provisional Application 63087104 · Oct 2, 2020
Related Publication 20220108171A1 · Apr 7, 2022
References Cited (127)
US 11393182B2 · Jacquot · 2022 [cited by examiner]
US 20070011114A1 · Chen · 2007 [cited by examiner]
US 20190197395A1 · Kibune · 2019 [cited by examiner]
US 20200349416A1 · Yang · 2020 [cited by examiner]
US 20220051079A1 · Laszlo · 2022 [cited by examiner]
US 20220058478A1 · Kuo · 2022 [cited by examiner]
US 20220121902A1 · Hann · 2022 [cited by examiner]
US 20220180199A1 · Xu · 2022 [cited by examiner]
Xia et al. (“Transferring Ensemble Representations Using Deep Convolutional Neural Networks for Small-Scale Image Classification”, vol. 7, 2019 pp. 168175-168186) (Year: 2019). [cited by examiner]
Macko et al. (“Improving Neural Architecture Search Image Classifiers via Ensemble Learning”, Nov. 20, 2019 ) (Year: 2019). [cited by examiner]
Li et al. (“Predicting the Number of Nearest Neighbor for kNN Classifier”, Nov. 20, 2019 ) (Year: 2019). [cited by examiner]
Raschka et al. (“Model Evaluation, Model Selection, and Algorithm Selection in Machine Learning” 2018) (Year: 2018). [cited by examiner]
Li et al. (“Machine Learning Based Online Performance Prediction for Runtime Parallelization and Task Scheduling”, 2009 IEEE International Symposium on Performance Analysis of Systems and Software, Boston, MA, USA, 2009… [cited by examiner]
Acharya et al, “Transfer learning with cluster ensembles” ICML, 2012, 11 pages. [cited by applicant]
Altman, “An introduction to kernel and nearest-neighbor nonparametric regression” The American Statistician, 12 pages. [cited by applicant]
Argyriou et al, “Multi-task feature learning” NIPS, 2006, 8 pages. [cited by applicant]
Bachman et al, “Learning with pseudo-ensembles” NIPS, 2014, 9 pages. [cited by applicant]
Barbu et al, “Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models” ICML, 2019, 11 pages. [cited by applicant]
Ben-David et al, “A theory of learning from different domains” Machine Learning, 2010, 25 pages. [cited by applicant]
Ben-David et al, “Analysis of representations for domain adaptation” NIPS, 2006, 8 pages. [cited by applicant]
Bengio, “Deep learning of representations for unsupervised and transfer learning” JMLR, 2012, 21 pages. [cited by applicant]
Berg et al, “Birdsnap: Large-scale fine-grained visual categorization of birds” CVPR, 2014, 8 pages. [cited by applicant]
Blitzer et al, “Learning bounds for domain adaptation” NIPS, 2007, 8 pages. [cited by applicant]
Bossard et al, “Food-101-mining discriminative components with random forests” ECCV, 2014, 16 pages. [cited by applicant]
Caruana et al, “Ensemble selection from libraries of models” ICML, 2004, 9 pages. [cited by applicant]
Caruana et al, “Multitask learning” Machine Learning, 1997, 255 pages. [cited by applicant]
Chen et al, “A closer look at few-shot classification” arXiv, 2019, 17 pages. [cited by applicant]
Cheng et al, “Remote sensing image scene classification: Benchmark and state of the art” arXiv, 2017, 17 pages. [cited by applicant]
Cheung et al, “Superposition of many models into one” NIPS, 2019, 18 pages. [cited by applicant]
Chollet, “Xception: Deep learning with depthwise separable convolutions” IEEE, 2017, 8 pages. [cited by applicant]
Cimpoi et al, “Describing textures in the wild” CVPR, 2014, 8 pages. [cited by applicant]
Dai et al, “Boosting for transfer learning” ICML, 2007, 8 pages. [cited by applicant]
Deng et al, “ImageNet: A large-scale hierarchical image database” IEEE, 2009, 8 pages. [cited by applicant]
Dhillon et al, “A baseline for few-shot image classification” arXiv, 2020, 20 pages. [cited by applicant]
Djolonga et al, “On robustness and transferability of convolutional neural networks” arXiv, 2020, 24 pages. [cited by applicant]
Donahue et al, “Decaf: a deep convolutional activation feature for generic visual recognition” ICML, 2014, 9 pages. [cited by applicant]
Du et al, “Hypothesis transfer learning via transformation functions” NIPS, 2017, 11 pages. [cited by applicant]
Dvornik et al, “Diversity with cooperation: ensemble methods for few-shot classification” arXiv, 2019, 12 pages. [cited by applicant]
Dvornik et al, “Selecting relevant features from a universal representation for few-shot classification” arXiv, 2020, 24 pages. [cited by applicant]
Eigen et al, “Learning factored representations in a deep mixture of experts” arXiv, 2013, 8 pages. [cited by applicant]
Evgeniou et al, “Learning multiple tasks with kernel methods” JMLR, 2005, 23 pages. [cited by applicant]
Fedus et al, “Hyperbolic discounting and learning over multiple horizons” arXiv, 2019, 28 pages. [cited by applicant]
Fei-Fei et al, “Learning generative visual models from few training examples: an incremental bayesian approach tested on 101 object categories” CVPR, 2007, 12 pages. [cited by applicant]
Fellbaum, “Wordnet” encyclopedia of applied linguistics, 2012, 8 pages. [cited by applicant]
Fort et al, “Deep ensembles: a loss landscape perspective” arXiv, 2019, 14 pages. [cited by applicant]
Fukushima, “Neocognitron: a self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position” Biological cybernetics, 1980, 10 pages. [cited by applicant]
Geiger et al, “are we ready for autonomous driving? The kitti vision benchmark suite” IEEE, 2012, 8 pages. [cited by applicant]
github.com [online], “dSprites: dis-entanglement testing sprites dataset” 2017, retrieved on Feb. 9, 2022, retrieved from URL <https://github.com/deepmind/dsprites-dataset/, 4 pages. [cited by applicant]
Glorot et al, “Deep sparse rectifier neural networks” ICAIS, 2011, 9 pages. [cited by applicant]
He et al, “AutoML: A survey of the state-of-the-art” arXiv, 2019, 17 pages. [cited by applicant]
He et al, “Deep residual learning for image recognition” IEEE, 2016, 9 pages. [cited by applicant]
He et al, “Identity mappings in deep residual networks” arXiv, 2016, 15 pages. [cited by applicant]
Helber et al, “EuroSAT: A novel dataset and deep learning benchmark for land use and land cover classification ” arXiv, 2019, 9 pages. [cited by applicant]
Hendrycks et al, “Benchmarking neural network robustness to common corruptions and perturbations” arXiv, 2019, 16 pages. [cited by applicant]
Hendrycks et al, “Natural adversarial examples” arXiv, 2019, 12 pages. [cited by applicant]
Hendrycks et al, “The many faces of robustness: a critical analysis of out-of-distribution generalization” arXiv, 2020, 18 pages. [cited by applicant]
Hinton et al, “Distilling the knowledge in a neural network” arXiv, 2015, 9 pages. [cited by applicant]
Houlsby et al, “Parameter-efficient transfer learning for NLP” arXiv, 2019, 13 pages. [cited by applicant]
Jacob et al, “Clustered multi-task learning: a convex formulation” NIPS, 2008, 8 pages. [cited by applicant]
Jacobs et al, “Learning piecewise control strategies in a modular neural network architecture” IEEE, 1993, 9 pages. [cited by applicant]
Johnson et al, “CLEVR: a diagnostic dataset for compositional language and elementary visual reasoning” CVPR, 2017, 10 pages. [cited by applicant]
kaggle.com [online], “diabetic retinopathy detection” 2015, retrieved on Feb. 9, 2022, retrieved from URL <https://www.kaggle.com/c/diabetic-retinopathy-detection/data>, 2 pages. [cited by applicant]
Kim et al, “Tree-guided group lasso for multi-response regression with structured sparsity, with an application to eqtl mapping” Annals of Applied Statistics, 2012, 24 pages. [cited by applicant]
Kolesnikov et al, “Big transfer (BiT): General visual representation learning” ECCV, 2020, 17 pages. [cited by applicant]
Krause et al, “3D object representations for fine-grained categorization” IEEE, 2013, 8 pages. [cited by applicant]
Krizhevsky, “Learning multiple layers of features from tiny images” Technical report, 2009, 60 pages. [cited by applicant]
Kuzborskij et al, “Stability and hypothesis transfer learning” ICML, 2013, 9 pages. [cited by applicant]
Laine et al, “Temporal ensembling for semi-supervised learning” arXiv, 2017, 13 pages. [cited by applicant]
Lakshminarayanan et al, “Simple and scalable predictive uncertainty estimation using deep ensembles” NIPS, 2017, 12 pages. [cited by applicant]
LeCun et al, “Backpropagation applied to handwritten zip code recognition” Neural computation, 1989, 11 pages. [cited by applicant]
LeCun et al, “Learning methods for generic object recognition with invariance to pose and lighting” IEEE, 2004, 8 pages. [cited by applicant]
Lee et al, “Why M heads are better than one: training a diverse ensemble of deep networks” arXiv, 2015, 18 pages. [cited by applicant]
Liu et al, “Representation learning using multi-task deep neural networks for semantic classification and information retrieval” Proceedings of the Annual Conference of the North American Chapter of the Association for … [cited by applicant]
Long et al, “Deep transfer learning with joint adaptation networks” ICML, 2017, 10 pages. [cited by applicant]
Long et al, “Learning transferable features with deep adaptation networks” arXiv, 2015, 9 pages. [cited by applicant]
Lounici et al, “Taking advantage of sparsity in multi-task learning” arXiv, 2009, 21 pages. [cited by applicant]
Maji et al, “Fine-grained visual classification of aircraft” arXiv, 2013, 6 pages. [cited by applicant]
Mansour et al, “Domain adaptation: Learning bounds and algorithms” arXiv, 2009, 16 pages. [cited by applicant]
Misra et al, “Cross-stitch networks for multi-task learning” IEEE, 2016, 10 pages. [cited by applicant]
Netzer et al, “Reading digits in natural images with unsupervised feature learning” NIPS, 2011, 9 pages. [cited by applicant]
Neyshabur et al, “What is being transferred in transfer learning?” arXiv, 2020, 34 pages. [cited by applicant]
Ngiam et al, “Domain adaptive transfer learning with specialist models” arXiv, 2018, 10 pages. [cited by applicant]
Nilsback et al, “A visual vocabulary for flower classification” IEEE, 2006, 8 pages. [cited by applicant]
Oquab et al, “Learning and transferring mid-level image representations using convoloution neural networks” IEEE, 2014, 8 pages. [cited by applicant]
Pan et al, “A survey on transfer learning” IEEE, 2009, 15 pages. [cited by applicant]
Pan et al, “Domain adaptation via transfer component analysis” IEEE, 2011, 12 pages. [cited by applicant]
Pardoe et al, “Boosting for regression transfer” ICML, 2010, 8 pages. [cited by applicant]
Parkhi et al, “Cats and dogs” IEEE, 2012, 8 pages. [cited by applicant]
phytorch.org [online], “Pytorch Hub” 2019, retrieved on Feb. 9, 2022, retrieved from URL <https://pytorch.org/hub/, 2 pages. [cited by applicant]
Puigcerver et al, “Scalable transfer learning with expert models” arXiv, 2020, 27 pages. [cited by applicant]
Raghu et al, “Transfusion: understanding transfer learning for medical imaging” NIPS, 2019, 11 pages. [cited by applicant]
Razavian et al, “CNN feature off-the-shelf: an astounding baseline for recognition” IEEE, 2014, 8 pages. [cited by applicant]
Real et al, “Regularized evolution for image classifier architecture search” AAAI, 2019, 10 pages. [cited by applicant]
Real et al, “YouTube-BoundingBoxes: A large high-precision human-annotated data set for object detection in video” CVPR, 2017, 10 pages. [cited by applicant]
Rebuffi et al, “Efficient parametrization of multi-domain deep neural networks” IEEE, 2018, 9 pages. [cited by applicant]
Rebuffi et al, “Learning multiple visual domains with residual adapters” NIPS, 2017, 11 pages. [cited by applicant]
Recht et al, “Do ImageNet classifiers generalize to ImageNet?” ICML, 2019, 12 pages. [cited by applicant]
Rosenfeld et al, “Incremental learning through deep adaptation” arXiv, 2017, 10 pages. [cited by applicant]
Rusu et al, “Progressive neural networks” arXiv, 2016, 14 pages. [cited by applicant]
Seni et al, “Ensemble methods in data mining: improving accuracy through combining predictions” Morgan and Claypool Publishers, 2010, 126 pages. [cited by applicant]
Shankar et al, “Do image classifiers generalize across time?” IMLR, 2019, 23 pages. [cited by applicant]
Shazeer et al, “Outrangeously large neural networks: the sparsely-grated mixture-of-experts layer” arXiv, 2017, 19 pages. [cited by applicant]
Stickland et al, “Diverse ensembles improve calibration ”arXiv, 2020, 6 pages. [cited by applicant]
Sun et al, “Revisiting unreasonable effectiveness of data in deep learning era” ICCV, 2017, 10 pages. [cited by applicant]
Sun et al, “Transfer learning with part-based ensembles” Multiple Classifier Systems, 2013, 12 pages. [cited by applicant]
Szegedy et al, “Going deeper with convolutions” IEEE, 2015, 9 pages. [cited by applicant]
Szegedy et al, “Rethinking the inception architecture for computer vision” IEEE, 2016, 9 pages. [cited by applicant]
Tan et al, “A survey on deep transfer learning” arXiv, 2018, 10 pages. [cited by applicant]
Tzeng et al, “Deep domain confusion: Maximizing for domain invariance” arXiv, 2014, 9 pages. [cited by applicant]
Veeling et al, “Rotation equivariant CNNs for digital pathology” arXiv, 2018, 8 pages. [cited by applicant]
Wan et al, “Bi-weighting domain adaptation for cross-language text classification” IJCAI, 2011, 6 pages. [cited by applicant]
Wang, “Theoretical guarantees of transfer learning” arXiv, 2018, 11 pages. [cited by applicant]
Webb et al, “To ensemble or not ensemble: when does end-to-end training fail?” CVPR, 2019, 16 pages. [cited by applicant]
Weiss et al, “A survey of transfer learning” Journal of Big data, 2016, 40 pages. [cited by applicant]
Wen et al, “Batchensemble: an alternative approach to efficient ensemble and lifelong learning” ICLR, 2020, 20 pages. [cited by applicant]
Wenzel et al, “Hyperparameter ensembles for robustness and uncertainty quantification” arXiv, 2020, 31 pages. [cited by applicant]
Wistuba et al, “Automatic frankensteining: creating complex ensembles autonomously” ICDM, 2017, 9 pages. [cited by applicant]
Wu et al, “Group normalization” ECCV, 2018, 17 pages. [cited by applicant]
Xavier-Junior et al, “A novel evolutionary algorithm for automated machine learning focusing on classifier ensembles” BRACIS, 2018, 6 pages. [cited by applicant]
Xie et al, “Self-tanning with noisy student improves imagenet classification” arXiv, 2019, 13 pages. [cited by applicant]
Xu et al, “A unified framework for metric transfer learning” IEEE, 2017, 14 pages. [cited by applicant]
Yalniz et al, “Billion-scale semi-supervised learning for image classification” arXiv, 2019, 12 pages. [cited by applicant]
Yan et al, “Neural data server: A large-scale search engine for transfer learning data”, 2020, 10 pages. [cited by applicant]
Yosinski et al, “How transferable are features in deep neural networks?” NIPS, 2014, 14 pages. [cited by applicant]
Zhai et al, “A large-scale study of representation learning with the visual task adaptation benchmark” arXiv, 2020, 33 pages. [cited by applicant]
Zhai et al, “The visual task adaptation benchmark” arXiv, 2019, 35 pages. [cited by applicant]
Zhang et al, “Facial landmark detection by deep multi- task learning” ECCV, 2014, 15 pages. [cited by applicant]