IP Library Granted Patent US 12,657,259
Granted Patent B1
US 12,657,259 · App. 17/127,680 · Granted Jun 16, 2026

Neural network training technique

Inventors: Zhiding Yu (Santa Clara, CA); Wuyang Chen (College Station, TX); Shalini De Mello (San Francisco, CA); Sifei Liu (San Diego, CA); Jose Manuel Alvarez Lopez (Mountain View, CA); Anima Anandkumar (Santa Clara, CA)
Assignee: NVIDIA Corporation
G06F18/2148G06F18/217G06F18/22G06N3/045G06N3/08G06V10/95
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,259
App. No.
17/127,680
Granted
Jun 16, 2026
Kind
B1
Abstract

Apparatuses, systems, and techniques to train neural networks to perform image processing tasks. In at least one embodiment, one or more second neural networks are used to train one or more first neural networks based, at least in part, on a first object type in one or more images and a second object type in the one or more images, in parallel.

Claims (48)

1 . One or more processors, comprising:

circuitry to train one or more first neural networks based, at least in part, on a comparison between one or more features identified, in one or more portions of synthetic data, by the one or more first neural networks and one or more features identified, in the one or more portions of synthetic data, by one or more pre-trained second neural networks, wherein the one or more features identified by the one or more first neural networks include a first object type in one or more images, the one or more features identified by the one or more pre-trained second neural networks include a second object type in the one or more images, and wherein the second object type is selected to be a negative sample of the first object type to reduce overfitting during training of the one or more first neural networks on positive samples of the first object type.

2 . The one or more processors of claim 1 , wherein the synthetic data includes one or more synthetic images.

3 . The one or more processors of claim 1 , wherein the one or more pre-trained second neural networks are pre-trained based on non-synthetic data.

4 . The one or more processors of claim 1 , wherein the one or more pre-trained second neural networks are frozen during training of the first one or more neural networks.

5 . The one or more processors of claim 1 , the circuitry to at least:

use the one or more first neural networks to generate a first embedding of an image;

use the one or more pre-trained second neural networks to generate a second embedding of the image; and

generate a loss value component indicative of a distance between the first and second embeddings.

6 . The one or more processors of claim 1 , the circuitry to at least:

use the one or more first neural networks to generate a first embedding of a first image;

use the one or more pre-trained second neural networks to generate a second

embedding of a second image; and

generate a loss value component based, at least in part, on a distance between the first and second embeddings, wherein the loss value component is reduced based on the distance.

7 . The one or more processors of claim 1 , wherein the one or more features identified by the one or more first neural networks include a first object type in one or more images, the one or more features identified by the one or more pre-trained second neural networks include a second object type in the one or more images, and wherein the second object type is selected to be different than the first object type.

8 . The one or more processors of claim 1 , wherein the one or more first neural networks and the one or more pre-trained second neural networks generate output based, at least in part, on an attention matrix.

9 . A system, comprising:

one or more processors to train one or more first neural networks based, at least in part, on a comparison between one or more features identified, in one or more portions of synthetic data, by the one or more first neural networks and one or more features identified, in the one or more portions of synthetic data, by one or more pre-trained second neural networks, wherein the one or more features identified by the one or more first neural networks include a first object type in one or more images, the one or more features identified by the one or more pre-trained second neural networks include a second object type in the one or more images, and wherein the second object type is selected to be a negative sample of the first object type to reduce overfitting during training of the one or more first neural networks on positive samples of the first object type.

10 . The system of claim 9 , wherein the one or more processors are to select the second object type based, at least in part, on one or more differences between the second object type and the first object type.

11 . The system of claim 9 , wherein the one or more pre-trained second neural networks are pre-trained based on images of real objects.

12 . The system of claim 9 , wherein the one or more pre-trained second neural networks are frozen during a training of the first one or more neural networks.

13 . The system of claim 9 , wherein the one or more processors to cause the system to at least generate a loss value component based on a distance between an embedding of an object of a first type by the one or more first neural networks and an embedding of the object of the first type by the one or more pre-trained second neural networks.

14 . The system of claim 9 , wherein the one or more processors to cause the system to at least generate a loss value component reduced based on distance between a first embedding and a second embedding, wherein the one or more features identified by the one or more first neural networks include a first object type in one or more images, wherein the one or more features identified by the one or more pre-trained second neural networks include a second object type in the one or more images, wherein the first embedding is generated by the one or more first neural networks based on a first image comprising a depiction of an object of the first object type, and wherein the second embedding is generated by the one or more pre-trained second neural networks based on a second image comprising a depiction of an object of the second object type.

15 . The system of claim 9 , wherein the one or more first neural networks and the one or more pre-trained second neural networks generate output based, at least in part, on an attention matrix.

16 . A non-transitory machine-readable medium having stored thereon a set of instructions, which cause one or more processors to at least:

train one or more first neural networks based, at least in part, on a comparison between one or more features identified, in one or more portions of synthetic data, by the one or more first neural networks and one or more features identified, in the one or more portions of synthetic data, by one or more pre-trained second neural networks, wherein the one or more features identified by the one or more first neural networks include a first object type in one or more images, the one or more features identified by the one or more pre-trained second neural networks include a second object type in the one or more images, and wherein the second object type is selected to be a negative sample of the first object type to reduce overfitting during training of the one or more first neural networks on positive samples of the first object type.

17 . The non-transitory machine-readable medium of claim 16 , having stored thereon a further set of instructions, which cause one or more processors to at least:

use the one or more pre-trained second neural networks to train the one or more first neural networks based, at least in part, on a first object type in one or more images and a second object type in the one or more images; and

select the second object type based, at least in part, on one or more differences between the second object type and the first object type.

18 . The non-transitory machine-readable medium of claim 16 , wherein the one or more pre-trained second neural networks are pre-trained based on images of real objects and frozen during training of the one or more first neural networks.

19 . The non-transitory machine-readable medium of claim 16 , having stored thereon a further set of instructions, which cause one or more processors to at least:

generate, by the one or more first neural networks, a first embedding of an image generate, by the one or more pre-trained second neural networks, a second embedding of the image; and

generate a loss value based, at least in part, on a distance between the first and second embeddings.

20 . The non-transitory machine-readable medium of claim 16 , having stored thereon a further set of instructions, which cause one or more processors to at least:

use the one or more pre-trained second neural networks to train the one or more first neural networks based, at least in part, on a first object type in one or more images and a second object type in the one or more images;

generate, by the one or more first neural networks, a first embedding of a first image of the one or more images, the first image comprising the first object type;

generate, by the one or more pre-trained second neural networks, a second embedding of a second image of the one or more images, the second image comprising the second object type;

and generate a loss value based, at least in part, on a distance between the first and second embeddings.

21 . The non-transitory machine-readable medium of claim 16 , wherein the one or more first neural networks are trained based, at least in part, on a contrastive loss signal.

22 . The non-transitory machine-readable medium of claim 16 , wherein the one or more first neural networks and the one or more pre-trained second neural networks generate output based, at least in part, on an attention matrix.

23 . A computing device, comprising:

one or more processors to perform an image processing task based, at least in part, on one or more first neural networks trained using one or more pre-trained second neural networks, based, at least in part on a comparison between one or more features identified, in one or more portions of synthetic data, by the one or more first neural networks and one or more features identified, in the one or more portions of synthetic data, by one or more pre-trained second neural networks, wherein the one or more features identified by the one or more first neural networks include a first object type in one or more images, the one or more features identified by the one or more pre-trained second neural networks include a second object type in the one or more images, and wherein the second object type is selected to be a negative sample of the first object type to reduce overfitting during training of the one or more first neural networks on positive samples of the first object type.

24 . The computing device of claim 23 , wherein the image processing task comprises recognition of an object.

25 . The computing device of claim 23 , wherein the one or more processors to use the one or more images in parallel to generate a loss signal for training the one or more first neural networks.

26 . The computing device of claim 23 , wherein the one or more pre-trained second neural networks are pre-trained based on images of real objects and frozen during a training of the one or more first neural networks.

27 . The computing device of claim 23 , wherein the one or more processors to cause the computing device to at least generate a loss value component increased in proportion to a distance between an embedding of an object of a first type by the one or more first neural networks and an embedding of the object of the first type by the one or more pre-trained second neural networks.

28 . The computing device of claim 23 , wherein the one or more processors to cause the computing device to at least generate a loss value component reduced in proportion to a distance between a first embedding and a second embedding.

29 . The computing device of claim 23 , wherein the one or more first neural networks and the one or more pre-trained second neural networks generate output based, at least in part, on an attention matrix.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2020
From: YU, ZHIDING; CHEN, WUYANG; DE MELLO, SHALINI; LIU, SIFEI; ALVAREZ LOPEZ, JOSE MANUEL; ANANDKUMAR, ANIMA
To: NVIDIA CORPORATION
Reel/Frame 054743/0503 →
References Cited (121)
US 11093789B2 · Wang · 2021 [cited by examiner]
US 11195051B2 · Huang · 2021 [cited by examiner]
US 11238300B2 · Karianakis · 2022 [cited by examiner]
US 11367268B2 · Yang · 2022 [cited by examiner]
US 20160078339A1 · Li · 2016 [cited by examiner]
US 20180307969A1 · Shibahara · 2018 [cited by applicant]
US 20190030371A1 · Han · 2019 [cited by examiner]
US 20190220738A1 · Flank · 2019 [cited by applicant]
US 20200097742A1 · Ratnesh Kumar · 2020 [cited by examiner]
US 20200134506A1 · Wang · 2020 [cited by examiner]
US 20200226421A1 · Almazan · 2020 [cited by examiner]
US 20200334538A1 · Meng · 2020 [cited by examiner]
US 20210034983A1 · Asano et al. · 2021 [cited by applicant]
US 20210279595A1 · Sridhar · 2021 [cited by examiner]
CN 106170800A · 2016 [cited by applicant]
Li, “Learning without Forgetting,” arxiv in 1606.09282, Feb. 14, 2017. (Year: 2017). [cited by examiner]
Meng, “Conditional Teacher-student Learning,” arxiv 1904.12399, Apr. 28, 2019. (Year: 2019). [cited by examiner]
Bhardwaj, “Dream Distillation: A Data-Independent Model Compression Framework”, arxiv 1905.07072 May 17, 2019 (Year: 2019). [cited by examiner]
Chen et al., “A Simple Framework for Contrastive Learning of Visual Representations, ” Jul. 1, 2020, 20 pages. [cited by applicant]
Chen et al., “Automated Synthetic-to-Real Generalization,” Jul. 14, 2020, 11 pages. [cited by applicant]
Chen et al., “Improved Baselines with Momentum Contrastive Learning,” Mar. 9, 2020, 3 pages. [cited by applicant]
Chen et al., “Learning Semantic Segmentation from Synthetic Data: A Geometrically Guided Input-Output Adaptation Approach,” CVPR, 2019, 10 pages. [cited by applicant]
Chen et al., “Rethinking Atrous Convolution for Semantic Image Segmentation,” Dec. 5, 2017, 14 pages. [cited by applicant]
Chen et al., “Road: Reality Oriented Adaptation for Semantic Segmentation of Urban Scenes,” CVPR, 2018, 10 pages. [cited by applicant]
Cordts et al., “The Cityscapes Dataset for Semantic Urban Scene Understanding,” IEEE Conference on ComputerVision and Pattern Recognition, 2016, 11 pages. [cited by applicant]
Cubuk et al., “Randaugment: Practical Automated Data Augmentation with a Reduced Search Space,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 14-19, 2020, 10 pages. [cited by applicant]
Deng et al. “Imagenet: A Large-Scale Hierarchical Image Database,” ICLR, 2009, 8 pages. [cited by applicant]
Dosovitskiy et al., “Carla: An Open Urban Driving Simulator,” CoRL, 2017, 16 pages. [cited by applicant]
Dosovitskiy et al., “FlowNet: Learning Optical Flow with Convolutional Networks,” ICCV, 2015, 9 pages. [cited by applicant]
Gan et al., “Learning Attributes Equals Multi-Source Domain Generalization,” CVPR, 2016, 11 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” CVPR, 2016, 9 pages. [cited by applicant]
He et al., “Momentum Contrast for Unsupervised Visual Representation Learning,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, 10 pages. [cited by applicant]
Hjelm et al., “Learning Deep Representations by Mutual Information Estimation and Maximization,” Oct. 3, 2018, 21 pages. [cited by applicant]
Hoffman et al., “CyCADA: Cycle-Consistent Adversarial Domain Adaptation,” International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Hénaff et al., “Data-Efficient Image Recognition with Contrastive Predictive Coding,” Dec. 6, 2019, 15 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Kuhnke et al., “Deep Head Pose Estimation Using Synthetic Images and Partial Adversarial Domain Adaption for Continuous Label Spaces,” Proceedings of the IEEE International Conference on Computer Vision, 2019, 10 pages. [cited by applicant]
Li et al., “Deeper, Broader and Artier Domain Generalization,” ICCV, 2017, 9 pages. [cited by applicant]
Li et al., “Learning to Generalize: Metalearning for Domain Generalization,” AAAI, 2018, 8 pages. [cited by applicant]
Li et al., “Learning without Forgetting,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(12): Feb. 14, 2017, 13 pages. [cited by applicant]
Lin et al., “Microsoft COCO: Common Objects in Context,” European Conference on Computer Vision, Jul. 5, 2014, 14 pages. [cited by applicant]
Liu et al., “Feature-level Frankenstein: Eliminating Variations for Discriminative Recognition,” CVPR, 2019, 10 pages. [cited by applicant]
Liu et al., “Learning Towards Minimum Hyperspherical Energy,” Advances in Neural Information Processing Systems, 2018, 12 pages. [cited by applicant]
Long et al., “Fully Convolutional Networks for Semantic Segmentation,” CVPR, 2015, 10 pages. [cited by applicant]
Mayer et al., “A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation,” CVPR, 2016, 9 pages. [cited by applicant]
Misra et al., “Self-Supervised Learning of Pretext-Invariant Representations,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, 11 pages. [cited by applicant]
Muandet et al., “Domain Generalization via Invariant Feature Representation,” International Conference on Machine Learning, 2013, 9 pages. [cited by applicant]
Oord et al., “Representation Learning with Contrastive Predictive Coding,” Jul. 10, 2018, 13 pages. [cited by applicant]
Pan et al., “Two at Once: Enhancing Learning and Generalization Capacities via IBN-Net,” ECCV, 2018, 16 pages. [cited by applicant]
Park et al., “Contrastive Learning for Unpaired Image-to-Image Translation,” Aug. 20, 2020, 29 pages. [cited by applicant]
Peng et al., “VisDA: The Visual Domain Adaptation Challenge,” Nov. 29, 2017, 17 pages. [cited by applicant]
Richter et al., “Playing for Benchmarks,” ICCV, 2017, 10 pages. [cited by applicant]
Richter et al., “Playing for Data: Ground Truth from Computer Games,” ECCV, 2016, 17 pages. [cited by applicant]
Ros et al., “The Synthia Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes,” CVPR, 2016, 10 pages. [cited by applicant]
Saito et al., “Adversarial Dropout Regularization,” ICLR, 2018, 15 pages. [cited by applicant]
Savva et al., “Habitat: A Platform for Embodied AI Research,” ICCV, 2019, 9 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Tian et al., “Contrastive Multiview Coding,” Oct. 21, 2019, 21 pages. [cited by applicant]
Tian et al., “What Makes for Good Views for Contrastive Learning,” Dec. 18, 2020, 24 pages. [cited by applicant]
Tsai et al., “Learning to Adapt Structured Output Space for Semantic Segmentation,” CVPR, 2018, 10 pages. [cited by applicant]
U.S. Appl. No. 16/859,360, “Neural Network Training Technique,” filed Apr. 27, 2020. [cited by applicant]
Wang et al., “Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere,” Nov. 10, 2020, 41 pages. [cited by applicant]
Wu et al., “3D Shapenets: A Deep Representation for Volumetric Shapes,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 1912-1920. [cited by applicant]
Wu et al., “Unsupervised Feature Learning via Nonparametric Instance Discrimination,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Yue et al., “Domain Randomization and Pyramid Consistency: Simulation-to-Real Generalization without Accessing Target Domain Data,” Proceedings of the IEEE International Conference on Computer Vision, 2019, 11 pages. [cited by applicant]
Zenke et al., “Continual Learning through Synaptic Intelligence,” ICML, 2017, 9 pages. [cited by applicant]
Zhang et al., “Self-produced Guidance for Weakly-Supervised Object Localization,” Proceedings of the European Conference on Computer Vision, 2018, 17 pages. [cited by applicant]
Zhou et al., “Learning Deep Features for Discriminative Localization,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, 9 pages. [cited by applicant]
Zou et al., “Confidence Regularized Self-Training,” ICCV, 2019, 10 pages. [cited by applicant]
Zou et al., “Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training,” ECCV, 2018, 17 pages. [cited by applicant]
Geirhos et al., “Imagenet-trained CNNs are Biased Towards Texture; Increasing Shape Bias Improves Accuracy and Robustness,” ICLR, 2019, 22 pages. [cited by applicant]
Kirkpatrick et al., “Overcoming Catastrophic Forgetting in Neural Networks,” Proceedings of the National Academy of Sciences, 114(13): 2017, 6 pages. [cited by applicant]
Real et al., “YouTube-BoundingBoxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video,” CVPR, 2017, 10 pages. [cited by applicant]
Scott, “Multivariate Density Estimation: Theory, Practice, and Visualization,” John Wiley & Sons, 2015, 10 pages. [cited by applicant]
Tarvainen et al., “Mean Teachers are Better Role Models: Weight-averaged Consistency Targets Improve Semi-supervised Deep Learning Results,”Neural Information Processing Systems, 2017, 10 pages. [cited by applicant]
Tian et al., “Contrastive Representation Distillation,” 2020, 19 pages. [cited by applicant]
Andrychowicz et al., “Learning to Learn by Gradient Descent by Gradient Descent,” Nov. 16, 2016, 17 pages. [cited by applicant]
Bae, Ji-Hoon, et al. “Densely distilled flow-based knowledge transfer in teacher-student framework for image classification.” IEEE Transactions on Image Processing 29 (2020): 5698-5710. (Year: 2020), 14 pages. [cited by applicant]
Baik et al., “Learning to Forget for Meta-Learning,” Jun. 13, 2019, 17 Pages. [cited by applicant]
Bhardwaj et al., “Dream Distillation: A Data-Independent Model Compression Framework,” ICML Workshop, May 17, 2019, 4 pages. [cited by applicant]
Chen et al., “Learning to Learn Without Gradient Descent by Gradient Descent,” Proceedings of the 34th International Conference on Machine Learning, vol. 70, 2017, 9 pages. [cited by applicant]
Chen et al., “Road: Reality Oriented Adaptation for Semantic Segmentation of Urban Scenes,” Apr. 7, 2018, 10 pages. [cited by applicant]
Coumans et al., “PyBullet, A Python Module for Physics Simulation for Games,” Robotics and Machine Learning, retrieved http://pybullet.org, 2016, 10 pages. [cited by applicant]
Deng et al., “Imagenet: A Large-Scale Hierarchical Image Database,” CVPR, 2009, 8 pages. [cited by applicant]
Dundar et al., “Domain Stylization: A Fast Covariance Matching Framework Towards Domain Adaptation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(7): Jul. 2021, 13 pages. [cited by applicant]
Finn et al., “Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks,” Jul. 18, 2017, 13 Pages. [cited by applicant]
Gaidon et al., “Virtual Worlds as Proxy for Multi-Object Tracking Analysis,” CVPR, 2016, 10 pages. [cited by applicant]
Ghiasi et al., Dropblock: A regularization method for convolutional networks. In Advances in Neural Information Processing Systems, 2018, 11 pages. [cited by applicant]
Hoffman et al., “FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation,” Dec. 8, 2016, 9 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2021/028404, mailed Sep. 24, 2021, filed Apr. 21, 2021, 18 pages. [cited by applicant]
Johnson-Roberson et al., “Driving in the matrix: Can virtual worlds replace human-generated annotations for real world tasks?” Oct. 6, 2016, 8 pages. [cited by applicant]
Jung et al., “Less-forgetful Learning for Domain Expansion in Deep Neural Networks,” Nov. 16, 2017, 8 Pages. [cited by applicant]
Li et al., “Bidirectional Learning for Domain Adaptation of Semantic Segmentation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, 10 pages. [cited by applicant]
Li et al., “Learning to Generalize: Meta-Learning for Domain Generalization,” The Thirty-Second AAAI Conference on Artificial Intelligence, 2018, 8 pages. [cited by applicant]
Li et al., “Learning to Optimize,” Jun. 6, 2016, 9 pages. [cited by applicant]
Li et al., “Learning without Forgetting,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Feb. 14, 2017, 13 pages. [cited by applicant]
Long et al., “Learning Transferable Features with Deep Adaptation Networks,” May 27, 2015, 9 pages. [cited by applicant]
Lopez-Paz et al., “Gradient Episodic Memory for Continual Learning,” Advances in Neural Information Processing Systems, 2017, 10 pages. [cited by applicant]
Meng et al., “Conditional Teacher-Student Learning,” Apr. 28, 2019, 5 pages. [cited by applicant]
Nayak, Gaurav Kumar, et al. “Zero-shot knowledge distillation in deep networks.” International conference on machine learning. PMLR, 2019. (Year: 2019), 9 pages. [cited by applicant]
Office Action for Chinese Application No. 202180006001.5, mailed Apr. 30, 2025, 37 pages. [cited by applicant]
Pinheiro, “Unsupervised Domain Adaptation with Similarity Learning,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Real et al., “Youtube-boundingboxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, 10 pages. [cited by applicant]
Ren et al., “Learning to Reweight Examples for Robust Deep Learning,” May 5, 2019, 13 Pages. [cited by applicant]
Rumelhart et al., “Learning Representations by Back-Propagating Errors,” Nature, 323(6088): Oct. 9, 1986, 4 pages. [cited by applicant]
Saito et al., “Adversarial Dropout Regularization,” Nov. 7, 2017, 14 pages. [cited by applicant]
Saito et al., “Maximum Classifier Discrepancy for Unsupervised Domain Adaptation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Sankaranarayanan et al., “Generate to Adapt: Aligning Domains using Generative Adversarial Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Shafahi et al., “Adversarially Robust Transfer Learning,” ICLR, Feb. 21, 2020, 14 pages. [cited by applicant]
Shin et al., “Continual Learning with Deep Generative Replay,” Advances in Neural Information Processing Systems, 2017, 10 pages. [cited by applicant]
Sun et al., “Deep Coral: Correlation Alignment for Deep Domain Adaptation,” European Conference on Computer Vision, Jul. 6, 2016, 7 pages. [cited by applicant]
Thrun, “Lifelong Learning Algorithms,” Learning to Learn, 1998, 29 pages. [cited by applicant]
Tzeng et al., “Deep Domain Confusion: Maximizing for Domain Invariance,” Dec. 10, 2014, 9 pages. [cited by applicant]
Wichrowska et al., “Learned Optimizers that Scale and Generalize,” Proceedings of the 34th International Conference on Machine Learning, vol. 70, 2017, 10 pages. [cited by applicant]
Williams, “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning,” Machine Learning, 8(3-4): 1992, pp. 229-256. [cited by applicant]
Wu et al., “DCAN: Dual Channel-wise Alignment Networks for Unsupervised Scene Adaptation,” Proceedings of the European Conference on Computer Vision, 2018, 17 pages. [cited by applicant]
Xu et al., “Learning an Adaptive Learning Rate Schedule,” Sep. 20, 2019, 6 pages. [cited by applicant]
Yao et al., “Adversarial Feature Alignment: Avoid Catastrophic Forgetting in Incremental Task Lifelong Learning,” Oct. 24, 2019, 13 pages. [cited by applicant]
Yin et al., “Dreaming to Distill: Data-free Knowledge Transfer via DeepInversion,” Dec. 18, 2019, 15 pages. [cited by applicant]
Zhu et al., “Penalizing Top Performers: Conservative Loss for Semantic Segmentation Adaptation,” Proceedings of the European Conference on Computer Vision, 2018, 16 pages. [cited by applicant]