IP Library Granted Patent US 12,731,024
Granted Patent B2
US 12,731,024 · App. 16/859,360 · Granted Sep 8, 2026

Neural network training technique

Inventors: Zhiding Yu (Santa Clara, CA); Wuyang Chen (College Station, TX); Anima Anandkumar (Santa Clara, CA)
Assignee: NVIDIA Corporation
G06N3/08G06N3/044G06N3/045G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,024
App. No.
16/859,360
Granted
Sep 8, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to train one or more neural networks. In at least one embodiment, one or more neural networks are trained based, at least in part, on inferencing output from one or more second neural networks.

Claims (47)

1 . One or more processors, comprising:

circuitry to use one or more first neural networks to identify one or more objects in one or more synthetic images based, at least in part, on:

identifying a plurality of layers of the one or more first neural networks based on one or more inputs to the plurality of layers;

a scaling factor, computed based on calculating a loss corresponding to an inferencing response of one or more second neural networks in response to an input of one or more second synthetic images, wherein the one or more second neural networks are trained on non-synthetic images; and

one or more hyperparameters generated using the one or more second neural networks and adjusted for the one or more first neural networks using the scaling factor in a policy corresponding to the plurality of layers of the one or more first neural networks, wherein the one or more first neural networks are trained using the one or more hyperparameters while one or more parameters of the one or more second neural networks are frozen.

2 . The one or more processors of claim 1 , wherein during training, a learning rate of the one or more first neural networks is adjusted according to the scaling factor in the policy associated with the plurality of layers of the one or more first neural networks, based, at least in part, on an inferencing output of the one or more second neural networks.

3 . The one or more processors of claim 2 , wherein the learning rate is adjusted according to the scaling factor in the policy for a region of the one or more first neural networks, the region comprising the plurality of layers of the one or more first neural networks grouped based, at least in part, on input resolution.

4 . The one or more processors of claim 1 , wherein the circuitry adjusts training of the one or more first neural networks based, at least in part, on an inferencing output of the one or more second neural networks in response to the input of the one or more second synthetic images.

5 . The one or more processors of claim 4 , wherein the circuitry is to compute the scaling factor and adjust a learning rate of the one or more first neural networks based, at least in part, on the scaling factor associated with the plurality of layers of the one or more first neural networks.

6 . The one or more processors of claim 1 , wherein the one or more second neural networks are trained to perform an image processing task equivalent to an image processing task performed by the one or more first neural networks.

7 . The one or more processors of claim 1 , wherein the one or more hyperparameters of the one or more second neural networks are not adjusted based on the input of the one or more second synthetic images.

8 . The one or more processors of claim 1 , wherein the one or more first neural networks and the one or more second neural networks have equivalent structures.

9 . A system, comprising:

one or more processors to use one or more first neural networks to identify one or more objects in one or more synthetic images based, at least in part, on:

identifying a plurality of layers of the one or more first neural networks based on one or more inputs to the plurality of layers;

a scaling factor, computed based on calculating a loss corresponding to an inferencing response of one or more second neural networks in response to an input of one or more second synthetic images, wherein the one or more second neural networks are trained on non-synthetic images; and

one or more hyperparameters, generated using the one or more second neural networks and adjusted for the one or more first neural networks using the scaling factor in a policy corresponding to the plurality of layers of the one or more first neural networks, wherein the one or more first neural networks are trained using the one or more hyperparameters while one or more parameters of the one or more second neural networks are frozen.

10 . The system of claim 9 , wherein during training, a learning rate of the one or more first neural networks is adjusted based, at least in part, on an inferencing output of the one or more second neural networks, and wherein the learning rate is further adjusted according to the scaling factor in the policy associated with the plurality of layers of the one or more first neural networks.

11 . The system of claim 10 , wherein the one or more processors are to adjust the learning rate according to the scaling factor in the policy for a region of the one or more first neural networks.

12 . The system of claim 11 , wherein the region comprises the plurality of layers of the one or more first neural networks grouped according to input resolution.

13 . The system of claim 9 , wherein the one or more processors are to adjust training of the one or more first neural networks based, at least in part, on an inferencing output of the one or more second neural networks in response to the input of the one or more second synthetic images.

14 . The system of claim 9 , wherein the one or more processors are to compute a learning rate scaling factor based, at least in part, on a divergence factor, the divergence factor computed, at least in part, based on an inferencing output of the one or more second neural networks.

15 . The system of claim 9 , wherein the one or more second neural networks are frozen during training of the one or more first neural networks.

16 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

use one or more first neural networks to identify one or more objects in one or more synthetic images based, at least in part, on:

identifying a plurality of layers of the one or more first neural networks based on one or more inputs to the plurality of layers;

a scaling factor, computed based on calculating a loss corresponding to an inferencing response of one or more second neural networks in response to an input of one or more second synthetic images, wherein the one or more second neural networks are trained on non-synthetic images; and

one or more hyperparameters generated using the one or more second neural networks and adjusted for the one or more first neural networks using the scaling factor in a policy corresponding to the plurality of layers of the one or more first neural networks, wherein the one or more first neural networks are trained using the one or more hyperparameters while one or more parameters of the one or more second neural networks are frozen.

17 . The non-transitory machine-readable medium of claim 16 , having stored thereon the set of instructions which if performed by one or more processors, cause the one or more processors to at least:

compute an adjustment to a learning rate of the one or more first neural networks, based, at least in part, on an inferencing output of the one or more second neural networks, wherein the learning rate is adjusted according to the scaling factor in the policy associated with the plurality of layers of the one or more first neural networks.

18 . The non-transitory machine-readable medium of claim 17 , wherein the adjustment to the learning rate is calculated for a region of the one or more first neural networks, the region comprising the plurality of layers of the one or more first neural networks grouped based, at least in part, on input resolution.

19 . The non-transitory machine-readable medium of claim 18 , wherein the adjustment to the learning rate is calculated based, at least in part, by a long short-term memory (“LSTM”) module.

20 . The non-transitory machine-readable medium of claim 16 , wherein an inferencing output of the one or more second neural networks is based, at least in part, on an input comprising the one or more second synthetic images.

21 . The non-transitory machine-readable medium of claim 16 , wherein the one or more second neural networks are trained to perform an image processing task equivalent to an image processing task performed by the one or more first neural networks.

22 . The non-transitory machine-readable medium of claim 16 , wherein the one or more hyperparameters of the one or more second neural networks are not adjusted based on the input derived from the one or more second synthetic images.

23 . The non-transitory machine-readable medium of claim 16 , wherein the one or more first neural networks, once trained, are usable to perform an inferencing task independently of the one or more second neural networks.

24 . A computing device, comprising:

one or more processors to perform an image processing task based, at least in part, on one or more first neural networks to identify one or more objects in one or more synthetic images based, at least in part, on:

identifying a plurality of layers of the one or more first neural networks based on one or more inputs to the plurality of layers;

a scaling factor, computed based on calculating a loss corresponding to an inferencing response of one or more second neural networks in response to an input of one or more second synthetic images, wherein the one or more second neural networks are trained on non-synthetic images; and

one or more hyperparameters generated using the one or more second neural networks and adjusted for the one or more first neural networks using the scaling factor in a policy corresponding to the plurality of layers of the one or more first neural networks, wherein the one or more first neural networks are trained using the one or more hyperparameters while one or more parameters of the one or more second neural networks are frozen.

25 . The computing device of claim 24 , wherein the image processing task comprises at least one of recognition or classification.

26 . The computing device of claim 24 , the one or more processors to compute an adjustment to a learning rate of the one or more first neural networks, based, at least in part, on an inferencing output of the one or more second neural networks, wherein the learning rate is adjusted according to the scaling factor in the policy associated with the plurality of layers of the one or more first neural networks.

27 . The computing device of claim 26 , wherein the adjustment to the learning rate is calculated for a region of the one or more first neural networks, the region comprising the plurality of layers of the one or more first neural networks grouped based, at least in part, on input resolution; and wherein the adjustment to the learning rate is according to the scaling factor in the policy associated with the plurality of layers of the one or more first neural networks.

28 . The computing device of claim 24 , wherein an inferencing output of the one or more second neural networks is based, at least in part, on an input comprising the more second synthetic images.

29 . The computing device of claim 24 , wherein the one or more first neural networks and the one or more second neural networks are each trained to perform the image processing task.

30 . The computing device of claim 29 , wherein the one or more second neural networks are trained only on non-synthetic images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 28, 2020
From: YU, ZHIDING; CHEN, WUYANG; ANANDKUMAR, ANIMA
To: NVIDIA CORPORATION
Reel/Frame 052773/0776 →
Continuity (1)
Related Publication 20210334644A1 · Oct 28, 2021
References Cited (83)
US 20160078339A1 · Li · 2016 [cited by examiner]
US 20180307969A1 · Shibahara · 2018 [cited by examiner]
US 20190030371A1 · Han · 2019 [cited by examiner]
US 20190220738A1 · Flank · 2019 [cited by examiner]
US 20200134506A1 · Wang · 2020 [cited by examiner]
US 20200334538A1 · Meng · 2020 [cited by examiner]
US 20210034983A1 · Asano · 2021 [cited by examiner]
US 20210279595A1 · Sridhar · 2021 [cited by examiner]
CN 106170800A · 2016 [cited by applicant]
CN 109063565A · 2018 [cited by applicant]
Z. Li and D. Hoiem, “Learning without Forgetting,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, No. 12, pp. 2935-2947, Dec. 1, 2018 (Year: 2018). [cited by examiner]
Z. Meng, J. Li, Y. Zhao and Y. Gong, “Conditional Teacher-student Learning,” ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 6445-6449 (Year: 2019). [cited by examiner]
Bhardwaj et al., “Dream Distillation: A Data-Independent Model Compression Framework”, May 17, 2019 (Year: 2019). [cited by examiner]
Bae, Ji-Hoon, et al. “Densely distilled flow-based knowledge transfer in teacher-student framework for image classification.” IEEE Transactions on Image Processing 29 (2020): 5698-5710. (Year: 2020). [cited by examiner]
Nayak, Gaurav Kumar, et al. “Zero-shot knowledge distillation in deep networks.” International conference on machine learning. PMLR, 2019. (Year: 2019). [cited by examiner]
Andrychowicz et al., “Learning to Learn by Gradient Descent by Gradient Descent,” Nov. 16, 2016, 17 pages. [cited by applicant]
Chen et al., “Automated Synthetic-to-Real Generalization,” Jul. 14, 2020, 11 pages. [cited by applicant]
Chen et al., “Learning to Learn Without Gradient Descent by Gradient Descent,” Proceedings of the 34th International Conference on Machine Learning, vol. 70, 2017, 9 pages. [cited by applicant]
Chen et al., “ROAD: Reality Oriented Adaptation for Semantic Segmentation of Urban Scenes,” Apr. 7, 2018, 10 pages. [cited by applicant]
Cordts et al., “The Cityscapes Dataset for Semantic Urban Scene Understanding,” IEEE Conference on ComputerVision and Pattern Recognition, 2016, 11 pages. [cited by applicant]
Coumans et al., :PyBullet: A Python Module for Physics Simulation for Games, Robotics and Machine Learning, retrieved http://pybullet.org, 2016-2020, 10 pages. [cited by applicant]
Deng et al., “Imagenet: A Large-Scale Hierarchical Image Database,” CVPR, 2009, 8 pages. [cited by applicant]
Dosovitskiy et al., “FlowNet: Learning Optical Flow with Convolutional Networks,” ICCV, 2015, 9 pages. [cited by applicant]
Dundar et al., “Domain Stylization: A Fast Covariance Matching Framework Towards Domain Adaptation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(7): Jul. 2021, 13 pages. [cited by applicant]
Gaidon et al., “Virtual Worlds as Proxy for Multi-Object Tracking Analysis,” CVPR, 2016, 10 pages. [cited by applicant]
Gan et al., “Learning Attributes Equals Multi-Source Domain Generalization,” CVPR, 2016, 11 pages. [cited by applicant]
Ghiasi et al., Dropblock: A regularization method for convolutional networks. In Advances in Neural Information Processing Systems, 2018, 11 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” CVPR, 2016, 9 pages. [cited by applicant]
Hoffman et al., “CyCADA: Cycle-Consistent Adversarial Domain Adaptation,” International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Hoffman et al., “FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation,” Dec. 8, 2016, 9 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2021/028404, mailed Sep. 24, 2021, filed Apr. 21, 2021, 18 pages. [cited by applicant]
Johnson-Roberson et al., “Driving in the matrix: Can virtual worlds replace human-generated annotations for real world tasks?” Oct. 6, 2016, 8 pages. [cited by applicant]
Li et al., “Bidirectional Learning for Domain Adaptation of Semantic Segmentation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, 10 pages. [cited by applicant]
Li et al., “Deeper, Broader and Artier Domain Generalization,” ICCV, 2017, 9 pages. [cited by applicant]
Li et al., “Learning to Generalize: Meta-Learning for Domain Generalization,” The Thirty-Second AAAI Conference on Artificial Intelligence, 2018, 8 pages. [cited by applicant]
Li et al., “Learning to Optimize,” Jun. 6, 2016, 9 pages. [cited by applicant]
Li et al., “Learning without Forgetting,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Feb. 14, 2017, 13 pages. [cited by applicant]
Lin et al., “Microsoft COCO: Common Objects in Context,” European Conference on Computer Vision, Jul. 5, 2014, 14 pages. [cited by applicant]
Long et al., “Fully Convolutional Networks for Semantic Segmentation,” CVPR, 2015, 10 pages. [cited by applicant]
Long et al., “Learning Transferable Features with Deep Adaptation Networks,” May 27, 2015, 9 pages. [cited by applicant]
Lopez-Paz et al., “Gradient Episodic Memory for Continual Learning,” Advances in Neural Information Processing Systems, 2017, 10 pages. [cited by applicant]
Mayer et al., “A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation,” CVPR, 2016, 9 pages. [cited by applicant]
Meng et al., “Conditional Teacher-Student Learning,” Apr. 28, 2019, 5 pages. [cited by applicant]
Muandet et al., “Domain Generalization via Invariant Feature Representation,” International Conference on Machine earning, 2013, 9 pages. [cited by applicant]
Pan et al., “Two at Once: Enhancing Learning and Generalization Capacities via IBN-Net,” ECCV, 2018, 16 pages. [cited by applicant]
Peng et al., “VisDA: The Visual Domain Adaptation Challenge,” Nov. 29, 2017, 17 pages. [cited by applicant]
Pinheiro, “Unsupervised Domain Adaptation with Similarity Learning,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Real et al., “Youtube-boundingboxes: A Large High-Precision Human-Annotated Data Set for Object Detection in Video,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, 10 pages. [cited by applicant]
Richter et al., “Playing for Benchmarks,” ICCV, 2017, 10 pages. [cited by applicant]
Richter et al., “Playing for Data: Ground Truth from Computer Games,” ECCV, 2016, 17 pages. [cited by applicant]
Ros et al., “The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes,” CVPR, 2016, 10 pages. [cited by applicant]
Rumelhart et al., “Learning Representations by Back-Propagating Errors,” Nature, 323(6088): Oct. 9, 1986, 4 pages. [cited by applicant]
Saito et al., “Adversarial Dropout Regularization,” Nov. 7, 2017, 14 pages. [cited by applicant]
Saito et al., “Maximum Classifier Discrepancy for Unsupervised Domain Adaptation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Sankaranarayanan et al., “Generate to Adapt: Aligning Domains using Generative Adversarial Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Savva et al., “Habitat: A Platform for Embodied AI Research,” ICCV, 2019, 9 pages. [cited by applicant]
Shafahi et al., “Adversarially Robust Transfer Learning,” ICLR, Feb. 21, 2020, 14 pages. [cited by applicant]
Shin et al., “Continual Learning with Deep Generative Replay,” Advances in Neural Information Processing Systems, 2017, 10 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Sun et al., “Deep CORAL: Correlation Alignment for Deep Domain Adaptation,” European Conference on Computer Vision, Jul. 6, 2016, 7 pages. [cited by applicant]
Thrun, “Lifelong Learning Algorithms,” Learning to Learn, 1998, 29 pages. [cited by applicant]
Tsai et al., “Learning to Adapt Structured Output Space for Semantic Segmentation,” CVPR, 2018, 10 pages. [cited by applicant]
Tzeng et al., “Deep Domain Confusion: Maximizing for Domain Invariance,” Dec. 10, 2014, 9 pages. [cited by applicant]
Wichrowska et al., “Learned Optimizers that Scale and Generalize,” Proceedings of the 34th International Conference on Machine Learning, vol. 70, 2017, 10 pages. [cited by applicant]
Williams, “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning,” Machine Learning, 8(3-4): 1992, pp. 229-256. [cited by applicant]
Wu et al., “DCAN: Dual Channel-wise Alignment Networks for Unsupervised Scene Adaptation,” Proceedings of the European Conference on Computer Vision, 2018, 17 pages. [cited by applicant]
Xu et al., “Learning an Adaptive Learning Rate Schedule,” Sep. 20, 2019, 6 pages. [cited by applicant]
Yao et al., “Adversarial Feature Alignment: Avoid Catastrophic Forgetting in Incremental Task Lifelong Learning,” Oct. 24, 2019, 13 pages. [cited by applicant]
Yin et al., “Dreaming to Distill: Data-free Knowledge Transfer via DeepInversion,” Dec. 18, 2019, 15 pages. [cited by applicant]
Yue et al., “Domain Randomization and Pyramid Consistency: Simulation-to-Real Generalization without Accessing Target Domain Data,” Proceedings of the IEEE International Conference on Computer Vision, 2019, 11 pages. [cited by applicant]
Zhu et al., “Penalizing Top Performers: Conservative Loss for Semantic Segmentation Adaptation,” Proceedings of the European Conference on Computer Vision, 2018, 16 pages. [cited by applicant]
Zou et al., “Confidence Regularized Self-Training,” ICCV, 2019, 10 pages. [cited by applicant]
Zou et al., “Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training,” ECCV, 2018, 17 pages. [cited by applicant]
Baik et al., “Learning to Forget for Meta-Learning,” Jun. 13, 2019, 17 Pages. [cited by applicant]
Finn et al., “Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks,” Jul. 18, 2017, 13 Pages. [cited by applicant]
Jung et al., “Less-forgetful Learning for Domain Expansion in Deep Neural Networks,” Nov. 16, 2017, 8 Pages. [cited by applicant]
Ren et al., “Learning to Reweight Examples for Robust Deep Learning,” May 5, 2019, 13 Pages. [cited by applicant]
Office Action for Chinese Application No. 202180006001.5, mailed Apr. 30, 2025, 37 pages. [cited by applicant]
Office Action for Chinese Application No. 202180006001.5, mailed Dec. 15, 2025, 20 pages. [cited by applicant]
Li et al., “Learning to Optimize Neural Nets,” Nov. 30, 2017, 10 Pages. [cited by applicant]
Decision of Rejection for Chinese Application No. 202180006001.5, mailed Mar. 18, 2026, 17 pages. [cited by applicant]