IP Library › Granted Patent US 12,406,023
Granted Patent B1
US 12,406,023 · App. 17/141,005 · Granted Sep 2, 2025

Neural network training method

Inventors: Jose Manuel Alvarez Lopez (Mountain View, CA); Akshay Chawla (Santa Clara, CA); Pavlo Molchanov (Mountain View, CA); Hongxu Yin (San Jose, CA)
Assignee: NVIDIA Corporation
G06F18/2148G06N3/045G06N3/08G06V20/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,023
App. No.
17/141,005
Granted
Sep 2, 2025
Kind
B1
Abstract

Apparatuses, systems, and techniques to generate images of objects. In at least one embodiment, one or more neural networks are trained to identify one or more objects within one or more images, and the one or more neural networks are used to generate an image of one or more objects.

Claims (49)

1. A processor, comprising:

one or more circuits to:

use one or more neural networks to identify one or more objects within one or more images and to cause an image of the one or more objects to be generated; and

use one or more second neural networks to detect the one or more objects based, at least in part, on a distance between output of the one or more neural networks and the one or more second neural networks.

2. The processor of claim 1 , the one or more circuits to train the one or more second neural networks to detect the one or more objects based, at least in part, on the one or more neural networks, the one or more second neural networks compact relative to the one or more first neural networks.

3. The processor of claim 1 , the one or more circuits to generate the image by modification of an input image based, at least in part, on a gradient derived from output of the one or more neural networks, wherein the input image is initialized as noise.

4. The processor of claim 1 , wherein the image is generated based, at least in part, on an object detection loss and a regularization term, the object detection loss corresponding to an object detection loss used to pre-train the one or more neural networks.

5. The processor of claim 1 , wherein the image is generated based, at least in part, on one or more augmented versions of an input image, the one or more augmented versions of the input image comprise three or more of flipped, jittered, contrast adjusted, brightness adjusted, or cutout images.

6. The processor of claim 1 , the one or more circuits to identify a plurality of tiles within the image, each of the plurality of tiles associated with an object category during generation of the image.

7. The processor of claim 1 , wherein the one or more second neural networks are trained to detect the one or more objects based, at least in part, on a mimic loss which is applied to minimize distance between output of the one or more neural networks and the one or more second neural networks.

8. The processor of claim 1 , wherein a category and bounding box of an object of the one or more objects is determined based, at least in part, on a false positive identification of the category.

9. A system, comprising:

one or more processors to:

use one or more neural networks to detect one or more objects within one or more images and to cause an image of the one or more objects to be generated; and

use one or more second neural networks to detect the one or more objects based, at least in part, on a distance between output of the one or more neural networks and the one or more second neural networks.

10. The system of claim 9 , wherein the one or more second neural networks are trained to detect the one or more objects based, at least in part, on the one or more neural networks.

11. The system of claim 9 , the one or more processors to generate the image by successive applications of a gradient derived from output of the one or more neural networks.

12. The system of claim 9 , wherein the image is generated based, at least in part, on an object detection loss calculated from output of the one or more neural networks.

13. The system of claim 9 , wherein the image is generated based, at least in part, on a regularization term.

14. The system of claim 9 , the one or more processors to generate depictions of the one or more objects in the image by at least identifying a plurality of tiles within the image and associating each of the plurality of tiles with an object category.

15. The system of claim 9 , wherein the one or more second neural networks are trained to detect the one or more objects based, at least in part, on a mimic loss indicative of distance between output of the one or more neural networks and the one or more second neural networks.

16. The system of claim 9 , wherein a category and bounding box of an object of the one or more objects is determined based, at least in part, on a false-positive identification of the category.

17. A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

use one or more neural networks to detect one or more objects within one or more images and to cause an image of the one or more objects to be generated; and

use one or more second neural networks to detect the one or more objects based, at least in part, on a distance between output of the one or more neural networks and the one or more second neural networks.

18. The machine-readable medium of claim 17 , having stored thereon a further set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

train the one or more second neural networks to detect the one or more objects based, at least in part, on the one or more neural networks and the image of the one or more objects.

19. The machine-readable medium of claim 17 , having stored thereon a further set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

generate the image by successive applications of a gradient derived from output of the one or more neural networks.

20. The machine-readable medium of claim 17 , wherein the image is generated based, at least in part, on an object detection loss calculated from output of the one or more neural networks.

21. The machine-readable medium of claim 17 , having stored thereon a further set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

identify a plurality of tiles within the image and associating each of the plurality of tiles with an object category.

22. The machine-readable medium of claim 17 , having stored thereon a further set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

train the one or more second neural networks to detect the one or more objects based, at least in part, on a mimic loss function to minimize distance between output of the one or more neural networks and the one or more second neural networks.

23. The machine-readable medium of claim 17 , wherein a category of an object of the one or more objects is determined based, at least in part, on a false-positive identification of the category.

24. A method, comprising:

training one or more first neural networks to detect one or more objects within one or more images;

using the one or more first neural networks to generate an image of the one or more objects; and

training one or more second neural networks to detect the one or more objects based, at least in part, on a distance between output of the one or more first neural networks and the one or more second neural networks.

25. The method of claim 24 , further comprising:

attempting to detect an object in a region of the image;

assigning a category indicator to the region, based at least in part on a false positive identification of the object in the region; and

adjusting the image to increase confidence in detection of the object in the region.

26. The method of claim 24 , further comprising:

generating a plurality of augmented versions of an input image, the plurality of augmented versions of the input image comprising three or more of flipped, jittered, contrast adjusted, brightness adjusted, or cutout images.

27. The method of claim 24 , further comprising: identifying a plurality of tiles in the image; and

generating an object in each of the plurality of tiles of the image using the one or more first neural networks.

28. The method of claim 24 , wherein one or more second neural networks are trained to detect the one or more objects based, at least in part, on a mimic loss which is applied to minimize the distance between output of the one or more neural networks and the one or more second neural networks.

29. The method of claim 24 , wherein image detection comprises identification of a region of an image estimated to contain an object and a classification of the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2021
From: ALVAREZ LOPEZ, JOSE MANUEL; CHAWLA, AKSHAY; MOLCHANOV, PAVLO; YIN, HONGXU
To: NVIDIA CORPORATION
Reel/Frame 055368/0043 →
References Cited (75)
US 20170185872A1 · Chakraborty · 2017 [cited by examiner]
US 20180174046A1 · Xiao · 2018 [cited by examiner]
US 20190147610A1 · Frossard · 2019 [cited by examiner]
US 20190295261A1 · Kang · 2019 [cited by examiner]
US 20200020093A1 · Frei · 2020 [cited by examiner]
US 20200027002A1 · Hickson · 2020 [cited by examiner]
US 20200202502A1 · Tsymbalenko · 2020 [cited by examiner]
US 20200265255A1 · Li · 2020 [cited by examiner]
US 20200293828A1 · Wang · 2020 [cited by examiner]
US 20200327450A1 · Huang · 2020 [cited by examiner]
US 20200401856A1 · Kim · 2020 [cited by examiner]
US 20210004681A1 · Tate · 2021 [cited by examiner]
US 20210019544A1 · Park · 2021 [cited by examiner]
US 20210042558A1 · Choi · 2021 [cited by examiner]
US 20210089841A1 · Mithun · 2021 [cited by examiner]
US 20210177296A1 · Saalbach · 2021 [cited by examiner]
US 20210225038A1 · Holzer · 2021 [cited by examiner]
US 20210265017A1 · Dutta · 2021 [cited by examiner]
US 20210279640A1 · Tu · 2021 [cited by examiner]
US 20210327029A1 · Chen · 2021 [cited by examiner]
US 20210350555A1 · Fischetti · 2021 [cited by examiner]
US 20220148241A1 · Park · 2022 [cited by examiner]
US 20220254137A1 · Tu · 2022 [cited by examiner]
Terrance De Vries et al. , “Improved Regularization of Convolutional Neural Networks with Cutout,” Nov. 29, 2017, Computer Vision and Pattern Recognition, arXiv:1708.04552,pp. 1-4. [cited by examiner]
Mark Everingham et al. ,“The PASCAL Visual Object Classes (VOC) Challenge,” Sep. 9, 2009, International Journal of Computer Vision vol. 88;pp. 303-315. [cited by examiner]
Daniel Mas Montserrat,“Training Object Detection And Recognition CNN Models Using Data Augmentation,” Jan. 2017,IS&T International Symposium on Electronic Imaging 2017,vol. 29 | Article ID: art00005, pp. 27-32. [cited by examiner]
Ba et al., “Do deep nets really need to be deep?” Advances in Neural Information Processing Systems, 2014, 9 pages. [cited by applicant]
Bhardwaj et al., “Dream distillation: A Data-Independent Model Compression Framework,” May 17, 2019, 4 pages. [cited by applicant]
Brock et al., “Large Scale GAN Training for High Fidelity Natural Image Synthesis,” Sep. 28, 2018, 29 pages. [cited by applicant]
Chen et al., “Data-Free Learning of Student Networks,” Proceedings of the IEEE International Conference on Computer Vision, 2019, 10 pages. [cited by applicant]
Chen et al., “Learning Efficient Object Detection Models with Knowledge Distillation,” Advances in Neural Information Processing Systems, 2017, 10 pages. [cited by applicant]
Deng et al. “Imagenet: A Large-Scale Hierarchical Image Database,” ICLR, 2009, 8 pages. [cited by applicant]
Devries et al., “Improved Regularization of Convolutional Neural Networks with Cutout,” Nov. 29, 2017, 8 pages. [cited by applicant]
Dosovitskiy et al., “Inverting Visual Representations with Convolutional Networks,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, 9 pages. [cited by applicant]
Everingham et al., “The Pascal Visual Object Classes (VOC) Challenge,” International Journal of Computer Vision, 38(2): 2010, 34 pages. [cited by applicant]
Fredrikson et al., “Model Inversion Attacks that Exploit Confidence Information and Basic Countermeasures,” Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security, 2015, 12 pages. [cited by applicant]
Girshick et al., “Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, 8 pages. [cited by applicant]
Girshick, “Fast R-CNN,” In IEEE Conference on Computer Vision and Pattern Recognition, 2015, 9 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition, ” CVPR, 2016, 9 pages. [cited by applicant]
Hinton et al., “Distilling the Knowledge in a Neural Network,” Mar. 9, 2015, 9 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Jocher et al., “Ultralytics/yolov3: [email protected]:0.95 on coco2014,” retrieved from https://zenodo.org/record/3785397# YOXZszNJGUI, May 4, 2020, 5 pages. [cited by applicant]
Karras et al., “Analyzing and Improving the Image Quality of Stylegan,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, 10 pages. [cited by applicant]
Karras et al., “Progressive Growing of GANs for Improved Quality, Stability, and Variation,” Machine Learning, Oct. 2017, 26 pages. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks,” In Advances in Neural Information Processing Systems, 2012, 9 pages. [cited by applicant]
Li et al., “Mimicking Very Efficient Network for Object Detection,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, 9 pages. [cited by applicant]
Lin et al., “Microsoft COCO: Common Objects in Context,” European Conference on Computer Vision, Jul. 5, 2014, 14 pages. [cited by applicant]
Liu et al., “SSD: Single Shot Multibox Detector,” European Conference on Computer Vision, Springer, Dec. 29, 2016, 17 pages. [cited by applicant]
Lopes et al., “Data-Free Knowledge Distillation for Deep Neural Networks,” Nov. 23, 2017, 8 pages. [cited by applicant]
Mahendran et al., “Understanding Deep Image Representations by Inverting Them,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, 9 pages. [cited by applicant]
Mehta et al., “Object Detection at 200 Frames Per Second,” Proceedings of the European Conference on Computer Vision, 2018, 15 pages. [cited by applicant]
Micaelli et al., “Zero-Shot Knowledge Transfer via Adversarial Belief Matching,” Advances in Neural Information Processing Systems, 2019, 11 pages. [cited by applicant]
Mordvintsev et al., Inceptionism: Going Deeper Into Neural Networks, retrieved from https://ai.googleblog.com/2015/06/inceptionism-going-deeper-into-neural.html, Jun. 17, 2015, 6 pages. [cited by applicant]
Nayak et al., “Zero-Shot Knowledge Distillation in Deep Networks,” May 20, 2019, 17 pages. [cited by applicant]
Nguyen et al., “Synthesizing the Preferred Inputs for Neurons in Neural Networks via Deep Generator Networks,” Advances in Neural Information Processing Systems, 2016, 9 pages. [cited by applicant]
Radford et al., “Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks,” Nov. 19, 2015, 15 pages. [cited by applicant]
Redmon et al., “YOLO9000: Better, Faster, Stronger,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, 9 pages. [cited by applicant]
Redmon et al., “YOLOv3: An Incremental Improvement”, Apr. 8, 2018, 6 pages. [cited by applicant]
Redmon et al., “You Only Look Once: Unified, Realtime Object Detection, ” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, 10 pages. [cited by applicant]
Ren et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” Advances in Neural Information Processing Systems, 2015, 9 pages. [cited by applicant]
Rezatofighi et al., “Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, 9 pages. [cited by applicant]
Richter et al., “Playing for Data: Ground Truth from Computer Games,” ECCV, 2016, 17 pages. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” International Journal of Computer Vision, 115(3): 2015, 42 pages. [cited by applicant]
Shmelkov et al., “Incremental Learning of Object Detectors without Catastrophic Forgetting,” Proceedings of the EEE International Conference on Computer Vision, 2017, 10 pages. [cited by applicant]
Simonyan et al., “Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps,” Dec. 20, 2013, 8 pages. [cited by applicant]
Simonyan et al., Very Deep Convolutional Networks for Large-Scale Image Recognition, Dec. 23, 2014, 13 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Tan et al., “Efficientnet: Rethinking Model Scaling for Convolutional Neural Networks,” In International Conference on Machine Learning, 2019, 10 pages. [cited by applicant]
Wang et al., “Distilling Object Detectors with Fine-Grained Feature Imitation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, 10 pages. [cited by applicant]
Wei et al., “Quantization Mimic: Towards Very Tiny CNN for Object Detection,” Proceedings of the European Conference on Computer Vision, 2018, 17 pages. [cited by applicant]
Wu et al., “Group Normalization,” In European Conference on Computer Vision, 2018, 17 pages. [cited by applicant]
Yin et al., “Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversion,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, 10 pages. [cited by applicant]
Yu et al., “BDD100K: A Diverse Driving Video Database with Scalable Annotation Tooling,” May 12, 2018, 16 pages. [cited by applicant]
Zhu et al., “Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks,” In Proceedings of the IEEE International Conference on Computer Vision, 2017, 10 pages. [cited by applicant]
Cited By (4)
US 12,597,176 US 12,665,916 US 12,682,593 US 12,718,576