IP Library › Granted Patent US 12,026,845
Granted Patent B2
US 12,026,845 · App. 17/007,079 · Granted Jul 2, 2024

Image generation using one or more neural networks

Inventor: Siddhant Pardeshi (Pune, IN)
Assignee: NVIDIA Corporation
G06T19/20G06N3/045G06N3/084G06T7/0002G06T7/70G06T19/006G06N5/046G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,026,845
App. No.
17/007,079
Filed
Aug 31, 2020
Granted
Jul 2, 2024
Kind
B2
Art Unit
2635
USPC
382/100
Abstract

Apparatuses, systems, and techniques are presented to generate augmented images. In at least one embodiment, one or more neural networks are used to modify one or more first objects in an image based at least in part upon a modification to be made to one or more second objects in the image.

Claims (42)

1. A processor, comprising:

one or more circuits to use one or more neural networks to modify one or more first objects in one or more images based, at least in part, upon one or more modifications to one or more second objects in the one or more images having one or more relationships to the one or more first objects.

2. The processor of claim 1 , wherein the one or more circuits are further to use the one or more neural networks are further to determine at least one relationship between the one or more first objects and the one or more second objects before determining to modify the one or more first objects, the at least one relationship including at least one of a logical relationship or a physical relationship.

3. The processor of claim 2 , wherein the one or more circuits are further to use the neural networks to identify features, position, and state information for objects in the one or more images, the objects including the one or more first objects and the one or more second objects, and wherein the one or more neural networks include a plurality of variational autoencoders (VAEs) trained to encode the features, position, state information, and at least one relationship into a latent space.

4. The processor of claim 3 , wherein the one or more neural networks include a generative adversarial network (GAN) for generating an output image based on image content of the one or more images and using the latent space as a constraint to cause the output image to include the one or more modifications to the one or more first objects and the one or more second objects.

5. The processor of claim 1 , wherein the one or more circuits are further to use the neural networks to detect one or more anomalies in the one or more images after the one or more first objects and the one or more second objects are modified, and cause the one or more images to be regenerated to attempt to remove the one or more anomalies.

6. The processor of claim 1 , wherein the one or more modifications to the one or more second objects includes a modification to at least one of object position, orientation, or state, and wherein the one or more modifications are determined based at least in part upon an input reference image including at least one object of at least one similar object class to the one or more second objects.

7. A system comprising:

one or more processors to use one or more neural networks to modify one or more first objects in one or more images based, at least in part, upon one or more modifications to one or more second objects in the one or more images having one or more relationships to the one or more first objects.

8. The system of claim 7 , wherein the one or more processors are further to use the one or more neural networks are further to determine at least one relationship between the one or more first objects and the one or more second objects before determining to modify the one or more first objects, the at least one relationship including at least one of a logical relationship or a physical relationship.

9. The system of claim 8 , wherein the one or more processors are further to use the neural networks to identify features, position, and state information for objects in the one or more, the objects including the one or more first objects and the one or more second objects, and wherein the one or more neural networks include a plurality of variational autoencoders (VAEs) trained to encode the features, position, state information, and at least one relationship into a latent space.

10. The system of claim 9 , wherein the one or more neural networks include a generative adversarial network (GAN) for generating an output image based on image content of the one or more and using the latent space as a constraint to cause the output image to include the one or more modifications to the one or more first objects and the one or more second objects.

11. The system of claim 7 , wherein the one or more processors are further to use the neural networks to detect one or more anomalies in the one or more images after the one or more first objects and the one or more second objects are modified, and cause the one or more images to be regenerated to attempt to remove the one or more anomalies.

12. The system of claim 7 , wherein the one or more modifications to the one or more second objects includes a modification to at least one of object position, orientation, or state, and wherein the one or more modifications are determined based at least in part upon an input reference image including at least one object of at least one similar object class to the one or more second objects.

13. A method comprising:

using one or more neural networks to modify one or more first objects in one or more images based, at least in part, upon one or more modifications to one or more second objects in the one or more images having one or more relationships to the one or more first objects image.

14. The method of claim 13 , further comprising:

determining at least one relationship between the one or more first objects and the one or more second objects before determining to modify the one or more first objects, the at least one relationship including at least one of a logical relationship or a physical relationship.

15. The method of claim 14 , further comprising:

identifying features, position, and state information for objects in the one or more images, the objects including the one or more first objects and the one or more second objects, and wherein the one or more neural networks include a plurality of variational autoencoders (VAEs) trained to encode the features, position, state information, and at least one relationship into a latent space.

16. The method of claim 15 , wherein the one or more neural networks include a generative adversarial network (GAN) for generating an output image based on image content of the one or more images and using the latent space as a constraint to cause the output image to include the one or more modifications to the one or more first objects and the one or more second objects.

17. The method of claim 13 , further comprising:

detecting one or more anomalies in the one or more images after the one or more first objects and the one or more second objects are modified, and cause the one or more images to be regenerated to attempt to remove the one or more anomalies.

18. The method of claim 13 , wherein the one or more modifications to the one or more second objects includes a modification to at least one of object position, orientation, or state, and wherein the one or more modifications are determined based at least in part upon an input reference image including at least one object of at least one similar object class to the one or more second objects.

19. A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

use one or more neural networks to modify one or more first objects in one or more images based, at least in part, upon one or more modifications to one or more second objects in the one or more images having one or more relationships to the one or more first objects.

20. The machine-readable medium of claim 19 , wherein the set of instructions if performed further causes the one or more processors to:

determine at least one relationship between the one or more first objects and the one or more second objects before determining to modify the one or more first objects, the at least one relationship including at least one of a logical relationship or a physical relationship.

21. The machine-readable medium of claim 20 , wherein the set of instructions if performed further causes the one or more processors to:

identifying features, position, and state information for objects in the one or more images, the objects including the one or more first objects and the one or more second objects, and wherein the one or more neural networks include a plurality of variational autoencoders (VAEs) trained to encode the features, position, state information, and at least one relationship into a latent space.

22. The machine-readable medium of claim 21 , wherein the one or more neural networks include a generative adversarial network (GAN) for generating an output image based on image content of the one or more images and using the latent space as a constraint to cause the output image to include the one or more modifications to the one or more first objects and the one or more second objects.

23. The machine-readable medium of claim 19 , wherein the set of instructions if performed further causes the one or more processors to:

detect one or more anomalies in the one or more images after the one or more first objects and the one or more second objects are modified, and cause the one or more images to be regenerated to attempt to remove the one or more anomalies.

24. The machine-readable medium of claim 19 , wherein the one or more modifications to the one or more second objects includes a modification to at least one of object position, orientation, or state, and wherein the one or more modifications are determined based at least in part upon an input reference image including at least one object of at least one similar object class to the one or more second objects.

25. An image generation system, comprising:

one or more processors to use one or more neural networks to modify one or more first objects in one or more images based, at least in part, upon one or more modifications to one or more second objects in the one or more images having one or more relationships to the one or more first objects; and

memory for storing network parameters for the one or more neural networks.

26. The image generation system of claim 25 , wherein the one or more processors are further to use the one or more neural networks to determine at least one relationship between the one or more first objects and the one or more second objects before determining to modify the one or more first objects, the at least one relationship including at least one of a logical relationship or a physical relationship.

27. The image generation system of claim 26 , wherein the one or more processors are further to use the neural networks to identify features, position, and state information for objects in the one or more images, the objects including the one or more first objects and the one or more second objects, and wherein the one or more neural networks include a plurality of variational autoencoders (VAEs) trained to encode the features, position, state information, and at least one relationship into a latent space.

28. The image generation system of claim 27 , wherein the one or more neural networks include a generative adversarial network (GAN) for generating an output image based on image content of the one or more images and using the latent space as a constraint to cause the output image to include the one or more modifications to the one or more first objects and the one or more second objects.

29. The image generation system of claim 25 , wherein the one or more processors are further to use the neural networks to detect one or more anomalies in the one or more images after the one or more first objects and the one or more second objects are modified, and cause the one or more images to be regenerated to attempt to remove the one or more anomalies.

30. The image generation system of claim 25 , wherein the one or more modifications to the one or more second objects includes a modification to at least one of object position, orientation, or state, and wherein the one or more modifications are determined based at least in part upon an input reference image including at least one object of at least one similar object class to the one or more second objects.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2020
From: PARDESHI, SIDDHANT
To: NVIDIA CORPORATION
Reel/Frame 053714/0707 →
Continuity (1)
Related Publication 20220068037A1 · Mar 3, 2022
Cited By (34)
US 12,210,800 US 12,249,045 US 12,260,530 US 12,277,652 US 12,288,279 US 12,299,858 US 12,307,600 US 12,333,691 US 12,333,692 US 12,347,005 US 12,347,080 US 12,347,124 US 12,394,166 US 12,395,722 US 12,423,855 US 12,450,810 US 12,456,243 US 12,456,274 US 12,462,420 US 12,462,519 US 12,469,194 US 12,482,172 US 12,488,523 US 12,499,574 US 12,505,520 US 12,505,596 US 12,536,625 US 12,597,186 US 12,614,301 US 12,646,188 US 12,657,902 US 12,669,914 US 12,699,852 US 12,731,309