IP Library › Granted Patent US 11,783,510
Granted Patent B2
US 11,783,510 · App. 17/002,279 · Granted Oct 10, 2023

View generation using one or more neural networks

Inventors: Siddhant Pardeshi (Pune, IN); Pranit P. Kothari (Pune, IN); Vinayak Vilas Gaikwad (Pune, IN)
Assignee: NVIDIA Corporation
G06T9/002G06N3/045G06T1/20G06T2200/04G06T2210/61
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,783,510
App. No.
17/002,279
Granted
Oct 10, 2023
Kind
B2
Abstract

Apparatuses, systems, and techniques are presented to generate image or video content representing at least one point of view. In at least one embodiment, one or more neural networks are used to generate one or more images of one or more objects from a first point of view based at least in part upon one or more images of the one or more objects from a second point of view.

Claims (39)

1. A processor, comprising:

one or more circuits to use one or more neural networks to indicate whether a second point of view of one or more objects within one or more images can be generated based, at least in part, on a first point of view of the one or more objects.

2. The processor of claim 1 , wherein the one or more circuits are further to perform instance segmentation to identify features for two or more objects in the one or more images.

3. The processor of claim 2 , wherein the one or more neural networks include at least one variational autoencoder to encode the features of the two or more objects into a latent space.

4. The processor of claim 3 , wherein the one or more neural networks include a generative network for generating the second point of view, the generative network accepting as input at least the latent space and an indication of the first point of view.

5. The processor of claim 1 , wherein the one or more images correspond to one or more frames of video content, and wherein the one or more circuits are further to use at least two passes to generate the second point of view, the one or more circuits further to store information for two or more objects identified during a first pass to a cache for use in extrapolating image content for the two or more objects in the second point of view to be generated in a second pass.

6. The processor of claim 1 , wherein an indication of the second point of view is received from a user and determined to satisfy a minimum generation probability criterion.

7. A system comprising:

one or more processors to use one or more neural networks to indicate whether a second point of view of one or more objects within one or more images can be generated based, at least in part, on a first point of view of the one or more objects.

8. The system of claim 7 , wherein the one or more processors are further to perform instance segmentation to identify features for two or more objects in the one or more images.

9. The system of claim 8 , wherein the one or more neural networks include at least one variational autoencoder to encode the features of the two or more objects into a latent space.

10. The system of claim 9 , wherein the one or more neural networks include a generative network for generating the second point of view, the generative network accepting as input at least the latent space and an indication of the first point of view.

11. The system of claim 7 , wherein the one or more images correspond to one or more frames of video content, and wherein the one or more processors are further to use at least two passes to generate the second point of view, the one or more processors further to store information for two or more objects identified during a first pass to a cache for use in extrapolating image content for the two or more objects in the second point of view to be generated in a second pass.

12. The system of claim 7 , wherein an indication of the second point of view is received from a user and determined to satisfy a minimum generation probability criterion.

13. A method comprising:

using one or more neural networks to indicate whether a second point of view of one or more objects within one or more images can be generated based, at least in part, on a first point of view of the one or more objects.

14. The method of claim 13 , further comprising: performing instance segmentation to identify features for two or more objects in the one or more images.

15. The method of claim 14 , wherein the one or more neural networks include at least one variational autoencoder to encode the features of the two or more objects into a latent space.

16. The method of claim 15 , wherein the one or more neural networks include a generative network for generating the second point of view, the generative network accepting as input at least the latent space and an indication of the first point of view.

17. The method of claim 13 , wherein the one or more images correspond to one or more frames of video content, and further comprising:

using at least two passes to generate the second point of view, further storing information for two or more objects identified during a first pass to a cache for use in extrapolating image content for the two or more objects in the second point of view to be generated in a second pass.

18. The method of claim 13 , wherein an indication of the second point of view is received from a user and determined to satisfy a minimum generation probability criterion.

19. A non-transitory computer-readable storage medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

use one or more neural networks to indicate whether a second point of view of one or more objects within one or more images can be generated based, at least in part, on a first point of view of the one or more objects.

20. The non-transitory computer-readable storage medium of claim 19 , wherein the instructions if performed further cause the one or more processors to:

perform instance segmentation to identify features for two or more objects in the one or more images.

21. The non-transitory computer-readable storage medium of claim 20 , wherein the one or more neural networks include at least one variational autoencoder to encode the features of the two or more objects into a latent space.

22. The non-transitory computer-readable storage medium of claim 21 , wherein the one or more neural networks include a generative network for generating the second point of view, the generative network accepting as input at least the latent space and an indication of the first point of view.

23. The non-transitory computer-readable storage medium of claim 19 , wherein the one or more images correspond to one or more frames of video content, and wherein the instructions if performed further cause the one or more processors to:

use at least two passes to generate the second point of view, further storing information for two or more objects identified during a first pass to a cache for use in extrapolating image content for the two or more objects in the second point of view to be generated in a second pass.

24. The non-transitory computer-readable storage medium of claim 19 , wherein an indication of the second point of view is received from a user and determined to satisfy a minimum generation probability criterion.

25. An image augmentation system, comprising:

one or more processors to use one or more neural networks to indicate whether a second point of view of one or more objects within one or more images can be generated based, at least in part, on a first point of view of the one or more objects; and

memory for storing network parameters for the one or more neural networks.

26. The image augmentation system of claim 25 , wherein the one or more processors are further to perform instance segmentation to identify features for two or more objects in the one or more images.

27. The image augmentation system of claim 26 , wherein the one or more neural networks include at least one variational autoencoder to encode the features of the two or more objects into a latent space.

28. The image augmentation system of claim 27 , wherein the one or more neural networks include a generative network for generating the second point of view, the generative network accepting as input at least the latent space and an indication of the first point of view.

29. The image augmentation system of claim 25 , wherein the one or more images correspond to one or more frames of video content, and wherein the one or more processors are further to use at least two passes to generate the second point of view, the one or more processors further to store information for two or more objects identified during a first pass to a cache for use in extrapolating image content for the two or more objects in the second point of view to be generated in a second pass.

30. The image augmentation system of claim 25 , wherein an indication of the second point of view is received from a user and determined to satisfy a minimum generation probability criterion.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2020
From: PARDESHI, SIDDHANT; KOTHARI, PRANIT P.; GAIKWAD, VINAYAK VILAS
To: NVIDIA CORPORATION
Reel/Frame 053624/0340 →
Continuity (1)
Related Publication 20220067982A1 · Mar 3, 2022