Generating images based on coordinates associated with objects depicted in the image using neural networks
Apparatuses, systems, and techniques to generate an image of one or more objects. In at least one embodiment, an image of one or more objects is generated using a neural network based on, for example, a representation of a scene.
1 . A system, comprising:
at least one processor; and
at least one memory comprising instructions that, in response to execution by the at least one processor, cause the system to at least:
generate a set of data indicating a plurality of rays that intersect with a plurality of objects within a scene based, at least in part, on a representation of the scene that comprises a plurality of latent codes comprising one or more latent codes for individual ones of the plurality of objects;
determine a plurality of coordinates associated with the plurality of objects in the scene based, at least in part, on where the plurality of rays intersect with the plurality of objects; and
generate an image based, at least in part, on a set of color and density values generated by at least one neural network based, at least in part, on the plurality of coordinates and the plurality of latent codes.
2 . The system of claim 1 , wherein the image is a first image depicting the plurality of objects in the scene from a first viewpoint associated with the plurality of rays, and the at least one memory comprises further instructions that, in response to execution by the at least one processor, cause the system to at least:
use at least the representation of the scene to generate a second image that depicts the plurality of objects in the scene from a second viewpoint.
3 . The system of claim 1 , wherein the at least one memory comprises further instructions that, in response to execution by the at least one processor, cause the system to at least:
calculate the plurality of latent codes based, at least in part, on one or more images of the scene, wherein the plurality of latent codes correspond to the plurality of objects.
4 . The system of claim 1 , wherein a first coordinate of the plurality of coordinates is along a first ray of the plurality of rays and is within a volume of a first object of the plurality of objects.
5 . The system of claim 1 , wherein the at least one memory comprises further instructions that, in response to execution by the at least one processor, cause the system to at least:
generate the plurality of rays in reference to a camera pose.
6 . The system of claim 1 , wherein the representation comprises a scene graph corresponding to the scene.
7 . A method, comprising:
obtaining a representation of a plurality of objects in a scene;
generating a set of matrices indicating a plurality of rays that intersect the plurality of objects in the scene based, at least in part, on where the plurality of rays intersect the representation, and wherein the plurality of rays are associated with a first viewpoint of the scene;
determining a set of coordinates along the plurality of rays, indicating where the plurality of rays intersect the plurality of objects based, at least in part, on the set of matrices; and
using a neural network to generate an image of the scene from the first viewpoint based, at least in part, on the set of coordinates and a plurality of latent codes associated with the plurality of objects, wherein the plurality of latent codes comprises one or more latent codes for individual ones of the plurality of objects.
8 . The method of claim 7 , further comprising:
generating a second set of matrices based, at least in part, on a second viewpoint in the scene;
determining a second set of coordinates based, at least in part, on the second set of matrices, the image being a first image; and
using the neural network to generate a second image of the scene from the second viewpoint based, at least in part, on the second set of coordinates and the plurality of latent codes.
9 . The method of claim 7 , further comprising:
determining the plurality of latent codes based, at least in part, on one or more images of the plurality of objects.
10 . The method of claim 7 , wherein the set of coordinates includes a plurality of coordinates that are within a plurality of volumes of the plurality of objects.
11 . The method of claim 7 , wherein the image is generated using one or more graphics processing units (GPUs).
12 . The method of claim 7 , wherein the representation indicates at least one or more poses of the plurality of objects in the scene.
13 . A non-transitory computer-readable medium comprising instructions that, when performed by at least one processor of a computing device, cause the computing device to at least:
obtain a representation of a scene indicating at least a plurality of first poses of a plurality of objects;
determine a plurality of rays that intersect the plurality of objects based, at least in part, on where the plurality of rays intersect the representation of the scene;
determine a set of coordinates based, at least in part, on the plurality of rays, wherein the set of coordinates indicate where the plurality of rays intersect the plurality of objects in the scene;
use a neural network to generate first data representing the plurality of objects based, at least in part, on the set of coordinates and a plurality of latent codes associated with the plurality of objects, wherein the plurality of latent codes comprises one or more latent codes determined for individual ones of the plurality of objects; and
generate a first image based at least in part on the first data, the first image comprising one or more first visual representations of the plurality of objects as viewed from a first viewpoint.
14 . The non-transitory computer-readable medium of claim 13 , comprising further instructions that when performed by the at least one processor of the computing device, cause the computing device to at least:
modify the representation of the scene to indicate at least one or more second poses of the plurality of objects;
use the neural network to generate second data representing the plurality of objects based, at least in part, on the representation of the scene; and
generate a second image based at least in part on the second data.
15 . The non-transitory computer-readable medium of claim 13 , comprising further instructions that when performed by the at least one processor of the computing device, cause the computing device to at least:
obtain a plurality of images depicting the plurality of objects; and
determine the plurality of latent codes based, at least in part, on the plurality of images and the neural network.
16 . The non-transitory computer-readable medium of claim 13 , comprising further instructions that when performed by the at least one processor of the computing device, cause the computing device to at least:
use at least the neural network to determine one or more pixel colors corresponding to the plurality of rays; and
generate the first image from at least the plurality of pixel colors.
17 . The non-transitory computer-readable medium of claim 13 , wherein the plurality of rays are indicated at least in part by one or more matrices.
18 . The non-transitory computer-readable medium of claim 13 , wherein each coordinate of the set of coordinates corresponds to a respective object of the plurality of objects.
19 . The non-transitory computer-readable medium of claim 13 , wherein the neural network is a decoder.
20 . The non-transitory computer-readable medium of claim 13 , wherein:
the plurality of objects are of a particular category; and
the neural network is trained in connection with one or more images depicting objects of one or more categories.
21 . The non-transitory computer-readable medium of claim 13 , comprising further instructions that when performed by the at least one processor of the computing device, cause the computing device to at least:
generate a second image based at least in part on the representation of the scene, the second image comprising one or more second visual representations of the plurality of objects as viewed from a second viewpoint.