Shared latent spaces for volumetric rendering
Systems and methods described herein support enhanced computer vision capabilities which may be applicable to, for example, autonomous vehicle operation. An example method includes An example method includes training a shared latent space and a first decoder based on first image data that includes multiple images, and training the shared latent space and a second decoder based on second image data that includes multiple images. The method also includes generating a volumetric embedding that is representative of a novel viewing frame the first scene. Further, the method includes decoding, with the first decoders, the shared latent space with the volumetric embedding, and generating the novel viewing frame of the first scene based on the output of the first decoder.
1 . A method comprising:
training a shared latent space and a first decoder based on first image data that includes multiple images, wherein each image has a different viewing frame of a first scene;
training the shared latent space and a second decoder based on second image data that includes multiple images, wherein each image has a different viewing frame of a second scene;
generating a volumetric embedding that is representative of a novel viewing frame the first scene;
decoding, with the first decoder, the shared latent space with the volumetric embedding; and
generating the novel viewing frame of the first scene based on the output of the first decoder.
2 . The method of claim 1 , wherein the volumetric embedding is a concatenation of an origin embedding and a depth embedding.
3 . The method of claim 1 , wherein the novel viewing frame includes a predicted depth map of the first scene.
4 . The method of claim 3 , wherein the predicted depth map is used to control at least one function of a vehicle.
5 . The method of claim 1 , wherein the novel viewing frame includes a bitmap of a novel image from a perspective of the novel viewing frame.
6 . The method of claim 1 , further comprising:
generating a second volumetric embedding that is representative of a second novel viewing frame the second scene;
decoding, with the second decoder, the shared latent space with the second volumetric embedding; and
generating the second novel viewing frame of the second scene based on the output of the second decoder.
7 . The method of claim 6 , wherein the second novel viewing frame includes a second predicted depth map of the second scene.
8 . A system comprising:
a preprocessing platform, comprising at least one processor and memory, configured to:
train a shared latent space and a first decoder based on first image data that includes multiple images, wherein each image has a different viewing frame of a first scene;
train the shared latent space and a second decoder based on second image data that includes multiple images, wherein each image has a different viewing frame of a second scene;
a computer vision platform configured to:
generate a volumetric embedding that is representative of a novel viewing frame the first scene;
decode, with the first decoder, the shared latent space with the volumetric embedding; and
generate the novel viewing frame of the first scene based on the output of the first decoder.
9 . The system of claim 8 , wherein the volumetric embedding is a concatenation of an origin embedding and a depth embedding.
10 . The system of claim 8 , wherein the novel viewing frame includes a predicted depth map of the first scene.
11 . The system of claim 10 , wherein the predicted depth map is used to control at least one function of a vehicle.
12 . The system of claim 8 , wherein the novel viewing frame includes a bitmap of a novel image from a perspective of the novel viewing frame.
13 . The system of claim 8 , wherein the computer vision platform is further configured to:
generate a second volumetric embedding that is representative of a second novel viewing frame the second scene;
decode, with the second decoder, the shared latent space with the second volumetric embedding; and
generate the second novel viewing frame of the second scene based on the output of the second decoder.
14 . The system of claim 13 , wherein the second novel viewing frame includes a second predicted depth map of the second scene.
15 . A non-transitory computer readable medium comprising instructions that, when executed, cause a system to:
train a shared latent space and a first decoder based on first image data that includes multiple images, wherein each image has a different viewing frame of a first scene;
train the shared latent space and a second decoder based on second image data that includes multiple images, wherein each image has a different viewing frame of a second scene;
generate a volumetric embedding that is representative of a novel viewing frame the first scene;
decode, with the first decoder, the shared latent space with the volumetric embedding; and
generate the novel viewing frame of the first scene based on the output of the first decoder.
16 . The computer readable medium of claim 15 , wherein the volumetric embedding is a concatenation of an origin embedding and a depth embedding.
17 . The computer readable medium of claim 15 , wherein the novel viewing frame includes a predicted depth map of the first scene.
18 . The computer readable medium of claim 17 , wherein the predicted depth map is used to control at least one function of a vehicle.
19 . The computer readable medium of claim 15 , wherein the novel viewing frame includes a bitmap of a novel image from a perspective of the novel viewing frame.
20 . The computer readable medium of claim 15 , wherein the instructions further cause the system to:
generate a second volumetric embedding that is representative of a second novel viewing frame the second scene;
decode, with the second decoder, the shared latent space with the second volumetric embedding; and
generate the second novel viewing frame of the second scene based on the output of the second decoder.