Systems and methods for generating an image using interpolation of features
System, methods, and other embodiments described herein relate to generating an image by interpolating features estimated from a learning model. In one embodiment, a method includes sampling three-dimensional (3D) points of a light ray that crosses a frustum space associated with a single-view camera, the 3D points reflecting depth estimates derived from data that the single-view camera generates for a scene. The method also includes deriving feature values for the 3D points using tri-linear interpolation across feature planes of the frustum space, the feature planes being estimated by a learning model. The method also includes inferring an image in two dimensions (2D) by translating the feature values and compositing the data with volumetric rendering for the scene. The method also includes executing a control task by a controller using the image.
1 . A prediction system comprising:
a processor; and
a memory storing instructions that, when executed by the processor, cause the processor to:
sample three-dimensional (3D) points of a light ray that crosses a frustum space associated with a single-view camera, the 3D points reflecting depth estimates derived from data that the single-view camera generates for a scene;
derive feature values for the 3D points using tri-linear interpolation;
generate an image using the tri-linear interpolation across feature planes of the frustum space with the feature values, the feature planes being estimated by a multi-layer perceptron (MLP), and the light ray is associated with generated grid points from adjacent planes within the frustum space, selective objects over the frustum space are unviewable, and the adjacent planes include a dense area within the scene having a density disparity from an increased radiance associated with deriving fine and coarse details for the feature values;
infer the image in two dimensions (2D) using the MLP by translating the feature values for the 3D points and compositing the data with a volumetric rendering for the scene, the feature values associated with one of the fine and the coarse details; and
execute a control task by a controller using the image.
2 . The prediction system of claim 1 , wherein the instructions to derive the feature values further include instructions to generate the image as a multi-plane image (MPI) by mixing information directly among the feature planes and between the feature planes using the tri-linear interpolation, wherein the information expands a capacity of the MPI.
3 . The prediction system of claim 2 , wherein the information describes denser locations within the scene for a driving environment and the denser locations represent mixed colors.
4 . The prediction system of claim 2 further including instructions to acquire the data by a vehicle for a target view and wherein the MPI is associated with an area of the target view.
5 . The prediction system of claim 1 , wherein the instructions to derive the feature values further include instructions to generate the image as a multi-plane image (MPI) to acquire information among the feature planes and between the feature planes using the tri-linear interpolation, wherein the information describes regions within the scene that attract attention according to the density disparity.
6 . The prediction system of claim 5 , wherein the density disparity represents an increase in a radiance density between the feature planes associated with the light ray.
7 . The prediction system of claim 5 , wherein the instructions to infer the image further include instructions to map the feature values to a red-green-blue (RGB) space using the MLP associated with the volumetric rendering.
8 . The prediction system of claim 1 , wherein the instructions to derive the feature values further include instructions to generate at least eight grid points from the adjacent planes for the light ray associated with the frustum space and a target view; and
wherein first objects within the frustum space are visible by the single-view camera and second objects outside the frustum space are unviewable by the single-view camera.
9 . The prediction system of claim 1 , wherein the feature values include classes describing objects within the scene and the tri-linear interpolation acquires information about the scene from the adjacent planes of the frustum space.
10 . A non-transitory computer-readable medium comprising:
instructions that when executed by a processor cause the processor to:
sample three-dimensional (3D) points of a light ray that crosses a frustum space associated with a single-view camera, the 3D points reflecting depth estimates derived from data that the single-view camera generates for a scene;
derive feature values for the 3D points using tri-linear interpolation;
generate an image using the tri-linear interpolation across feature planes of the frustum space with the feature values, the feature planes being estimated by a multi-layer perceptron (MLP), and the light ray is associated with generated grid points from adjacent planes within the frustum space, selective objects over the frustum space are unviewable, and the adjacent planes include a dense area within the scene having a density disparity from an increased radiance associated with deriving fine and coarse details for the feature values;
infer the image in two dimensions (2D) using the MLP by translating the feature values for the 3D points and compositing the data with a volumetric rendering for the scene, the feature values associated with one of the fine and the coarse details; and
execute a control task by a controller using the image.
11 . The non-transitory computer-readable medium of claim 10 , wherein the instructions to derive the feature values further include instructions to generate the image as a multi-plane image (MPI) by mixing information directly among the feature planes and between the feature planes using the tri-linear interpolation, wherein the information expands a capacity of the MPI.
12 . A method comprising:
sampling three-dimensional (3D) points of a light ray that crosses a frustum space associated with a single-view camera, the 3D points reflecting depth estimates derived from data that the single-view camera generates for a scene;
deriving feature values for the 3D points using tri-linear interpolation;
generate an image using the tri-linear interpolation across feature planes of the frustum space with the feature values, the feature planes being estimated by a multi-layer perceptron (MLP), and the light ray is associated with generated grid points from adjacent planes within the frustum space, selective objects over the frustum space are unviewable, and the adjacent planes include a dense area within the scene having a density disparity from an increased radiance associated with deriving fine and coarse details for the feature values;
inferring the image in two dimensions (2D) using the MLP by translating the feature values for the 3D points and compositing the data with a volumetric rendering for the scene, the feature values associated with one of the fine and the coarse details; and
executing a control task by a controller using the image.
13 . The method of claim 12 , wherein deriving the feature values further includes generating the image as a multi-plane image (MPI) by mixing information directly among the feature planes and between the feature planes using the tri-linear interpolation, wherein the information expands a capacity of the MPI.
14 . The method of claim 13 , wherein the information describes denser locations within the scene for a driving environment and the denser locations represent mixed colors.
15 . The method of claim 13 further including acquiring the data by a vehicle for a target view and wherein the MPI is associated with an area of the target view.
16 . The method of claim 12 , wherein deriving the feature values further includes generating the image as a multi-plane image (MPI) to acquire information among the feature planes and between the feature planes using the tri-linear interpolation, wherein the information describes regions within the scene that attract attention according to the density disparity.
17 . The method of claim 16 , wherein the density disparity represents an increase in a radiance density between the feature planes associated with the light ray.
18 . The method of claim 16 , wherein inferring the image further includes mapping the feature values to a red-green-blue (RGB) space using the MLP associated with the volumetric rendering.
19 . The method of claim 12 , wherein deriving the feature values further includes generating at least eight grid points from the adjacent planes for the light ray associated with the frustum space and a target view; and
wherein first objects within the frustum space are visible by the single-view camera and second objects outside the frustum space are unviewable by the single-view camera.
20 . The method of claim 12 , wherein the feature values include classes describing objects within the scene and the tri-linear interpolation acquires information about the scene from the adjacent planes of the frustum space.