Distance estimation using a geometrical distance aware machine learning model
Techniques and systems are provided for generating depth information for an image. For instance, a process can include obtaining one or more images of an environment. The process can further include generating a set of features for the one or more images. The process can also include combining the set of features with one or more distance maps to generate combined feature distance information, wherein the one or more distance maps indicate distances based on relative height above a ground level. The process can further include generating depth information of the environment based on the combined feature distance information, and outputting the depth information of the environment.
1 . An apparatus for generating depth information, comprising:
at least one memory comprising instructions; and
at least one processor coupled to the at least one memory and configured to:
obtain one or more images of an environment;
generate a set of features for the one or more images;
combine the set of features with one or more distance maps to generate combined feature distance information, wherein the one or more distance maps indicate distances based on a height of features above a ground plane in the one or more images, and wherein the one or more distance maps are predetermined;
generate depth information of the environment based on the combined feature distance information; and
output the depth information of the environment.
2 . The apparatus of claim 1 , wherein the one or more images are obtained from a monocular camera.
3 . The apparatus of claim 2 , wherein the one or more distance maps are based on camera properties of the monocular camera.
4 . The apparatus of claim 1 , wherein the at least one processor is further configured to generate the depth information of the environment using a machine learning model.
5 . The apparatus of claim 1 , wherein the at least one processor is further configured to combine the set of features with one or more distance maps by concatenating or multiplying one or more values of the one or more distance maps to the set of features.
6 . The apparatus of claim 1 , wherein the at least one processor is further configured to select a distance map based on an estimated height of a feature of the set of features.
7 . The apparatus of claim 1 , wherein the at least one processor is further configured to generate, based on the depth information of the environment, at least one of a segmented depth map of the environment, a segmented birds-eye-view (BEV) of the environment, or three-dimensional location information for one or more objects in the environment.
8 . The apparatus of claim 1 , wherein the at least one processor is further configured to:
perform a transformation operation based on the depth information of the environment and the set of features to generate a transformed set of features; and
output the transformed set of features, wherein the depth information of the environment is output implicitly in the transformed set of features.
9 . A method for generating depth information, comprising:
obtaining one or more images of an environment;
generating a set of features for the one or more images;
combining the set of features with one or more distance maps to generate combined feature distance information, wherein the one or more distance maps indicate distances based on a height of features above a ground plane in the one or more images, and wherein the one or more distance maps are predetermined;
generating depth information of the environment based on the combined feature distance information; and
outputting the depth information of the environment.
10 . The method of claim 9 , wherein the one or more images are obtained from a monocular camera.
11 . The method of claim 10 , wherein the one or more distance maps are based on camera properties of the monocular camera.
12 . The method of claim 9 , further comprising generating the depth information of the environment using a machine learning model.
13 . The method of claim 9 , further comprising combining the set of features with one or more distance maps by concatenating or multiplying one or more values of the one or more distance maps to the set of features.
14 . The method of claim 9 , further comprising selecting a distance map based on an estimated height of a feature of the set of features.
15 . The method of claim 9 , further comprising generating, based on the depth information of the environment, at least one of a segmented depth map of the environment, a segmented birds-eye-view (BEV) of the environment, or three-dimensional location information for one or more objects in the environment.
16 . The method of claim 9 , further comprising:
performing a transformation operation based on the depth information of the environment and the set of features to generate a transformed set of features; and
outputting the transformed set of features, wherein the depth information of the environment is output implicitly in the transformed set of features.
17 . A non-transitory computer-readable medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to:
obtain one or more images of an environment;
generate a set of features for the one or more images;
combine the set of features with one or more distance maps to generate combined feature distance information, wherein the one or more distance maps indicate distances based on a height of features above a ground plane in the one or more images, and wherein the one or more distance maps are predetermined;
generate depth information of the environment based on the combined feature distance information; and
output the depth information of the environment.
18 . The non-transitory computer-readable medium of claim 17 , wherein the one or more images are obtained from a monocular camera.
19 . The non-transitory computer-readable medium of claim 18 , wherein the one or more distance maps are based on camera properties of the monocular camera.
20 . The non-transitory computer-readable medium of claim 17 , wherein instructions further cause the at least one processor to generate the depth information of the environment using a machine learning model.
21 . The non-transitory computer-readable medium of claim 17 , wherein instructions further cause the at least one processor to combine the set of features with one or more distance maps by concatenating or multiplying one or more values of the one or more distance maps to the set of features.
22 . The non-transitory computer-readable medium of claim 17 , wherein instructions further cause the at least one processor to select a distance map based on an estimated height of a feature of the set of features.
23 . The non-transitory computer-readable medium of claim 17 , wherein instructions further cause the at least one processor to generate, based on the depth information of the environment, at least one of a segmented depth map of the environment, a segmented birds-eye-view (BEV) of the environment, or three-dimensional location information for one or more objects in the environment.
24 . The non-transitory computer-readable medium of claim 17 , wherein instructions further cause the at least one processor to:
perform a transformation operation based on the depth information of the environment and the set of features to generate a transformed set of features; and
output the transformed set of features, wherein the depth information of the environment is output implicitly in the transformed set of features.
25 . An apparatus for generating depth information comprising:
means for obtaining one or more images of an environment;
means for generating a set of features for the one or more images;
means for combining the set of features with one or more distance maps to generate combined feature distance information, wherein the one or more distance maps indicate distances based on a height of features above a ground plane in the one or more images, and wherein the one or more distance maps are predetermined;
means for generating depth information of the environment based on the combined feature distance information; and
means for outputting the depth information of the environment.
26 . The apparatus of claim 25 , wherein the one or more images are obtained from a monocular camera.
27 . The apparatus of claim 26 , wherein the one or more distance maps are based on camera properties of the monocular camera.
28 . The apparatus of claim 26 , further comprising means for generating the depth information of the environment using a machine learning model.
29 . The apparatus of claim 26 , further comprising means for combining the set of features with one or more distance maps by concatenating or multiplying one or more values of the one or more distance maps to the set of features.
30 . The apparatus of claim 1 , wherein the locations of the features in the one or more images comprise a distance of a feature from a bottom or a top of an image.