Retrofit vision assist with monocular depth estimation
Systems, methods, and other embodiments described herein relate to retrofit vision assist systems for vehicles based on monocular depth estimation. In one embodiment, a method for operating a depth system of a vehicle includes acquiring a camera image from a monocular camera that is retrofitted to a vehicle. The method also includes estimating depth information of features in the camera image to output a depth map including three-dimensional data associated with the camera image. The method also includes generating an annotated image based on the three-dimensional data. The method further includes displaying the annotated image to aid a driver in understanding a surrounding environment of the vehicle.
1 . A system comprising:
a processor;
a memory communicably coupled to the processor and storing:
a depth estimation module including instructions that, when executed by the processor, cause the processor to:
calibrate a monocular camera retrofitted to a vehicle by receiving information regarding one or more dimensions of the vehicle, constructing a bounding box to relate the one or more dimensions with the monocular camera, and defining a location of the monocular camera relative to the vehicle using the bounding box;
acquire a camera image from the monocular camera and additional images from additional cameras on the vehicle;
estimate depth information of features including objects in the camera image and the additional images to by performing depth estimation using the bounding box to identify locations of the objects relative to the vehicle and output depth maps including three-dimensional data associated with the camera image and the additional images;
generate an annotated image including a three-dimensional representation of a scene surrounding the vehicle based on the three-dimensional data; and
display the annotated image to aid a driver in understanding the scene by superimposing colors onto the objects according to the depth information.
2 . The system of claim 1 , wherein the instructions to estimate the depth information of the features include instructions to perceive objects in a field-of-view of the monocular camera and determine distances of the objects in the field-of-view from the monocular camera using a depth model that performs monocular depth estimation.
3 . The system of claim 1 , wherein the instructions to estimate the depth information of the features include instructions to perceive objects in a field-of-view of the monocular camera and determine distances of the objects in the field-of-view from the monocular camera, and wherein the instructions to generate the annotated image include instructions to colorize the objects in the field-of-view according to a risk of impact between the vehicle and the objects, wherein the risk of impact is calculated based on travel information of the vehicle, wherein the travel information includes at least one of a location, speed, acceleration, and heading of the vehicle.
4 . The system of claim 1 , wherein the instructions to estimate depth information of features include instructions to i) output a depth map including three-dimensional data associated with the scene depicted by the camera image, ii) perceive objects in a field-of-view of the monocular camera, and iii) determine distances of the objects in the field-of-view from the monocular camera according to the depth map, and wherein the instructions to generate the annotated image include instructions to highlight objects in the field-of-view that are within a distance threshold to the monocular camera.
5 . The system of claim 1 , wherein the additional cameras are retrofitted to different areas of the vehicle, and wherein the instructions to estimate depth information of features in the additional images to output depth maps include instructions to identify relative locations of the additional cameras with respect to the vehicle.
6 . A non-transitory computer-readable medium including instructions that, when executed by a processor, cause the processor to:
calibrate a monocular camera retrofitted to a vehicle by receiving information regarding one or more dimensions of the vehicle, constructing a bounding box to relate the one or more dimensions with the monocular camera, and defining a location of the monocular camera relative to the vehicle using the bounding box;
acquire a camera image from the monocular camera and additional images from additional cameras on the vehicle;
estimate depth information of features including objects in the camera image and the additional images by performing depth estimation using the bounding box to identify locations of the objects relative to the vehicle and output depth maps including three-dimensional data associated with the camera image and the additional images;
generate an annotated image including a three-dimensional representation of a scene surrounding the vehicle based on the three-dimensional data; and
display the annotated image to aid a driver in understanding the scene by superimposing colors onto the objects according to the depth information.
7 . The non-transitory computer-readable medium of claim 6 , wherein the instructions to estimate depth information of the features include instructions to perceive objects in a field-of-view of the monocular camera and determine distances of the objects in the field-of-view from the monocular camera using a depth model that performs monocular depth estimation.
8 . The non-transitory computer-readable medium of claim 6 , wherein the instructions to estimate depth information of features include instructions to perceive objects in a field-of-view of the monocular camera and determine distances of the objects in the field-of-view from the monocular camera, and wherein the instructions to generate the annotated image include instructions to colorize the objects in the field-of-view according to a risk of impact between the vehicle and the objects in the field-of-view, wherein the risk of impact is calculated based on travel information of the vehicle, wherein the travel information includes at least one of a location, speed, acceleration, and heading of the vehicle.
9 . The non-transitory computer-readable medium of claim 6 , wherein the instructions to estimate depth information of features include instructions to i) output a depth map including three-dimensional data associated with the scene depicted by the camera image, ii) perceive objects in a field-of-view of the monocular camera, and iii) determine distances of the objects in the field-of-view from the monocular camera according to the depth map, and wherein the instructions to generate the annotated image include instructions to highlight objects in the field-of-view that are within a distance threshold to the monocular camera.
10 . A method comprising:
calibrating a monocular camera retrofitted to a vehicle by receiving information regarding one or more dimensions of the vehicle, constructing a bounding box to relate the one or more dimensions with the monocular camera, and defining a location of the monocular camera relative to the vehicle using the bounding box;
acquiring a camera image from the monocular camera and additional images from additional cameras on the vehicle;
estimating depth information of features including objects in the camera image and the additional images by performing depth estimation using the bounding box to identify locations of the objects relative to the vehicle and output depth maps including three-dimensional data associated with the camera image and the additional images;
generating an annotated image including a three-dimensional representation of a scene surrounding the vehicle based on the three-dimensional data; and
displaying the annotated image to aid a driver in understanding the scene by superimposing colors onto the objects according to the depth information.
11 . The method of claim 10 , wherein estimating the depth information of the features in the camera image includes perceiving objects in a field-of-view of the monocular camera and determining distances of the objects in the field-of-view from the monocular camera using a depth model that performs monocular depth estimation.
12 . The method of claim 10 , wherein estimating depth information of features in the camera image includes perceiving objects in a field-of-view of the monocular camera and determining distances of the objects in the field-of-view from the monocular camera, and wherein generating the annotated image includes colorizing the objects in the field-of-view according to a risk of impact between the vehicle and the objects in the field-of-view, wherein the risk of impact is calculated based on travel information of the vehicle, wherein the travel information includes at least one of a location, speed, acceleration, and heading of the vehicle.
13 . The method of claim 10 , wherein estimating depth information of features in the camera image includes i) outputting a depth map including three-dimensional data associated with the scene depicted by the camera image, ii) perceiving objects in a field-of-view of the monocular camera, and iii) determining distances of the objects in the field-of-view from the monocular camera according to the depth map, and wherein generating the annotated image includes colorizing the objects in the field-of-view according to the distances including highlighting objects in the field-of-view that are within a distance threshold to the monocular camera.
14 . The method of claim 10 , wherein the additional cameras are retrofitted to different areas of the vehicle, and wherein estimating depth information of features in the additional images to output depth maps includes perceiving relative locations of the additional cameras with respect to the vehicle.