Monocular depth estimation system
Systems and methods are directed to monocular depth estimation. A method includes capturing an image; encoding the image to form an encoded image; masking one or more dynamic objects in the encoded image; predicting a final prediction performance based on one or more unmasked static objects from the encoded image; decoding and projecting the final prediction performance; and producing an information based on Global Positioning System (GPS), Inertial Measurement Unit (IMU), and wheel encoder fusion.
1 . A method for estimating monocular depth, comprising:
capturing an image;
encoding the image to form an encoded image;
comparing a projected mask in one or more dynamic objects with a pre-defined mask in the encoded image to predict an overlap ratio;
updating a pre-defined one or more dynamic objects as a static object or a dynamic object based on the overlap ratio;
masking one or more dynamic objects in the encoded image;
predicting a final prediction performance based on one or more unmasked static objects from the encoded image;
decoding and projecting the final prediction performance;
producing an information based on Global Positioning System (GPS), Inertial Measurement Unit (IMU), and wheel encoder fusion; and
estimating the monocular depth based on the final prediction performance and the information.
2 . A method for estimating monocular depth, comprising:
capturing an image;
encoding the image to form an encoded image;
comparing a projected mask in one or more dynamic objects with a pre-defined mask in the encoded image to predict an overlap ratio;
updating a pre-defined one or more dynamic objects to either a static object or a dynamic object based on the overlap ratio;
predicting a final prediction performance based on one or more unmasked static objects from the encoded image; and
decoding and projecting the final prediction performance;
producing an information based on Global Positioning System (GPS), Inertial Measurement Unit (IMU), and wheel encoder fusion; and
estimating the monocular depth based on the final prediction performance and the information.
3 . The method of claim 1 , wherein masking one or more dynamic objects in the encoded image comprises masking the one or more dynamic objects with a self-detected mask on calculating loss in training.
4 . The method of claim 3 , wherein the self-detected mask is formed by comparing an estimated depth and a projected depth map.
5 . The method of claim 1 , wherein the one or more dynamic objects is masked from a final loss function.
6 . The method of claim 5 , wherein the final loss function is calculated to improve the final prediction performance.
7 . The method of claim 6 , wherein the final loss function comprises a re-projection loss, a smoothness loss, and a geometry consistency loss.
8 . The method of claim 7 , wherein the re-projection loss is summation of photometric loss and structural similarity (SSIM) difference.
9 . The method of claim 7 , wherein the geometry consistency loss comprises one or more re-projection weights.
10 . A system for estimating monocular depth, comprising:
a memory and a processor, wherein the memory is connected to the processor;
the memory stores a computer program; and
the processor implements the method of claim 1 when executing the computer program.
11 . A system for estimating monocular depth, comprising:
a processor, wherein the memory is connected to the processor;
the memory stores a computer program; and
the processor implements the method of claim 2 when executing the computer program.