Bird's eye view (BEV) semantic mapping systems and methods using monocular camera
Bird's eye view (BEV) semantic mapping systems and methods are provided. A method includes receiving an image captured by a monocular camera having a first point of view (POV) of an environment including a plurality of features. The method further includes processing, by an artificial neural network (ANN), the captured image to generate a semantic map for the captured image, the semantic map associated with a second POV different from the first POV. The features exhibit a uniform scale in the semantic map. Additional methods and associated systems are also provided.
1 . A method comprising:
receiving multiple images of an environment that are captured by one or more cameras, the environment comprising a plurality of features; and
processing the captured images by an artificial neural network (ANN), wherein the processing comprises:
generating, for each captured image:
a semantic map, wherein the plurality of features that are represented in the semantic map have a uniform scale in the semantic map; and
a corresponding occlusion mask identifying occluded image locations without a direct line of sight in the corresponding semantic map; and
generating, from the semantic maps and the occlusion masks, a combined semantic map of the environment;
wherein generating the combined semantic map comprises filling one or more of locations that are identified as occluded for at least one of the semantic maps but not for another one of the semantic maps, and using the other one of the semantic maps to perform the filling.
2 . The method of claim 1 ,
wherein at least two of the multiple images of the environment are captured at different points in time.
3 . The method of claim 1 , wherein in the filling, the at least one of the semantic maps and the other one of the semantic maps are captured at different points of time.
4 . The method of claim 1 , wherein:
the one or more cameras comprise a plurality of monocular cameras mounted on a mobile structure and having respective different views of the environment relative to the mobile structure; and
each of the semantic maps and the combined map comprises a bird's eye view (BEV) representation of the environment.
5 . The method of claim 1 , wherein:
the one or more cameras are mounted on a mobile structure; and
the method further comprises at least one of:
aligning, using a coordinate frame transformation matrix, the semantic map to a coordinate frame of the mobile structure; or
removing image drift from the semantic map.
6 . The method of claim 1 , wherein the one or more cameras are mounted on a boat, and the semantic map is an orthographic map to facilitate a boat docking maneuver for the boat.
7 . The method of claim 1 , wherein:
the semantic maps and the combined semantic map of the environment share a POV.
8 . The method of claim 7 , further comprising:
generating a human viewable representation of the environment, the human viewable representation having the same POV as the combined semantic map, wherein the generating comprises processing one or more of:
the combined semantic map;
one of more of the multiple images;
one or more simulated images of the environment;
one or more images provided by a drone;
one or more images provided by a satellite; and/or
one or more photos.
9 . The method of claim 8 , wherein:
the features represented in the combined semantic map and the human viewable map have the uniform scale in the combined semantic map and the human viewable representation.
10 . The method of claim 7 , further comprising:
training the ANN using a plurality of simulated images; and/or
training the ANN by:
generating, by a simulator, a simulated semantic map of the environment, and
comparing the combined semantic map against the simulated semantic map.
11 . A system comprising:
one or more cameras configured to capture multiple images of an environment comprising a plurality of features; and
an artificial neural network (ANN) configured to process the captured images, wherein the processing comprises
generating, for each captured image:
a semantic map, wherein the plurality of features that are represented in the semantic map have a uniform scale in the semantic map; and
a corresponding occlusion mask identifying occluded image locations without a direct line of sight in the corresponding semantic map; and
generating, from the semantic maps and the occlusion masks, a combined semantic map of the environment;
wherein generating the combined semantic map comprises filling one or more of locations that are identified as occluded for at least one of the semantic maps but not for another one of the semantic maps, and using the other one of the semantic maps to perform the filling.
12 . The system of claim 11 , wherein:
at least two of the multiple images of the environment are captured at different points in time.
13 . The system of claim 1 , wherein in the filling, the at least one of the semantic maps and the other one of the semantic maps are captured at different points of time.
14 . The system of claim 12 , wherein:
the one or more cameras comprise a plurality of monocular cameras mounted on a mobile structure and having respective different views of the environment relative to the mobile structure; and
each of the semantic maps and the combined map comprises a bird's eye view (BEV) representation of the environment.
15 . The system of claim 11 , wherein:
the one or more cameras are configured to be mounted on a mobile structure; and
the system is configured to perform at least one of:
aligning the semantic map to a coordinate frame of the mobile structure using a coordinate frame transformation matrix; and/or
remove image drift from the semantic map.
16 . The system of claim 11 , wherein the one or more cameras are mounted on a boat, and the semantic map is an orthographic map to facilitate a boat docking maneuver for the boat.
17 . The system of claim 11 , wherein
the semantic maps and the combined semantic map of the environment share a POV.
18 . The system of claim 17 , wherein:
the ANN is configured to generate a human viewable representation of the environment, the human viewable representation having the same POV as the combined semantic map, wherein the ANN is configured to generate the human viewable representation by processing one or more of:
the combined semantic map;
the image;
the additional images;
one or more simulated images of the environment;
one or more images provided by a drone;
one or more images provided by a satellite; and/or
one or more photos.
19 . The system of claim 18 , wherein:
the features represented in the combined semantic map and the human viewable map have the uniform scale in the combined semantic map and the human viewable representation.
20 . The system of claim 17 , wherein:
the ANN is trained using a plurality of simulated images; and/or
the ANN is trained by comparing the combined semantic map against a simulated semantic map of the environment.