IP Library Granted Patent US 12,633,126
Granted Patent B2
US 12,633,126 · App. 18/815,712 · Granted May 19, 2026

Bird's eye view (BEV) semantic mapping systems and methods using monocular camera

Inventors: Mark Johnson (Vannes Cedez, FR); James Ross (Fareham, GB); Richard Bowden (Surrey, GB); Oscar Mendez Maldonado (Surrey, GB)
Assignee: Raymarine UK Limited
G06V20/56G06T5/77G06T7/10G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,633,126
App. No.
18/815,712
Granted
May 19, 2026
Kind
B2
Abstract

Bird's eye view (BEV) semantic mapping systems and methods are provided. A method includes receiving an image captured by a monocular camera having a first point of view (POV) of an environment including a plurality of features. The method further includes processing, by an artificial neural network (ANN), the captured image to generate a semantic map for the captured image, the semantic map associated with a second POV different from the first POV. The features exhibit a uniform scale in the semantic map. Additional methods and associated systems are also provided.

Claims (73)

1 . A method comprising:

receiving multiple images of an environment that are captured by one or more cameras, the environment comprising a plurality of features; and

processing the captured images by an artificial neural network (ANN), wherein the processing comprises:

generating, for each captured image:

a semantic map, wherein the plurality of features that are represented in the semantic map have a uniform scale in the semantic map; and

a corresponding occlusion mask identifying occluded image locations without a direct line of sight in the corresponding semantic map; and

generating, from the semantic maps and the occlusion masks, a combined semantic map of the environment;

wherein generating the combined semantic map comprises filling one or more of locations that are identified as occluded for at least one of the semantic maps but not for another one of the semantic maps, and using the other one of the semantic maps to perform the filling.

2 . The method of claim 1 ,

wherein at least two of the multiple images of the environment are captured at different points in time.

3 . The method of claim 1 , wherein in the filling, the at least one of the semantic maps and the other one of the semantic maps are captured at different points of time.

4 . The method of claim 1 , wherein:

the one or more cameras comprise a plurality of monocular cameras mounted on a mobile structure and having respective different views of the environment relative to the mobile structure; and

each of the semantic maps and the combined map comprises a bird's eye view (BEV) representation of the environment.

5 . The method of claim 1 , wherein:

the one or more cameras are mounted on a mobile structure; and

the method further comprises at least one of:

aligning, using a coordinate frame transformation matrix, the semantic map to a coordinate frame of the mobile structure; or

removing image drift from the semantic map.

6 . The method of claim 1 , wherein the one or more cameras are mounted on a boat, and the semantic map is an orthographic map to facilitate a boat docking maneuver for the boat.

7 . The method of claim 1 , wherein:

the semantic maps and the combined semantic map of the environment share a POV.

8 . The method of claim 7 , further comprising:

generating a human viewable representation of the environment, the human viewable representation having the same POV as the combined semantic map, wherein the generating comprises processing one or more of:

the combined semantic map;

one of more of the multiple images;

one or more simulated images of the environment;

one or more images provided by a drone;

one or more images provided by a satellite; and/or

one or more photos.

9 . The method of claim 8 , wherein:

the features represented in the combined semantic map and the human viewable map have the uniform scale in the combined semantic map and the human viewable representation.

10 . The method of claim 7 , further comprising:

training the ANN using a plurality of simulated images; and/or

training the ANN by:

generating, by a simulator, a simulated semantic map of the environment, and

comparing the combined semantic map against the simulated semantic map.

11 . A system comprising:

one or more cameras configured to capture multiple images of an environment comprising a plurality of features; and

an artificial neural network (ANN) configured to process the captured images, wherein the processing comprises

generating, for each captured image:

a semantic map, wherein the plurality of features that are represented in the semantic map have a uniform scale in the semantic map; and

a corresponding occlusion mask identifying occluded image locations without a direct line of sight in the corresponding semantic map; and

generating, from the semantic maps and the occlusion masks, a combined semantic map of the environment;

wherein generating the combined semantic map comprises filling one or more of locations that are identified as occluded for at least one of the semantic maps but not for another one of the semantic maps, and using the other one of the semantic maps to perform the filling.

12 . The system of claim 11 , wherein:

at least two of the multiple images of the environment are captured at different points in time.

13 . The system of claim 1 , wherein in the filling, the at least one of the semantic maps and the other one of the semantic maps are captured at different points of time.

14 . The system of claim 12 , wherein:

the one or more cameras comprise a plurality of monocular cameras mounted on a mobile structure and having respective different views of the environment relative to the mobile structure; and

each of the semantic maps and the combined map comprises a bird's eye view (BEV) representation of the environment.

15 . The system of claim 11 , wherein:

the one or more cameras are configured to be mounted on a mobile structure; and

the system is configured to perform at least one of:

aligning the semantic map to a coordinate frame of the mobile structure using a coordinate frame transformation matrix; and/or

remove image drift from the semantic map.

16 . The system of claim 11 , wherein the one or more cameras are mounted on a boat, and the semantic map is an orthographic map to facilitate a boat docking maneuver for the boat.

17 . The system of claim 11 , wherein

the semantic maps and the combined semantic map of the environment share a POV.

18 . The system of claim 17 , wherein:

the ANN is configured to generate a human viewable representation of the environment, the human viewable representation having the same POV as the combined semantic map, wherein the ANN is configured to generate the human viewable representation by processing one or more of:

the combined semantic map;

the image;

the additional images;

one or more simulated images of the environment;

one or more images provided by a drone;

one or more images provided by a satellite; and/or

one or more photos.

19 . The system of claim 18 , wherein:

the features represented in the combined semantic map and the human viewable map have the uniform scale in the combined semantic map and the human viewable representation.

20 . The system of claim 17 , wherein:

the ANN is trained using a plurality of simulated images; and/or

the ANN is trained by comparing the combined semantic map against a simulated semantic map of the environment.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2025
From: JOHNSON, MARK; ROSS, JAMES; BOWDEN, RICHARD; MALDONADO, OSCAR MENDEZ
To: FLIR BELGIUM BVBA
Reel/Frame 070999/0309 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2025
From: FLIR BELGIUM BVBA
To: RAYMARINE UK LIMITED
Reel/Frame 071149/0656 →
Continuity (5)
Continuation PCTUS2023063373 · Feb 27, 2023
Continuation PCTUS2023063369 · Feb 27, 2023
Provisional Application 63396210 · Aug 8, 2022
Provisional Application 63314990 · Feb 28, 2022
Related Publication 20240420483A1 · Dec 19, 2024
References Cited (70)
US 6549660B1 · Lipson et al. · 2003 [cited by applicant]
US 7738707B2 · Wiedemann et al. · 2010 [cited by applicant]
US 8315433B2 · Hsu et al. · 2012 [cited by applicant]
US 8854463B2 · Imamura · 2014 [cited by applicant]
US 10467500B1 · Bao · 2019 [cited by examiner]
US 10678256B2 · Schulter et al. · 2020 [cited by applicant]
US 12211265B2 · Ross et al. · 2025 [cited by applicant]
US 20190050648A1 · Stojanovic · 2019 [cited by examiner]
US 20190096125A1 · Schulter · 2019 [cited by examiner]
US 20200369351A1 · Behrendt · 2020 [cited by applicant]
US 20220398775A1 · Streem · 2022 [cited by examiner]
CN 108052940 · 2018 [cited by applicant]
CN 108876707 · 2022 [cited by applicant]
WO WO2021178603 · 2022 [cited by applicant]
WO WO2023164705 · 2023 [cited by applicant]
WO WO2023164707 · 2023 [cited by applicant]
Roddick T, Cipolla R. Predicting semantic map representations from images using pyramid occupancy networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition 2020 (pp. 11138-11147). (Ye… [cited by examiner]
Abbas et al. “A Geometric Approach to Obtain a Bird's Eye View from an Image”. In: CoRR abs/1905.02231 (2019). [cited by applicant]
Arnold et al. “A comparative study of methods for transductive transfer learning.” In Seventh IEEE International Conference on Data Mining Workshops (ICDMW 2007), pp. 77-82, 2007. [cited by applicant]
Bailey et al. “Simultaneous localization and mapping: part ii”. In: IEEE Robotics Automation Magazine 13.3 (2006), pp. 108-117. [cited by applicant]
Bousmalis et al. “Unsupervised pixel-level domain adaptation with generative adversarial networks.” CoRR, abs/1612.05424, 2016. [cited by applicant]
Bruls et al. “The right (angled) perspective: improving the understanding of road scenes using boosted inverse perspective mapping.” 2019. [cited by applicant]
Cherian et al. “Sem-GAN: Semantically-Consistent Image-To-Image Translation,CoRR” abs/1908.04409, 2018. [cited by applicant]
Chopra et al. “DLID: Deep learning for domain adaptation by interpolating between domains.” In ICML Workshop on Challenges in Representation Learning, 2013. [cited by applicant]
Durrant-Whyte et al. “Simultaneous localization and mapping: part i”. In: IEEE Robotics Automation Magazine 13.2 (2006), pp. 99-110. [cited by applicant]
Engel et al. “Semi-dense visual odometry for a monocular camera,” 2013 IEEE International Conference on Computer Vision, 2013, pp. 1449-1456. [cited by applicant]
Evangelidis et al. “Parametric image alignment using enhanced correlation coefficient maximization.” IEEE Transactions on Pattern Analysis and Machine Intelligence, Institute of Electrical and Electronics Engineers, 200… [cited by applicant]
Ganin et al. “Domain-adversarial training of neural networks,” 2016. [cited by applicant]
Ghifary et al. “Domain adaptive neural networks for object recognition.” CoRR, abs/1409.6041, 2014. [cited by applicant]
Ghifary et al. “Domain generalization for object recognition with multi-task autoencoders.” CoRR, abs/1508.07680, 2015. [cited by applicant]
Ghifary et al. “Deep reconstruction-classification networks for unsupervised domain adaptation.” CoRR, abs/1607.03516, 2016. [cited by applicant]
Gretton et al. “A kernel two-sample test. J. Mach.” Learn. Res., 13(null):723-773, Mar. 2012. [cited by applicant]
Guo et al. “Beyond the line of sight: labeling the underlying surfaces”. English (US). In: Computer Vision, ECCV 2012—12th European Conference on Computer Vision, Proceedings. Part 5.Oct. 2012, pp. 761-774. [cited by applicant]
Hou et al. “Convolutional neural networkbased image representation for visual loop closure detection,” CoRR, vol. abs/1504.05241, 2015. [cited by applicant]
Hu et al. “FIERY: future instance prediction in bird's-eye view from surround monocular cameras”. [cited by applicant]
Huang et al. “Arbitrary style transfer in real-time with adaptive instance normalization.” CoRR, abs/1703.06868, 2017. [cited by applicant]
Kendall et al. “Convolutional networks for realtime 6-dof camera relocalization,” CoRR, vol. abs/1505.07427, 2015. [cited by applicant]
Kim et al. “Learning to discover cross-domain relations with generative adversarial networks.” CoRR, abs/1703.05192, 2017. [cited by applicant]
Leonard et al. “Simultaneous map building and localization for an autonomous mobile robot,” Proceedings IROS '91: IEEE/RSJ International Workshop on Intelligent Robots and Systems, 1991, pp. 1442-1447. [cited by applicant]
Li et al. “Revisiting batch normalization for practical domain adaptation.” CoRR, abs/1603.04779, 2016. [cited by applicant]
Long et al. “Learning transferable features with deep adaptation networks.” In Proceedings of the 32nd International Conference on International Conference on Machine Learning—vol. 37, ICML'15, p. 97-105. JMLR.org, 2015. [cited by applicant]
Lu et al. “Monocular semantic occupancy grid mapping with convolutional variational autoencoders,” CoRR, vol. abs/1804.02176, 2018. [cited by applicant]
Lu et al. “Monocular semantic occupancy grid mapping with convolutional variational encoder-decoder networks”. In: IEEE Robotics and Automation Letters 4.2 (Apr. 2019), pp. 445-452. [cited by applicant]
Mani et al. “MonoLayout: Amodal scene layout from a single image.” 2020. [cited by applicant]
Mur-Artal et al. “ORB-SLAM: a versatile and accurate monocular SLAM system.” IEEE Transactions on Robotics, 31(5):1147-1163, 2015. [cited by applicant]
Mur-Artal et al. “ORB-SLAM2: an open-source SLAM system for monocular, stereo and RGB-D cameras.” IEEE Transactions on Robotics, 33(5):1255-1262, 2017. [cited by applicant]
Pan et al. “Cross-view semantic segmentation for sensing surroundings”. In: IEEE Robotics and Automation Letters 5.3 (Jul. 2020), pp. 4867-4873. [cited by applicant]
Philion et al. “Lift, splat, shoot: encoding images from arbitrary camera rigs by implicitly unprojecting to 3D”. In: CoRR abs/2008.05711 (2020). [cited by applicant]
Ragot et al. “Benchmark of visual slam algorithms: ORB-SLAM2 vs RTAB-Map,” 2019 Eighth International Conference on Emerging Security Technologies (EST), 2019, pp. 1-6. [cited by applicant]
Rao et al. R1-cyclegan: “Reinforcement learning aware simulation-to-real.” CoRR, abs/2006.09001, 2020. [cited by applicant]
Regmi et al. “Cross-view image synthesis using conditional GANs.” 2018. [cited by applicant]
Roddick et al. “Orthographic feature transform for monocular 3D object detection.” CoRR, abs/1811.08188, 2018. [cited by applicant]
Roddick et al. “Predicting Semantic Map Representations From Images Using Pyramid Occupancy Networks”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 13, 2020 (Jun. 13, 2020), pp.… [cited by applicant]
Ross et al., “BEV-SLAM: Building a Globally-Consistent World Map using Monocular Vision,” IEEE RJS Int. Conf. on Intelligent Robots and Systems (IROS) Oct. 23-27, 2022. [cited by applicant]
Saha et al. “Enabling spatio-temporal aggregation in birds-eyeview vehicle estimation”. In: ICRA (2021). [cited by applicant]
Saha et al. “Translating Images into Maps”, 2022 International Conference on Robotics and Automation (ICRA) Oct. 3, 2021 (Oct. 3, 2021), pp. 9200-9206 Retrieved from the Internet: URL: https://arxiv.org/pdf/2110.00966vl… [cited by applicant]
Saha et al. “The Pedestrian next to the Lamppost” Adaptive Object Graphs for Better Instantaneous Mapping, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 18, 2022 (Jun. 18, 2022),… [cited by applicant]
Schulter et al. “Learning to look around objects for top-view representations of outdoor scenes”. In: CoRR abs/1803.10870 (2018). [cited by applicant]
Sengupta et al. “Automatic dense visual semantic mapping from street-level imagery”. In: 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems. 2012, pp. 857-862. [cited by applicant]
Sun et al. “Return of frustratingly easy domain adaptation.” CoRR, abs/1511.05547, 2015. [cited by applicant]
Tighe et al. “Scene parsing with object instances and occlusion ordering”. In: 2014 IEEE Conference on Computer Vision and Pattern Recognition. 2014, pp. 3748-3755. [cited by applicant]
Tzeng et al. “Deep domain confusion: Maximizing for domain invariance.” CoRR, abs/1412.3474, 2014. [cited by applicant]
Wang et al. “Deep visual domain adaptation: A survey.” In Neurocomputing, 2018. [cited by applicant]
Weiss et al. “A survey of transfer learning.” Journal of Big Data, 2016. [cited by applicant]
Winkelbauer et al. “Learning to localize in new environments from synthetic training data,” CoRR, vol. abs/2011.04539, 2020. [cited by applicant]
Yi et al. “Unsupervised dual learning for image-to-image translation.” CoRR, abs/1704.02510, 2017. [cited by applicant]
Zhai et al. “Predicting ground-level scene layout from aerial imagery”. In: CoRR abs/1612.02709 (2016). [cited by applicant]
Zhu et al. “Unpaired image-to-image translation using cycle-consistent adversarial networks.” CoRR, abs/1703.10593, 2017. [cited by applicant]
Zhu et al. “Generative adversarial frontal view to bird view synthesis.” 2019. [cited by applicant]
Zhuang et al. “Supervised representation learning: Transfer learning with deep autoencoders.” In Proceedings of the 24th International Conference on Artificial Intelligence, IJCAI'15, p. 4119-4125. AAAI Press, 2015. [cited by applicant]