IP Library Granted Patent US 12,561,906
Granted Patent B2
US 12,561,906 · App. 18/154,219 · Granted Feb 24, 2026

Method for generating at least one ground truth from a bird's eye view

Inventors: Denis Tananaev (Sindelfingen, DE); Ze Guo (Berlin, DE)
Assignee: ROBERT BOSCH GMBH
G06T17/05G01S17/89G06T5/20G06T7/10G06T2207/20036
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,906
App. No.
18/154,219
Granted
Feb 24, 2026
Kind
B2
Abstract

A method for generating at least one image from a bird's eye view. The method includes: a) carrying out a sensor data point cloud compression; b) carrying out a point cloud filtering in a camera perspective; c) carrying out an object completion; and d) carrying out a bird's eye view segmentation and generating an elevation map.

Claims (62)

1 . A method for generating at least one representation from a bird's eye view, the method comprising the following steps:

a) carrying out a sensor data point cloud compression;

b) carrying out a point cloud filtering in a perspective of a camera;

c) carrying out an object completion; and

d) carrying out a bird's eye view segmentation and generating an elevation map;

wherein the method includes at least one of the following features (I)-(III):

(I) the cloud compression includes at least one of:

(i) compressing data points from respective frames of a plurality of respective points in time into a single frame of data points; and

(ii) removing respective data points based on the respective data points corresponding to dynamic objects;

(II) the point cloud filtering includes at least one of:

(i) removing data points that correspond to a location that is outside a viewing frustum of the camera; and

(ii) applying the data points to a matrix having extrinsic parameters of pose of the camera in space and intrinsic parameters of the camera; and

(III) the object completion includes at least one of:

(i) a neural network completing partially visible objects to which at least a part of the point cloud corresponds; and

(ii) carrying out a morphological operation that fills in a missing portion of an incomplete object to which the at least the part of the point cloud corresponds.

2 . The method according to claim 1 , wherein the representation is also at least based on sensor data obtained from at least one active surroundings sensor, wherein the at least one active surroundings sensor includes a LiDAR sensor and/or a radar sensor.

3 . The method according to claim 1 , wherein, in step b), at least one camera parameter is used to back-project sensor data points onto a current image area.

4 . The method according to claim 1 , wherein, in step c), at least one cuboid box is projected onto a current bird's eye view region.

5 . The method according to claim 1 , wherein in step d), all 3D points for a valid bird's eye view region are collected.

6 . The method according to claim 1 , wherein the cloud compression includes the compressing of the data points from the respective frames of the plurality of respective points in time into the single frame of data points.

7 . The method according to claim 1 , wherein the cloud compression includes the removing of the respective data points based on the respective data points corresponding to the dynamic objects.

8 . The method according to claim 7 , wherein, in step a), the removing is performed based on a splitting of a sensor data point cloud is split into static and dynamic object points on the basis of semantic information.

9 . The method according to claim 1 , wherein the point cloud filtering includes the removing of the data points that correspond to the location that is outside the viewing frustum of the camera.

10 . The method according to claim 1 , wherein the point cloud filtering includes the applying of the data points to the matrix having the extrinsic parameters of pose of the camera in space and the intrinsic parameters of the camera.

11 . The method according to claim 1 , wherein the object completion includes the neural network completing the partially visible objects to which the at least the part of the point cloud corresponds.

12 . The method according to claim 1 , wherein the object completion includes the carrying out of the morphological operation that fills in the missing portion of the incomplete object to which the at least the part of the point cloud corresponds.

13 . The method according to claim 12 , wherein, in step c), the morphological operation is applied to at least one object of a current bird's eye view region.

14 . The method according to claim 12 , wherein the morphological operation is a dilation.

15 . The method according to claim 1 , wherein:

the cloud compression includes the compressing of the data points from the respective frames of the plurality of respective points in time into the single frame of data points;

the respective points in time include a keyframe point in time and non-keyframe points in time that are at least one of before and after the keyframe point in time; and

the cloud compression further includes performing for each of the non-keyframe points in time, the removing of the respective data points based on the respective data points corresponding to the dynamic objects, while keeping respective data points corresponding to dynamic objects in the keyframe.

16 . A non-transitory machine-readable storage medium on which is stored a computer program for generating at least one representation from a bird's eye view, the computer program, when executed by a computer, causing the computer to perform a method having the following steps:

a) carrying out a sensor data point cloud compression;

b) carrying out a point cloud filtering in a perspective of a camera;

c) carrying out an object completion; and

d) carrying out a bird's eye view segmentation and generating an elevation map;

wherein the method includes at least one of the following features (I)-(III);

(I) the cloud compression includes at least one of:

(i) compressing data points from respective frames of a plurality of respective points in time into a single frame of data points; and

(ii) removing respective data points based on the respective data points corresponding to dynamic objects;

(II) the point cloud filtering includes at least one of:

(i) removing data points that correspond to a location that is outside a viewing frustum of the camera; and

(ii) applying the data points to a matrix having extrinsic parameters of pose of the camera in space and intrinsic parameters of the camera; and

(III) the object completion includes at least one of:

(i) a neural network completing partially visible objects to which at least a part of the point cloud corresponds; and

(ii) carrying out a morphological operation that fills in a missing portion of an incomplete object to which the at least the part of the point cloud corresponds.

17 . An object recognition system for a vehicle, the system configured to generate at least one representation from a bird's eye view, the system configured to

a) carry out a sensor data point cloud compression;

b) carry out a point cloud filtering in a perspective of a camera;

c) carry out an object completion; and

d) carry out a bird's eye view segmentation and generate an elevation map;

wherein the object recognition system includes at least one of the following features (I)-(III):

(I) the cloud compression includes at least one of:

(i) compressing data points from respective frames of a plurality of respective points in time into a single frame of data points; and

(ii) removing respective data points based on the respective data points corresponding to dynamic objects;

(II) the point cloud filtering includes at least one of:

(i) removing data points that correspond to a location that is outside a viewing frustum of the camera; and

(ii) applying the data points to a matrix having extrinsic parameters of pose of the camera in space and intrinsic parameters of the camera; and

(III) the object completion includes at least one of:

(i) a neural network completing partially visible objects to which at least a part of the point cloud corresponds; and

(ii) carrying out a morphological operation that fills in a missing portion of an incomplete object to which the at least the part of the point cloud corresponds.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2023
From: TANANAEV, DENIS; GUO, ZE
To: ROBERT BOSCH GMBH
Reel/Frame 062840/0769 →
Priority Claims (2)
DE 10 2022 200 503.1 · Jan 18, 2022 · national
DE 10 2022 214 330.2 · Dec 22, 2022 · national
Continuity (1)
Related Publication 20230230317A1 · Jul 20, 2023
References Cited (23)
US 11354913B1 · Houston · 2022 [cited by examiner]
US 20190266418A1 · Xu · 2019 [cited by examiner]
US 20190371052A1 · Kehl · 2019 [cited by examiner]
US 20200025931A1 · Liang · 2020 [cited by examiner]
US 20210063578A1 · Wekel · 2021 [cited by examiner]
US 20210146952A1 · Vora · 2021 [cited by examiner]
US 20210201569A1 · Marschner · 2021 [cited by examiner]
US 20210278852A1 · Urtasun · 2021 [cited by examiner]
US 20210342608A1 · Smolyanskiy et al. · 2021 [cited by applicant]
US 20210365697A1 · Vaquero Gomez · 2021 [cited by examiner]
US 20220012466A1 · Taghavi · 2022 [cited by examiner]
US 20220114764A1 · Beijbom · 2022 [cited by examiner]
US 20220214457A1 · Liang · 2022 [cited by examiner]
US 20220317305A1 · Chou · 2022 [cited by examiner]
US 20230105331A1 · Cheng · 2023 [cited by examiner]
US 20230136860A1 · Wang · 2023 [cited by examiner]
Cao, et al.: “Multi-View Frustum Pointnet for Object Detection in Autonomous Driving,” 2019 IEEE International Conference on Image Processing (ICIP) IEEE, (2019), pp. 3896-3899; doi: 10.1109/ICIP.2019.8803572. [cited by applicant]
Imad, et al.: “Transfer Learning Based Semantic Segmentation for 3D Object Detection from Point Cloud,” Sensors, 21, (2021), pp. 1-15. doi:10.3390/s21123964. [cited by applicant]
Meyer, et al.: “Sensor Fusion for Joint 3D Object Detection and Semantic Segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, (2019), pp. 1-8; https://openaccess.th… [cited by applicant]
Qi, et al.: “Offboard 3D Object Detection from Point Cloud Sequences,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2021), pp. 6134-6144; https://openaccess. thecvf.com/content/CVP… [cited by applicant]
Wen, et al . . . : “Fast and Accurate 3D Object Detection for Lidar-Camera-Based Autonomous Vehicles Using One Shared Voxel-Based Backbone,” IEEE Access 9, (2021), pp. 22080-22089. doi:10.1109/ACCESS.2021.3055491. [cited by applicant]
Yang, et al.: “IPOD: Intensive Pint-based Oject Dtector for Pint Cloud,” arXiv:1812.05276, 2018, pp. 1-9; doi:10.48550/arXiv.1812.05276. [cited by applicant]
Higgins, Sean: “How dynamic object removal helps you capture active sites with ease”, Blog, (2021), pp. 1-8, with English translation; https://www.navvis.com/blog/how-dynamic-object-removal-helps-you-capture-active-site… [cited by applicant]