IP Library Granted Patent US 12,517,252
Granted Patent B2
US 12,517,252 · App. 17/971,069 · Granted Jan 6, 2026

Three-dimensional object detection

Inventors: Ming Liang (Toronto, CA); Bin Yang (Toronto, CA); Shenlong Wang (Toronto, CA); Wei-Chiu Ma (Toronto, CA); Raquel Urtasun (Toronto, CA)
Assignee: AURORA OPERATIONS, INC.
G01S17/89G01S7/4817G01S17/931G05D1/0231G06F18/253G06N3/02G06N3/04G06N3/08G06V10/764G06V10/803G06V20/56G06V20/64G06T2207/30261
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,517,252
App. No.
17/971,069
Granted
Jan 6, 2026
Kind
B2
Abstract

Generally, the disclosed systems and methods implement improved detection of objects in three-dimensional (3D) space. More particularly, an improved 3D object detection system can exploit continuous fusion of multiple sensors and/or integrated geographic prior map data to enhance effectiveness and robustness of object detection in applications such as autonomous driving. In some implementations, geographic prior data (e.g., geometric ground and/or semantic road features) can be exploited to enhance three-dimensional object detection for autonomous vehicle applications. In some implementations, object detection systems and methods can be improved based on dynamic utilization of multiple sensor modalities. More particularly, an improved 3D object detection system can exploit both LIDAR systems and cameras to perform very accurate localization of objects within three-dimensional space relative to an autonomous vehicle. For example, multi-sensor fusion can be implemented via continuous convolutions to fuse image data samples and LIDAR feature maps at different levels of resolution.

Claims (54)

1 . A method, comprising:

obtaining image data and point cloud data, the image data and the point cloud data descriptive of an environment of an autonomous vehicle, the image data associated with a camera view;

for a respective portion of the environment:

determining one or more datapoints of the point cloud data associated with the respective portion;

projecting the one or more datapoints of the point cloud data into the camera view;

determining one or more portions of the image data respectively corresponding to the projected one or more datapoints of the point cloud; and

generating, based at least in part on the one or more portions of the image data, an output feature associated with the respective portion of the environment; and

generating an output feature map based at least in part on the output feature, the output feature map corresponding to a two-dimensional overhead view of the environment, the output feature map comprising a fusion of image data features and point cloud data features.

2 . The method of claim 1 , comprising:

obtaining a bird's-eye-view point cloud map associated with the respective portion, the bird's-eye-view point cloud map generated using a bird's-eye-view projection of the point cloud data.

3 . The method of claim 2 , wherein the output feature map is characterized by a data structure also associated with the bird's-eye-view point cloud map.

4 . The method of claim 3 , comprising:

generating a dense bird's-eye-view feature map using the output feature map and the bird's-eye-view point cloud map.

5 . The method of claim 1 , wherein the output feature map is generated using a learnable continuous convolution operator.

6 . The method of claim 1 , comprising:

controlling the autonomous vehicle based on the output feature map.

7 . The method of claim 6 , comprising:

performing, in a point-cloud metric space, and based at least in part on the output feature map, at least one of object detection or motion planning.

8 . The method of claim 6 , wherein the output feature map is fused with point-cloud data convolution layers.

9 . One or more non-transitory computer-readable media storing instructions that are executable by one or more processors to cause an autonomous vehicle control system to perform operations, the operations comprising:

obtaining image data and point cloud data, the image data and the point cloud data descriptive of an environment of an autonomous vehicle, the image data associated with a camera view;

for a respective portion of the environment:

determining one or more datapoints of the point cloud data associated with the respective portion;

projecting the one or more datapoints of the point cloud data into the camera view;

determining one or more portions of the image data respectively corresponding to the projected one or more datapoints of the point cloud; and

generating, based at least in part on the one or more portions of the image data, an output feature associated with the respective portion of the environment; and

generating an output feature map based at least in part on the output feature, the output feature map corresponding to a two-dimensional overhead view of the environment, the output feature map comprising a fusion of image data features and point cloud data features.

10 . The one or more non-transitory computer-readable media of claim 9 , wherein the operations comprise:

obtaining a bird's-eye-view point cloud map associated with the respective portion, the bird's-eye-view point cloud map generated using a bird's-eye-view projection of the point cloud data.

11 . The one or more non-transitory computer-readable media of claim 10 , wherein the output feature map is characterized by a data structure also associated with the bird's-eye-view point cloud map.

12 . The one or more non-transitory computer-readable media of claim 11 , wherein the operations comprise:

generating a dense bird's-eye-view feature map using the output feature map and the bird's-eye-view point cloud map.

13 . The one or more non-transitory computer-readable media of claim 9 , wherein the output feature map is generated using a learnable continuous convolution operator.

14 . The one or more non-transitory computer-readable media of claim 9 , wherein the operations comprise:

controlling the autonomous vehicle based on the output feature map.

15 . The one or more non-transitory computer-readable media of claim 14 , wherein the operations comprise:

performing, in a point-cloud metric space, and based at least in part on the output feature map, at least one of object detection or motion planning.

16 . The one or more non-transitory computer-readable media of claim 9 , wherein the output feature map is fused with point-cloud data convolution layers.

17 . An autonomous vehicle control system for controlling an autonomous vehicle, the autonomous vehicle control system comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions that are executable by the one or more processors to cause the autonomous vehicle control system to perform operations, the operations comprising:

obtaining image data and point cloud data, the image data and the point cloud data descriptive of an environment of an autonomous vehicle, the image data associated with a camera view;

for a respective portion of the environment:

determining one or more datapoints of the point cloud data associated with the respective portion;

projecting the one or more datapoints of the point cloud data into the camera view;

determining one or more portions of the image data respectively corresponding to the projected one or more datapoints of the point cloud; and

generating, based at least in part on the one or more portions of the image data, an output feature associated with the respective portion of the environment; and

generating an output feature map based at least in part on the output feature, the output feature map corresponding to a two-dimensional overhead view of the environment, the output feature map comprising a fusion of image data features and point cloud data features.

18 . The autonomous vehicle control system of claim 17 , wherein the operations comprise:

obtaining a bird's-eye-view point cloud map associated with the respective portion, the bird's-eye-view point cloud map generated using a bird's-eye-view projection of the point cloud data.

19 . The autonomous vehicle control system of claim 18 , wherein the operations comprise:

generating a dense bird's-eye-view feature map using the output feature map and the bird's-eye-view point cloud map.

20 . The autonomous vehicle control system of claim 17 , wherein the output feature map comprises a fusion of image data features and point cloud data features, and wherein the operations comprise:

performing, in a point-cloud metric space, and based at least in part on the output feature map, at least one of object detection or motion planning.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: UATC, LLC
To: AURORA OPERATIONS, INC.
Reel/Frame 067733/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2024
From: UBER TECHNOLOGIES, INC.
To: UATC, LLC
Reel/Frame 067073/0861 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2022
From: MA, WEI-CHIU; WANG, SHENLONG; YANG, BIN
To: UBER TECHNOLOGIES, INC.
Reel/Frame 061576/0809 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2022
From: URTASUN, RAQUEL
To: UATC, LLC
Reel/Frame 061576/0817 →
EMPLOYMENT AGREEMENT Recorded Oct 28, 2022
From: LIANG, MING
To: UBER TECHNOLOGIES, INC.
Reel/Frame 061795/0963 →