IP Library › Granted Patent US 12,546,882
Granted Patent B2
US 12,546,882 · App. 18/076,723 · Granted Feb 10, 2026

Camera-radar sensor fusion using local attention mechanism

Inventors: Jyh-Jing Hwang (Mountain View, CA); Henrik Kretzschmar (Mountain View, CA); Dragomir Anguelov (San Francisco, CA)
Assignee: Waymo LLC
G01S13/867G01S7/417G01S13/89G06T7/194G06T7/50G06V10/80G06V10/82G06T2207/10028G06T2207/20084G06T2207/20221G06T2207/30252G06V20/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,546,882
App. No.
18/076,723
Granted
Feb 10, 2026
Kind
B2
Abstract

Methods, computer systems, and apparatus, including computer programs encoded on computer storage media, for processing sensor data. In one aspect, a method includes obtaining image data representing a camera sensor measurement of a scene; obtaining radar data representing a radar sensor measurement of the scene; generating a feature representation of the image data; generating a respective initial depth estimate for each of a subset of the plurality of pixels; generating a feature representation of the radar data; for each of the subset of the plurality of pixels, generating a respective adjusted depth estimate for the pixel using the initial depth estimate for the pixel and the radar feature vectors for a corresponding subset of the plurality of radar reflection points; generating a fused point cloud that includes a plurality of three-dimensional data points; and processing the fused point cloud to generate an output that characterizes the scene.

Claims (70)

1 . A method comprising:

obtaining image data representing a camera sensor measurement of a scene captured by a camera sensor, the image data comprising a plurality of two-dimensional pixels;

obtaining radar data representing a radar sensor measurement of the scene captured by a radar sensor, the radar data comprising a plurality of two-dimensional radar reflection points;

processing the image data using a first neural network to generate as output a feature representation of the image data that comprises a respective image feature vector for each of the pixels;

generating a respective initial depth estimate for each of a subset of the plurality of pixels;

processing the radar data using a second neural network to generate as output a feature representation of the radar data that comprises a radar feature vector for each of the plurality of radar reflection points;

for each of the subset of the plurality of pixels, generating a respective adjusted depth estimate for the pixel using the initial depth estimate for the pixel and the radar feature vectors for one or more corresponding radar reflection points;

generating a respective elevation estimate for each of a subset of the plurality of radar reflection points;

generating a fused point cloud that includes (i) a first plurality of three-dimensional data points, wherein each three-dimensional data point in the first plurality of three-dimensional data points corresponds to a respective one of the subset of pixels in the image data and has a depth value that is equal to the respective adjusted depth estimate for the corresponding pixel, and (ii) a second plurality of three-dimensional data points, wherein each three-dimensional data point in the second plurality of three-dimensional data points corresponds to a respective one of the subset of the plurality of radar reflection points and has an elevation value that is equal to the respective elevation estimate for the corresponding radar reflection point; and

processing the fused point cloud using an output neural network to generate a network output that characterizes the scene.

2 . The method of claim 1 , wherein the output neural network comprises an object detection neural network and the network output is an object detection output that identifies objects that are located in the scene.

3 . The method of claim 1 , wherein the fused point cloud further comprises, for each of the first plurality of three-dimensional data points, corresponding feature information derived from the image data.

4 . The method of claim 1 , wherein generating the fused point cloud comprises transforming the subset of the plurality of pixels in the image data into three-dimensional coordinates in accordance with the respective adjusted depth estimates.

5 . The method of claim 1 , wherein generating the respective initial depth estimates comprises processing the image data using a depth prediction neural network that is configured to process the image data to generate as output the respective initial depth estimate of each of the plurality of pixels.

6 . The method of claim 1 , wherein the first and second neural networks each comprise a respective convolutional neural network.

7 . The method of claim 1 , further comprising generating a first mask that assigns each of the pixels to be either a foreground pixel or a background pixel, wherein the subset of the plurality of pixels includes only the foreground pixels.

8 . The method of claim 1 , wherein generating the respective adjusted depth estimate comprises, for each pixel in the subset of the plurality of pixels:

generating a plurality of candidate three-dimensional positions along a ray that is cast from the camera sensor to a three-dimensional position having spatial coordinates of the pixel and a depth equal to the initial depth estimate, wherein each candidate three-dimensional position specifies a respective candidate initial depth estimate for the pixel.

9 . The method of claim 8 , further comprising:

determining, as the one or more corresponding radar reflection points to be used to generate the respective adjusted depth estimate for the pixel, one or more radar reflection points that spatially map to each of the plurality of candidate three-dimensional positions by using the respective candidate initial depth estimate that is specified by the candidate three-dimensional position.

10 . The method of claim 8 , further comprising generating a second mask that assigns each of the plurality of radar reflection points to be either a foreground radar reflection point or a background radar reflection point, and wherein the one or more corresponding radar reflection points to be used to generate the respective adjusted depth estimate for the pixel include only the foreground radar reflection points.

11 . The method of claim 8 , wherein generating the adjusted depth estimate for the pixel further comprises, for each pixel in the subset of the plurality of pixels:

processing, using a fusion neural network, a fusion network input comprising the respective initial depth estimate for the pixel to generate a fusion network output that specifies the respective adjusted depth estimate for the pixel, wherein the fusion neural network is configured to generate the respective adjusted depth estimate for the pixel at least in part by applying an attention mechanism over the radar feature vectors for the one or more corresponding radar reflection points by using the image feature vector for the pixel to generate a query to be used in the attention mechanism.

12 . The method of claim 11 , wherein the fusion neural network is further configured to, for each pixel in the subset of the plurality of pixels:

use the image feature vector for the pixel to generate the query for the pixel to be used in the attention mechanism;

use the radar feature vector for each radar reflection point in the corresponding subset of the plurality of radar reflection points to generate a respective key for the radar reflection point to be used in the attention mechanism;

use the respective initial depth estimate and the plurality of candidate initial depth estimates for the pixel to generate respective values for the pixel to be used in the attention mechanism;

determine a corresponding attention weight for each respective value for the pixel by computing a product between the query for the pixel and each of the respective keys; and

generate the respective adjusted depth estimate for the pixel based on determining a weighted sum of the respective values for the pixel weighted by the corresponding attention weights for the respective values.

13 . The method of claim 1 , wherein generating the respective elevation estimate for each of the subset of the plurality of radar reflection points comprises generating the respective elevation estimate based on a height of the radar sensor, a terrain of the scene, or both.

14 . The method of claim 1 , wherein the fused point cloud further comprises, for each of the second plurality of three-dimensional data points, corresponding feature information derived from the radar data.

15 . The method of claim 1 , further comprising training the output neural network by:

obtaining a fused point cloud that includes a first plurality of three-dimensional data points that correspond to image pixels and a second plurality of three-dimensional data points that correspond to radar reflection points;

determining, in accordance with a predetermined dropout probability, whether to mask out feature information for either the first or the second plurality of three-dimensional data points included in the fused point cloud;

in response to a positive determination, generating a masked fused point cloud by masking out the feature information for either the first or the second plurality of three-dimensional data points; and

processing the masked fused point cloud using the output neural network in accordance with current values of output network parameters to generate a training network output for a given machine learning task.

16 . The method of claim 15 , further comprising:

determining an update to the current values of the output network parameters by determining a gradient with respect to the output network parameters of a loss function that includes a first term that depends on a difference between the training network output and a ground truth network output associated with the fused point cloud.

17 . The method of claim 15 , wherein masking out the feature information for the first plurality of three-dimensional data points included in the fused point cloud comprises, for each three-dimensional data point in the first plurality of three-dimensional data points:

replacing the feature information associated with data point with predetermined numeric values.

18 . The method of claim 17 , wherein the predetermined numeric value is zero.

19 . The method of claim 15 , wherein determining, in accordance with the predetermined dropout probability, whether to mask out feature information for either the first or the second plurality of three-dimensional data points included in the fused point cloud comprises:

sampling, with uniform randomness, a number between zero and one; and

determining whether the sampled number is greater than the predetermined dropout probability.

20 . The method of claim 16 , wherein obtaining the fused point cloud comprises generating the fused point cloud from the image and radar data using the first neural network, the second neural network, the depth prediction neural network, and the fusion neural network, and wherein the method further comprises:

determining a respective update to current values of the first, second, depth prediction, and fusion network parameters based on the determined gradient of the loss function.

21 . The method of claim 16 , wherein the loss function further includes a second term that depends on a difference between the initial depth estimates generated by using the depth prediction neural network and ground truth depth values.

22 . The method of claim 21 , further comprising determining the ground truth depth values using:

point cloud data representing a LiDAR sensor measurement of the scene captured by a LiDAR sensor, or

ground truth object detection labels associated with the image data.

23 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising:

obtaining image data representing a camera sensor measurement of a scene captured by a camera sensor, the image data comprising a plurality of two-dimensional pixels;

obtaining radar data representing a radar sensor measurement of the scene captured by a radar sensor, the radar data comprising a plurality of two-dimensional radar reflection points;

processing the image data using a first neural network to generate as output a feature representation of the image data that comprises a respective image feature vector for each of the pixels;

generating a respective initial depth estimate for each of a subset of the plurality of pixels;

processing the radar data using a second neural network to generate as output a feature representation of the radar data that comprises a radar feature vector for each of the plurality of radar reflection points;

for each of the subset of the plurality of pixels, generating a respective adjusted depth estimate for the pixel using the initial depth estimate for the pixel and the radar feature vectors for one or more corresponding radar reflection points;

generating a respective elevation estimate for each of a subset of the plurality of radar reflection points;

generating a fused point cloud that includes (i) a first plurality of three-dimensional data points, wherein each three-dimensional data point in the first plurality of three-dimensional data points corresponds to a respective one of the subset of pixels in the image data and has a depth value that is equal to the respective adjusted depth estimate for the corresponding pixel, and (ii) a second plurality of three-dimensional data points, wherein each three-dimensional data point in the second plurality of three-dimensional data points corresponds to a respective one of the subset of the plurality of radar reflection points and has an elevation value that is equal to the respective elevation estimate for the corresponding radar reflection point; and

processing the fused point cloud using an output neural network to generate a network output that characterizes the scene.

24 . One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

obtaining image data representing a camera sensor measurement of a scene captured by a camera sensor, the image data comprising a plurality of two-dimensional pixels;

obtaining radar data representing a radar sensor measurement of the scene captured by a radar sensor, the radar data comprising a plurality of two-dimensional radar reflection points;

processing the image data using a first neural network to generate as output a feature representation of the image data that comprises a respective image feature vector for each of the pixels;

generating a respective initial depth estimate for each of a subset of the plurality of pixels;

processing the radar data using a second neural network to generate as output a feature representation of the radar data that comprises a radar feature vector for each of the plurality of radar reflection points;

for each of the subset of the plurality of pixels, generating a respective adjusted depth estimate for the pixel using the initial depth estimate for the pixel and the radar feature vectors for one or more corresponding radar reflection points;

generating a respective elevation estimate for each of a subset of the plurality of radar reflection points;

generating a fused point cloud that includes (i) a first plurality of three-dimensional data points, wherein each three-dimensional data point in the first plurality of three-dimensional data points corresponds to a respective one of the subset of pixels in the image data and has a depth value that is equal to the respective adjusted depth estimate for the corresponding pixel, and (ii) a second plurality of three-dimensional data points, wherein each three-dimensional data point in the second plurality of three-dimensional data points corresponds to a respective one of the subset of the plurality of radar reflection points and has an elevation value that is equal to the respective elevation estimate for the corresponding radar reflection point; and

processing the fused point cloud using an output neural network to generate a network output that characterizes the scene.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2022
From: HWANG, JYH-JING; KRETZSCHMAR, HENRIK; ANGUELOV, DRAGOMIR
To: WAYMO LLC
Reel/Frame 062084/0923 →
Continuity (2)
Continuation In Part 17569385 · Jan 5, 2022
Related Publication 20230213643A1 · Jul 6, 2023
References Cited (62)
US 12007728B1 · Mohta · 2024 [cited by examiner]
US 20200218979A1 · Kwon · 2020 [cited by examiner]
US 20200286247A1 · Niesen · 2020 [cited by examiner]
US 20200301013A1 · Banerjee · 2020 [cited by examiner]
US 20220026557A1 · Arbabian · 2022 [cited by examiner]
US 20220357441A1 · Ansari · 2022 [cited by examiner]
Ouyang, Z., Feng, Y., He, Z., Hao, T., Dai, T., & Xia, S. T. (Jul. 2019). Attentiondrop for convolutional neural networks. In 2019 IEEE International Conference on Multimedia and Expo (ICME) (pp. 1342-1347). IEEE. (Year… [cited by examiner]
Ba et al., “Layer normalization,” CoRR, Jul. 21, 2016, arXiv:1607.06450, 14 pages. [cited by applicant]
Bijelic et al., “Seeing through fog without seeing fog: Deep multimodal sensor fusion in unseen adverse weather,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 11682… [cited by applicant]
Brazil et al., “M3D-RPN: Monocular 3D region proposal network for object. detection,” roceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 9287-9296. [cited by applicant]
Casser et al., “Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos,” Proceedings of the AAAI conference on artificial intelligence, Jul. 17, 2019, 33(01):8001-8008. [cited by applicant]
Chen et al., “Monocular 3D object detection for autonomous driving,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 2147-2156. [cited by applicant]
Chen et al., “Monopair: Monocular 3D object detection using pairwise spatial relationships,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 12093-12102. [cited by applicant]
Chen et al., “Multi-view 3D object detection network for autonomous driving,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 1907-1915. [cited by applicant]
Ding et al., “Learning depth-guided convolutions for monocular 3D object detection,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2020, pp. 1000-1001. [cited by applicant]
Graham et al., “Submanifold sparse convolutional networks,” CoRR, arXiv:1706.01307, Jun. 6, 2017, 10 pages. [cited by applicant]
Huang et al., “EPNet: Enhancing point features with image semantics for 3D object detection,” European Conference on Computer Vision, Nov. 16, 2020, pp. 35-52. [cited by applicant]
Ioffe et al., “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” Proceedings of the 32nd International Conference on Machine Learning, 2015, 37:448-456. [cited by applicant]
Kim et al., “GRIF Net: Gated region of interest fusion network for robust 3D object detection from radar point cloud and monocular image,” IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Oct.… [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization,” CoRR, Dec. 22, 2014, arXiv:1412.6980, 15 pages. [cited by applicant]
Lang et al., “Pointpillars: Fast encoders for object detection from point clouds,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 12697-12705. [cited by applicant]
Li et al., “RTM3D: Real-time monocular 3D detection from object keypoints for autonomous driving,” European Conference on Computer Vision, Dec. 3, 2020, 11 pages. [cited by applicant]
Liang et al., “Deep continuous fusion for multi-sensor 3D object detection,” Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 641-656. [cited by applicant]
Lim et al., “Radar and camera early fusion for vehicle detection in advanced driver assistance systems,” Machine Learning for Autonomous Driving Workshop at the 33rd Conference on Neural Information Processing Systems, … [cited by applicant]
Lin et al., “Feature pyramid networks for object detection,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2117-2125. [cited by applicant]
Lin et al., “Focal loss for dense object detection,” Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2980-2988. [cited by applicant]
Liu et al., “Reinforced axial refinement network for monocular 3D object detection,” European Conference on Computer Vision, Nov. 19, 2020, 17 pages. [cited by applicant]
Ma et al., “Rethinking pseudo-LiDAR representation,” European Conference on Computer Vision, Nov. 28, 2020, 21 pages. [cited by applicant]
Major et al., “Vehicle detection with automotive radar using deep learning on range-azimuth-doppler tensors,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 0-0. [cited by applicant]
Manhardt et al., “ROI-10D: Monocular lifting of 2D detection to 6D pose and metric shape,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 2069-2078. [cited by applicant]
Mousavian et al., “3D bounding box estimation using deep learning and geometry,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 7074-7082. [cited by applicant]
Nabati et al., “Centerfusion: Center-based radar and camera fusion for 3D object detection,” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2021, pp. 1527-1536. [cited by applicant]
Nobis et al., “Radar voxel fusion for 3D object detection,” Applied Sciences, Jun. 17, 2021, 11(12):5598. [cited by applicant]
Park et al., “Is pseudo-lidar needed for monocular 3D object detection?,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 3142-3152. [cited by applicant]
Piergiovanni et al., “4D-Net for learned multimodal alignment,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 15435-15445. [cited by applicant]
Qi et al., “Frustum pointnets for 3D object detection from RGB-D data,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 918-927. [cited by applicant]
Qi et al., “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” Advances in Neural Information Processing Systems 30, 2017, 10 pages. [cited by applicant]
Reading et al., “Categorical depth distribution network for monocular 3D object detection,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 8555-8564. [cited by applicant]
Roddick et al., “Orthographic feature transform for monocular 3D object detection,” CoRR, Nov. 20, 2018, arXiv:1811.08188, 10 pages. [cited by applicant]
Ronneberger et al., “U-net: Convolutional networks for biomedical image segmentation,” International Conference on Medical image computing and computer-assisted intervention, Nov. 18, 2015, pp. 234-241. [cited by applicant]
Schumann et al., “Semantic segmentation on radar point clouds,” 2018 21st International Conference on Information Fusion (FUSION), Jul. 10-13, 2018, pp. 2179-2186. [cited by applicant]
Sheeny et al., “Radiate: A radar dataset for automotive perception in bad weather,” 2021 IEEE International Conference on Robotics and Automation (ICRA), May 30, 2021, 7 pages. [cited by applicant]
Shi et al., “Distance-normalized united representation for monocular 3D object detection,” European Conference on Computer Vision, 2020, pp. 91-107. [cited by applicant]
Shi et al., “Pointrenn: 3D object proposal generation and detection from point cloud,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 770-779. [cited by applicant]
Simonelli et al., “Disentangling monocular 3D object detection,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 1991-1999. [cited by applicant]
Simonelli et al., “Towards generalization across depth for monocular 3D object detection,” European Conference on Computer Vision, Nov. 17, 2020, pp. 767-782. [cited by applicant]
Srivastava et al., “Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research, 2014, 15(1): 1929-1958. [cited by applicant]
Srivastava et al., “Learning 2D to 3D lifting for object detection in 3D for autonomous vehicles,” 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Nov. 3-8, 2019, pp. 4504-4511. [cited by applicant]
Sun et al., “RSN: Range Sparse Net for Efficient, Accurate LiDAR 3D Object Detection,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 5725-5734. [cited by applicant]
Sun et al., “Scalability in perception for autonomous driving: Waymo Open Dataset,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 2446-2454. [cited by applicant]
Vaswani et al., “Attention is all you need,” Advances in Neural Information Processing Systems 30, 2017, 11 pages. [cited by applicant]
Vladimir, et al., “Hitnet: Hierarchical iterative tile refinement network for real-time stereo matching,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 14362-14372. [cited by applicant]
Vora et al., “Pointpainting: Sequential fusion for 3D object detection,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 4604-4612. [cited by applicant]
Wang et al., “Frustum convnet: Sliding frustums to aggregate local pointwise features for amodal 3D object detection,” 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Nov. 3-8, 2019, 8 p… [cited by applicant]
Wang et al., “Pointaugmenting: Cross-modal augmentation for 3D object detection,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 11794-11803. [cited by applicant]
Wang et al., “Pseudo-lidar from visual depth estimation: Bridging the gap in 3D object detection for autonomous driving,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, p… [cited by applicant]
Weng et al., “Monocular 3D object detection with pseudo-lidar point cloud,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 0-0. [cited by applicant]
Yan et al., “Second: Sparsely embedded convolutional detection,” Sensors, Oct. 6, 2018, 18(10):3337. [cited by applicant]
You et al., “Pseudo-liDAR++: Accurate depth for 3D object detection in autonomous driving,” CoRR, Jun. 14, 2019, arXiv:1906.06310, 22 pages. [cited by applicant]
Zhou et al., “End-to-end multi-view fusion for 3D object detection in LiDAR point clouds,” Proceedings of the Conference on Robot Learning, 2020, 100:923-932. [cited by applicant]
Zhou et al., “IoU loss for 2D/3D object detection,” 2019 International Conference on 3D Vision (3DV), Sep. 16-19, 2019, 10 pages. [cited by applicant]
Zhou et al., “Objects as points,” CoRR, Apr. 16, 2019, arXiv:1904.07850, 12 pages. [cited by applicant]