IP Library › Granted Patent US 12,736,626
Granted Patent B2
US 12,736,626 · App. 18/351,816 · Granted Sep 15, 2026

Systems and methods for camera-to-radar knowledge distillation

Inventors: Sirajum Munir (Pittsburgh, PA); Wenpeng Wang (Pittsburgh, PA)
Assignee: Robert Bosch GmbH
G01S7/417G01S13/9027G01S13/931G06V10/82G01S2013/9322G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,736,626
App. No.
18/351,816
Granted
Sep 15, 2026
Kind
B2
Abstract

A workflow is described herein for training a radar data processing model with a reduced requirement for annotated training data. The training workflow leverages an image sensor, such as a camera, that is synchronized with a radar sensor to capture synchronous image and radar point cloud data. The training workflow utilizes a self-supervised knowledge distillation process to pre-train the radar data processing model on the unannotated synchronous image and radar point cloud data, using a pre-trained image model. Subsequently, the radar data processing model is fine-tuned with a limited set of annotated radar point cloud data, thereby greatly reducing the human annotation burden.

Claims (57)

1 . A method for training a first neural network to perform a radar data processing task, the method comprising:

receiving, with a processor, a plurality of training data pairs, each respective training data pair in the plurality of training data pairs including a respective image and a respective radar point cloud, which were captured synchronously with one another of a same scene;

training, with the processor, in a first phase based on the plurality of training data pairs, the first neural network to extract features from radar point clouds using a second neural network that is pre-trained to extract features from images;

receiving, with the processor, a plurality of annotated radar point clouds, the annotated radar point clouds having labels corresponding to a radar data processing task; and

further training, with the processor, in a second phase based on the plurality of annotated radar point clouds, the first neural network to perform the radar data processing task.

2 . The method according to claim 1 , wherein:

the first neural network comprises an encoder followed by a decoder,

the training the first neural network in the first phase comprises training the encoder of the first neural network using the second neural network and based on the plurality of training data pairs; and

the training the first neural network in the second phase comprises training the decoder of the first neural network based on the plurality of annotated radar point clouds.

3 . The method according to claim 2 , wherein parameters of the encoder of the first neural network are frozen during training the decoder of the first neural network in the second phase.

4 . The method according to claim 1 , wherein parameters of the second neural network are frozen during the training the first neural network in the first phase.

5 . The method according to claim 1 , the training the first neural network in the first phase comprising:

determining, using the first neural network, a first feature output based on a respective radar point cloud in the plurality of training data pairs;

determining, using the second neural network, a second feature output based on a respective image in the plurality of training data pairs that corresponds to the respective radar point cloud; and

refining the first neural network based on the first feature output and the second feature output.

6 . The method according to claim 5 , the training the first neural network in the first phase comprising:

determining a mapping between (i) points of radar point clouds in the plurality of training data pairs and (ii) pixels of images in the plurality of training data pairs.

7 . The method according to claim 6 , the training the first neural network in the first phase comprising:

determining a contrastive loss based on the first feature output and the second feature output and based on the mapping; and

refining the first neural network based on the contrastive loss.

8 . The method according to claim 7 , the training the first neural network in the first phase comprising:

segmenting the respective radar point cloud into superpoints, each superpoint including a subset of points from the respective radar point cloud; and

segmenting the respective image into superpixels, each superpixel including a subset of pixels from the respective image.

9 . The method according to claim 8 , the training the first neural network in the first phase comprising:

determining, for each respective superpoint in the respective radar point cloud, a respective superpoint feature output;

determining, for each respective superpixel in the respective image, a respective superpixel feature output; and

determining the contrastive loss based on the respective superpoint feature output of each superpoint in the respective radar point cloud and the respective superpixel feature output of each superpixel in the respective image.

10 . The method according to claim 9 , the training the first neural network in the first phase comprising:

matching each superpoint in the respective radar point cloud to a respective superpixel in the respective image; and

determining the contrastive loss based on the respective superpoint feature output and the respective superpixel feature output of each matched superpoint and superpixel.

11 . The method according to claim 9 , the determining the respective superpoint feature output further comprising:

determining the respective superpoint feature output as a weighted average of features in the first feature output for points of the respective superpoint, the average being weighted depending on an average distance of each point with each other point in the respective superpoint.

12 . The method according to claim 9 , the determining the respective superpixel feature output further comprising:

determining the respective superpixel feature output as an average of features in the second feature output for pixels of the respective superpixel.

13 . The method according to claim 8 , the segmenting the respective image into superpixels further comprising:

defining each superpixel of the respective image by applying at least one of an image segmentation algorithm and a pixel clustering algorithm to the respective image.

14 . The method according to claim 13 , the segmenting the respective radar point cloud into superpoints further comprising:

matching each point in the respective radar point cloud to a respective superpixel in the respective image; and

defining each superpoint of the respective radar point cloud as a group of radar points that match to a same superpixel in the respective image.

15 . The method according to claim 8 , the segmenting the respective radar point cloud into superpoints further comprising:

defining each superpoint of the respective radar point cloud by applying a point clustering algorithm to the respective radar point cloud.

16 . The method according to claim 15 , the segmenting the respective image into superpixels further comprising:

matching each superpixel in the respective image to a respective superpoint in the respective radar point cloud; and

defining each superpixel of the respective image as a group of pixels that match to a same superpoint in the respective radar point cloud.

17 . The method according to claim 15 , the segmenting the respective image into superpixels further comprising:

initially defining superpixels of the respective image by applying at least one of an image segmentation algorithm and a pixel clustering algorithm to the respective image;

matching each initially defined superpixel in the respective image to a respective superpoint in the respective radar point cloud; and

defining each superpixel of the respective image as a group of initially defined superpixels that match to a same superpoint in the respective radar point cloud.

18 . The method according to claim 15 , the applying the point clustering algorithm to the respective radar point cloud further comprising:

generating a respective combined radar point cloud by combining the respective radar point cloud with at least one further radar point cloud that was captured at an immediately previous or subsequent time compared to a time at which the respective radar point cloud was captured.

19 . The method according to claim 18 , the applying the point clustering algorithm to the respective radar point cloud further comprising:

defining each superpoint of the respective radar point cloud by applying a point clustering algorithm to the respective combined radar point cloud.

20 . A non-transitory computer-readable medium for training a first neural network to perform a radar data processing task, the non-transitory computer-readable medium storing program instructions that, when executed by a processor, cause the processor to:

receive a plurality of training data pairs, each respective training data pair in the plurality of training data pairs including a respective image and a respective radar point cloud, which were captured synchronously with one another of a same scene;

train, in a first phase based on the plurality of training data pairs, the first neural network to extract features from radar point clouds using a second neural network that is pre-trained to extract features from images;

receive a plurality of annotated radar point clouds, the annotated radar point clouds having labels corresponding to a radar data processing task; and

further train, in a second phase based on the plurality of annotated radar point clouds, the first neural network to perform the radar data processing task.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2023
From: MUNIR, SIRAJUM; WANG, WENPENG
To: ROBERT BOSCH GMBH
Reel/Frame 064625/0134 →
Continuity (1)
Related Publication 20250020774A1 · Jan 16, 2025
References Cited (18)
US 11971955B1 · Chakraborty · 2024 [cited by examiner]
US 12189026B1 · Alferdaous Alazem · 2025 [cited by examiner]
US 20210027113A1 · Goldstein · 2021 [cited by examiner]
US 20220196798A1 · Chen · 2022 [cited by examiner]
US 20220261593A1 · Yu · 2022 [cited by examiner]
US 20220357441A1 · Ansari · 2022 [cited by examiner]
US 20240096105A1 · Zhao · 2024 [cited by examiner]
US 20240104913A1 · Redford · 2024 [cited by examiner]
US 20240185719A1 · Wei · 2024 [cited by examiner]
R. Achanta, A. Shaji, K. Smith, A. Lucchi, P. Fua, and S. Susstrunk. SLIC Superpixels Compared to State-of-the-Art Superpixel Methods. IEEE transactions on pattern analysis and machine intelligence,vol. 34(11), pp. 2274… [cited by applicant]
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, nuScenes: A Multimodal dataset for autonomous driving, In Proceedings of the IEEE/CVF conference on compute… [cited by applicant]
C. Choy, J. Gwak, and S. Savarese. 4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3075-3084, 2019. [cited by applicant]
P. F. Felzenszwalb and D. P. Huttenlocher. Efficient Graph-Based Image Segmentation. International Journal of Computer Vision, 59(2), pp. 167-181, 2004. [cited by applicant]
H. Huang, L. Lin, R. Tong, H. Hu, Q. Zhang, Y. Iwamoto, X. Han, Y.W. Chen, and J. Wu. UNET 3+: A Full-Scale Connected UNET for Medical Image Segmentation. In ICASSP 2020-2020 IEEE International Conference on Acoustics, … [cited by applicant]
A. Popov, P. Gebhardt, K. Chen, R. Oldja, H. Lee, S. Murray, R. Bhargava, and N. Smolyanskiy. NVRadarNet: Real-Time Radar obstacle and Free Space Detection for Autonomous Driving. arXiv preprint arXiv:2209.14499v2, 2023. [cited by applicant]
C. Sautier, G. Puy, S. Gidaris, A. Boulch, A. Bursuc, and R. Marlet. Image-to-Lidar Self-Supervised Distillation for Autonomous Driving Data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco… [cited by applicant]
S. Xie, J. Gu, D. Guo, C. R. Qi, L. Guibas, and O. Litany. PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding. In European conference on computer vision, pp. 574-591. Springer, 2020. [cited by applicant]
Z. Zhang, R. Girdhar, A. Joulin, and I. Misra. Self-Supervised Pretraining of 3D Features on any Point-Cloud. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10252-10263, 2021. [cited by applicant]