IP Library Granted Patent US 12,196,859
Granted Patent B2
US 12,196,859 · App. 17/185,414 · Granted Jan 14, 2025

Label transfer between data from multiple sensors

Inventors: Sarah Najmark (Palo Alto, CA); Sean Kirmani (Mountain View, CA)
Assignee: Google LLC
G01S17/89G01S17/93G06F18/2155H04L67/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,196,859
App. No.
17/185,414
Granted
Jan 14, 2025
Kind
B2
Abstract

A method includes receiving first sensor data captured by a first sensor. The method further includes receiving a plurality of labels or predictions corresponding to the first sensor data. The method also includes receiving second sensor data captured by a second sensor. The method further includes determining time-synchronized sensor data comprising a subset of the first sensor data and a subset of the second sensor data. The method additionally includes determining, based on the plurality of labels or predictions and the time-synchronized sensor data, a plurality of pseudo-labels corresponding to the second sensor data. The method also includes generating a training data set comprising at least the subset of the second sensor data and one or more pseudo-labels from the plurality of pseudo-labels.

Claims (48)

1. A method comprising:

receiving first sensor data captured by a first sensor, wherein the first sensor data includes two-dimensional image data and depth data;

receiving a plurality of labels or predictions defining coordinates of one or more two-dimensional bounding boxes of the first sensor data;

receiving second sensor data captured by a second sensor, wherein the second sensor data includes point cloud data;

determining time-synchronized sensor data comprising a subset of the first sensor data and a subset of the second sensor data;

determining, based on the plurality of labels or predictions, the time-synchronized sensor data, and the depth data, a plurality of pseudo-labels corresponding to the second sensor data, wherein the plurality of pseudo-labels define coordinates of one or more three-dimensional bounding boxes based on the one or more two-dimensional bounding boxes and the depth data; and

generating a training data set comprising at least the subset of the second sensor data and one or more pseudo-labels from the plurality of pseudo-labels.

2. The method of claim 1 , wherein the first sensor is a red green blue depth (RGB-D) camera, wherein the second sensor is a LIDAR sensor.

3. The method of claim 1 , wherein the second sensor has a wider field of view than the first sensor.

4. The method of claim 1 , wherein the first sensor data represents an area smaller than the second sensor data, and wherein determining the plurality of pseudo-labels corresponding to the second sensor data comprises:

determining, based on the first sensor and the second sensor, cropped second sensor data corresponding to the area represented by the first sensor data.

5. The method of claim 1 , wherein the plurality of labels or predictions comprise labels determined by using a human-assisted labeling process.

6. The method of claim 1 , wherein the plurality of labels or predictions comprise predictions determined by using an automated prediction process run on a computing device.

7. The method of claim 1 , wherein the plurality of predictions or labels indicate one or more object locations in the first sensor data, and wherein the plurality of pseudo-labels indicate one or more object locations in the second sensor data.

8. The method of claim 1 , wherein the time-synchronized sensor data comprising a subset of the first sensor data and a subset of the second sensor data is based on a threshold timestamp difference between the first sensor data and the second sensor data.

9. The method of claim 1 , wherein the subset the first sensor data is an improper subset of the first sensor data.

10. The method of claim 1 , wherein the subset of the second sensor data is an improper subset of the second sensor data.

11. The method of claim 1 , wherein the method further comprises before generating the training data set, validating, based at least on the subset of the second sensor data and a quality measure associated with the plurality of pseudo-labels corresponding to the second sensor data, one or more pseudo-labels from the plurality of pseudo-labels and wherein the training data set is generated based on the one or more validated pseudo-labels.

12. The method of claim 11 , wherein the second sensor is a LIDAR sensor, wherein the second sensor data includes point cloud data, and wherein the quality measure is a threshold number of points within the point cloud data.

13. The method of claim 1 , wherein the method further comprises determining, based at least on the plurality of pseudo-labels, confidence scores, wherein generating the training data set is based on the confidence scores.

14. The method of claim 13 , wherein the method further comprises:

determining, based on the confidence scores, a high confidence subset of the second sensor data and a high confidence subset of the plurality of pseudo-labels; and

generating a high confidence training data set comprising at least the high confidence subset of the second sensor data and the high confidence subset of the plurality of pseudo-labels, wherein the high confidence training data set is to be used in supervised machine learning.

15. A system comprising:

a first sensor;

a second sensor; and

a computing device configured to:

receive first sensor data captured by a first sensor, wherein the first sensor data includes two-dimensional image data and depth data;

receive a plurality of labels or predictions defining coordinates of one or more two-dimensional bounding boxes of the first sensor data;

receive second sensor data captured by a second sensor, wherein the second sensor data includes point cloud data;

determine time-synchronized sensor data comprising a subset of the first sensor data and a subset of the second sensor data;

determine, based on the plurality of labels or predictions, the time-synchronized sensor data, and the depth data, a plurality of pseudo-labels corresponding to the second sensor data, wherein the plurality of pseudo-labels define coordinates of one or more three-dimensional bounding boxes based on the one or more two-dimensional bounding boxes and the depth data; and

generate a training data set comprising at least the subset of the second sensor data and one or more pseudo-labels from the plurality of pseudo-labels.

16. The system of claim 15 , wherein the computing device is further configured to:

receive additional second sensor data captured by the second sensor;

receive a second plurality of predictions corresponding to the additional second sensor data, wherein the second plurality of predictions was generated by a machine learning model trained on the plurality of pseudo-labels, wherein the second plurality of predictions were generated by applying the machine learning model to the additional second sensor data;

receive additional first sensor data captured by the first sensor;

determine additional time-synchronized sensor data comprising a subset of the additional second sensor data and a subset of the additional first sensor data;

determine, based on the second plurality of predictions and the additional time-synchronized sensor data, an additional plurality of pseudo-labels corresponding to the additional first sensor data; and

generate an additional training data set comprising at least the subset of the additional first sensor data and an additional one or more pseudo-labels from the additional plurality of pseudo-labels.

17. A non-transitory computer readable medium comprising program instructions executable by at least one processor to cause the at least one processor to perform functions comprising:

receiving first sensor data captured by a first sensor, wherein the first sensor data includes two-dimensional image data and depth data;

receiving a plurality of labels or predictions defining coordinate of one or more two-dimensional bounding boxes of the first sensor data;

receiving second sensor data captured by a second sensor, wherein the second sensor data includes point cloud data;

determining time-synchronized sensor data comprising a subset of the first sensor data and a subset of the second sensor data;

determining, based on the plurality of labels or predictions, the time-synchronized sensor data and the depth data, a plurality of pseudo-labels corresponding to the second sensor data, wherein the plurality of pseudo-labels define coordinates of one or more three-dimensional bounding boxes based on the one or more two-dimensional bounding boxes and the depth data; and

generating a training data set comprising at least the subset of the second sensor data and one or more pseudo-labels from the plurality of pseudo-labels.

18. The method of claim 1 , wherein the one or more two-dimensional bounding boxes are two-dimensional rectangles, and the one or more three-dimensional bounding boxes are three-dimensional rectangles.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2025
From: GOOGLE LLC
To: GDM HOLDING LLC
Reel/Frame 071109/0342 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2023
From: X DEVELOPMENT LLC
To: GOOGLE LLC
Reel/Frame 064658/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2021
From: NAJMARK, SARAH; KIRMANI, SEAN
To: X DEVELOPMENT LLC
Reel/Frame 055433/0414 →
Continuity (1)
Related Publication 20220268939A1 · Aug 25, 2022
References Cited (29)
US 8290305B2 · Minear et al. · 2012 [cited by applicant]
US 8605998B2 · Samples et al. · 2013 [cited by applicant]
US 9098773B2 · Huang et al. · 2015 [cited by applicant]
US 9282321B2 · Sandrew et al. · 2016 [cited by applicant]
US 9904867B2 · Fathi et al. · 2018 [cited by applicant]
US 10452949B2 · Jia et al. · 2019 [cited by applicant]
US 10663594B2 · Tsishkou et al. · 2020 [cited by applicant]
US 20090232388A1 · Minear et al. · 2009 [cited by applicant]
US 20130286017A1 · Sanjuan et al. · 2013 [cited by applicant]
US 20170103510A1 · Wang et al. · 2017 [cited by applicant]
US 20190026957A1 · Gausebeck · 2019 [cited by applicant]
US 20200210887A1 · Jain · 2020 [cited by examiner]
Yalniz “Billion-scale semi-supervised learning for image classification”, arXiv:1905.00546v1 [cs.CV] May 2, 2019 (Year: 2019). [cited by examiner]
Li et al., “Real-time 3D object proposal generation and classification using limited processing resources,” Robotics and Autonomous Systems, 2020, 12 pages, vol. 130. [cited by applicant]
Çiçek et al., “3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation,” arXiv:1606.06650v1, Jun. 21, 2016, 8 pages. [cited by applicant]
Huang et al., “PerspectiveNet: 3D Object Detection from a Single RGB Image via Perspective Points,” arXiv:1912.07744v1, Dec. 16, 2019, 13 pages. [cited by applicant]
Lin et al., “A novel point cloud registration using 2D image features,” EURASIP Journal on Advances in Signal Processing, 2017, 11 pages, vol. 5. [cited by applicant]
Mousavian et al., “3D Bounding Box Estimation Using Deep Learning and Geometry,” arXiv:1612.0049v2, Apr. 10, 2017, 10 pages. [cited by applicant]
Qi et al., “Deep Hough Voting for 3D Object Detection in Point Clouds,” arXiv:1904.09664v2, Aug. 22, 2019, 14 pages. [cited by applicant]
Shi et al., “PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud,” arXiv:1812.04244v2, May 16, 2019, 10 pages. [cited by applicant]
Tang et al., “Transferable Semi-supervised 3D Object Detection from RGB-D Data,” arXiv:1904.10300v1, Apr. 23, 2019, 16 pages. [cited by applicant]
Wang et al., “Fusing Bird's Eye View LIDAR Point Cloud and Front View Camera Image for Deep Object Detection,” arXiv:1711.06703v3, Feb. 14, 2018, 12 pages. [cited by applicant]
Wang et al., “LDLS: 3D Object Segmentation through Label Diffusion from 2D Images,” IEEE Robotics and Automation Letters, Preprint Version, Accepted May 2019, arXiv:1910.13955.v1, Oct. 30, 2019, 8 pages. [cited by applicant]
Wang et al., “SGPN: Similarity Group Proposal Network for 3D Point Cloud Instance Segmentation,” arXiv.1711.08588v2, May 30, 2019, 13 pages. [cited by applicant]
Yi et al., “GSPN: Generative Shape Proposal Network for 3D Instance Segmentation in Point Cloud,” arXiv:1812.03320v1, Dec. 8, 2018, 13 pages. [cited by applicant]
Zhou et al., “VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection,” arXiv1711.06396v1, Nov. 17, 2017, 10 pages. [cited by applicant]
Ding et al., “Semantic Segmentation of Indoor 3D Point Cloud with SLENet,” The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, ISPRS Geospatial Week 2019, The Netherlands, … [cited by applicant]
Piewak et al., “Boosting LiDAR-Based Semantic Labeling by Cross-modal Training Data Generation,” ECCV 2018 Workshops, LNCS 11134, 2019, pp. 497-513. [cited by applicant]
Xie et al., “Semantic Instance Annotation of Street Scenes by 3D and 2D Label Transfer,” 2016 IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 3688-3697. [cited by applicant]