IP Library › Granted Patent US 12,299,916
Granted Patent B2
US 12,299,916 · App. 17/545,987 · Granted May 13, 2025

Three-dimensional location prediction from images

Inventors: Longlong Jing (Mountain View, CA); Ruichi Yu (Mountain View, CA); Jiyang Gao (Foster City, CA); Henrik Kretzschmar (Mountain View, CA); Kang Li (Sammamish, WA); Ruizhongtai Qi (Mountain View, CA); Hang Zhao (Sunnyvale, CA); Alper Ayvaci (Santa Clara, CA); Xu Chen (Livermore, CA); Dillon Cower (Woodinville, WA); Congcong Li (Cupertino, CA)
Assignee: Waymo LLC
G06T7/70G06T7/50G06V10/40G06V10/806G06N20/00G06T2207/10016G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,916
App. No.
17/545,987
Granted
May 13, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for predicting three-dimensional object locations from images. One of the methods includes obtaining a sequence of images that comprises, at each of a plurality of time steps, a respective image that was captured by a camera at the time step; generating, for each image in the sequence, respective pseudo-lidar features of a respective pseudo-lidar representation of a region in the image that has been determined to depict a first object; generating, for a particular image at a particular time step in the sequence, image patch features of the region in the particular image that has been determined to depict the first object; and generating, from the respective pseudo-lidar features and the image patch features, a prediction that characterizes a location of the first object in a three-dimensional coordinate system at the particular time step in the sequence.

Claims (53)

1. A method performed by one or more computers, the method comprising:

obtaining a temporal sequence of images that comprises, at each of a plurality of time steps, a respective image that was captured by a camera at the time step;

generating, for each image in the temporal sequence, respective pseudo-lidar features of a respective pseudo-lidar representation of a region in the image that has been determined to depict a first object by processing the region in the image using a first neural network, wherein the pseudo-lidar features represent one or more pixels within the region in the image as a point in a three-dimensional coordinate system based on an initial depth estimate for the image;

generating, for a particular image at a particular time step in the temporal sequence, image patch features of the region in the particular image that has been determined to depict the first object by processing the region in the particular image using a second neural network, wherein the image patch features are generated from intensity values of pixels in the image; and

generating, from the respective pseudo-lidar features and the image patch features, a prediction that characterizes a location of the first object in the three-dimensional coordinate system at the particular time step in the temporal sequence by processing the respective pseudo-lidar features and the image patch features using a third neural network, wherein generating, from the respective pseudo-lidar features and the image patch features, a prediction that characterizes the first object at the particular time step in the temporal sequence comprises:

combining the respective pseudo-lidar features that represent one or more pixels within the region in the image as a point in the three-dimensional coordinate system based on the initial depth estimate for the image and the image patch features to generate combined features; and

processing the combined features using the third neural network to generate the prediction.

2. The method of claim 1 , wherein the prediction includes an updated depth estimate that estimates a depth of a specified point on the first object at the particular time step in the temporal sequence, wherein the updated depth estimate is a predicted distance from the specified point on the first object to the camera at the particular time step.

3. The method of claim 1 , wherein the prediction specifies a three-dimensional region that corresponds to a predicted location of the first object at the particular time step relative to the camera.

4. The method of claim 1 , wherein the third neural network is a decoder neural network.

5. The method of claim 4 , wherein combining the respective pseudo-lidar features and the image patch features comprises concatenating the respective pseudo-lidar features and the image patch features.

6. The method of claim 1 , wherein generating image patch features of the region in the image at the particular time step in the temporal sequence comprises:

processing the image using an image feature extraction neural network to generate image features for the image; and

selecting, as the image patch features, a subset of the image features that correspond to the region in the image.

7. The method of claim 1 , further comprising:

generating, for each image in the temporal sequence, an initial depth estimate that assigns a respective estimated depth value to each pixel in the image; and

generating, for each image in the temporal sequence, the respective pseudo-lidar representation using the initial depth estimate for the image.

8. The method of claim 7 , wherein generating, for each image in the temporal sequence, an initial depth estimate that assigns a respective estimated depth value to each pixel in the image comprises:

processing the image using a depth estimation neural network to generate the initial depth estimate for the image.

9. The method of claim 8 , wherein generating the pseudo-lidar representation comprises:

mapping each pixel that is within the region in the image that has been determined to depict the first object to the three-dimensional coordinate system based on the estimated depth value for the pixel in the initial depth estimate for the image and properties of the camera.

10. The method of claim 9 , wherein the properties of the camera include the horizontal and vertical focal lengths of the camera.

11. The method of claim 1 , wherein generating respective pseudo-lidar features of each of the pseudo-lidar representations comprises:

processing the pseudo-lidar representation using a pseudo-lidar feature extraction neural network to generate the pseudo-lidar features for the pseudo-lidar representation.

12. A method performed by one or more computers, the method comprising:

obtaining a temporal sequence of images that comprises, at each of a plurality of time steps, a respective image that was captured by a camera at the time step;

generating, for each image in the temporal sequence, an initial depth estimate that assigns a respective estimated depth value to each pixel in the image;

obtaining object tracklet data for a first object that identifies, for each of the images in the temporal sequence, a respective two-dimensional bounding box in the image that has been determined to depict the first object;

generating, for each image in the temporal sequence, a respective pseudo-lidar representation of the two-dimensional bounding box in the image from the initial depth estimate for the image;

generating respective pseudo-lidar features of each of the pseudo-lidar representations by processing the pseudo-lidar representation using a first neural network, wherein the pseudo-lidar features represent one or more pixels within the region in the image as a point in the three-dimensional coordinate system based on the initial depth estimate for the image;

generating image patch features of the two-dimensional bounding box in the last image in the temporal sequence by processing the two-dimensional bounding box in the last image using a second neural network, wherein the image patch features are generated from intensity values of pixels in the image; and

generating, from the respective pseudo-lidar features and the image patch features, a prediction that characterizes a location of the first object in the three-dimensional coordinate system at the last time step in the temporal sequence by processing the respective pseudo-lidar features and the image patch features using a third neural network, wherein generating, from the respective pseudo-lidar features and the image patch features, a prediction that characterizes the first object at the particular time step in the temporal sequence comprises:

combining the respective pseudo-lidar features that represent one or more pixels within the region in the image as a point in a three-dimensional coordinate system based on an initial depth estimate for the image and the image patch features to generate combined features; and

processing the combined features using the third neural network to generate the prediction.

13. A system comprising one or more computers and one or more storage devices storing instructions then when executed by the one or more computers cause the one or more computers to perform operations comprising:

obtaining a temporal sequence of images that comprises, at each of a plurality of time steps, a respective image that was captured by a camera at the time step;

generating, for each image in the temporal sequence, respective pseudo-lidar features of a respective pseudo-lidar representation of a region in the image that has been determined to depict a first object by processing the region in the image using a first neural network, wherein the pseudo-lidar features represent one or more pixels within the region in the image as a point in a three-dimensional coordinate system based on an initial depth estimate for the image;

generating, for a particular image at a particular time step in the temporal sequence, image patch features of the region in the particular image that has been determined to depict the first object by processing the region in the particular image using a second neural network; and

generating, from the respective pseudo-lidar features and the image patch features, a prediction that characterizes a location of the first object in the three-dimensional coordinate system at the particular time step in the temporal sequence by processing the respective pseudo-lidar features and the image patch features using a third neural network, wherein generating, from the respective pseudo-lidar features and the image patch features, a prediction that characterizes the first object at the particular time step in the temporal sequence comprises:

combining the respective pseudo-lidar features that represent one or more pixels within the region in the image as a point in the three-dimensional coordinate system based on the initial depth estimate for the image and the image patch features to generate combined features; and

processing the combined features using the third neural network to generate the prediction.

14. The system of claim 13 , wherein the prediction includes an updated depth estimate that estimates a depth of a specified point on the first object at the particular time step in the temporal sequence, wherein the updated depth estimate is a predicted distance from the specified point on the first object to the camera at the particular time step.

15. The system of claim 13 , wherein the prediction specifies a three-dimensional region that corresponds to a predicted location of the first object at the particular time step relative to the camera.

16. The system of claim 13 , wherein the third neural network is a decoder neural network.

17. The system of claim 16 , wherein combining the respective pseudo-lidar features and the image patch features comprises concatenating the respective pseudo-lidar features and the image patch features.

18. The system of claim 13 , wherein generating image patch features of the region in the image at the particular time step in the temporal sequence comprises:

processing the image using an image feature extraction neural network to generate image features for the image; and

selecting, as the image patch features, a subset of the image features that correspond to the region in the image.

19. The system of claim 13 , the operations further comprising:

generating, for each image in the temporal sequence, an initial depth estimate that assigns a respective estimated depth value to each pixel in the image; and

generating, for each image in the temporal sequence, the respective pseudo-lidar representation using the initial depth estimate for the image.

20. The system of claim 19 , wherein generating the pseudo-lidar representation of the region in the image from the initial depth estimate for the image comprises:

mapping each pixel that is within the region in the image that has been determined to depict the first object to the three-dimensional coordinate system based on the estimated depth value for the pixel in the initial depth estimate for the image and properties of the camera.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2022
From: JING, LONGLONG; YU, RUICHI; GAO, JIYANG; KRETZSCHMAR, HENRIK; LI, KANG; QI, RUIZHONGTAI; ZHAO, HANG; AYVACI, ALPER; CHEN, XU; COWER, DILLON; LI, CONGCONG
To: WAYMO LLC
Reel/Frame 058675/0001 →
Continuity (2)
Provisional Application 63122899 · Dec 8, 2020
Related Publication 20220180549A1 · Jun 9, 2022
References Cited (75)
US 11409304B1 · Cai · 2022 [cited by examiner]
US 11436743B2 · Guizilini · 2022 [cited by examiner]
US 11927668B2 · Fontijne · 2024 [cited by examiner]
US 20150254834A1 · Chandraker · 2015 [cited by examiner]
US 20210133461A1 · Ren · 2021 [cited by examiner]
US 20210255304A1 · Fontijne · 2021 [cited by examiner]
US 20220357441A1 · Ansari · 2022 [cited by examiner]
CN 111386550 · 2020 [cited by applicant]
CN 111602141 · 2020 [cited by applicant]
CN 111915663 · 2020 [cited by applicant]
Ma, X., Liu, S., Xia, Z., Zhang, H., Zeng, X., & Ouyang, W. (2020). Rethinking Pseudo-LiDAR Representation. arXiv.Org. https://doi.org/10.48550/arxiv.2008.04582 (Year: 2020). [cited by examiner]
Wang, Y., Chao, W.-L., Garg, D., Hariharan, B., Campbell, M., & Weinberger, K. Q. (2019). Pseudo-LiDAR From Visual Depth Estimation. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 8437-8445.… [cited by examiner]
Alhashim et al., “High quality monocular depth estimation via transfer learning,” CoRR, Dec. 2018, arXiv:1812.11941, 12 pages. [cited by applicant]
Beery et al., “Context r-cnn: Long term temporal context for per-camera object detection,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 13075-13085. [cited by applicant]
Behle et al., “Semantickitti: A dataset for semantic scene understanding of lidar sequences,” Proceedings of the IEEE International Conference on Computer Vision, Apr. 2, 2019, pp. 9297-9307. [cited by applicant]
Bewley et al., “Simple online and realtime tracking,” IEEE International Conference on Image Processing (ICIP), Sep. 25-28, 2016, pp. 3464-3468. [cited by applicant]
Brazil et al., “M3d-rpn: Monocular 3d region proposal network for object detection,” Proceedings of the IEEE International Conference on Computer Vision, Jul. 13, 2019, pp. 9287-9296. [cited by applicant]
Caesar et al., “nuscenes: A multimodal dataset for autonomous driving,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 1, 2020, pp. 11621-11631. [cited by applicant]
Cao et al., “Circle marker based distance measurement using a single camera,” Lecture Notes on Software Engineering, Nov. 1, 2013, 1(4):376-380. [cited by applicant]
Chen et al., “Monopair: Monocular 3d object detection using pairwise spatial relationships,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 1, 2020, pp. 12093-12102. [cited by applicant]
Chiu et al., “Probabilistic 3D multi-object tracking for autonomous driving,” CoRR, Jan. 16, 2020, arXiv:2001.05673, 8 pages. [cited by applicant]
Ding et al., “Learning depth-guided convolutions for monocular 3D object detection,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Jun. 14, 2020, pp. 1000-1001. [cited by applicant]
Dosovitskiy et al., “Flownet: Learning optical flow with convolutional networks,” Proceedings of the IEEE international conference on computer vision, 2015, pp. 2758-2766. [cited by applicant]
Eigen et al., “Depth map prediction from a single image using a multi-scale deep network,” Advances in neural information processing systems, 2014, pp. 2366-2374. [cited by applicant]
Fadadu et al., “SMulti-view fusion of sensor data for improved perception and prediction in autonomous driving,” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2022, pp. 2349-23… [cited by applicant]
Feichtenhofer et al., “Convolutional two-stream network fusion for video action recognition,” Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1933-1941. [cited by applicant]
Feng et al., “A new distance detection algorithm for images in deflecting angle,” 2016 2nd IEEE International Conference on Computer and Communications (ICCC), Oct. 2016, pp. 746-750. [cited by applicant]
Fu et al., “Deep ordinal regression network for monocular depth estimation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 2002-2011. [cited by applicant]
Geiger et al., “Vision meets robotics: The KITTI dataset,” The International Journal of Robotics Research, 2013, 32(11):1231-1237. [cited by applicant]
Godard et al., “Digging into self-supervised monocular depth estimation,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 3828-3838. [cited by applicant]
Godard et al., “Unsupervised monocular depth estimation with left-right consistency,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 270-279. [cited by applicant]
Gokce et al., “Vision-based detection and distance estimation of micro unmanned aerial vehicles,” Sensors, 2015,15(9):23805-23846. [cited by applicant]
Guizilini et al., “Semantically-guided representation learning for self-supervised monocular depth,” CoRR, Feb. 27, 2020, arXiv:2002.12319, 14 pages. [cited by applicant]
Hazirbas et al., “Fusenet: Incorporating depth into semantic segmentation via fusion-based cnn architecture,” Asian conference on computer vision, Mar. 10, 2017, pp. 213-228. [cited by applicant]
He et al., “Deep residual learning for image recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770-778. [cited by applicant]
Hu et al., “Joint monocular 3d vehicle detection and tracking,” Proceedings of the IEEE international conference on computer vision, 2019, pp. 5390-5399. [cited by applicant]
Hung et al., “SoDA: Multi-object tracking with soft data association,” CoRR, Aug. 18, 2020, arXiv:2008.07725, 13 pages. [cited by applicant]
Huo et al., “Learning depth-guided convolutions for monocular 3D object detection,” 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2020, pp. 4306-4315. [cited by applicant]
Jorgensen et al., “Monocular 3D object detection and box fitting trained end-toend using intersection-over-union loss,” CoRR, Jun. 19, 2019, arXiv:1906.08070, 10 pages. [cited by applicant]
Lang et al., “Pointpillars: Fast encoders for object detection from point clouds,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 12697-12705. [cited by applicant]
Li et al., “Rtm3d: Real-time monocular 3d detection from object keypoints for autonomous driving,” CoRR, Jan. 10, 2020, arXiv:2001.03343, 11 pages. [cited by applicant]
Liang et al., “Deep continuous fusion for multi-sensor 3D object detection,” Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 641-656. [cited by applicant]
Liang et al., “Multi-task multi-sensor fusion for 3d object detection,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 7345-7353. [cited by applicant]
Liu et al., “Flownet3D: Learning scene flow in 3D point clouds,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 529-537. [cited by applicant]
Liu et al., “Looking fast and slow: Memoryguided mobile video object detection,” CoRR, Mar. 25, 2019, arXiv:1903.10172, 10 pages. [cited by applicant]
Liu et al., “Smoke: Singlestage monocular 3D object detection via keypoint estimation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 996-997. [cited by applicant]
Luo et al., “Every pixel counts++: Joint learning of geometry and motion with 3D holistic understanding,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Oct. 1, 2020, 42(10):2624-2641. [cited by applicant]
Ma et al., “Accurate monocular 3D object detection via color-embedded 3D reconstruction for autonomous driving,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 6851-6860. [cited by applicant]
Ma et al., “Rethinking pseudo-lidar representation,” CoRR, Aug. 11, 2020, arXiv:2008.04582, 21 pages. [cited by applicant]
Poggi et al., “On the uncertainty of self-supervised monocular depth estimation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 3227-3237. [cited by applicant]
Qi et al., “Frustum pointnets for 3D object detection from RGB-D Data,” Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 918-927. [cited by applicant]
Rezaei et al., “Robust vehicle detection and distance estimation under challenging lighting conditions,” IEEE Transactions on Intelligent Transportation Systems, Oct. 5, 2015, 16(5):2723-2743. [cited by applicant]
Roberts et al., “Accurate marker based distance measurement with single camera,” 2015 International Conference on Image and Vision Computing New Zealand (IVCNZ), Nov. 23-24, 2015, 6 pages. [cited by applicant]
Simonelli et al., “Disentangling monocular 3D object detection,” Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 1991-1999. [cited by applicant]
Simonyan et al., “Two-stream convolutional networks for action recognition in videos,” Advances in Neural Information Processing Systems 27, 2014, 9 pages. [cited by applicant]
Sun et al., “Scalability in perception for autonomous driving: Waymo open dataset,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 2446-2454. [cited by applicant]
Tuohy et al., “Distance determination for an automobile environment using inverse perspective mapping in openCV,” IET Irish Signals and Systems Conference (ISSC 2010), Jun. 23-24, 2010, 6 pages. [cited by applicant]
Vianney et al., “Refinedmpl: Refined monocular pseudolidar for 3D object detection in autonomous driving,” CoRR, Nov. 21, 2019, arXiv:1911.09712, 10 pages. [cited by applicant]
Vora et al., “Pointpainting: Sequential fusion for 3D object detection,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 4604-4612. [cited by applicant]
Wang et al., “Centernet3d: An anchor free object detector for autonomous driving,” CoRR, Jul. 13, 2020, arXiv:2007.07214, 13 pages. [cited by applicant]
Wang et al., “Pseudo-LiDAR from visual depth estimation: Bridging the gap in 3D object detection for autonomous driving,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern. Recognition (CVPR), 2019, … [cited by applicant]
Wang et al., “Task-aware monocular depth estimation for 3D object detection,” AAAI-20 Technical Tracks 7, Apr. 3, 2020, 34(7):12257-12264. [cited by applicant]
Wang et al., “Towards real-time multi-object tracking,” CoRR, Sep. 27, 2019, arXiv:1909.12605, 17 pages. [cited by applicant]
Weng et al., “3D multi-object tracking: A baseline and new evaluation metrics,” CoRR, Jul. 9, 2019, arXiv:1907.03961, 9 pages. [cited by applicant]
Weng et al., “GNN3DMOT: Graph neural network for 3D multi-object tracking with 2D-3D multi-feature learning,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 6499-6508. [cited by applicant]
Weng et al., “Monocular 3D object detection with pseudo-LiDAR point cloud,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 0-0. [cited by applicant]
Xiao et al., “Audiovisual slowfast networks for video recognition,” CoRR, Jan. 23, 2020, arXiv:2001.08740, 14 pages. [cited by applicant]
Yang et al., “RadarNet: Exploiting radar for robust perception of dynamic objects,” CoRR, Jul. 28, 2020, arXiv:2007.14366, 16 pages. [cited by applicant]
Yin et al., “Centerbased 3D object detection and tracking,” CoRR, Jun. 19, 2020, arXiv:2006.11275, 12 pages. [cited by applicant]
Zhang et al., “FairMot: On the fairness of detection and re-identification in multiple object tracking,” CoRR, Apr. 4, 2020, arXiv:2004.01888, 19 pages. [cited by applicant]
Zhou et al., “Objects as Points,” CoRR, Apr. 16, 2019, arXiv:1904.07850, 12 pages. [cited by applicant]
Zhou et al., “Tracking objects as points,” Computer Vision—ECCV, Aug. 21, 2020, pp. 474-490. [cited by applicant]
Zhu et al., “Flow-guided feature aggregation for video object detection,” Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 408-417. [cited by applicant]
Zhu et al., “Learning object-specific distance from a monocular image,” Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 3839-3848. [cited by applicant]
Office Action in Chinese Appln. No. 202111493822.4, dated May 31, 2024, 16 pages (with English translation). [cited by applicant]