IP Library Granted Patent US 12,373,984
Granted Patent B2
US 12,373,984 · App. 18/614,254 · Granted Jul 29, 2025

Multi-modal 3-D pose estimation

Inventors: Jingxiao Zheng (San Jose, CA); Xinwei Shi (Cupertino, CA); Alexander Gorban (Scotts Valley, CA); Junhua Mao (Palo Alto, CA); Andre Liang Cornman (San Francisco, CA); Yang Song (San Jose, CA); Ting Liu (Los Angeles, CA); Ruizhongtai Qi (Mountain View, CA); Yin Zhou (San Jose, CA); Congcong Li (Cupertino, CA); Dragomir Anguelov (San Francisco, CA)
Assignee: Waymo LLC
G06T7/73G06F18/214G06F18/251G06V20/58G06T2207/10028G06T2207/20081G06T2207/20084G06T2207/30196G06T2207/30261
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,984
App. No.
18/614,254
Granted
Jul 29, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for estimating a 3-D pose of an object of interest from image and point cloud data. In one aspect, a method includes obtaining an image of an environment; obtaining a point cloud of a three-dimensional region of the environment; generating a fused representation of the image and the point cloud; and processing the fused representation using a pose estimation neural network and in accordance with current values of a plurality of pose estimation network parameters to generate a pose estimation network output that specifies, for each of multiple keypoints, a respective estimated position in the three-dimensional region of the environment.

Claims (43)

1. A method performed by one or more computers, wherein the method comprises:

generating a fused representation of an image of an environment and a point cloud of a three-dimensional region of the environment, wherein the point cloud comprises a plurality of data points, and wherein the generating comprises:

generating, based on the image and for each of a plurality of keypoints, a score for each of a plurality of locations in the image that represents a likelihood that the keypoint is located at the location in the image; and

generating, for each of the plurality of data points in the point cloud, a respective feature vector that includes at least some of the scores generated based on the image;

processing the fused representation to generate, for each of the plurality of keypoints, an estimated position in the three-dimensional region of the environment; and

controlling an agent based at least on the estimated position for each of the plurality of keypoints.

2. The method of claim 1 , wherein controlling the agent based at least on the estimated position for each of the plurality of keypoints comprises:

generating, based at least on the estimated position for each of the plurality of keypoints, a planning decision for the agent; and

controlling the agent to implement the planning decision by transmitting one or more electronic signals that have been generated in accordance with the planning decision to one or more control units of the agent.

3. The method of claim 1 , wherein the agent comprises a vehicle, and wherein the environment is an environment in a vicinity of the vehicle.

4. The method of claim 1 , wherein the plurality of keypoints collectively define an estimated pose of each of one or more pedestrians in the environment.

5. The method of claim 3 , wherein the vehicle comprises an autonomous vehicle, or a semi-autonomous vehicle.

6. The method of claim 2 , wherein the planning decision plans a future trajectory of the vehicle.

7. The method of claim 1 , wherein the agent comprises a robot.

8. The method of claim 1 , wherein processing the fused representation to generate the estimated position in the three-dimensional region of the environment comprises:

processing the fused representation using a pose estimation neural network that has been trained on training data comprising both (i) labeled point cloud data that associates each of multiple point clouds with corresponding human assigned keypoints and (ii) unlabeled point cloud data for which human assigned keypoints are unavailable.

9. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

generating a fused representation of an image of an environment and a point cloud of a three-dimensional region of the environment, wherein the point cloud comprises a plurality of data points, and wherein the generating comprises:

generating, based on the image and for each of a plurality of keypoints, a score for each of a plurality of locations in the image that represents a likelihood that the keypoint is located at the location in the image; and

generating, for each of the plurality of data points in the point cloud, a respective feature vector that includes at least some of the scores generated based on the image;

processing the fused representation to generate, for each of the plurality of keypoints, an estimated position in the three-dimensional region of the environment; and

controlling an agent based at least on the estimated position for each of the plurality of keypoints.

10. The system of claim 9 , wherein controlling the agent based at least on the estimated position for each of the plurality of keypoints comprises:

generating, based at least on the estimated position for each of the plurality of keypoints, a planning decision for the agent; and

controlling the agent to implement the planning decision by transmitting one or more electronic signals that have been generated in accordance with the planning decision to one or more control units of the agent.

11. The system of claim 9 , wherein the agent comprises a vehicle, and wherein the environment is an environment in a vicinity of the vehicle.

12. The system of claim 9 , wherein the plurality of keypoints collectively define an estimated pose of each of one or more pedestrians in the environment.

13. The system of claim 11 , wherein the vehicle comprises an autonomous vehicle, or a semi-autonomous vehicle.

14. The system of claim 10 , wherein the planning decision plans a future trajectory of the vehicle.

15. The system of claim 9 , wherein the agent comprises a robot.

16. The system of claim 9 , wherein processing the fused representation to generate the estimated position in the three-dimensional region of the environment comprises:

processing the fused representation using a pose estimation neural network that has been trained on training data comprising both (i) labeled point cloud data that associates each of multiple point clouds with corresponding human assigned keypoints and (ii) unlabeled point cloud data for which human assigned keypoints are unavailable.

17. One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

generating a fused representation of an image of an environment and a point cloud of a three-dimensional region of the environment, wherein the point cloud comprises a plurality of data points, and wherein the generating comprises:

generating, based on the image and for each of a plurality of keypoints, a score for each of a plurality of locations in the image that represents a likelihood that the keypoint is located at the location in the image; and

generating, for each of the plurality of data points in the point cloud, a respective feature vector that includes at least some of the scores generated based on the image;

processing the fused representation to generate, for each of the plurality of keypoints, an estimated position in the three-dimensional region of the environment; and

controlling an agent based at least on the estimated position for each of the plurality of keypoints.

18. The non-transitory computer-readable storage media of claim 17 , wherein controlling the agent based at least on the estimated position for each of the plurality of keypoints comprises:

generating, based at least on the estimated position for each of the plurality of keypoints, a planning decision for the agent; and

controlling the agent to implement the planning decision by transmitting one or more electronic signals that have been generated in accordance with the planning decision to one or more control units of the agent.

19. The non-transitory computer-readable storage media of claim 17 , wherein the agent comprises a vehicle, and wherein the environment is an environment in a vicinity of the vehicle.

20. The non-transitory computer-readable storage media of claim 17 , wherein the plurality of keypoints collectively define an estimated pose of each of one or more pedestrians in the environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2024
From: ZHENG, JINGXIAO; SHI, XINWEI; GORBAN, ALEXANDER; MAO, JUNHUA; CORNMAN, ANDRE LIANG; SONG, YANG; LIU, TING; QI, RUIZHONGTAI; ZHOU, YIN; LI, CONGCONG; ANGUELOV, DRAGOMIR
To: WAYMO LLC
Reel/Frame 067306/0297 →
Continuity (3)
Continuation 17505900 · Oct 20, 2021
Provisional Application 63114448 · Nov 16, 2020
Related Publication 20250037303A1 · Jan 30, 2025
References Cited (32)
US 10366502B1 · Li · 2019 [cited by applicant]
US 10408939B1 · Kim · 2019 [cited by examiner]
US 20180348374A1 · Laddha · 2018 [cited by examiner]
US 20190108651A1 · Gu · 2019 [cited by examiner]
US 20200005485A1 · Xu et al. · 2020 [cited by applicant]
US 20200057442A1 · Deiters et al. · 2020 [cited by applicant]
US 20200184721A1 · Ge et al. · 2020 [cited by applicant]
US 20200302160A1 · Hashimoto · 2020 [cited by examiner]
US 20210166150A1 · Wang et al. · 2021 [cited by applicant]
US 20220043441A1 · Huang · 2022 [cited by examiner]
US 20220343639A1 · Li · 2022 [cited by examiner]
Akhter et al., “Pose-conditioned joint angle limits for 3d human pose reconstruction,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 1446-1455. [cited by applicant]
Bin et al., “Simple baselines for human pose estimation and tracking,” Proceedings of the European conference on computer vision, 2018, pp. 466-481. [cited by applicant]
Chen et al., “3d human pose estimation=2d pose estimation+matching,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 5759-5767. [cited by applicant]
Chen et al., “Unsupervised 3d pose estimation with geometric self-supervision,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 5707-5717. [cited by applicant]
Extended Search Report in European Appln. No. 21892516.2, dated Sep. 9, 2024, 14 pages. [cited by applicant]
Furst et al., “HPERL: 3d human pose estimation from RGB and LiDAR,” 2020 25th International Conference on Pattern Recognition, May 2021, 7 pages. [cited by applicant]
Ge et al., “Hand pointnet: 3d hand pose estimation using point sets,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 8417-8426. [cited by applicant]
Ge et al., “Point-to-point regression pointnet for 3D hand pose estimation,” Proceedings of the European Conference on Computer Vision, 2018, pp. 475-491. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2021/049838, dated Dec. 24, 2021, 10 pages. [cited by applicant]
Kim et al., “PedX: Benchmark Dataset for Metric 3D Pose Estimation of Pedestrians in Complex Urban Intersections,” CoRR, Sep. 10, 2018, arXiv:1809.03605v1, 8 pages. [cited by applicant]
Kocabas et al., “Self-supervised learning of 3d human pose using multi-view geometry,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 1077-1086. [cited by applicant]
Qi et al., “Frustum pointnets for 3d object detection from RGB-D data,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 918-927. [cited by applicant]
Sun et al., “Deep high-resolution representation learning for human pose estimation,” Proceedings of the IEEE conference on computer vision and pattern recognition, 2019, pp. 5693-5703. [cited by applicant]
Sun et al., “Integral Human Pose Regression,” CoRR, Sep. 18, 2018, arXiv:1711.08229v4, 17 pages. [cited by applicant]
Tripathi et al., “Posenet3d: Unsupervised 3D human shape and pose estimation,” eprint arXiv:2003.03473, Mar. 2020, 17 pages. [cited by applicant]
Vora et al., “PointPainting: Sequential Fusion for 3D Object Detection,” CoRR, Nov. 22, 2019, arXiv:1911.10150v2, 11 pages. [cited by applicant]
Zhang et al., “Weakly supervised adversarial learning for 3d human pose estimation from point clouds,” IEEE Transactions on Visualization and Computer Graphics, May 2020, 26(5):1851-1859. [cited by applicant]
Zheng et al., “Deep learning-based human pose estimation: A survey,” CoRR, Dec. 2020, arxiv.org/abs/2012.13392, 25 pages. [cited by applicant]
Zhou et al., “Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised Approach,” 2017 IEEE International Conference on Computer Vision (ICCV), Oct. 22, 2017, pp. 398-407. [cited by applicant]
Zhou et al., “Towards 3d human pose estimation in the wild: A weakly-supervised approach,” Proceedings of the IEEE International Conference on Computer Vision, Oct. 2017, 10 pages. [cited by applicant]
Zimmermann et al., “3D Human Pose Estimation in RGBD Images for Robotic Task Learning,” CoRR, Mar. 13, 2018, arXiv:1803.02622v2, 7 pages. [cited by applicant]