IP Library › Granted Patent US 12,646,195
Granted Patent B2
US 12,646,195 · App. 17/750,785 · Granted Jun 2, 2026

Object pose tracking from video images

Inventors: Yunzhi Lin (Marietta, GA); Jonathan Tremblay (Redmond, WA); Stephen Walter Tyree (University City, MO); Stanley Thomas Birchfield (Sammamish, WA)
Assignee: NVIDIA Corporation
G06T7/70G06T7/277G06T2207/10016G06T2207/10024G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,195
App. No.
17/750,785
Granted
Jun 2, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to determined a pose of an object from a plurality of images. In at least one embodiment, the pose of an object is determined from at least two images of a video sequence using one or more neural networks, in which the neural network produces a distribution of pose information that is filtered to determine the current pose.

Claims (45)

1 . A computer-implemented method comprising:

generating a first distribution of pose information for an object in an image using a trained network from the image;

generating a second distribution of prior pose information of the object in a prior image, wherein the second distribution comprises an uncertainty estimate corresponding to a prior cuboid generated in response to the prior image;

determining a cuboid having dimensions centered around the object;

applying a first filter to the first distribution to generate updated pose information corresponding to the object;

applying a second filter to the second distribution to generate relative cuboid dimensions; and

identifying a current pose of the object based, at least in part, on the updated pose information and the relative cuboid dimensions.

2 . The computer-implemented method of claim 1 , wherein the current pose or a next pose is a 6-degree of freedom pose.

3 . The computer-implemented method of claim 1 , wherein the image and the prior image are two-dimensional images.

4 . The computer-implemented method of claim 1 , wherein the first filter comprises a Bayesian filter, and the current pose is determined at least in part by applying the first filter to the first distribution of the pose information.

5 . The computer-implemented method of claim 1 , wherein the first distribution of the pose information and the second distribution of the prior pose information includes a center heatmap and a keypoint heatmap.

6 . The computer-implemented method of claim 1 , wherein the trained network is used to calculate a center track offset and a keypoint track offset.

7 . The computer-implemented method of claim 1 , wherein the cuboid comprises a bounding cuboid for the object.

8 . The computer-implemented method of claim 1 , wherein the trained network determines the first distribution of the pose information using a third distribution of additional pose information older than the prior pose information.

9 . The computer-implemented method of claim 1 , wherein the first filter comprises a Kalman filter, and wherein an uncertainty is determined by applying the first filter to the first distribution of the pose information.

10 . A system comprising one or more circuits to:

generate a first distribution of pose information for an object in an image using a trained network from the image;

generate a second distribution of prior pose information of the object in a prior image, wherein the second distribution comprises an uncertainty estimate corresponding to a prior cuboid generated in response to the prior image;

determine a cuboid having dimensions centered around the object;

apply a first filter to the first distribution to generate updated pose information corresponding to the object;

apply a second filter to the second distribution to generate relative cuboid dimensions; and

identify a current pose of the object based, at least in part, on the updated pose information and the relative cuboid dimensions.

11 . The system of claim 10 , wherein the current pose or a next pose is a 6-degree of freedom pose.

12 . The system of claim 10 , wherein the image and the prior image are two-dimensional images.

13 . The system of claim 10 , wherein the first filter comprises a Bayesian filter, and wherein the current pose is determined at least in part by applying the first filter to the first distribution of the pose information.

14 . The system of claim 10 , wherein the first distribution of the pose information and the second distribution of the prior pose information includes a center heatmap and a keypoint heatmap.

15 . The system of claim 10 , wherein the trained network is used to calculate a center track offset and a keypoint track offset.

16 . The system of claim 10 , wherein the cuboid comprises a bounding cuboid for the object.

17 . The system of claim 10 , wherein the trained network determines the first distribution of the pose information using a third distribution of additional pose information older than the prior pose information.

18 . The system of claim 10 , wherein the first filter comprises a Kalman filter, and wherein an uncertainty is determined by applying the first filter to the first distribution of the pose information.

19 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

generate a first distribution of pose information for an object in an image using a trained network from the image;

generate a second distribution of prior pose information of the object in a prior image, wherein the second distribution comprises an uncertainty estimate corresponding to a prior cuboid generated in response to the prior image;

determine a cuboid having dimensions centered around the object;

apply a first filter to the first distribution to generate updated pose information corresponding to the object;

apply a second filter to the second distribution to generate relative cuboid dimensions; and

identify a current pose of the object based, at least in part, on the updated pose information and the relative cuboid dimensions.

20 . The non-transitory machine-readable medium of claim 19 , wherein the current pose or a next pose is a 6-degree of freedom pose.

21 . The non-transitory machine-readable medium of claim 19 , wherein the image and the prior image are two-dimensional images.

22 . The non-transitory machine-readable medium of claim 19 , wherein the first filter comprises a Bayesian filter, and wherein the current pose is determined at least in part by applying the first filter to the first distribution of the pose information.

23 . The non-transitory machine-readable medium of claim 19 , wherein the first distribution of the pose information and the second distribution of the prior pose information includes a center heatmap and a keypoint heatmap.

24 . The non-transitory machine-readable medium of claim 19 , wherein the trained network is used to calculate a center track offset and a keypoint track offset.

25 . The non-transitory machine-readable medium of claim 19 , wherein the cuboid comprises a bounding cuboid for the object.

26 . The non-transitory machine-readable medium of claim 19 , wherein the trained network determines the first distribution of the pose information using a third distribution of additional pose information older than the prior pose information.

27 . The non-transitory machine-readable medium of claim 19 , wherein the first filter comprises a Kalman filter, and wherein an uncertainty is determined by applying the first filter to the first distribution of the pose information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2022
From: LIN, YUNZHI; TREMBLAY, JONATHAN; TYREE, STEPHEN WALTER; BIRCHFIELD, STANLEY THOMAS
To: NVIDIA CORPORATION
Reel/Frame 059997/0844 →
Continuity (1)
Related Publication 20240005547A1 · Jan 4, 2024
References Cited (41)
US 20220277472A1 · Birchfield · 2022 [cited by examiner]
US 20230281864A1 · Guo · 2023 [cited by examiner]
Majcher, Deep Quaternion Pose Proposals for 6D Object Pose Tracking, 2021 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW) (Year: 2021). [cited by examiner]
Braso, The Center of Attention: Center-Keypoint Grouping via Attention for Multi-Person Pose Estimation, 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (Year: 2021). [cited by examiner]
Tremblay, Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects, 2nd Conference on Robot Learning (CoRL 2018), Zurich, Switzerland (Year: 2018). [cited by examiner]
Sei, Accuracy Enhancement in Face-pose Estimation Network Using Incrementally Updated Face-shape Parameters, 2020 17th International Conference on Ubiquitous Robots (UR), Jun. 22-26, 2020. Kyoto, Japan (Year: 2020). [cited by examiner]
Janabi-Sharifi, A Kalman-Filter-Based Method for Pose Estimation in Visual Servoing, IEEE Transactions on Robotics, vol. 26, No. 5, Oct. 2010 (Year: 2010). [cited by examiner]
Li, Semantic Scene Models for Visual Localization Under Large Viewpoint Changes, 2018 15th Conference on Computer and Robot Vision (Year: 2018). [cited by examiner]
Ahmadyan et al., “Objectron: A Large Scale Dataset of Object-Centric Videos in the Wild with Pose Annotations,” Dec. 18, 2020, 11 pages. [cited by applicant]
Bewley et al., “Simple Online and Realtime Tracking,” IEEE, Jul. 7, 2017, 5 pages. [cited by applicant]
Chen et al., “Learning Canonical Shape Space for Category-level 6D Object Pose and Size Estimation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, 10 pages. [cited by applicant]
Chen et al., “Real-Time Multiple People Tracking With Deeply Learned Candidate Selection and Person Re-Identification,” IEEE, Sep. 12, 2018, 6 pages. [cited by applicant]
Choi et al., “RGB-D Object Tracking: A Particle Filter Approach on GPU,” International Conference on Intelligent Robots and Systems, 2013, 8 pages. [cited by applicant]
Deng et al., PoseRBPF: A Rao-Blackwellized Particle Filter for 6D Object Pose Tracking, May 22, 2019, 10 pages. [cited by applicant]
Dosovitskiy et al., “FlowNet: Learning Optical Flow with Convolutional Networks,” ICCV, 2015, 9 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Issac et al., “Depth-Based Object Tracking Using a Robust Gaussian Filter,” IEEE, Feb. 19, 2016, 8 pages. [cited by applicant]
Li et al., “DeepIM: Deep Iterative Matching for 6D Pose Estimation,” Sep. 8, 2014, 16 pages. [cited by applicant]
Pauwels et al., “SimTrack: A Simulation-based Framework for Scalable Real-time Object Pose Detection and Tracking,” IEEE, Dec. 17, 2015, 9 pages. [cited by applicant]
Peng et al., “PVNet: Pixel-Wise Voting Network for 6DoF Pose Estimation,” CVPR, 2019, 10 pages. [cited by applicant]
Rad et al., “BB8: A Scalable, Accurate, Robust to Partial Occlusion Method for Predicting the 3D Poses of Challenging Objects without Using Depth,” ICCV, 2017, 9 pages. [cited by applicant]
Sharma et al., “Beyond Pixels: Leveraging Geometry and Shape Cues for Online Multi-Object Tracking,” IEEE, Feb. 26, 2018, 9 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Standard No. J3016-201806, dated Jun. … [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Sundermeyer et al., “Implicit 3D Orientation Learning for 6D Object Detection from RGB Images,” ECCV, 2018, 17 pages. [cited by applicant]
Tan et al., “A Versatile Learning-based 3D Temporal Tracker: Scalable, Robust, Online,” IEEE, Dec. 1, 2015, 9 pages. [cited by applicant]
Tang et al., “Multiple People Tracking by Lifted Multicut and Person Re-identification,” IEEE, Jul. 1, 2017, 10 pages. [cited by applicant]
Tian et al., “Shape Prior Deformation for Categorical 6D Object Pose and Size Estimation,” Jul. 16, 2020, 21 pages. [cited by applicant]
Tjaden et al., “Real-Time Monocular Pose Estimation of 3D Objects using Temporally Consistent Local Color Histograms,” IEEE, Oct. 1, 2017, 9 pages. [cited by applicant]
Tremblay et al., “Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects,” Conference on Robot Learning, Sep. 27, 2018, 11 pages. [cited by applicant]
Wang et al., “Densefusion: 6D Object Pose Estimation by Iterative Dense Fusion,” Jan. 15, 2019, 11 pages. [cited by applicant]
Wang et al., “Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation,” Jun. 23, 2019, 12 pages. [cited by applicant]
Wen et al., “se(3)-TrackNet: Data-driven 6D Pose Tracking by Calibrating Image Residuals in Synthetic Domains,” IEEE, Jul. 27, 2020, 7 pages. [cited by applicant]
Wojke et al., “Simplw Online and Realtime Tracking With a Deep Association Metric,” IEEE, Mar. 21, 2017, 5 pages. [cited by applicant]
Wüthrich et al., “Probabilistic Object Tracking using a Range Camera,” IEEE, May 1, 2015, 8 pages. [cited by applicant]
Xiang et al., “PoseCNN: A Convolutional Neural Network for 6D Object Pose Estimation in Cluttered Scenes,” May 26, 2018, 10 pages. [cited by applicant]
Xu et al., “Consistent Online Multi-object Trackingwith Part-Based Deep Network,” Nov. 23, 2018, 13 pages. [cited by applicant]
Xu et al., “Spatial-Temporal Relation Networks for Multi-Object Tracking,” IEEE, Apr. 25, 2019, 11 pages. [cited by applicant]
Zeng et al., “Multi-view Self-Supervised Deep Learning for 6D Pose Estimation in the Amazon Picking Challenge,” IEEE International Conference on Robotics and Automation, May 7, 2017, 8 pages. [cited by applicant]
Zhou et al., “Tracking Objects as Points,” In “Proceedings of the European Conference on Computer Vision (ECCV),” Aug. 21, 2020, 22 pages. [cited by applicant]
Zhu et al., “Online Multi-Object Tracking with Dual Matching Attention Networks,” Feb. 2, 2019, 17 pages. [cited by applicant]