IP Library › Granted Patent US 12,380,594
Granted Patent B2
US 12,380,594 · App. 17/921,784 · Granted Aug 5, 2025

Pose estimation method and apparatus

Inventors: Slobodan Ilic (Munich, DE); Roman Kaskman (Munich, DE); Ivan Shugurov (Munich, DE); Sergey Zakharov (San Francisco, CA)
Assignee: SIEMENS AKTIENGESELLSCHAFT
G06T7/73G06T7/13G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,594
App. No.
17/921,784
Granted
Aug 5, 2025
Kind
B2
Abstract

Various embodiments of the teachings herein include a computer implemented pose estimation method for providing poses of objects of interest in a scene. The scene comprises a visual representation of the objects of interest in an environment. The method comprising: conducting for each one of the objects of interest a pose estimation; determining edge data of the object of interest from the visual representation representing the edges of the respective object of interest; determining keypoints of the respective object of interest by a previously trained artificial neural keypoint detection network, wherein the artificial neural keypoint detection network utilizes the determined edge data of the respective object of interest Oi as input and provides the keypoints of the respective object of interest as output; and estimating the pose of the respective object of interest based on the respective object's keypoints provided by the artificial neural keypoint detection network.

Claims (41)

1. A computer implemented pose estimation method for providing poses of objects of interest in a scene, wherein the scene comprises a visual representation of the objects of interest in an environment, the method comprising:

conducting for each one of the objects of interest a pose estimation;

determining edge data of the object of interest from the visual representation representing the edges of the respective object of interest;

determining keypoints of the respective object of interest by a previously trained artificial neural keypoint detection network, wherein the artificial neural keypoint detection network utilizes the determined edge data of the respective object of interest Oi as input and provides the keypoints of the respective object of interest as output;

estimating the pose of the respective object of interest based on the respective object's keypoints provided by the artificial neural keypoint detection network.

2. A method according to claim 1 , wherein training the previously trained artificial neural keypoint detection network is based on synthetic models which correspond to potential objects of interest.

3. A method according to claim 1 , wherein, in case poses are required for more than one object of interest, the pose estimation is conducted simultaneously for all objects of interest.

4. A method according to claim 1 , wherein generating the visual representation of the scene includes creating a sequence of images of the scene, representing the one or more objects of interest from different perspectives.

5. A method according to claim 4 , wherein a structure-from-motion approach processes the visual representation to provide a sparse reconstruction at least of the one or more objects of interest for the determination of the edge data.

6. A method according to claim 5 , wherein, for the determination of the edge data, the provided sparse reconstruction is further approximated with lines wherein the edge data are obtained by sampling points along the lines.

7. A method according to claim 1 , wherein:

training of the artificial neural keypoint detection network, during which KDN-parameters which define the artificial neural keypoint detection network are adjusted, is based on the creation of one or more synthetic scenes;

for each synthetic scene

in a preparation step

one or more known synthetic models of respective one or more potential objects of interest are provided;

the synthetic scene is generated by arranging the provided synthetic models in a virtual environment, such that poses of the provided synthetic models are known;

each synthetic model is represented in the synthetic scene by an edge-like representation of the respective synthetic model; and

known keypoints for each one of the provided synthetic models are provided;

in an adjustment step one or more training loops to adjust the KDN-parameters are executed;

wherein in each training loop

an edge-like representation of the generated synthetic scene comprising the known edge-like representations of the provided synthetic models is provided as input training data to the artificial neural keypoint detection network; and

the artificial neural keypoint detection network is trained utilizing the provided input training data and the provided known keypoints, representing an aspired output of the artificial neural keypoint detection network when processing the input training data.

8. A Method according to claim 7 , wherein for the generation of the known synthetic scene, a virtual surface is generated and the provided one or more known synthetic models are randomly arranged on the virtual surface.

9. A method according to claim 7 , wherein the edge-like representations for the corresponding one or more provided synthetic models of the generated synthetic scene are generated by:

rendering the generated synthetic scene to generate rendered data, preferably rendering each one of the provided synthetic models of the synthetic scene;

obtaining a dense point cloud model of the generated synthetic scene and by applying a perspective-based sampling approach on the rendered data; and

reconstructing prominent edges of the dense point cloud model to the respective edge-like generate representation of the generated synthetic scene.

10. A method according to claim 9 , wherein for each known perspective, features of a respective synthetic model which are not visible in the particular perspective are removed from the respective edge-like representation.

11. A method according to claim 7 , wherein:

the known object keypoints of a respective synthetic model comprise at least a first keypoint and one or more further keypoints with k=2, 3, . . . for the respective synthetic model;

the keypoints are calculated such that the first keypoint corresponds to the centroid of the respective synthetic model and the one or more further keypoints are obtained via a farthest-point-sampling approach.

12. A method according to claim 7 , wherein for each synthetic scene-Byas, the network is trained simultaneously for all the provided synthetic models arranged in the respective synthetic scene.

13. An apparatus for providing pose information of one or more objects of interest in a scene, an apparatus comprising:

a computer with a memory and a processor;

wherein the memory stores a set of instructions, the set of instructions causing the processor to implement a pose estimation method for providing poses of objects of interest in a scene, wherein the scene comprises a visual representation of the objects of interest oi in environment, the method comprising:

conducting for each one of the objects of interest pose estimation;

determining edge data of the object of interest from the visual representation representing the edges of the respective object of interest;

determining keypoints of the respective object of interest by previously trained artificial neural keypoint detection network, wherein the artificial neural keypoint detection network utilizes the determined edge data of the respective object of interest oi as input and provides the keypoints of the respective object of interest as output;

estimating the pose of the respective object of interest based on the respective object's keypoints provided by the artificial neural keypoint detection network.

14. An apparatus according to claim 13 , comprising a camera system for imaging the scene comprising the one or more objects of interest to generate the visual representation.

15. An apparatus according to claim 14 , wherein the relative positioning of the camera system and the scene is adjustable.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2023
From: ILIC, SLOBODAN; KASKMAN, ROMAN; SHUGUROV, IVAN; ZAKHAROV, SERGEY
To: SIEMENS AKTIENGESELLSCHAFT
Reel/Frame 063579/0052 →
Priority Claims (1)
EP 20172252 · Apr 30, 2020 · regional
Continuity (1)
Related Publication 20230169677A1 · Jun 1, 2023
References Cited (19)
US 20190156086A1 · Plummer · 2019 [cited by examiner]
US 20190355150A1 · Tremblay · 2019 [cited by examiner]
US 20200306980A1 · Choi · 2020 [cited by examiner]
CN 110853036A · 2020 [cited by examiner]
WO 2020086217 · 2020 [cited by applicant]
Search Report for International Application No. PCT/EP2021/061363, 13 pages, Aug. 12, 2021. [cited by applicant]
Search Report for EP Application No. 20172252.7, 10 pages, Oct. 9, 2020. [cited by applicant]
Pereira, Nuno et al:“Masked Fusion: Mask-based 6D Object Pose Detection”; Arxiv.Org; Cornell University Library; 201, Olin Library Cornell University Ithaca; NY 14853; XP081534671, Nov. 18, 2019. [cited by applicant]
Kingma, Diederik P. et al. “Adam: A Method for Stochastic Optimization” 3rd International Conference for Learning Representations (ICLR), San Diego, 2015, arXiv:1412.6980. [cited by applicant]
Birdal 2017: Birdal T, Ilic S; “A point sampling algorithm for 3d matching of irregular geometries”; IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); 6871-6878, 2017. [cited by applicant]
S. Umeyama, “Least-Squares Estimation of Transformation Parameters Between Two Point Patterns,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 13, No. 4, Apr. 1, 1991. [cited by applicant]
Dai 2017: Dai A, Chang AX, Savva M, Halber M, Funkhouser T, Nießner M; “Scannet: Richly-annotated 3d reconstructions of indoor scenes”; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 5828… [cited by applicant]
Drummond 2002: Drummond T, Cipolla R; “Real-Time Visual Tracking of Complex Structures”; IEEE Trans. Pattern Anal. Mach. Intell.; 24, 932-946; doi: 10.1109/TPAMI.2002.1017620. [cited by applicant]
COLMAP: Schonberger JL, Frahm JM; “Structure-from-motion re-visited”; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; 4104-4113, 2016. [cited by applicant]
Peng, Sida et al:“PVNet: Pixel-Wise Voting Network for 6DoF Pose Estimation”; 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE, Jun. 15, 2019. [cited by applicant]
Hofer 2017: Hofer M, Maurer M, Bischof H; “Efficient 3D scene abstraction using line segments”; Computer Vision and Image Understanding; 157, 167-178. [cited by applicant]
Qi et al., “PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space,” Advances in Neural Information Processing Systems 30 (NIPS 2017), 10 pgs. [cited by applicant]
Schwarz, Max et al:“RGB-D object recognition and pose estimation based on pre-trained convolutional neural network features”; 2015 IEEE International Conference on Robotics And Automation (ICRA); pp. 1329-1335; ISBN: 97… [cited by applicant]
“SSD: Single Shot MultiBox Detector” Wei Liu; Dragomir Anguelov; et al UNC Chapel Hill; Zoox Inc. ; Google Inc.; University of Michigan; Ann-Arbor. [cited by applicant]