IP Library Granted Patent US 12,620,121
Granted Patent B2
US 12,620,121 · App. 18/456,761 · Granted May 5, 2026

Method of recognizing position and attitude of object, and non-transitory computer-readable storage medium

Inventors: Masaki Hayashi (Matsumoto, JP); Hirokazu Kasahara (Okaya, JP); Guoyi Fu (Richmond Hill, CA); Zhongzhen Luo (Vaughan, CA)
Assignee: SEIKO EPSON CORPORATION
G06T7/70G06V10/771G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,121
App. No.
18/456,761
Granted
May 5, 2026
Kind
B2
Abstract

A method of the present disclosure includes (a) generating an input image by imaging a scene containing the M objects by a camera, (b) obtaining a feature map showing feature amounts relating to the N keypoints from the input image using a learned machine learning model with the input image as input and the feature map as output, (c) obtaining three-dimensional coordinates of the N keypoints belonging to each of the M objects using the feature map, and (d) determining positions and attitudes of one or more objects of the M objects using the three-dimensional coordinates of the N keypoints belonging to each of the M objects, wherein (c) includes (c1) obtaining M×N keypoints having undetermined correspondence relationships with the M objects and determining the three-dimensional coordinates of the M×N keypoints, and (c2) grouping the M×N keypoints to the N keypoints belonging to each of the M objects.

Claims (51)

1 . A method of recognizing a position and an attitude of an object using first to Nth N keypoints set for the object, M being an integer of 1 or more and N being an integer of 2 or more, comprising:

(a) generating an input image by imaging a scene containing the M objects by a camera;

(b) obtaining a feature map showing feature amounts relating to the N keypoints from the input image using a learned machine learning model with the input image as input and the feature map as output;

(c) obtaining three-dimensional coordinates of the N keypoints belonging to each of the M objects using the feature map; and

(d) determining positions and attitudes of one or more objects of the M objects using the three-dimensional coordinates of the N keypoints belonging to each of the M objects, wherein

(c) includes:

(c1) obtaining M×N keypoints having undetermined correspondence relationships with the M objects and determining the three-dimensional coordinates of the M×N keypoints; and

(c2) grouping the M×N keypoints to the N keypoints belonging to each of the M objects,

the feature map used at (c2) contains N directional vector maps as maps in which vectors indicating directions from a plurality of pixels belonging to a same object to an object keypoint are assigned to the plurality of pixels with each of the N keypoints as the object keypoint, and

(c2) includes:

(c2-1) selecting one ith keypoint from M ith keypoints and selecting one jth keypoint from M jth keypoints;

(c2-2) calculating a first degree of conformance indicating a degree of coincidence of directions of a first vector obtained from a jth directional vector map and indicating a direction from a pixel position of the ith keypoint toward the jth keypoint and a second vector indicating a direction from a pixel position expressed by the three-dimensional coordinates of the ith keypoint to a pixel position expressed by the three-dimensional coordinates of the jth keypoint, i and j being integers from 1 to N different from each other; and

(c2-3) repeating (c2-1) and (c2-2) and performing the grouping of the M×N keypoints according to the first degree of conformance,

wherein (c2-2) further includes:

(2a) calculating a second degree of conformance indicating a degree of coincidence of directions of a third vector obtained from an ith directional vector map and indicating a direction from a pixel position of the jth keypoint toward the ith keypoint and a fourth vector indicating a direction from a pixel position expressed by the three-dimensional coordinates of the jth keypoint to a pixel position expressed by the three-dimensional coordinates of the ith keypoint; and

(2b) calculating an integrated degree of conformance by integration of the first degree of conformance and the second degree of conformance, and

(c2-3) further executes the grouping according to the integrated degree of conformance,

wherein the feature map used at (c2) further contains a field map showing whether pixels belong to a same object, and

(c2-3) further includes:

(3a) estimating that the ith keypoint and the jth keypoint do not belong to a same object when the integrated degree of conformance is lower than a threshold;

(3b) estimating whether the ith keypoint and the jth keypoint belong to a same object using the field map when the integrated degree of conformance is equal to or higher than the threshold;

(3c) adjusting the integrated degree of conformance to a first value when estimated that the ith keypoint and the jth keypoint do not belong to a same object and adjusting the integrated degree of conformance to a second value higher than the first value when estimated that the ith keypoint and the jth keypoint belong to a same object;

(3d) selecting one arbitrary keypoint set including N keypoints from the first keypoint to the Nth keypoint from the M×N keypoints;

(3e) calculating a set degree of conformance for the keypoint set by adding the integrated degrees of conformance for N (N−1)/2 keypoint pairs respectively formed by two arbitrary keypoints contained in the keypoint set;

(3f) repeating (3d), ( 3 e ) and obtaining the set degrees of conformance for a plurality of the keypoint sets; and

(3g) settling the grouping relating to the keypoint set in descending order of the set degree of conformance.

2 . A non-transitory computer-readable storage medium storing a computer program for controlling a processor to execute processing of recognizing a position and an attitude of an object using first to Nth N keypoints set for the object, M being an integer of 1 or more and N being an integer of 2 or more, the computer program for controlling the processor to execute:

(a) processing of generating an input image by imaging a scene containing M objects by a camera;

(b) processing of obtaining a feature map showing feature amounts relating to the N keypoints from the input image using a learned machine learning model with the input image as input and the feature map as output;

(c) processing of obtaining three-dimensional coordinates of the N keypoints belonging to each of the M objects using the feature map; and

(d) processing of determining positions and attitudes of one or more objects of the M objects using the three-dimensional coordinates of the N keypoints belonging to each of the M objects, wherein

(c) includes:

(c1) processing of obtaining M×N keypoints having undetermined correspondence relationships with the M objects and determining the three-dimensional coordinates of the M×N keypoints; and

(c2) processing of grouping the M×N keypoints to the N keypoints belonging to each of the M objects, the feature map used at (c2) contains N directional vector maps as maps in which vectors indicating directions from a plurality of pixels belonging to a same object to an object keypoint are assigned to the plurality of pixels with each of the N keypoints as the object keypoint, and

(c2) includes:

(c2-1) processing of selecting one ith keypoint from M ith keypoints and selecting one jth keypoint from M jth keypoints;

(c2-2) processing of calculating a first degree of conformance indicating a degree of coincidence of directions of a first vector obtained from a jth directional vector map and indicating a direction from a pixel position of the ith keypoint toward the jth keypoint and a second vector indicating a direction from a pixel position expressed by the three-dimensional coordinates of the ith keypoint to a pixel position expressed by the three-dimensional coordinates of the jth keypoint, i and j being integers from 1 to N different from each other; and

(c2-3) processing of repeating (c2-1) and (c2-2) and performing the grouping of the M×N keypoints according to the first degree of conformance,

wherein (c2-2) further includes:

(2a) calculating a second degree of conformance indicating a degree of coincidence of directions of a third vector obtained from an ith directional vector map and indicating a direction from a pixel position of the jth keypoint toward the ith keypoint and a fourth vector indicating a direction from a pixel position expressed by the three-dimensional coordinates of the jth keypoint to a pixel position expressed by the three-dimensional coordinates of the ith keypoint; and

(2b) calculating an integrated degree of conformance by integration of the first degree of conformance and the second degree of conformance, and

(c2-3) further executes the grouping according to the integrated degree of conformance,

wherein the feature map used at (c2) further contains a field map showing whether pixels belong to a same object, and

(c2-3) further includes:

(3a) estimating that the ith keypoint and the jth keypoint do not belong to a same object when the integrated degree of conformance is lower than a threshold;

(3b) estimating whether the ith keypoint and the jth keypoint belong to a same object using the field map when the integrated degree of conformance is equal to or higher than the threshold;

(3c) adjusting the integrated degree of conformance to a first value when estimated that the ith keypoint and the jth keypoint do not belong to a same object and adjusting the integrated degree of conformance to a second value higher than the first value when estimated that the ith keypoint and the jth keypoint belong to a same object;

(3d) selecting one arbitrary keypoint set including N keypoints from the first keypoint to the Nth keypoint from the M×N keypoints;

(3e) calculating a set degree of conformance for the keypoint set by adding the integrated degrees of conformance for N (N−1)/2 keypoint pairs respectively formed by two arbitrary keypoints contained in the keypoint set;

(3f) repeating (3d), (3e) and obtaining the set degrees of conformance for a plurality of the keypoint sets; and

(3g) settling the grouping relating to the keypoint set in descending order of the set degree of conformance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2023
From: HAYASHI, MASAKI; KASAHARA, HIROKAZU; FU, GUOYI; LUO, ZHONGZHEN
To: SEIKO EPSON CORPORATION
Reel/Frame 064724/0465 →
Priority Claims (1)
JP 2022-135489 · Aug 29, 2022 · national
Continuity (1)
Related Publication 20240070896A1 · Feb 29, 2024
References Cited (19)
US 20210390731A1 · Wang · 2021 [cited by examiner]
US 20220277472A1 · Birchfield · 2022 [cited by examiner]
US 20250095192A1 · Pan · 2025 [cited by examiner]
WO WO2022001106A1 · 2022 [cited by examiner]
Dou, J., Qin, Q., & Tu, Z. (2021). Multi-modal image registration based on local self-similarity and bidirectional matching. Pattern Recognition and Image Analysis, 31(1), 7-17. (Year: 2021). [cited by examiner]
Z. Cao, G. Hidalgo, T. Simon, S.-E. Wei and Y. Sheikh, “OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, No. 1, … [cited by examiner]
X. Zhang, X. Zhang, F. Ji and Q. Yu, “Known Landing Area Rough Locating from Far Distance Based on DDM-SIFT,” 2008 Congress on Image and Signal Processing, Sanya, China, 2008, pp. 686-690, doi: 10.1109/CISP.2008.473. (Y… [cited by examiner]
N. Bold, C. Zhang and T. Akashi, “3D Point Cloud Retrieval With Bidirectional Feature Match,” in IEEE Access, vol. 7, pp. 164194-164202, 2019, doi: 10.1109/ACCESS.2019.2952157 (Year: 2019). [cited by examiner]
W. Gao and R. Tedrake, “kPAM-SC: Generalizable Manipulation Planning using KeyPoint Affordance and Shape Completion,” 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi'an, China, 2021, pp. 6527-65… [cited by examiner]
W. Gao and R. Tedrake, “kPAM 2.0: Feedback Control for Category-Level Robotic Manipulation,” in IEEE Robotics and Automation Letters, vol. 6, No. 2, pp. 2962-2969, Apr. 2021, doi: 10.1109/LRA.2021.3062315. (Year: 2021). [cited by examiner]
Z. Cao, G. Hidalgo, T. Simon, S.-E. Wei and Y. Sheikh, “OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, No. 1, … [cited by examiner]
Manuelli, L., Gao, W., Florence, P., Tedrake, R. (2022). KPAM: KeyPoint Affordances for Category-Level Robotic Manipulation. ISRR 2019. Springer Proceedings in Advanced Robotics, vol. 20. Springer, Cham. https://doi.org… [cited by examiner]
Z. Luo, W. Xue, J. Chae and G. Fu, “SKP: Semantic 3D Keypoint Detection for Category-Level Robotic Manipulation,” in IEEE Robotics and Automation Letters, vol. 7, No. 2, pp. 5437-5444, Apr. 2022, doi: 10.1109/LRA.2022.3… [cited by examiner]
Lauer, J., Zhou, M., Ye, S. et al. Multi-animal pose estimation, identification and tracking with DeepLabCut. Nat Methods 19, 496-504 (2022). https://doi.org/10.1038/s41592-022-01443-0 (Year: 2022). [cited by examiner]
Gao, Wei and Russ Tedrake: “kPAM 2.0: Feedback Control for Category-Level Robotic Manipulation”; IEEE Robotics and Automation Letters; arXiv:2102.06279v1 [cs.RO]; Feb. 11, 2021; (8pp). [cited by applicant]
Xu, Ruinian; Fu-Jen Chu; Chao Tang et al.: “An Affordance Keypoint Detection Network for Robot Manipulation”; IEEE Robotics and Automation Letters; vol. 6; Issue: 2; Apr. 2021; (8pp). [cited by applicant]
Manuelli, Lucas; Wei Gao; Peter Florence et al.: “kPAM: KeyPoint Affordances for Category-Level Robotic Manipulation”; arXiv:1903.06684v2 [cs.RO]; Oct. 29, 2019; (26pp). [cited by applicant]
Gao, Wei and Russ Tedrake: “kPAM-SC: Generalizable Manipulation Planning using KeyPoint Affordance and Shape Completion”; arXiv: 1909.06980v1 [cs.RO]; Sep. 16, 2019; (7pp). [cited by applicant]
Luo, Zhongzhen; Wenjie Xue; Julia Chae et al. “SKP: Semantic 3D Keypoint Detection for Category-Level Robotic Manipulation”; IEEE Robotics and Automation Letters; vol. 7; No. 2; Apr. 2022; pp. 5437-5444; (8pp). [cited by applicant]