IP Library Granted Patent US 12,387,346
Granted Patent B2
US 12,387,346 · App. 18/295,998 · Granted Aug 12, 2025

Object pose neural network system

Inventors: Mrinal Kalakrishnan (Palo Alto, CA); Adrian Ling Hin Li (San Francisco, CA); Nicolas Hudson (San Mateo, CA)
Assignee: Google LLC
G06T7/246G06T7/11G06T7/60G06T7/73G06V10/42G06V10/758G06V10/82G06V20/10G06V30/186G06V30/19173G06V30/194G06T2207/10004G06T2207/20084G06T2207/30244
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,346
App. No.
18/295,998
Granted
Aug 12, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium for predicting object pose. In one aspect, a method includes receiving an image of an object having one or more feature points; providing the image as an input to a neural network subsystem trained to receive images of objects and to generate an output including a heat map for each feature point; applying a differentiable transformation on each heat map to generate respective one or more feature coordinates for each feature point; providing the feature coordinates for each feature point as input to an object pose solver configured to compute a predicted object pose for the object, wherein the predicted object pose for the object specifies a position and an orientation of an object; and receiving, at the output of the object pose solver, a predicted object pose for the object in the image.

Claims (33)

1. A computer-implemented method comprising:

receiving an image;

providing the image as an input to a neural network system trained to generate an output comprising a plurality of heat maps corresponding to a plurality of feature points, each heat map representing a likelihood, for each image region of a plurality of image regions in the image, that the image region corresponds to a respective one of the plurality of feature points;

applying a differentiable transformation on each of the plurality of heat maps to generate a plurality of feature coordinates for the plurality of feature points, wherein applying the differentiable transformation on each heat map comprises:

applying a soft argmax function to one or more values of each respective heat map to generate a location of a measure of central tendency of values in the heat map; and

generating a single feature point location within the image region corresponding to the heat map based on the location of the measure of central tendency of values in the heat map; and

providing the plurality of feature coordinates for the plurality of feature points.

2. The method of claim 1 , wherein the image is of a robot, and wherein the method further comprises controlling the robot based on the plurality of feature coordinates for the plurality of feature points.

3. The method of claim 2 , wherein the plurality of feature points corresponds to a plurality of points of the robot.

4. The method of claim 1 , further comprising generating a measure of variance for each feature location from the plurality of feature coordinates of the measure of central tendency of values in the respective heat maps.

5. The method of claim 1 , further comprising:

applying a respective curve fitting procedure to each heat map, each curve fitting procedure for a heat map generating a function based on values in the heat map; and

computing the plurality of feature coordinates based on respective locations of peaks of the function output by the curve fitting procedure.

6. The method of claim 1 , wherein the image is of an object, wherein the method further comprises:

inputting the plurality of feature coordinates for the plurality of feature points into an object pose solver configured to compute a predicted object pose for the object, wherein the predicted object pose for the object specifies a position and an orientation of the object; and

outputting the predicted object pose for the object.

7. The method of claim 6 , wherein the object pose solver uses a least squares regression analysis procedure.

8. The method of claim 6 , further comprising obtaining a ground truth pose for the object by:

obtaining respective ground truth object poses of the object based on pose information contained in each of one or more other images;

computing one or more respective relative displacements of the object using kinematic data associated with the image and respective kinematic data associated with each of the one or more other images; and

computing an interpolated ground truth object pose of the object based on a respective ground truth pose of the object in each of the one or more other images and the one or more respective relative displacements.

9. The method of claim 1 , wherein the plurality of feature points each correspond to a visually distinguishable portion of a robotic arm.

10. A computer-implemented method comprising:

receiving an image, wherein the image is of an object;

providing the image as an input to a neural network system trained to generate an output comprising a plurality of heat maps corresponding to a plurality of feature points, each heat map representing a likelihood, for each image region of a plurality of image regions in the image, that the image region corresponds to a respective one of the plurality of feature points;

applying a differentiable transformation on each of the plurality of heat maps to generate a plurality of feature coordinates for the plurality of feature points;

providing the plurality of feature coordinates for the plurality of feature points;

obtaining a ground truth pose for the object by:

obtaining respective ground truth object poses of the object based on pose information contained in each of one or more other images;

computing one or more respective relative displacements of the object using kinematic data associated with the image and respective kinematic data associated with each of the one or more other images; and

computing an interpolated ground truth object pose of the object based on a respective ground truth pose of the object in each of the one or more other images and the one or more respective relative displacements;

inputting the plurality of feature coordinates for the plurality of feature points into an object pose solver configured to compute a predicted object pose for the object, wherein the predicted object pose for the object specifies a position and an orientation of the object; and

outputting the predicted object pose for the object.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2025
From: GOOGLE LLC
To: GDM HOLDING LLC
Reel/Frame 071465/0754 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2024
From: X DEVELOPMENT LLC
To: GOOGLE LLC
Reel/Frame 068927/0386 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2024
From: KALAKRISHNAN, MRINAL; LI, ADRIAN LING HIN; HUDSON, NICOLAS
To: X DEVELOPMENT LLC
Reel/Frame 066109/0992 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2023
From: X DEVELOPMENT LLC
To: GOOGLE LLC
Reel/Frame 064658/0001 →
Continuity (3)
Continuation 17114083 · Dec 7, 2020
Division 15410702 · Jan 19, 2017
Related Publication 20240078683A1 · Mar 7, 2024
References Cited (15)
US 8385687B1 · Blais-Morin · 2013 [cited by examiner]
US 20150213328A1 · Mase · 2015 [cited by examiner]
US 20170083752A1 · Saberian · 2017 [cited by examiner]
US 20170124415A1 · Choi · 2017 [cited by examiner]
US 20180060701A1 · Krishnamurthy · 2018 [cited by examiner]
US 20180165548A1 · Wang · 2018 [cited by examiner]
US 20180342050A1 · Fitzgerald · 2018 [cited by examiner]
US 20190114743A1 · Lund · 2019 [cited by examiner]
WO WO2014200742A1 · 2014 [cited by examiner]
WO WO2015069824A2 · 2015 [cited by examiner]
WO WO2016132371A1 · 2016 [cited by examiner]
WO WO2018045031A1 · 2018 [cited by examiner]
WO WO2018156133A1 · 2018 [cited by examiner]
WO WO2020163908A1 · 2020 [cited by examiner]
Bach S, Binder A, Montavon G, Klauschen F, Müller KR, Samek W. On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation. PLoS One. Jul. 10, 2015;10(7):e0130140. doi: 10.1371/jou… [cited by examiner]