IP Library Granted Patent US 10,891,484
Granted Patent B2
US 10,891,484 · App. 16/269,183 · Granted Jan 12, 2021

Selectively downloading targeted object recognition modules

Inventors: Nareshkumar Rajkumar (San Jose, CA); Stefan Hinterstoisser (Mountain View, CA); Max Bajracharya (Millbrae, CA)
Assignee: X DEVELOPMENT LLC
G06K9/00664B25J9/1697G06K9/00979G06K9/00993G06K9/628G06T1/0014H04W4/025H04W4/60H04W4/80Y10S901/47
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,891,484
App. No.
16/269,183
Granted
Jan 12, 2021
Kind
B2
Abstract

Methods, apparatus, systems, and computer-readable media are provided for downloading targeted object recognition modules that are selected from a library of candidate targeted object recognition modules based on various signals. In some implementations, an object recognition client may be operated to facilitate object recognition for a robot. It may download targeted object recognition module(s). Each targeted object recognition module may facilitate inference of an object type or pose of an observed object. The targeted object module(s) may be selected from a library of targeted object recognition modules based on various signals, such as a task to be performed by the robot. The object recognition client may obtain vision data capturing at least a portion of an environment in which the robot operates. The object recognition client may determine, based on the vision data and the downloaded object recognition module(s), information about an observed object in the environment.

Claims (44)

1. A method implemented using one or more processors, comprising:

operating, by one or more processors integral with a robot, an object recognition client to facilitate object recognition for the robot;

determining an assigned task to be performed by the robot;

determining one or more expected objects or object types associated with the assigned task;

based on the assigned task and subsequent to the determining, downloading, to the robot by the object recognition client, from a remote computing system, one or more targeted object recognition modules, wherein each targeted object recognition module is usable by the object recognition client to calculate, at the robot, an object type or pose of an observed object, and wherein the one or more targeted object recognition modules are selected from a library of targeted object recognition modules that is remote from the robot based at least in part on one or more of the expected objects or object types associated with the assigned task to be performed by the robot;

obtaining, by the object recognition client, from one or more vision sensors integral with the robot, vision data capturing, from a perspective of the robot, at least a portion of an area in which the robot is operating;

determining, at the robot by the object recognition client, based on the vision data and the one or more downloaded targeted object recognition modules, two or more conflicting inferences about an object in the area that is detected in the vision data captured by the one or more vision sensors integral with the robot;

disambiguating, by the object recognition client, between the two or more conflicting inferences based on canonical models associated with each of the two or more conflicting inferences; and

operating the robot to perform the assigned task.

2. The method of claim 1 , wherein the disambiguating comprises rendering the canonical models and comparing the rendered canonical models with the vision data.

3. The method of claim 1 , wherein the disambiguating comprises rendering the canonical models in poses inferred based on the vision data.

4. The method of claim 1 , wherein the canonical models comprise computer-aided designs.

5. The method of claim 1 , wherein the one or more targeted object recognition modules are selected from the library further based at least in part on available resources of the robot.

6. The method of claim 5 , wherein the available resources of the robot include one or more attributes of a wireless signal available to the robot.

7. The method of claim 1 , further comprising selecting, by the object recognition client, the one or more targeted object recognition modules from the library.

8. The method of claim 1 , wherein the one or more targeted object recognition modules are selected from the library by the remote computing system based on one or more of the expected objects or types associated with the assigned task.

9. A system comprising one or more processors and memory storing instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to perform the following operations:

operating, by one or more processors integral with a robot, an object recognition client to facilitate object recognition for the robot;

determining an assigned task to be performed by the robot;

determining one or more expected objects or object types associated with the assigned task;

based on the assigned task and subsequent to the determining, downloading, to the robot by the object recognition client, from a remote computing system, one or more targeted object recognition modules, wherein each targeted object recognition module is usable by the object recognition client to calculate, at the robot, an object type or pose of an observed object, and wherein the one or more targeted object recognition modules are selected from a library of targeted object recognition modules that is remote from the robot based at least in part on one or more of the expected objects or object types associated with the assigned task to be performed by the robot;

obtaining, by the object recognition client, from one or more vision sensors integral with the robot, vision data capturing, from a perspective of the robot, at least a portion of an area in which the robot is operating;

determining, at the robot by the object recognition client, based on the vision data and the one or more downloaded targeted object recognition modules, two or more conflicting inferences about an object in the area that is detected in the vision data captured by the one or more vision sensors integral with the robot;

disambiguating, at the robot by the object recognition client, between the two or more conflicting inferences based on canonical models associated with each of the two or more conflicting inferences; and

operating the robot to perform the assigned task.

10. The system of claim 9 , wherein the disambiguating comprises rendering the canonical models and comparing the rendered canonical models with the vision data.

11. The system of claim 9 , wherein the disambiguating comprises rendering the canonical models in poses inferred based on the vision data.

12. The system of claim 9 , wherein the canonical models comprise computer-aided designs.

13. The system of claim 9 , wherein the one or more targeted object recognition modules are selected from the library further based at least in part on available resources of the robot.

14. The system of claim 13 , wherein the available resources of the robot include one or more attributes of a wireless signal available to the robot.

15. The system of claim 9 , further comprising instructions for selecting, by the object recognition client, the one or more targeted object recognition modules from the library.

16. The system of claim 9 , wherein the one or more targeted object recognition modules are selected from the library by the remote computing system based on one or more of the expected objects or types associated with the assigned task.

17. At least one non-transitory computer-readable medium comprising instructions that, in response to execution of the instructions by one or more processors, cause the one or more processors to perform the following operations:

operating, by one or more processors integral with a robot, an object recognition client to facilitate object recognition for the robot;

determining an assigned task to be performed by the robot;

determining one or more expected objects or object types associated with the assigned task;

based on the assigned task and subsequent to the determining, downloading, to the robot by the object recognition client, from a remote computing system, one or more targeted object recognition modules, wherein each targeted object recognition module is usable by the object recognition client to calculate, at the robot, an object type or pose of an observed object, and wherein the one or more targeted object recognition modules are selected from a library of targeted object recognition modules that is remote from the robot based at least in part on one or more of the expected objects or object types associated with the assigned task to be performed by the robot;

obtaining, by the object recognition client, from one or more vision sensors integral with the robot, vision data capturing, from a perspective of the robot, at least a portion of an area in which the robot is operating;

determining, at the robot by the object recognition client, based on the vision data and the one or more downloaded targeted object recognition modules, two or more conflicting inferences about an object in the area that is detected in the vision data captured by the one or more vision sensors integral with the robot;

disambiguating, by the object recognition client, between the two or more conflicting inferences based on canonical models associated with each of the two or more conflicting inferences; and

operating the robot to perform the assigned task.

18. The at least one non-transitory computer-readable medium of claim 17 , wherein the disambiguating comprises rendering the canonical models and comparing the rendered canonical models with the vision data.

19. The at least one non-transitory computer-readable medium of claim 17 , wherein the disambiguating comprises rendering the canonical models in poses inferred based on the vision data.

20. The at least one non-transitory computer-readable medium of claim 17 , wherein the canonical models comprise computer-aided designs.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2025
From: GOOGLE LLC
To: GDM HOLDING LLC
Reel/Frame 071109/0342 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2023
From: X DEVELOPMENT LLC
To: GOOGLE LLC
Reel/Frame 063992/0371 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2019
From: RAJKUMAR, NARESHKUMAR; HINTERSTOISSER, STEFAN; BAJRACHARYA, MAX
To: GOOGLE INC.
Reel/Frame 048254/0900 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2019
From: GOOGLE INC.
To: X DEVELOPMENT LLC
Reel/Frame 048278/0139 →