IP Library Granted Patent US 10,572,775
Granted Patent B2
US 10,572,775 · App. 15/832,705 · Granted Feb 25, 2020

Learning and applying empirical knowledge of environments by robots

Inventor: Alexa Greenberg (San Francisco, CA)
Assignee: X DEVELOPMENT LLC
G06K9/6262B25J9/163B25J9/1697G05B13/027G06K9/00664G06K9/32G06N3/04G06N3/0454G06N3/08G06N3/084G06T7/97G05B2219/33038G05B2219/39046G05B2219/39543G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,572,775
App. No.
15/832,705
Granted
Feb 25, 2020
Kind
B2
Abstract

Techniques described herein relate to generating a posteriori knowledge about where objects are typically located within environments to improve object location. In various implementations, output from vision sensor(s) of a robot may include visual frame(s) that capture at least a portion of an environment in which a robot operates/will operate. The visual frame(s) may be applied as input across a machine learning model to generate output that identifies potential location(s) of an object of interest. The robot's position/pose may be altered based on the output to relocate one or more of the vision sensors. One or more subsequent visual frames that capture at least a not-previously-captured portion of the environment may be applied as input across the machine learning model to generate subsequent output identifying the object of interest. The robot may perform task(s) that relate to the object of interest.

Claims (18)

1. A method implemented by one or more processors, comprising:

determining an object of interest;

receiving vision data, the vision data generated based on vision sensor output from one or more vision sensors of a vision component of a robot, the vision data including one or more visual frames that capture at least a portion of an environment in which a robot operates or will operate;

applying one or more of the visual frames as input across a machine learning model to generate output, wherein the output identifies one or more probabilities that a plurality of pixels representing one or more surfaces in the portion of the environment captured in the one or more visual frames that conceal, from a vantage point of the one or more vision sensors, an instance of the object of interest; and

altering a position or pose of the robot based on the output to relocate one or more of the vision sensors to have a direct view behind one or more of the surfaces;

wherein the machine learning model was trained using one or more annotated vision frames in which a portion of a surface of the one or more annotated vision frames is annotated as concealing another object of interest of a same type as the object of interest.

2. The method of claim 1 , wherein the machine learning model comprises a convolutional neural network.

3. The method of claim 1 , wherein the input applied across the machine learning model includes a reduced dimensionality embedding of the object of interest.

4. A method implemented by one or more processors, comprising:

determining an object of interest;

receiving vision data, the vision data generated based on vision sensor output from one or more vision sensors of a vision component of a robot, the vision data including at least one visual frame that captures at least a portion of an environment in which a robot operates or will operate;

applying the at least one visual frame as input across a machine learning model to generate output, wherein the output identifies one or more other portions of the environment that are outside of the portion of the environment captured by the at least one visual frame, wherein the one or more other portions of the environment potentially include an instance of the object of interest;

altering a position or pose of the robot based on the output to relocate one or more of the vision sensors to have a direct view of a given other portion of the one or more other portions of the environment;

obtaining, from one or more of the vision sensors, at least one subsequent visual frame that captures the given other portion of the environment;

applying the at least one subsequent visual frame as input across the machine learning model to generate subsequent output, wherein the subsequent output identifies the instance of the object of interest; and

operating the robot to perform one or more tasks that relate to the instance of the object of interest.

5. The method of claim 4 , wherein the machine learning model comprises a convolutional neural network.

6. The method of claim 4 , wherein the input applied across the machine learning model includes a reduced dimensionality embedding of the object of interest.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2025
From: GOOGLE LLC
To: GDM HOLDING LLC
Reel/Frame 071109/0342 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2023
From: X DEVELOPMENT LLC
To: GOOGLE LLC
Reel/Frame 063992/0371 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2017
From: GREENBERG, ALEXA
To: X DEVELOPMENT LLC
Reel/Frame 044318/0931 →
Continuity (1)
Related Publication 20190171911A1 · Jun 6, 2019
Cited By (2)
US 12,290,917 US 12,485,426