IP Library › Granted Patent US 12,372,962
Granted Patent B1
US 12,372,962 · App. 18/067,004 · Granted Jul 29, 2025

Automatic learning of relevant actions in mobile robots

Inventors: Nigel D. Stepp (Santa Monica, CA); Charles E. Martin (Santa Monica, CA); Heiko Hoffmann (Simi Valley, CA)
Assignee: HRL LABORATORIES, LLC
G05D1/0088G06N3/091
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,372,962
App. No.
18/067,004
Filed
Dec 15, 2022
Granted
Jul 29, 2025
Kind
B1
Art Unit
3666
USPC
701/27
Abstract

Described is a system and method for robotic perception and action. The system includes a mobile autonomous platform programmed to receive a visual input related to an environment and perform a movement sequence that affects the environment based on the visual input. A neural network is in communication with the mobile autonomous platform. The neural network includes a vision autoencoder trained, based on movements of the mobile autonomous platform, to encode visual features of an object in the visual input onto a visual latent space. In addition, the neural network includes an action autoencoder trained, based on movements of the mobile autonomous platform, to encode action features of the movement sequence onto an action latent space. The neural network also includes a mapping layer between the vision autoencoder and the action autoencoder. The mapping layer decodes the visual inputs into an action output of the mobile autonomous platform.

Claims (114)

1. A system for automatic learning of relevant actions in mobile autonomous platforms, comprising:

a mobile autonomous platform programmed to receive a visual input related to an environment and perform a movement sequence that affects the environment based on the visual input;

a neural network in communication with the mobile autonomous platform, the neural network comprising:

a vision autoencoder trained, based on movements of the mobile autonomous platform, to encode visual features of a training object in the visual input onto a visual latent space;

an action autoencoder trained, based on movements of the mobile autonomous platform by the mobile autonomous platform, to encode action features of the movement sequence onto an action latent space; and

a mapping layer between the vision autoencoder and the action autoencoder, wherein the mapping layer maps the visual latent space to the action latent space and, given a new object that shares visual features with the training object and given a new objective, translates the visual features of the new object to action features and associated action possibilities of the mobile autonomous platform to achieve the objective.

2. The system as set forth in claim 1 , wherein the mapping layer encodes a series of action output controls for causing the autonomous platform to reach an object based on visual inputs from the autonomous platform to the training object.

3. The system as set forth in claim 1 , wherein the visual input comprises visual features specifying action possibilities for reaching the training object in the environment.

4. The system as set forth in claim 1 , wherein the visual features comprise texture gradients and edge features of the training object.

5. A method for automatic learning of relevant actions in mobile autonomous platforms, the method comprising:

receiving, by a mobile autonomous platform, a visual input related to an environment;

performing, by the mobile autonomous platform, a movement sequence that affects the environment;

training a vision autoencoder of a neural network to encode visual features of a training object in the visual input onto a visual latent space;

training an action autoencoder of the neural network to encode action features of the movement sequence onto an action latent space;

learning a mapping from the visual latent space to the action latent space;

using the mapping and given an objective, translating visual features from an instance of a new object that shares visual features with the training object to action features that specify action possibilities for the mobile autonomous platform to achieve the objective.

6. The method as set forth in claim 5 , wherein the visual features comprise texture gradients and edge features of the training object.

7. The method as set forth in claim 5 , further comprising training each of the vision autoencoder and the action autoencoder with a model loss function.

8. The method as set forth in claim 5 , further comprising triggering training of a mapping layer of the neural network with a model loss function when the mobile autonomous platform physically contacts the training object in the environment.

9. The method as set forth in claim 8 , wherein model loss is computed separately for the vision autoencoder, the action autoencoder, and the mapping layer.

10. The method as set forth in claim 9 , further comprising determining model loss functions for vision, action, and mapping, respectively, according to the following:

L

V

(

x

,

y

)

=

1

N

⁢

∑

n

N

y

n

⁢

log

⁢

σ

⁡

(

x

n

)

+

(

1

-

y

n

)

⁢

log

⁡

(

1

-

σ

⁡

(

x

n

)

)

L

A

(

x

,

y

)

=

1

N

⁢

∑

n

N

❘

"\[LeftBracketingBar]"

x

n

-

y

n

❘

"\[RightBracketingBar]"

L

M

(

x

,

y

)

=

1

N

⁢

∑

n

N

⁢

(

x

n

-

y

n

)

2

,

where for each loss function, x and y correspond to a decoded image and an input image for L V , a decoded action and an input action for L A , a mapped vision and encoded action for L M , N is a number of elements in x or y, and n is an index of a specific element.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2023
From: STEPP, NIGEL D.; MARTIN, CHARLES E.; HOFFMANN, HEIKO
To: HRL LABORATORIES, LLC
Reel/Frame 063908/0841 →
Continuity (1)
Provisional Application 63290617 · Dec 16, 2021
References Cited (8)
US 11427210B2 · Rosman · 2022 [cited by examiner]
US 11447127B2 · Tawari · 2022 [cited by examiner]
US 11687087B2 · Choi · 2023 [cited by examiner]
US 20230222328A1 · Wei · 2023 [cited by examiner]
US 20240096077A1 · Luo · 2024 [cited by examiner]
Martin Lohmann, Jordi Salvador, Aniruddha Kembhavi, Roozbeh Mottaghi. Leaming About Objects by Learning to Interact with Them. 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada, … [cited by applicant]
Guang-Bin Huang, Yan-Qiu Chen and H. A. Babri, “Classification ability of single hidden layer feed forward neural networks,” in IEEE Transactions on Neural Networks, vol. 11, No. 3, pp. 799-801, May 2000. [cited by applicant]
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., & Chintala, S. (2019). “Pytorch: An imperative style, high-performance deep learning library”, Advances in neural information processing systems, 32… [cited by applicant]