IP Library Patent Application 18663491
Patent Application
App. No. 18/663,491

BRIDGING LANGUAGE AND ENVIRONMENTS WITH RENDERING FUNCTIONS AND VISION-LANGUAGE MODELS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/663,491
Abstract

A robot system includes: image encodings generated based on renderings of configurations, respectively of a robot, the configurations including at least a predetermined number of different poses of the robot in an environment; an encoding module configured to receive text descriptive of an action to be performed by the robot and to encode the text into a text encoding; a scoring module configured to generate scores for the configurations based on comparisons of (a) the text encoding with (b) the image encoding of the respective configuration; a selection module configured to select k of the configurations based on the scores, where k is an integer greater than or equal to 1; and an actuation module configured to actuate the robot based on the selected k of the configurations based on actuating the robot to achieve the action described in the text.

Claims (42)

1 . A robot system comprising:

image encodings generated based on renderings of configurations, respectively of a robot, the configurations including at least a predetermined number of different poses of the robot in an environment;

an encoding module configured to receive text descriptive of an action to be performed by the robot and to encode the text into a text encoding;

a scoring module configured to generate scores for the configurations based on comparisons of (a) the text encoding with (b) the image encoding of the respective configuration;

a selection module configured to select k of the configurations based on the scores,

where k is an integer greater than or equal to 1; and

an actuation module configured to actuate the robot based on the selected k of the configurations based on actuating the robot to achieve the action described in the text.

2 . The robot system of claim 1 wherein the scoring module is configured to generate the scores using cosine similarity.

3 . The robot system of claim 1 wherein the selection module is configured to select k of the configurations with the k highest scores.

4 . The robot system of claim 1 wherein the renderings include at least two different renderings of each configuration from different points of view.

5 . The robot system of claim 4 wherein the different points of view are on a same horizontal plane.

6 . The robot system of claim 1 further comprising a vision-language model (VLM) module and a projection module configured to finetune the selected k of the configurations,

wherein the actuation module is configured to actuate the robot based on the k finetuned selected configurations.

7 . The robot system of claim 1 wherein the projection module is configured to finetune the selected k configurations based on one of gradient ascent and projected gradient ascent.

8 . The robot system of claim 1 wherein the scoring module is configured to generate a score for one of the configurations based on (a) a first score for the one of the configurations generated based on a first comparison of the text encoding with a first image encoding of the one of the configurations generated based on a first point of view and (b) a second score for the one of the configurations generated based on a second comparison of the text encoding with a second image encoding of the one of the configurations generated based on a second point of view that is different than the first point of view.

9 . The robot system of claim 8 wherein the scoring module is configured to generate the score for the one of the configurations based on an average of the first score and the second score.

10 . The robot system of claim 1 wherein the encoding module is configured to encode the text using a vision-language model (VLM) text encoding algorithm.

11 . The robot system of claim 1 wherein the encoding module includes a neural network configured to encode the text.

12 . The robot system of claim 1 wherein each of the configurations includes three-dimensional coordinates of a portion of the robot in the environment.

13 . The robot system of claim 1 wherein each of the configurations includes angles of a joint of the robot in the environment.

14 . The robot system of claim 1 wherein each of the configurations includes three-dimensional coordinates of an object to be acted upon by the robot in the environment.

15 . The robot system of claim 1 wherein each of the configurations includes at least one dimension describing the orientation of an object to be acted upon by the robot in the environment.

16 . The robot system of claim 1 wherein the image encodings are generated using a vision-language model (VLM) image encoding algorithm based on the renderings of configurations.

17 . The robot system of claim 1 wherein the renderings are generated using the MuJoCo rendering algorithm.

18 . A training system comprising:

the robot system of claim 1 ;

a rendering module configured to generate the renderings based on the configurations, respectively; and

a second encoding module configured to encode the renderings into the image encodings, respectively.

19 . A robot system comprising:

image encodings generated based on renderings of configurations, respectively of a robot, the configurations including at least a predetermined number of different poses of the robot in an environment;

an encoding module configured to receive text descriptive of an action to be performed by the robot and to encode the text into a text encoding;

a scoring module configured to generate scores for the configurations based on comparisons of (a) the text encoding with (b) the image encoding of the respective configuration;

a selection module configured to select k of the configurations based on the scores,

where k is an integer greater than or equal to 1; and

an actuation module configured to actuate the robot based on a dot product of the k image encodings of the selected k of the configurations and actuating the robot to achieve the action described in the text.

20 . A method comprising:

receiving image encodings generated based on renderings of configurations, respectively of a robot, the configurations including at least a predetermined number of different poses of the robot in an environment;

receiving text descriptive of an action to be performed by the robot and to encode the text into a text encoding;

generating scores for the configurations based on comparisons of (a) the text encoding with (b) the image encoding of the respective configuration;

selecting k of the configurations based on the scores,

where k is an integer greater than or equal to 1; and

actuating the robot based on the selected k of the configurations based on actuating the robot to achieve the action described in the text.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2024
From: NAVER LABS CORPORATION
To: NAVER CORPORATION
Reel/Frame 068820/0495 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2024
From: CACHET, THEO; DANCE, CHRISTOPHER
To: NAVER CORPORATION; NAVER LABS CORPORATION
Reel/Frame 067406/0239 →