IP Library Granted Patent US 12,468,902
Granted Patent B2
US 12,468,902 · App. 18/172,969 · Granted Nov 11, 2025

Systems and methods for automated response to natural language instructions

Inventors: Divyansh Garg (Stanford, CA); Skanda Vaidyanath (Stanford, CA)
Assignee: The Board of Trustees of the Leland Stanford Junior University
G06F40/44G06F40/216G06F40/30G06N3/045G06N20/00G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,902
App. No.
18/172,969
Granted
Nov 11, 2025
Kind
B2
Abstract

Systems and methods for automated response to natural language instructions in accordance with embodiments of the invention are illustrated. One embodiment includes a method for training an agent, the method including sampling an instruction and observation pair from a dataset, predicting a skill code, using a skill predictor, based on the instruction and observation pair, predicting, for each of a plurality of timesteps, a set of one or more actions based on the predicted skill code and a state history using a policy, and updating the skill predictor and the policy based on a comparison of the predicted set of actions and the observation. In many embodiments, the trained agent can then be used to carry out natural language instructions.

Claims (27)

1 . A method for enabling a machine to act upon natural language instructions, comprising:

obtaining a plurality of instruction and observation pairs;

generating language embeddings for each instruction in the plurality of instruction and observation pairs using a language encoder;

generating observation embeddings for each observation in the plurality of instruction and observation pairs;

predicting a set of skill codes for each given pair in the plurality of instruction and observation pairs based on a given language embedding and a given observation embedding generated from the given pair using a skill predictor;

predicting an action to correctly resolve the instruction of the given pair using a policy based on the set of skill codes; and

controlling a device to perform the predicted action.

2 . The method of claim 1 , wherein the set of skill codes are human interpretable.

3 . The method of claim 1 , wherein vector quantization is used to translate the predicted set of skill codes into enumerated skill codes from a digital codebook.

4 . The method of claim 1 , wherein the controlled device is a robot.

5 . The method of claim 1 , wherein the controlled device is an autonomous vehicle.

6 . The method of claim 1 , wherein the controlled device is a virtual avatar.

7 . A system for enabling a machine to act upon natural language instructions, comprising:

a processor;

a controllable device; and

a memory, comprising a natural language processing application that configures the processor to:

obtain a plurality of instruction and observation pairs;

generate language embeddings for each instruction in the plurality of instruction and observation pairs using a language encoder;

generate observation embeddings for each observation in the plurality of instruction and observation pairs;

predict a set of skill codes for each given pair in the plurality of instruction and observation pairs based on a given language embedding and a given observation embedding generated from the given pair using a skill predictor;

predict an action to correctly resolve the instruction of the given pair using a policy based on the set of skill codes; and

control the controllable device to perform the predicted action.

8 . The system of claim 7 , wherein the set of skill codes are human interpretable.

9 . The system of claim 7 , wherein vector quantization is used to translate the predicted set of skill codes into enumerated skill codes from a digital codebook.

10 . The system of claim 7 , wherein the controlled device is a robot.

11 . The system of claim 7 , wherein the controlled device is an autonomous vehicle.

12 . The system of claim 7 , wherein the controlled device is a virtual avatar.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 7, 2025
From: GARG, DIVYANSH; VAIDYANATH, SKANDA
To: THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIVERSITY
Reel/Frame 071966/0586 →
CONFIRMATORY LICENSE Recorded Feb 12, 2025
From: STANFORD UNIVERSITY
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 070189/0812 →
Continuity (2)
Provisional Application 63268364 · Feb 22, 2022
Related Publication 20230267284A1 · Aug 24, 2023
References Cited (13)
US 11651162B2 · Gadde · 2023 [cited by examiner]
US 11989523B2 · Gadde · 2024 [cited by examiner]
US 20200342175A1 · Gadde · 2020 [cited by examiner]
US 20210086353A1 · Shah et al. · 2021 [cited by applicant]
US 20220147548A1 · Verma · 2022 [cited by examiner]
US 20230206004A1 · Gadde · 2023 [cited by examiner]
US 20240242034A1 · Gadde · 2024 [cited by examiner]
CN 112809689A · 2021 [cited by applicant]
JP 6921022B2 · 2021 [cited by applicant]
WO 2006119577A1 · 2006 [cited by applicant]
WO 2021231895A1 · 2021 [cited by applicant]
Mees O, Hermann L, Rosete-Beas E, Burgard W. Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. IEEE Robotics and Automation Letters. Jun. 3, 2022;7(3):7327-34. (Year… [cited by examiner]
Garg et al., “LISA: Learning Interpretable Skill Abstractions from Language”, Proceedings of the 39th International Conference on Machine Learning, Baltimore, Maryland, PMLR, 162, 2022, 22 pgs. [cited by applicant]