IP Library Granted Patent US 10,672,398
Granted Patent B1
US 10,672,398 · App. 16/568,506 · Granted Jun 2, 2020

Speech recognition biasing

Inventor: Daniel Alex Lam (San Francisco, CA)
Assignee: X Development LLC
G10L15/22B25J13/003G06N7/005G10L15/265G10L15/183G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,672,398
App. No.
16/568,506
Granted
Jun 2, 2020
Kind
B1
Abstract

Systems and methods are described include a robot and/or an associated computing system that can use various cues about an environment of the robot to apply a bias to increase the accuracy of speech transcription. In some implementations, audio data corresponding to a spoken instruction to a robot is received. Candidate transcriptions of the audio data are obtained. A respective action of the robot corresponding to each of the candidate transcriptions of the audio data is determined. One or more scores indicating characteristics of a potential outcome of performing the respective action corresponding to the candidate transcription of the audio data are determined for each of the candidate transcriptions of the audio data. A particular candidate transcription is selected from among the candidate transcriptions based at least on the one or more scores. The action determined for the particular candidate transcription is performed.

Claims (49)

1. A method comprising:

receiving audio data corresponding to a spoken instruction to a robot;

obtaining candidate transcriptions for the audio data;

accessing object characteristics data that indicates characteristics for a plurality of objects or types of objects;

using, for each of one or more of the candidate transcriptions, the object characteristics data to evaluate a potential effect of the robot performing an action corresponding to the candidate transcription of the audio data;

selecting a particular candidate transcription from among the candidate transcriptions based on the evaluation; and

causing the robot to perform the action determined for the particular candidate transcription.

2. The method of claim 1 , wherein the object characteristics data indicates physical characteristics for the plurality of objects or types of objects.

3. The method of claim 2 , wherein the physical characteristics comprise a weight for the respective objects in the plurality objects or the respective types of objects.

4. The method of claim 2 , wherein the physical characteristics comprise a fragility for the respective objects in the plurality objects or the respective types of objects.

5. The method of claim 1 , wherein the object characteristics data indicates values indicative of how manipulations of specific objects or types of objects can result in damage to the specific objects or types of objects.

6. The method of claim 1 , wherein the object characteristics data indicates, for a particular object or type of object, a set of allowable actions and/or a set of unallowable actions.

7. The method of claim 1 , wherein the object characteristics data indicates multiple attributes of each of the plurality of objects or types of objects;

wherein using the object characteristics data to evaluate a potential effect of the robot performing an action corresponding to the candidate transcription of the audio data comprises:

identifying an object mentioned in a particular candidate transcription;

evaluating multiple attributes for the identified object that are indicated by the object characteristics data; and

generating, based on evaluating the multiple attributes, one or more scores representing the potential effect of the robot performing the action corresponding to the particular transcription.

8. The method of claim 1 , comprising receiving context data that indicates (i) a location of the robot, and (ii) one or more objects within a threshold level of proximity to the location of the robot;

wherein evaluating the potential effect of the robot performing an action corresponding to the candidate transcription of the audio data is further based on the context data.

9. The method of claim 1 , further comprising:

determining, for each of the candidate transcriptions of the audio data, a confidence score that reflects a likelihood that candidate transcription is an accurate representation of the audio data, wherein the confidence scores are each biased on one or more scores indicating results of evaluation of the potential effect of the robot performing the action corresponding to the candidate transcription.

10. The method of claim 1 , further comprising:

computing a recognition score for each of the candidate transcriptions;

computing an impact score for each of the candidate transcriptions, each impact score corresponding to a predicted effect of performing an action indicated by the corresponding candidate transcription; and

combining, for each of the candidate transcriptions, the recognition score and the impact score to compute a confidence for the candidate transcription.

11. A system comprising:

one or more computers; and

one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

receiving audio data corresponding to a spoken instruction to a robot;

obtaining candidate transcriptions for the audio data;

accessing object characteristics data that indicates characteristics for a plurality of objects or types of objects;

using, for each of one or more of the candidate transcriptions, the object characteristics data to evaluate a potential effect of the robot performing an action corresponding to the candidate transcription of the audio data;

selecting a particular candidate transcription from among the candidate transcriptions based on the evaluation; and

causing the robot to perform the action determined for the particular candidate transcription.

12. The system of claim 11 , wherein the object characteristics data indicates physical characteristics for the plurality of objects or types of objects.

13. The system of claim 12 , wherein the physical characteristics comprise a weight for the respective objects in the plurality objects or the respective types of objects.

14. The system of claim 12 , wherein the physical characteristics comprise a fragility for the respective objects in the plurality objects or the respective types of objects.

15. The system of claim 11 , wherein the object characteristics data indicates values indicative of how manipulations of specific objects or types of objects can result in damage to the specific objects or types of objects.

16. One or more non-transitory computer-readable storage media encoded with computer program instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:

receiving audio data corresponding to a spoken instruction to a robot;

obtaining candidate transcriptions for the audio data;

accessing object characteristics data that indicates characteristics for a plurality of objects or types of objects;

using, for each of one or more of the candidate transcriptions, the object characteristics data to evaluate a potential effect of the robot performing an action corresponding to the candidate transcription of the audio data;

selecting a particular candidate transcription from among the candidate transcriptions based on the evaluation; and

causing the robot to perform the action determined for the particular candidate transcription.

17. The computer-readable storage media of claim 16 , wherein the object characteristics data indicates physical characteristics for the plurality of objects or types of objects.

18. The computer-readable storage media of claim 17 , wherein the physical characteristics comprise a weight for the respective objects in the plurality objects or the respective types of objects.

19. The computer-readable storage media of claim 17 , wherein the physical characteristics comprise a fragility for the respective objects in the plurality objects or the respective types of objects.

20. The computer-readable storage media of claim 16 , wherein the object characteristics data indicates values indicative of how manipulations of specific objects or types of objects can result in damage to the specific objects or types of objects.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 29, 2025
From: GOOGLE LLC
To: GDM HOLDING LLC
Reel/Frame 071109/0342 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2023
From: X DEVELOPMENT LLC
To: GOOGLE LLC
Reel/Frame 064658/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2020
From: LAM, DANIEL ALEX
To: X DEVELOPMENT LLC
Reel/Frame 051794/0233 →