IP Library › Granted Patent US 10,515,637
Granted Patent B1
US 10,515,637 · App. 15/709,119 · Granted Dec 24, 2019

Dynamic speech processing

Inventors: David William Devries (San Jose, CA); Rajesh Mittal (Milpitas, CA)
Assignee: AMAZON TECHNOLOGIES, INC.
G10L15/30G10L13/00G10L15/183G10L15/1815G10L15/22G06F17/241G06F17/2705G10L15/02G10L2015/025G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,515,637
App. No.
15/709,119
Filed
Sep 19, 2017
Granted
Dec 24, 2019
Kind
B1
Art Unit
2657
USPC
704/270.1
Abstract

Techniques for dynamically maintaining speech processing data on a local device for frequently input commands are described. A system determines a usage history associated with a user profile. The usage history represents at least a first command. The system determines the first command is associated with an input frequency that satisfies an input frequently threshold. The system also determines the first command is missing from first speech processing data stored by a device associated with the user profile. The system then generates second speech processing data specific to the first command and sends the second speech processing data to the device.

Claims (81)

1. A method implemented by a computing system, comprising:

receiving, from a first device associated with at least one user profile, first data;

determining that the first data corresponds to a first command;

determining a unique identifier (ID) associated with the at least one user profile;

determining an updated spoken command history by updating, based at least in part on the first data corresponding to the first command, a spoken command history stored by the computing system and associated with the unique ID;

determining, using the updated spoken command history, that the first command is associated with an input frequency satisfying an input frequency threshold;

determining that the first command is not included in a first speech-recognition language model stored by the first device;

generating a second speech-recognition language model including second data sufficient to determine text corresponding to the first command;

sending, to the first device, the second speech-recognition language model; and

sending, to the first device, a first instruction to delete the first speech-recognition language model.

2. The method of claim 1 , further comprising:

determining a portion of the second speech-recognition language model associated with a second command;

determining, using the updated spoken command history, that the second command is associated with a second input frequency falling below the input frequency threshold; and

sending, to the first device, a second instruction to delete the portion of the second speech-recognition language model associated with the second command.

3. The method of claim 1 , further comprising:

determining that the first command corresponds to a first intent;

generating natural language processing data usable to determine the first intent from text corresponding to the first command; and

sending the natural language processing data to the first device.

4. The method of claim 1 , further comprising:

receiving, from the first device, audio data corresponding to an utterance;

determining, using the updated spoken command history, that the utterance is associated with a second input frequency falling below the input frequency threshold;

performing speech processing on the input audio data to determine output data; and

sending the output data to the first device.

5. A computing system comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the computing system to:

receive, from a first device associated with a user profile, first data;

determine that the first data corresponds to a first command;

determine an updated spoken command history by updating, based at least in part on the first data corresponding to the first command, a spoken command history stored by the computing system and associated with the user profile;

determine, using the updated spoken command history, that the first command is associated with an input frequency satisfying an input frequency threshold;

determine that the first command is at least in part not processable using first speech processing data stored by the first device;

generate second speech processing data including second data specific to the first command; and

send the second speech processing data to the first device.

6. The computing system of claim 5 , wherein the at least one memory further includes additional instructions that, when executed by the at least one processor, further cause the system to:

send, to the first device, an instruction to delete the first speech processing data.

7. The computing system of claim 5 , wherein the at least one memory further includes additional instructions that, when executed by the at least one processor, further cause the system to:

determine a second command represented in the second speech processing data;

determine, using the updated spoken command history, that the second command is associated with a second input frequency falling below the input frequency threshold; and

send, to the first device, an instruction to delete only a portion of the second speech processing data associated with the second command.

8. The computing system of claim 5 , wherein the second speech processing data includes data for performing speech recognition.

9. The computing system of claim 5 , wherein the at least one memory further includes additional instructions that, when executed by the at least one processor, further cause the system to:

determine that the first command corresponds to a first intent,

wherein the second speech processing data corresponds to natural language processing data specific to the first intent.

10. The computing system of claim 5 , wherein the at least one memory further includes additional instructions that, when executed by the at least one processor, further cause the system to:

receive, from the first device, audio data corresponding to a second command;

determine, using the updated spoken command history, that the second command is associated with a second input frequency falling below the input frequency threshold;

perform speech processing on the audio data to determine output data; and

send the output data to the first device.

11. The computing system of claim 5 , wherein the at least one memory further includes additional instructions that, when executed by the at least one processor, further cause the system to:

receive, from the first device, natural language processing results;

perform text-to-speech processing on the natural language processing results to generate output audio data; and

send the output audio data to the first device.

12. The computing system of claim 5 , wherein the second speech processing data includes data for performing text-to-speech processing.

13. A method implemented by a computing system, comprising:

receiving, from a first device associated with a user profile, first data;

determining that the first data corresponds to a first command;

determining an updated spoken command history by updating, based at least in part on the first data corresponding to the first command, a spoken command history stored by the computing system and associated with the user profile;

determining, using the updated spoken command history, that the first command is associated with an input frequency satisfying an input frequency threshold;

determining that the first command is at least in part not processable using first speech processing data stored by the first device;

generating second speech processing data including second data specific to the first command; and

sending the second speech processing data to the first device.

14. The method of claim 13 , further comprising

sending, to the first device, an instruction to delete the first speech processing data.

15. The method of claim 13 , further comprising:

determining a second command represented in the second speech processing data;

determining, using the updated spoken command history, that the second command is associated with a second input frequency falling below the input frequency threshold; and

sending, to the first device, an instruction to delete only a portion of the second speech processing data associated with the second command.

16. The method of claim 13 , wherein the second speech processing data includes data for performing speech recognition.

17. The method of claim 13 , further comprising:

determining that the first command corresponds to a first intent,

wherein the second speech processing data corresponds to natural language processing data specific to the first intent.

18. The method of claim 13 , further comprising:

receiving, from the first device, audio data corresponding to a second command;

determining, using the updated spoken command history, that the second command is associated with a second input frequency falling below the input frequency threshold;

performing speech processing on the input audio data to determine output data; and

sending the output data to the first device.

19. The method of claim 13 , further comprising:

receiving, from the first device, natural language processing results;

performing text-to-speech processing on the natural language processing results to generate output audio data; and

sending the output audio data to the first device.

20. The method of claim 13 , wherein the second speech processing data includes data for performing text-to-speech processing.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2017
From: DEVRIES, DAVID WILLIAM; MITTAL, RAJESH KUMAR
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 043631/0377 →
Cited By (6)
US 12,217,740 US 12,400,663 US 12,417,768 US 12,579,969 US 12,700,405 US 12,738,263