IP Library › Granted Patent US 11,437,026
Granted Patent B1
US 11,437,026 · App. 16/672,834 · Granted Sep 6, 2022

Personalized alternate utterance generation

Inventors: Alireza Roshan Ghias (Seattle, WA); Chenlei Guo (Redmond, WA); Pragaash Ponnusamy (Seattle, WA); Clint Solomon Mathialagan (Kirkland, WA)
Assignee: Amazon Technologies, Inc.
G10L15/197G10L15/063G10L15/1815G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,437,026
App. No.
16/672,834
Granted
Sep 6, 2022
Kind
B1
Abstract

A system is provided for handling errors during automatic speech recognition by leveraging past inputs spoken by the user. The system may process a user input to determine an ASR hypothesis. The system may then determine an alternate representation of the user input based on the inputs provided by the user in the past, and whether the ASR hypothesis sufficiently matches one of the past inputs.

Claims (110)

1. A computer-implemented method comprising:

receiving audio data representing an input utterance, the audio data corresponding to a user profile;

performing ASR processing on the audio data to determine:

first ASR data representing a first hypothesis potentially corresponding to the input utterance, the first ASR data including first text data and a first score, and

second ASR data representing a second hypothesis potentially corresponding to the input utterance, the second ASR data including second text data and a second score;

determining data representing an uncertainty exists with respect to ASR processing, the data indicating that the first score and the second score are below a threshold confidence level;

in response to determining the data representing the uncertainty exists, receiving historical utterance data associated with the user profile, the historical utterance data comprising third ASR data representing a past utterance and frequency data indicating a number of times the past utterance was spoken within a time period;

processing, using a trained model, the first ASR data, the second ASR data and the historical utterance data to determine fourth ASR data potentially corresponding to the input utterance; and

generating output data using the fourth ASR data.

2. The computer-implemented method of claim 1 , further comprising:

processing the first ASR data to determine a first feature vector;

processing the second ASR data to determine a second feature vector;

determining an average feature vector using the first feature vector and the second feature vector;

processing the third ASR data to determine a third feature vector;

processing, using an attention model, the average feature vector and the third feature vector to determine a fourth feature vector; and

determining a concatenated vector using the fourth feature vector and the third feature vector;

wherein processing using the trained model comprises processing, using the trained model, the concatenated vector to determine the fourth ASR data.

3. The computer-implemented method of claim 1 , further comprising prior to receiving the audio data:

receiving first utterance pair data associated with a second user profile and a first historical utterance, the first utterance pair data including fifth ASR data that resulted in a speech processing error and sixth ASR data that resulted in successful speech processing;

receiving second utterance pair data associated with a third user profile and a second historical utterance, the second utterance pair data including seventh ASR data that resulted in a speech processing error and eighth ASR data that resulted in successful speech processing;

storing the first utterance pair data;

storing the second utterance pair data; and

generating the trained model using first utterance pair data and the second utterance pair data.

4. The computer-implemented method of claim 1 , further comprising:

determining a first time corresponding to receipt of the input utterance;

determining a second time corresponding to receipt of the past utterance; and

determining that the first time and the second time are within the time period,

wherein processing using the trained model comprises processing the first ASR data, the second ASR data, the historical utterance data and the second time to determine the fourth ASR data.

5. A computer-implemented method comprising:

receiving audio data corresponding to an input utterance associated with a user profile;

performing speech recognition processing on the audio data to determine first automatic speech recognition (ASR) data potentially corresponding to the input utterance, the first ASR data associated with a first confidence level;

performing speech recognition processing on the audio data to determine second ASR data potentially corresponding to the input utterance, the second ASR data associated with a second confidence level;

determining an uncertainty exists based on the first confidence level and the second confidence level satisfying a condition;

in response to determining the uncertainty exists, receiving historical utterance data associated with the user profile, the historical utterance data corresponding to a first past utterance, wherein the historical utterance data includes first data indicating a number of times the first past utterance was spoken within a time period;

processing, using a machine learning model, the first ASR data, the second ASR data, and the historical utterance data to determine third ASR data potentially corresponding to the input utterance; and

generating output data using the third ASR data.

6. The computer-implemented method of claim 5 , further comprising:

determining a first time indicating when the input utterance was received;

determining a second time indicating when the first past utterance was received;

determining that the first time and the second time are within the time period; and

wherein receiving the historical utterance data comprises receiving historical ASR data representing the first past utterance, and

wherein processing the first ASR data, the second ASR data and the historical ASR data comprises processing, using the machine learning model, the first ASR data, the second ASR data, the historical ASR data and the first data.

7. The computer-implemented method of claim 5 , further comprising:

processing the first ASR data to determine a first feature vector;

processing the second ASR data to determine a second feature vector;

processing the historical utterance data to determine a third feature vector;

determining an average feature vector using the first feature vector and the second feature vector;

processing, using an attention model, the average feature vector and the third feature vector to determine a fourth feature vector; and

determining a concatenated vector using the fourth feature vector and the third feature vector,

wherein processing the first ASR data, the second ASR data and the historical ASR data comprises processing, using the machine learning model, the concatenated vector to determine the third ASR data.

8. The computer-implemented method of claim 5 , further comprising:

processing profile data associated with the user profile to determine a first past utterance spoken with the time period;

determining, using the profile data, that the first past utterance resulted in successful speech processing; and

determining to receive the historical utterance data corresponding to the first past utterance.

9. The computer-implemented method of claim 5 , further comprising:

performing, using the first ASR data, natural language understanding (NLU) to determine a first NLU hypothesis;

performing, using the third ASR data, NLU to determine a second NLU hypothesis;

determining that the first NLU hypothesis results in an error; and

determining to generate the output data using the second NLU hypothesis.

10. The computer-implemented method of claim 5 , wherein determining that the uncertainty exists comprises:

determining a value representing a difference between the first confidence level and the second confidence level;

determining that the value meets a threshold value; and

determining that the uncertainty exists based on the value meeting the threshold value.

11. The computer-implemented method of claim 5 , further comprising prior to receiving the audio data:

receiving first utterance pair data associated with a second user profile, the first utterance pair data including first historical ASR data corresponding to a first historical utterance and second historical ASR data corresponding to the first historical utterance, wherein the first historical ASR data resulted in a speech processing error and the second historical ASR data resulted in successful speech processing;

receiving second utterance pair data associated with a third user profile, the second utterance pair data including third historical ASR data corresponding to a second historical utterance and fourth historical ASR data corresponding to the second historical utterance, wherein the third historical ASR data resulted in a speech processing error and the fourth historical ASR data resulted in successful speech processing;

storing training data comprising the first utterance pair data and the second utterance pair data; and

generating the machine learning model using the training data.

12. A system comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the system to:

receive audio data corresponding to an input utterance associated with a user profile;

perform speech recognition processing on the audio data to determine first automatic speech recognition (ASR) data potentially corresponding to the input utterance, the first ASR data associated with a first confidence level;

perform speech recognition processing on the audio data to determine second ASR data potentially corresponding to the input utterance, the second ASR data associated with a second confidence level;

determine an uncertainty exists based on the first confidence level and the second confidence level satisfying a condition;

in response to determining the uncertainty exists, receive historical utterance data associated with the user profile, the historical utterance data corresponding to a first past utterance, wherein the historical utterance data includes first data indicating a number of times the first past utterance was spoken within a time period;

process, using a machine learning model, the first ASR data, the second ASR data and the historical utterance data to determine third ASR data potentially corresponding to the input utterance; and

generate output data using the third ASR data.

13. The system of claim 12 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

determine a first time indicating when the input utterance was received;

determine a second time indicating when the first past utterance was received;

determine that the first time and the second time are within the time period; and

wherein the instructions to receive the historical utterance data further cause the system to receive historical ASR data representing the first past utterance, and

wherein the instructions to process the first ASR data, the second ASR data and the historical ASR data further cause the system to process, using the machine learning model, the first ASR data, the second ASR data, the historical ASR data and the first data.

14. The system of claim 12 , wherein the instructions that, when executed by the at least one processor, further causes the system to:

process the first ASR data to determine a first feature vector;

process the second ASR data to determine a second feature vector;

process the historical utterance data to determine a third feature vector;

determine an average feature vector using the first feature vector and the second feature vector;

process, using an attention model, the average feature vector and the third feature vector to determine a fourth feature vector; and

determine a concatenated vector using the fourth feature vector and the third feature vector,

wherein the instructions to process the first ASR data, the second ASR data and the historical ASR data further causes the system to process, using the machine learning model, the concatenated vector to determine the third ASR data.

15. The system of claim 12 , wherein the instructions that, when executed by the at least one processor, further causes the system to:

process profile data associated with the user profile to determine a first past utterance spoken with the time period;

determine, using the profile data, that the first past utterance resulted in successful speech processing; and

determine to receive the historical utterance data corresponding to the first past utterance.

16. The system of claim 12 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

perform, using the first ASR data, natural language understanding (NLU) to determine a first NLU hypothesis;

perform, using the third ASR data, NLU to determine a second NLU hypothesis;

determine that the first NLU hypothesis results in an error; and

determine to generate the output data using the second NLU hypothesis.

17. The system of claim 12 , wherein the instructions that, when executed by the at least one processor, cause the system to determine that an uncertainty exists further cause the system to:

determine a value representing a difference between the first confidence level and the second confidence level;

determine that the value meets a threshold value; and

determine that the uncertainty exists based on the value meeting the threshold value.

18. The system of claim 12 , wherein the instructions that, when executed by the at least one processor, further cause the system to, prior to receiving the audio data:

receive first utterance pair data associated with a second user profile, the first utterance pair data including first historical ASR data corresponding to a first historical utterance and second historical ASR data corresponding to the first historical utterance, wherein the first historical ASR data resulted in a speech processing error and the second historical ASR data resulted in successful speech processing;

receive second utterance pair data associated with a third user profile, the second utterance pair data including third historical ASR data corresponding to a second historical utterance and fourth historical ASR data corresponding to the second historical utterance, wherein the third historical ASR data resulted in a speech processing error and the fourth historical ASR data resulted in successful speech processing;

store training data comprising the first utterance pair data and the second utterance pair data; and

generate the machine learning model using the training data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 4, 2019
From: ROSHAN GHIAS, ALIREZA; GUO, CHENLEI; PONNUSAMY, PRAGAASH; MATHIALAGAN, CLINT SOLOMON
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 050905/0016 →
Cited By (3)
US 12,327,090 US 12,347,425 US 12,711,948