IP Library Granted Patent US 10,861,446
Granted Patent B2
US 10,861,446 · App. 16/215,105 · Granted Dec 8, 2020

Generating input alternatives

Inventors: Ravi Chandra Reddy Yasa (Ashland, MA); Sai Rahul Reddy Pulikunta (North Andover, MA); Eliav Kahan (Jamaica Plain, MA); Gregory Newell (Waltham, MA)
Assignee: Amazon Technologies, Inc.
G10L15/1815G06F16/313G06F16/334G06N20/00G10L15/22G10L15/26G10L2015/223G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,861,446
App. No.
16/215,105
Granted
Dec 8, 2020
Kind
B2
Abstract

Exemplary embodiments relate to a system for recovering a conversation between a user and the system when the system is unable to properly respond to a user's input. The system may process the user input and determine an error condition exists. The system may query one or more storage systems to identify candidate text data based on their semantic similarity to the user input. The storage systems may store data related to past frequently entered inputs and/or user-generated inputs. Alternative text data is selected from the candidate text data, and presented to the user for confirmation.

Claims (98)

1. A method comprising:

receiving first input audio data corresponding to a first user utterance;

performing automatic speech recognition (ASR) processing on the first input audio data to determine input text data;

performing natural language understanding (NLU) processing on the input text data to determine first intent data associated with the input text data;

determining an error condition associated with the NLU processing;

querying, using at least a portion of the input text data, a first storage storing first indexed data to determine a plurality of candidate text representations of an utterance, wherein the first indexed data is based on past utterances spoken by a plurality of users;

identifying alternative text data from the plurality of candidate text representations based at least on a semantic similarity between the alternative text data and the first user utterance;

generating output audio data based on the alternative text data, wherein the output audio data requests a confirmation from to proceed with the alternative text data;

receiving second input audio data corresponding to a second user utterance;

in response to receiving the second input audio data, determining a second intent and slot data associated with the alternative text data; and

generating output data associated with the second intent.

2. The method of claim 1 , further comprising:

identifying frequent utterance data from a second storage, wherein the frequent utterance data represents the past utterances spoken by the plurality of users;

generating the first indexed data using the frequent utterance data, wherein a first entry in the first indexed data represents a keyword and one or more results among the frequent utterance data containing the keyword; and

storing the first indexed data in the first storage.

3. The method of claim 1 , further comprising:

determining a ranked list of candidate text representations by processing the plurality of candidate text representations using a trained machine learning model, the ranked list of candidate text representations having semantic similarity with the first user utterance;

determining a filtered list of candidate text representations by comparing the ranked list of candidate text representations to at least one of device type data, the first intent data, or domain data corresponding to the first input audio data;

selecting the alternative text data from the filtered list of candidate text representations;

determining stored utterance data based on the alternative text data; and

retrieving, from a database, the second intent associated with the stored utterance data.

4. The method of claim 1 , further comprising:

identifying utterance text data from a second storage, wherein the utterance text data represents text representations of utterances created by a system user;

generating second indexed data using the utterance text data, wherein a second entry in the second indexed data represents a keyword and one or more results among the utterance text data containing the keyword;

storing the second indexed data in the second storage; and

querying, using the at least a portion of the input text data, the second storage to identify the plurality of candidate text representations.

5. A method comprising:

receiving input data;

performing natural language processing on the input data;

determining an error condition associated with the natural language processing;

identifying alternative text data based on a semantic similarity between the alternative text data and the input data;

determining first application data associated with the alternative text data; and

generating output data based on the first application data.

6. The method of claim 5 , further comprising:

identifying frequent utterance data from a first storage, wherein the frequent utterance data represents past utterances spoken by a plurality of users;

generating first indexed data using the frequent utterance data, wherein a first entry in the first indexed data corresponds to a keyword and results data containing the keyword; and

storing the first indexed data in a second storage.

7. The method of claim 6 , wherein identifying the alternative text data comprises:

querying the second storage using at least a portion of the input data to determine a plurality of candidate text representations of an utterance; and

identifying the alternative text data based on the plurality of candidate text representations.

8. The method of claim 5 , further comprising:

determining a ranked list of candidate text representations by processing the input data using a trained model;

determining a second list of candidate text representations by processing the ranked list of candidate text representations based at least on device type data, second application data corresponding to the input data, or domain data corresponding to the input data;

selecting the alternative text data from the second list of candidate text representations; and

generating the output data using the alternative text data.

9. The method of claim 5 , further comprising:

identifying utterance text data from a first storage, wherein the utterance text data represents text representations of utterances created by a system user;

generating second indexed data using the utterance text data, wherein a second entry in the second indexed data represents a keyword and one or more results containing the keyword; and

storing the second indexed data in a second storage.

10. The method of claim 5 , wherein the output data corresponds to a confirmation request to generate an output using the alternative text data, and the method further comprises:

receiving second input data;

determining that the second input data indicates a confirmation;

in response to the second input data indicating a confirmation, determining the first application data associated with the alternative text data; and

generating second output data associated with the first application data.

11. The method of claim 5 , wherein the input data is audio data and the method further comprises:

performing ASR processing on the input data to generate text data;

performing the natural language processing on the text data to determine second application data and a confidence score, the second application data corresponding to the text data and the confidence score associated with the second application data; and

determining the error condition based at least in part on the text data, the second application data, or the confidence score.

12. The method of claim 5 , further comprising:

querying a storage using at least a portion of the input data to determine a plurality of candidate text representations of an utterance, the storage storing data representing a mapping between input data and system actionable data; and

identifying the alternative text data based on the plurality of candidate text representations.

13. A system, comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

receive input data,

perform natural language processing on the input data,

determine an error condition associated with the natural language processing,

identify alternative text data based on a semantic similarity between the alternative text data and the input data,

determine first application data associated with the alternative text data, and

generate output data based on the first application data.

14. The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

identify frequent utterance data from a first storage, wherein the frequent utterance data represents past utterances spoken by a plurality of users,

generate first indexed data using the frequent utterance data, wherein a first entry in the first indexed data corresponds to a keyword and results data containing the keyword, and

store the first indexed data in a second storage.

15. The system of claim 14 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

query the second storage using at least a portion of the input data to determine a plurality of candidate text representations of an utterance, and

identify the alternative text data based on the plurality of candidate text representations.

16. The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine a ranked list of candidate text representations by processing the input data using a trained model,

determine a second list of candidate text representations by processing the ranked list of candidate text representations based at least on device type data, second application data corresponding to the input data, or domain data corresponding to the input data,

select the alternative text data from the second list of candidate text representations, and

generate the output data using the alternative text data.

17. The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

identify utterance text data from a first storage, wherein the utterance text data represents text representations of utterances created by a system user,

generate second indexed data using the utterance text data, wherein a second entry in the second indexed data represents a keyword and one or more results containing the keyword, and

store the second indexed data in a second storage.

18. The system of claim 13 , wherein the output data corresponds to a confirmation request to generate an output using the alternative text data, and the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive second input data,

determine that the second input data indicates a confirmation,

in response to the second input data indicating a confirmation, determine the first application data associated with the alternative text data, and

generate second output data associated with the first application data.

19. The system of claim 13 , wherein the input data is audio data and the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

perform ASR processing on the input data to generate text data,

perform the natural language processing on the text data to determine second application data and a confidence score, the second application data corresponding to the input data and the confidence score associated with the second application data, and

determine the error condition based at least in part on the text data, the second application data, or the confidence score.

20. The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

query a storage using at least a portion of the input data to determine a plurality of candidate text representations, the storage storing data representing a mapping between input data and system actionable data, and

identify the alternative text data based on the plurality of candidate text representations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2019
From: YASA, RAVI CHANDRA REDDY; PULIKUNTA, SAI RAHUL REDDY; KAHAN, ELIAV; NEWELL, GREGORY
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 051202/0974 →
Continuity (1)
Related Publication 20200184959A1 · Jun 11, 2020
Cited By (1)
US 12,393,782