IP Library › Granted Patent US 11,626,106
Granted Patent B1
US 11,626,106 · App. 16/800,622 · Granted Apr 11, 2023

Error attribution in natural language processing systems

Inventors: Qing Ping (Santa Clara, CA); Govindarajan Sundaram Thattai (Fremont, CA); Joel Joseph Chengottusseriyil (San Jose, CA); Feiyang Niu (Hayward, CA)
Assignee: Amazon Technologies, Inc.
G10L15/1815G10L13/00G10L15/02G10L15/197G10L15/22G10L15/30G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,626,106
App. No.
16/800,622
Granted
Apr 11, 2023
Kind
B1
Abstract

A system is provided for determining which component of a speech processing system is the cause of an undesired response to a user input. The system processes ASR data and NLU data to determine the component likely to cause the undesired response. Based on which component is the cause of the undesired response, the system performs an appropriate conversation recovery technique to confirm the speech processing results with the user.

Claims (130)

1. A computer-implemented method comprising:

receiving, from a device, audio data corresponding to an utterance;

processing the audio data using at least one automatic speech recognition (ASR) component to determine ASR data corresponding to the utterance;

processing the ASR data using at least one natural language understanding (NLU) component to determine NLU data corresponding to the utterance;

processing the ASR data and the NLU data to determine vector data representing features determined by at least one component from a plurality of components, the features corresponding to at least one of: a first ASR hypothesis, a first ASR score, a first NLU hypothesis, a first NLU score, a first intent, and a first slot value;

determining, using a trained classifier model and the vector data, an error corresponding to the utterance, the trained classifier model configured to process the features to determine the error, wherein the error is of an error type from a plurality of error types corresponding to at least one component from the plurality of components;

determining that the error type corresponds to intent data determined by the at least one NLU component, the intent data representing the first intent;

processing the first intent to determine a second intent similar to the first intent;

determining, using first data representing the utterance, second data representing an alternative utterance corresponding to the second intent;

determining third data including the second data and a request to confirm use of the alternative utterance for further processing;

processing, using text-to-speech (TTS) processing, the third data to determine output audio data; and

sending the output audio data to the device.

2. The computer-implemented method of claim 1 , further comprising:

receiving, from the device, second audio data corresponding to a second utterance;

processing the audio data using the at least one ASR component to determine second ASR data corresponding to the second utterance;

processing the second ASR data using the at least one NLU component to determine second NLU data corresponding to the second utterance;

processing the second ASR data and the second NLU data to determine second vector data representing features determined during natural language processing of the second audio data;

processing, using the trained classifier model, the second vector data to determine a second error type corresponding to the second utterance;

determining that the second error type corresponds to data determined by the at least one ASR component;

generating output data representing a request to repeat the second utterance;

processing, using TTS processing, the output data to determine second output audio data; and

sending the second output audio data to the device.

3. The computer-implemented method of claim 1 , further comprising:

processing the ASR data to determine an ASR slot score corresponding to a first portion of the ASR hypothesis representing a slot;

processing the ASR data to determine an ASR entity score corresponding to a second portion of the ASR hypothesis representing an intent;

processing the NLU data to determine an intent score corresponding to a NLU hypothesis representing the utterance;

processing the NLU data to determine an NLU slot score corresponding to the NLU hypothesis;

processing the NLU data to determine an NLU entity score corresponding to the NLU hypothesis; and

determining the vector data using at least the ASR slot score, the ASR entity score, the intent score, the NLU slot score, and the NLU entity score.

4. The computer-implemented method of claim 1 , further comprising:

determining dialog session data including fourth data representing a previous utterance and fifth data representing a previous system-generated response to the previous utterance;

processing the dialog session data to determine encoded dialog session data;

determining a dialog outcome corresponding to the previous utterance;

processing the first data to determine encoded utterance data;

processing the audio data to determine audio features; and

determining the vector data using at least the encoded dialog session data, the dialog outcome, encoded utterance data, and the audio features.

5. A computer-implemented method comprising:

receiving audio data corresponding to an utterance;

processing the audio data using a natural language processing (NLP) system comprising a plurality of components to determine NLP data;

determining, using the NLP data, feature data representing processing, corresponding to the utterance, by one or more of the plurality of components;

determining, using the feature data, a first error type corresponding to processing of the utterance, the first error type determined from a plurality of error types corresponding to the NLP system;

determining the first error type indicates a first component of the plurality of components; and

determining, based on the first error type, that the first component corresponds to an undesired response.

6. The computer-implemented method of claim 5 , further comprising:

determining that automatic speech recognition (ASR) data determined by the first component representing the utterance caused the undesired response;

determining that the first component is configured to perform automatic speech recognition;

generating output data representing a request to repeat the utterance;

processing, using text-to-speech (TTS) processing, the output data to determine output audio data; and

sending the output audio data to a device that received the audio data.

7. The computer-implemented method of claim 5 , further comprising:

determining that intent data determined by the first component corresponding to the utterance caused the undesired response;

determining that the first component is configured to perform intent classification;

generating output data representing an alternative utterance corresponding to the utterance;

processing, using TTS processing, the output data to determine output audio data; and

sending the output audio data to a device that received the audio data.

8. The computer-implemented method of claim 5 , further comprising:

determining that entity data determined by the first component corresponding to the utterance causes the undesired response;

determining that the first component is configured to perform entity recognition;

generating output data representing a request to confirm an entity name;

processing, using TTS processing, the output data to determine output audio data; and

sending the output audio data to a device that received the audio data.

9. The computer-implemented method of claim 5 , further comprising:

processing the NLP data to determine an ASR confidence score corresponding to an ASR hypothesis corresponding to the utterance;

processing the NLP data to determine an ASR slot score corresponding to a first portion of the ASR hypothesis representing a slot;

processing the NLP data to determine an ASR entity score corresponding to a second portion of the ASR hypothesis representing an intent; and

determining the feature data using at least the ASR confidence score, the ASR slot score, and the ASR entity score.

10. The computer-implemented method of claim 5 , further comprising:

processing the NLP data to determine an intent score corresponding to a NLU hypothesis corresponding to the utterance;

processing the NLP data to determine an NLU slot score corresponding to the NLU hypothesis;

processing the NLP data to determine an NLU entity score corresponding to the NLU hypothesis; and

determining the feature data using at least the intent score, the NLU slot score, and the NLU entity score.

11. The computer-implemented method of claim 5 , further comprising:

determining dialog session data representing a previous utterance and a previous system-generated response to the previous utterance;

processing the dialog session data to determine encoded dialog session data;

determining a dialog outcome corresponding to the previous utterance;

processing ASR data representing the utterance to determine encoded utterance data; and

determining the feature data using at least the encoded dialog session data, the dialog outcome, and the encoded utterance data.

12. The computer-implemented method of claim 5 , further comprising:

processing the audio data to determine audio features;

determining an audio utterance ratio representing a ratio of a first duration corresponding to the audio data and a second duration corresponding to a portion of the audio data representing the utterance;

determining a wakeword represented in the audio data; and

determining the feature data using at least the audio features, the audio utterance ratio, and the wakeword.

13. A system comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the system to:

receive audio data corresponding to an utterance;

process the audio data using a natural language processing (NLP) system comprising a plurality of components to determine NLP data;

determine, using the NLP data, feature data representing processing, corresponding to the utterance, by one or more of the plurality of components;

determine, using the feature data, a first error type corresponding to processing of the utterance, the first error type determined from a plurality of error types corresponding to the NLP system;

determine the first error type indicates a first component of the plurality of components; and

determine, based on the first error type, that the first component corresponds to an undesired response.

14. The system of claim 13 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

determine that automatic speech recognition (ASR) data determined by the first component representing the utterance caused the undesired response;

determine that the first component is configured to perform automatic speech recognition;

generate output data representing a request to repeat the utterance;

process, using text-to-speech (TTS) processing, the output data to determine output audio data; and

send the output audio data to a device that received the audio data.

15. The system of claim 13 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

determine that intent data determined by the first component corresponding to the utterance caused the undesired response;

determine that the first component is configured to perform intent classification;

generate output data representing an alternative utterance corresponding to the utterance;

process, using TTS processing, the output data to determine output audio data; and

send the output audio data to a device that received the audio data.

16. The system of claim 13 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

determine that entity data determined by the first component corresponding to the utterance causes the undesired response;

determine that the first component is configured to perform entity recognition;

generate output data representing a request to confirm an entity name;

process, using TTS processing, the output data to determine output audio data; and

send the output audio data to a device that received the audio data.

17. The system of claim 13 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

process the NLP data to determine an ASR confidence score corresponding to an ASR hypothesis corresponding to the utterance;

process the NLP data to determine an ASR slot score corresponding to a first portion of the ASR hypothesis representing a slot;

process the NLP data to determine an ASR entity score corresponding to a second portion of the ASR hypothesis representing an intent; and

determine the feature data using at least the ASR confidence score, the ASR slot score, and the ASR entity score.

18. The system of claim 13 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

process the NLP data to determine an intent score corresponding to a NLU hypothesis corresponding to the utterance;

process the NLP data to determine an NLU slot score corresponding to the NLU hypothesis;

process the NLP data to determine an NLU entity score corresponding to the NLU hypothesis; and

determine the feature data using at least the intent score, the NLU slot score, and the NLU entity score.

19. The system of claim 13 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

determining dialog session data representing a previous utterance and a previous system-generated response to the previous utterance;

processing the dialog session data to determine encoded dialog session data;

determining a dialog outcome corresponding to the previous utterance;

processing ASR data representing the utterance to determine encoded utterance data; and

determining the feature data using at least the encoded dialog session data, the dialog outcome, and the encoded utterance data.

20. The system of claim 13 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

process the audio data to determine audio features;

determine an audio utterance ratio representing a ratio of a first duration corresponding to the audio data and a second duration corresponding to a portion of the audio data representing the utterance;

determine a wakeword represented in the audio data; and

determine the feature data using at least the audio features, the audio utterance ratio, and the wakeword.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2020
From: PING, QING; THATTAI, GOVINDARAJAN SUNDARAM; CHENGOTTUSSERIYIL, JOEL JOSEPH; NIU, FEIYANG
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 051923/0046 →
Cited By (3)
US 12,443,388 US 12,541,648 US 12,567,402