Speech recognition using dialog history
Described herein is a system for rescoring automatic speech recognition hypotheses for conversational devices that have multi-turn dialogs with a user. The system leverages dialog context by incorporating data related to past user utterances and data related to the system generated response corresponding to the past user utterance. Incorporation of this data improves recognition of a particular user utterance within the dialog.
1 . A computer-implemented method, comprising:
receiving input audio data corresponding to a first utterance;
performing speech recognition using the input audio data to determine first data indicating at a least a first speech recognition hypothesis representing the first utterance;
encoding the first data to generate first encoded data;
receiving second encoded data corresponding to a past utterance;
processing the first encoded data and the second encoded data using a machine learning model to determine model output data indicating a second speech recognition hypothesis as representing the first utterance, the second speech recognition hypothesis being different from the first speech recognition hypothesis;
processing the second speech recognition hypothesis to generate second data indicating an action to be performed in response to the first utterance; and
generating, using the second data, output data responsive to the first utterance.
2 . The computer-implemented method of claim 1 , further comprising:
receiving third data corresponding to the past utterance; and
encoding the third data to generate the second encoded data.
3 . The computer-implemented method of claim 2 , further comprising:
determining fourth data representing content of a user input corresponding to the past utterance; and
including the fourth data in the third data.
4 . The computer-implemented method of claim 2 , further comprising:
determining fourth data representing content of a system response to the past utterance; and
including the fourth data in the third data.
5 . The computer-implemented method of claim 1 , further comprising:
determining third encoded data representing parts-of-speech of at least one of the first utterance or the past utterance,
wherein the machine learning model further processes the third encoded data to determine the model output data.
6 . The computer-implemented method of claim 1 , further comprising:
determining third encoded data representing a device corresponding to the first utterance,
wherein the machine learning model further processes the third encoded data to determine the model output data.
7 . The computer-implemented method of claim 1 , further comprising:
determining weight data based at least in part on the second encoded data,
wherein the machine learning model uses the weight data to determine the model output data.
8 . The computer-implemented method of claim 1 , further comprising:
determining third encoded data representing a topic of at least one of the first utterance or the past utterance,
wherein the machine learning model further processes the third encoded data to determine the model output data.
9 . The computer-implemented method of claim 1 , wherein the first data represents a plurality of speech recognition hypotheses.
10 . A system comprising:
at least one processor; and
at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
receive input audio data corresponding to a first utterance;
perform speech recognition using the input audio data to determine first data indicating at a least a first speech recognition hypothesis representing the first utterance;
encode the first data to generate first encoded data;
receive second encoded data corresponding to a past utterance;
process the first encoded data and the second encoded data using a machine learning model to determine model output data indicating a second speech recognition hypothesis representing the first utterance, the second speech recognition hypothesis being different from the first speech recognition hypothesis;
process the second speech recognition hypothesis to generate second data indicating an action to be performed in response to the first utterance; and
generate, using the second data, output data responsive to the first utterance.
11 . The system of claim 10 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
receive third data corresponding to the past utterance; and
encode the third data to generate the second encoded data.
12 . The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine fourth data representing content of a user input corresponding to the past utterance; and
include the fourth data in the third data.
13 . The system of claim 11 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine fourth data representing content of a system response to the past utterance; and
include the fourth data in the third data.
14 . The system of claim 10 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine third encoded data representing parts-of-speech of at least one of the first utterance or the past utterance,
wherein the machine learning model further processes the third encoded data to determine the model output data.
15 . The system of claim 10 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine third encoded data representing a device corresponding to the first utterance,
wherein the machine learning model further processes the third encoded data to determine the model output data.
16 . The system of claim 10 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine weight data based at least in part on the second encoded data,
wherein the machine learning model uses the weight data to determine the model output data.
17 . The system of claim 10 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:
determine third encoded data representing a topic of at least one of the first utterance or the past utterance,
wherein the machine learning model further processes the third encoded data to determine the model output data.
18 . The system of claim 10 , wherein the first data represents a plurality of speech recognition hypotheses.