Speech recognition platforms
A speech recognition platform configured to receive an audio signal that includes speech from a user and perform automatic speech recognition (ASR) on the audio signal to identify ASR results. The platform may identify: (i) a domain of a voice command within the speech based on the ASR results and based on context information associated with the speech or the user, and (ii) an intent of the voice command. In response to identifying the intent, the platform may perform a corresponding action, such as streaming audio to the device, setting a reminder for the user, purchasing an item on behalf of the user, making a reservation for the user or launching an application for the user. The speech recognition platform, in combination with the device, may therefore facilitate efficient interactions between the user and a voice-controlled device.
1. A system comprising:
a coordination component configured to receive an audio signal generated based at least in part on sound from an environment, the sound captured by a device that includes a microphone unit and a speaker and the audio signal including speech of a user in the environment;
a speech recognition component configured to perform speech recognition on the audio signal to generate speech-recognition results;
a natural language understanding (NLU) component configured to: identify, for different domains, multiple intents associated with the speech of the user based at least in part on the speech-recognition results and context information associated with the speech of the user, wherein identify the multiple intents includes to provide the speech-recognition results to the different domains, and to identify the multiple intents based at least in part on named entities identified from the different domains and one or more slots filled based at least in part on the context information, and rank the multiple intents for the different domains relative to one another;
a dialog component configured to select one of the domains based at least in part on the ranked intents and select one of the ranked intents that is associated with the selected domain; and
a response component configured to receive an indication of the selected domain and an indication of the selected intent and cause performance of an action corresponding to the selected intent, the action including providing an audio signal for output on the speaker of the device.
2. A system as recited in claim 1 , wherein at least one of coordination component or the speech recognition component is further configured to retrieve the context information associated with the speech and provide the context information to the NLU component for identifying the multiple intents.
3. A system as recited in claim 1 , wherein individual ones of the domains specify a set of related activities that the user may request the device to perform, the device configured to perform activities of the set of related activities based at least in part on receiving a command identified from speech of the user.
4. A system as recited in claim 1 , wherein individual ones of the intents are associated with one or more fields that, when associated with respective values, specify an action requested by the user, and the NLU component is configured to associated respective values to the one or more fields based at least in part on the context information associated with the speech of the user.
5. A system as recited in claim 1 , wherein:
the dialog component is further configured to engage in a dialog with the user by causing the speaker to output a first set of one or more questions, and wherein the dialog component is further configured to select the domain based at least in part a response from the user to the first set of one or more questions; and
the response component is further configured to engage in another dialog with the user by causing the speaker to output a second set of one or more questions, and wherein the dialog component is further configured to select the intent based at least in part a response from the user to the set of one or more questions.
6. A system comprising:
one or more processors; and
computer-readable media storing computer-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform acts comprising:
receiving an audio signal that includes speech from a user, the speech of the user received at a device that includes a speaker;
retrieving context information associated with the speech, the context information based at least in part on a previous interaction between the user and the device;
performing speech recognition on the audio signal to generate speech-recognition results;
identifying multiple intents, for different domains, associated with the speech based at least in part on the speech-recognition results and the context information, wherein identify the multiple intents includes to provide the speech-recognition results to the different domains, and to identify the multiple intents based at least in part on named entities identified from the different domains and one or more slots filled based at least in part on the context information;
selecting one of the domains based at least in part on the multiple intents for the different domains;
selecting, from the multiple intents, an intent that is associated with the selected domain; and
providing an audio signal for output on the speaker of the device based at least in part on the selected intent and the selected domain.
7. A system as recited in claim 6 , wherein individual ones of the domain specify set of related activities that the user may request the device to perform, the device configured to perform activities of the set of related activities based at least in part on receiving a command identified from speech of the user.
8. A system as recited in claim 7 , wherein individual ones of the domain is associated with multiple intents, and wherein individual ones of the intents of a respective domain is associated with a particular activity of the set of activities.
9. A system as recited in claim 6 , wherein the identifying comprises, for individual ones of the domains:
parsing the speech-recognition results to recognize one or more named entities associated with the domain;
associating a value with a field associated with the domain using the context information; and
identifying a respective intent from the domain based at least in part on the one or more recognized named entities and the value of the field.
10. A system as recited in claim 6 , the acts further comprising ranking the multiple intents prior to selecting the domain.
11. A system as recited in claim 6 , wherein selecting the domain comprises causing the speaker to output a question and receiving a response from the user to the question.
12. A system as recited in claim 6 , wherein selecting the intent comprises causing the speaker to output a question and receiving a response from the user to the question.
13. A system as recited in claim 6 , the acts further comprising at least one of: streaming audio to the device, setting a reminder for the user, ordering or purchasing an item on behalf of the user, making a reservation for the user, or launching an application for a user.
14. A method comprising:
under control of one or more computing systems configured with executable instructions,
receiving an audio signal generated by a device, the audio signal including speech of a user;
performing speech recognition on the speech to generate speech-recognition results;
identifying, based at least in part on the speech-recognition results, first potential intents of the speech, the first potential intents being associated with a first domain;
identifying, based at least in part on the speech-recognition results, second potential intents of the speech, the second potential intents being associated with a second domain, wherein identifying the first potential intents of the speech and the second potential intents of the speech include to provide the speech-recognition results to the first domain and to the second domain, and to identify the first potential intents of the speech and the second potential intents of the speech based at least in part on named entities identified from the first domain and the second domain and one or more slots filled by the first domain and the second domain;
selecting the first potential intents or the second potential intents; and
providing an audio signal for output on the device based at least in part on the selecting.
15. A method as recited in claim 14 , wherein the one or more computing systems form a portion of a network-accessible computing platform that is remote from an environment in which the device resides.
16. A method as recited in claim 14 , wherein the first or second potential intents are selected based at least in part on a dialog with the user that occurs at least partly subsequent to receiving the audio signal.
17. A method as recited in claim 14 , further comprising obtaining context information associated with the speech or with the user at least partly in response to receiving the audio signal, and wherein the first and second potential intents are identified based at least in part on the context information.
18. A method as recited in claim 17 , wherein the context information is based at least in part on previous speech of the user.
19. A method as recited in claim 17 , wherein the context information is based at least in part on a location of the user, preferences of the user, or information from an application called by the speech of the user.
20. A method as recited in claim 14 , further comprising selecting the first domain or the second domain prior to selecting the first potential intent or the second potential intent.
21. A method as recited in claim 20 , wherein the domain is selected based at least in part on a dialog with the user that occurs at least partly subsequent to receiving the audio signal.
22. A method as recited in claim 20 , wherein the first domain specifies a first set of related activities that the user may request the device to perform, the device configured to perform activities of the set of related activities in response to receiving a command identified from speech of the user, and the second domain specifies a second set of related activities that the user may request the device to perform, the device configured to perform activities of the set of related activities in response to receiving a command identified from speech of the user.
23. A method as recited in claim 20 , wherein the first potential intent comprises an activity of the first set of related activities and the second potential intent comprises an activity of the second set of related activities.