IP Library › Granted Patent US 9,959,869
Granted Patent B2
US 9,959,869 · App. 15/694,996 · Granted May 1, 2018

Architecture for multi-domain natural language processing

Inventors: Lambert Mathias (Arlington, MA); Ying Shi (Seattle, WA); Imre Attila Kiss (Arlington, MA); Ryan Paul Thomas (Redmond, WA); Frederic Johan Georges Deramat (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G10L15/22G06F17/277G06F17/278G06F17/279G10L13/08G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,959,869
App. No.
15/694,996
Granted
May 1, 2018
Kind
B2
Abstract

Features are disclosed for processing a user utterance with respect to multiple subject matters or domains, and for selecting a likely result from a particular domain with which to respond to the utterance or otherwise take action. A user utterance may be transcribed by an automatic speech recognition (“ASR”) module, and the results may be provided to a multi-domain natural language understanding (“NLU”) engine. The multi-domain NLU engine may process the transcription(s) in multiple individual domains rather than in a single domain. In some cases, the transcription(s) may be processed in multiple individual domains in parallel or substantially simultaneously. In addition, hints may be generated based on previous user interactions and other data. The ASR module, multi-domain NLU engine, and other components of a spoken language processing system may use the hints to more efficiently process input or more accurately generate output.

Claims (51)

1. A system comprising:

computer-readable memory storing executable instructions; and

one or more processors in communication with the computer-readable memory, wherein the one or more processors are programmed by the executable instructions to:

receive text data representing at least a portion of an utterance;

receive hint data associated with the utterance;

determine first ranking data using the text data and the hint data, wherein the first ranking data represents a relative rank of a first natural language understanding (“NLU”) component of a plurality of NLU components, wherein the first NLU component is associated with a first domain;

determine second ranking data using the text data and the hint data, wherein the second ranking data represents a relative rank of a second NLU component of the plurality of NLU components, wherein the second NLU component is associated with a second domain;

select the first NLU component based at least partly on the first ranking data and the second ranking data; and

generate a response to the utterance using the first NLU component and the text data.

2. The system of claim 1 , wherein the one or more processors are further programmed by the executable instructions to generate NLU result data using the text data and the first NLU component, wherein the NLU result data represents a plurality of named entities, wherein a first named entity of the plurality of named entities corresponds to an intent associated with the first domain, and wherein a second named entity of the plurality of named entities corresponds to information regarding the intent.

3. The system of claim 1 , wherein one or more processors are further programmed by the executable instructions to:

identify a plurality of tokens in the text data;

associate first token of the plurality of tokens with a first named entity;

associate a second token of the plurality of tokens with a second named entity; and

identify an intent that corresponds to the first named entity and the second named entity, wherein the intent is associated with the first domain.

4. The system of claim 1 , wherein the first NLU component being associated with the first domain comprises the first NLU component being configured to generate NLU result data comprising intents associated with a subject matter of the first domain.

5. The system of claim 1 , wherein the hint data indicates a likely domain with which the text data is associated.

6. The system of claim 1 , wherein the one or more processors are further programmed by the executable instructions to generate the hint data based on at least one of a previous utterance or a previous response.

7. The system of claim 1 , wherein the one or more processors are further programmed by the executable instructions to adjust at least one of the first ranking data or the second ranking data based at least partly on the hint data.

8. The system of claim 1 , wherein the first domain is associated with intents related to at least one of: phone dialing, shopping, getting directions, playing music, or performing a search, and wherein the second domain is different than the first domain.

9. The system of claim 1 , wherein one or more processors are further programmed by the executable instructions to:

receive audio data representing the utterance; and

generate the text data using the audio data and an automatic speech recognition (“ASR”) component.

10. The system of claim 1 , wherein executable instructions to generate the response comprise executable instructions to:

generate response text data using a natural language generation component; and

generate response audio data using the response text data and a text-to-speech component.

11. The system of claim 1 , wherein one or more processors are further programmed by the executable instructions to send the response to a remote device configured to present the response.

12. A computer-implemented method comprising:

under control of one or more computing devices configured with specific computer-executable instructions,

receiving text data representing at least a portion of an utterance;

receiving hint data associated with the utterance;

determining first ranking data using the text data and the hint data, wherein the first ranking data represents a relative rank of a first natural language understanding (“NLU”) component of a plurality of NLU components, wherein the first NLU component is associated with a first domain;

determining second ranking data using the text data and the hint data, wherein the second ranking data represents a relative rank of a second NLU component of the plurality of NLU components, wherein the second NLU component is associated with a second domain;

selecting the first NLU component based at least partly on the first ranking data and the second ranking data; and

generating a response to the utterance using the first NLU component and the text data.

13. The computer-implemented method of claim 12 , wherein receiving the hint data comprises receiving data indicating a likely domain with which the text data is associated.

14. The computer-implemented method of claim 12 , further comprising generating the hint data based on at least one of a previous utterance or a previous response.

15. The computer-implemented method of claim 12 , further comprising adjusting at least one of the first ranking data or the second ranking data based at least partly on the hint data.

16. The computer-implemented method of claim 12 , further comprising generating NLU result data using the text data and the first NLU component, wherein the NLU result data represents a plurality of named entities, wherein a first named entity of the plurality of named entities corresponds to an intent associated with the first domain, and wherein a second named entity of the plurality of named entities corresponds to information regarding the intent.

17. The computer-implemented method of claim 12 , further comprising:

identifying a plurality of tokens in the text data;

associating a first token of the plurality of tokens with a first named entity;

associating a second token of the plurality of tokens with a second named entity; and

identifying an intent that corresponds to the first named entity and the second named entity, wherein the intent is associated with the first domain.

18. The computer-implemented method of claim 12 , further comprising:

receiving audio data representing the utterance; and

generating the text data using the audio data and an automatic speech recognition (“ASR”) component.

19. The computer-implemented method of claim 12 , wherein generating the response comprises:

generating response text data using a natural language generation component; and

generating response audio data using the response text data and a text-to-speech component.

20. The computer-implemented method of claim 12 , further comprising sending the response to a remote device configured to present the response.

Continuity (4)
Continuation 15256176 · Sep 2, 2016
Continuation 14754598 · Jun 29, 2015
Continuation 13720909 · Dec 19, 2012
Related Publication 20180012597A1 · Jan 11, 2018