IP Library › Granted Patent US 11,176,936
Granted Patent B2
US 11,176,936 · App. 16/400,905 · Granted Nov 16, 2021

Architecture for multi-domain natural language processing

Inventors: Lambert Mathias (Arlington, MA); Ying Shi (Seattle, WA); Imre Attila Kiss (Arlington, MA); Ryan Paul Thomas (Redmond, WA); Frederic Johan Georges Deramat (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G10L15/22G06F40/284G06F40/295G06F40/35G06F40/40G06F40/56G10L13/08G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,176,936
App. No.
16/400,905
Filed
May 1, 2019
Granted
Nov 16, 2021
Kind
B2
Art Unit
2655
USPC
704/9
Abstract

Features are disclosed for processing a user utterance with respect to multiple subject matters or domains, and for selecting a likely result from a particular domain with which to respond to the utterance or otherwise take action. A user utterance may be transcribed by an automatic speech recognition (“ASR”) module, and the results may be provided to a multi-domain natural language understanding (“NLU”) engine. The multi-domain NLU engine may process the transcription(s) in multiple individual domains rather than in a single domain. In some cases, the transcription(s) may be processed in multiple individual domains in parallel or substantially simultaneously. In addition, hints may be generated based on previous user interactions and other data. The ASR module, multi-domain NLU engine, and other components of a spoken language processing system may use the hints to more efficiently process input or more accurately generate output.

Claims (60)

1. A system comprising:

computer-readable memory storing executable instructions; and

one or more processors in communication with the computer-readable memory,

 wherein the one or more processors are programmed by the executable instructions to:

receive interaction data representing one or more user interactions;

store history data regarding the one or more user interactions;

generate future hint data using the interaction data, wherein the future hint data represents at least a first natural language understanding (“NLU”) domain with which a future user utterance is likely to be associated, and wherein the first NLU domain is one of a plurality of NLU domains;

establish a natural language processing session subsequent to generating the future hint data;

receive natural language data representing a user utterance occurring during the natural language processing session;

access, in response to receiving the natural language data, the future hint data representing at least the first NLU domain;

determine, based at least partly on the future hint data, that the natural language data is associated with the first NLU domain;

generate first intent data based at least partly on the natural language data, wherein the first intent data represents a first intent associated with the first NLU domain;

generate second intent data based at least partly on the natural language data, wherein the second intent data represents a second intent associated with a second NLU domain;

rank the first intent data above the second intent data based at least partly on the future hint data; and

generate response data based at least partly on the first intent data.

2. The system of claim 1 , wherein the one or more processors are further programmed by the executable instructions to determine, based at least partly on the future hint data, that the natural language data is associated with the first intent, wherein the first intent is one of a plurality of intents associated with the first NLU domain.

3. The system of claim 1 , wherein the one or more user interactions includes a non-utterance interaction of a user.

4. The system of claim 1 , wherein the first NLU domain is associated with intents related to at least one of: phone dialing, shopping, getting directions, playing music, or performing a search.

5. The system of claim 1 , wherein the one or more processors are further programmed by the executable instructions to:

analyze the interaction data and the history data; and

determine, based at least partly on a result of analyzing the interaction data, that the future user utterance is likely to be associated with the first NLU domain.

6. The system of claim 1 , wherein the one or more processors are further programmed by the executable instructions to generate NLU result data using the natural language data and a first NLU component of a plurality of NLU components,

wherein the first NLU component is associated with the first NLU domain, and

wherein the NLU result data represents a named entity of a plurality of named entities corresponding to an executable action associated with the first intent.

7. The system of claim 1 , wherein the one or more processors are further programmed by the executable instructions to:

receive audio data representing the user utterance; and

generate the natural language data using the audio data and an automatic speech recognition (“ASR”) component.

8. The system of claim 1 , wherein the executable instructions to generate the response data comprise executable instructions to:

generate response text data using a natural language generation component; and

generate response audio data using the response text data and a text-to-speech component.

9. The system of claim 1 , wherein one or more processors are further programmed by the executable instructions to send the response data to a remote device configured to present a response using the response data.

10. A computer-implemented method comprising:

under control of one or more computing devices configured with specific computer-executable instructions,

receiving interaction data representing one or more user interactions;

storing history data regarding the one or more user interactions;

generating future hint data using the interaction data, wherein the future hint data represents at least a first natural language understanding (“NLU”) domain with which a future user utterance is likely to be associated, and wherein the first NLU domain is one of a plurality of NLU domains;

establishing a natural language processing session subsequent to generating the future hint data;

receiving natural language data representing a user utterance occurring during the natural language processing session;

accessing, in response to receiving the natural language data, the future hint data representing at least the first NLU domain;

determining, based at least partly on the future hint data, that the natural language data is associated with the first NLU domain;

generating first intent data based at least partly on the natural language data, wherein the first intent data represents a first intent associated with the first NLU domain;

generating second intent data based at least partly on the natural language data, wherein the second intent data represents a second intent associated with a second NLU domain;

ranking the first intent data above the second intent data based at least partly on the future hint data; and

generating response data based at least partly on the first intent data.

11. The computer-implemented method of claim 10 , further comprising determining, based at least partly on the future hint data, that the natural language data is associated with the first intent, wherein the first intent is one of a plurality of intents associated with the first NLU domain.

12. The computer-implemented method of claim 10 , wherein ranking the first intent data above the second intent data based at least partly on the future hint data comprises adjusting a rank of at least one of the first intent data or the second intent data.

13. The computer-implemented method of claim 10 , wherein receiving the interaction data comprises receiving data representing a non-utterance interaction of a user.

14. The computer-implemented method of claim 10 , wherein receiving the interaction data comprises receiving at least a portion of the interaction data in response to prompt.

15. The computer-implemented method of claim 10 , further comprising:

analyzing the interaction data and the history data; and

determining, based at least partly on a result of analyzing the interaction data, that the future user utterance is likely to be associated with the first NLU domain.

16. The computer-implemented method of claim 10 , further comprising generating NLU result data using the natural language data and a first NLU component of a plurality of NLU components,

wherein the first NLU component is associated with the first NLU domain, and

wherein the NLU result data represents a named entity of a plurality of named entities corresponding to an executable action associated with the first intent.

17. The computer-implemented method of claim 10 , further comprising:

receiving audio data representing the user utterance; and

generating the natural language data using the audio data and an automatic speech recognition (“ASR”) component.

18. The computer-implemented method of claim 10 , wherein the generating the response data comprises:

generating response text data using a natural language generation component; and

generating response audio data using the response text data and a text-to-speech component.

Continuity (6)
Continuation 15966400 · Apr 30, 2018
Continuation 15694996 · Sep 4, 2017
Continuation 15256176 · Sep 2, 2016
Continuation 14754598 · Jun 29, 2015
Continuation 13720909 · Dec 19, 2012
Related Publication 20190325873A1 · Oct 24, 2019
Cited By (2)
US 12,198,696 US 12,670,899