IP Library › Granted Patent US 10,515,625
Granted Patent B1
US 10,515,625 · App. 15/828,174 · Granted Dec 24, 2019

Multi-modal natural language processing

Inventors: Angeliki Metallinou (Mountain View, CA); Rahul Goel (Mountain View, CA); Vishal Ishwar (Tempe, AZ)
Assignee: Amazon Technologies, Inc.
G10L15/16G06F3/167G10L15/02G10L15/144G10L15/197G10L15/265G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,515,625
App. No.
15/828,174
Granted
Dec 24, 2019
Kind
B1
Abstract

Multi-modal natural language processing systems are provided. Some systems are context-aware systems that use multi-modal data to improve the accuracy of natural language understanding as it is applied to spoken language input. Machine learning architectures are provided that jointly model spoken language input (“utterances”) and information displayed on a visual display (“on-screen information”). Such machine learning architectures can improve upon, and solve problems inherent in, existing spoken language understanding systems that operate in multi-modal contexts.

Claims (66)

1. A computer-implemented method comprising:

as performed by a computing system comprising one or more computer processors configured to execute specific instructions,

receiving, from a computing device, audio data representing an utterance;

generating text data using the audio data;

receiving contextual data, wherein the contextual data represents content displayed by the computing device when the utterance occurred, and wherein the content is associated with domain data representing a first domain of a plurality of domains;

generating input data for a natural language understanding (“NLU”) model, wherein the input data comprises:

first vector data representing at least a portion of the text data; and

second vector data representing a context domain vector comprising a first element and a second element, wherein the first element corresponds to the first domain and the second element corresponds to a second domain of the plurality of domains, and wherein the second vector data indicates that content associated with the first domain was displayed when the utterance was made;

generating data representing an internal state of the model using the first vector data, the second vector data, and first parameter data, wherein the first parameter data represents a first portion of parameters of the NLU model;

generating intent score data using the data representing the internal state of the model and second parameter data representing a second portion of parameters of the NLU model, wherein the intent score data represents a confidence that the text data corresponds to intent data associated with the first domain; and

sending the intent data to the first domain.

2. The computer-implemented method of claim 1 , further comprising generating slot score data using the data representing the internal state of the NLU model and third parameter data representing a third portion of parameters of the NLU model, wherein the slot score data represents a confidence that the text data corresponds to a content slot associated with the intent data.

3. The computer-implemented method of claim 1 , further comprising:

generating second intent score data using the data representing the internal state of the model and third parameter data representing a third portion of parameters of the NLU model, wherein the second intent score data represents a second confidence that the text data corresponds to second intent data associated with a second domain distinct from the first domain; and

selecting the intent data based at least partly on an analysis of the intent score data with respect to the second intent score data.

4. The computer-implemented method of claim 1 , further comprising:

prior to receiving the audio data, generating display data using the first domain, wherein the display data represents an item to be displayed by the computing device, and wherein the item is associated with the first domain; and

sending the display data to the computing device.

5. The computer-implemented method of claim 1 , wherein generating the second vector data comprises generating data representing a value of the first element indicating that content associated with the first domain was displayed when the utterance was made.

6. The computer-implemented method of claim 5 , further comprising:

generating first hidden layer data using at least a portion of the text data and at least a portion of the first parameter data, wherein the NLU model comprises an artificial neural network;

generating long short term memory layer data using the first hidden layer data, the second vector data, and third parameter data representing values for a third portion of parameters of the NLU model; and

generating the intent score data and entity score data using the long short term memory data and at least a portion of the second parameter data.

7. The computer-implemented method of claim 5 , further comprising generating context content data comprising a fixed-length encoded representation of the item.

8. The computer-implemented method of claim 7 , further comprising:

generating first hidden layer data using at least a portion of the text data and at least a portion of the first parameter data, wherein the NLU model comprises an artificial neural network;

generating second hidden layer data using at least a portion of the context content data and third parameter data representing values for a third portion of parameters of the NLU model;

generating long short term memory layer data using the first hidden layer data, the second hidden layer data, the second vector data, and fourth parameter data representing values for a fourth portion of parameters of the NLU model; and

generating entity score data using the long short term memory data and at least a portion of the second parameter data.

9. The computer-implemented method of claim 8 , wherein the generating the second hidden layer data comprises using a convolutional neural network layer of the NLU model.

10. The computer-implemented method of claim 1 , wherein receiving the contextual data comprises receiving data representing at least one of: a background process of the computing device; an internal state of the computing device, or a capability of the computing device.

11. The computer-implemented method of claim 1 , further comprising:

obtaining training data representing a first plurality of utterances made during display of content and second plurality of utterances made without display of content;

generating first randomized training data comprising data representing a first utterance of the first plurality of utterances and first randomized contextual data, wherein the first randomized contextual data is based at least partly on randomizing an order of individual items of content displayed during the first utterance;

generating second randomized training data comprising data representing a second utterance of the second plurality utterances and second randomized contextual data, wherein the second randomized contextual data is based at least partly on content displayed during a randomly-selected utterance of the first plurality of utterances and not displayed during the second utterance; and

training the NLU model using the first randomized training data and the second randomized training data.

12. A system comprising:

computer-readable memory storing executable instructions; and

one or more processors in communication with the computer-readable memory and configured by the executable instructions to at least:

receive, from a computing device, audio data representing an utterance;

generate text data using the audio data;

receive contextual data, wherein the contextual data represents content displayed by the computing device when the utterance occurred, and wherein the content is associated with domain data representing a first domain of a plurality of domains;

generate input data for a natural language understanding (“NLU”) model, wherein the input data comprises:

first vector data representing at least a portion of the text data; and

second vector data representing a context domain vector comprising a first element and a second element, wherein the first element corresponds to the first domain and the second element corresponds to a second domain of the plurality of domains, and wherein the second vector data indicates that content associated with the first domain was displayed when the utterance was made;

generate data representing an internal state of the model using the first vector data, the second vector data, and first parameter data, wherein the first parameter data represents a first portion of parameters of the NLU model;

generate intent score data using the data representing the internal state of the model and second parameter data representing a second portion of parameters of the NLU model, wherein the intent score data represents a confidence that the text data corresponds to intent data associated with the first domain; and

send the intent data to the first domain.

13. The system of claim 12 , wherein the one or more processors are further configured by the executable instructions to generate slot score data using the data representing the internal state of the NLU model and third parameter data representing a third portion of parameters of the NLU model, wherein the slot score data represents a confidence that the text data corresponds to a content slot associated with the intent data.

14. The system of claim 12 , one or more processors are further configured by the executable instructions to:

generate second intent score data using the data representing the internal state of the model and third parameter data representing a third portion of parameters of the NLU model, wherein the second intent score data represents a second confidence that the text data corresponds to second intent data associated with a second domain distinct from the first domain; and

select the intent data based at least partly on an analysis of the intent score data with respect to the second intent score data.

15. The system of claim 12 , wherein data representing the first element comprises a value indicating that an item associated with the first domain was displayed when the utterance was made.

16. The system of claim 15 , wherein the one or more processors are further configured by the executable instructions to generate context content data comprising a fixed-length encoded representation of the item.

17. The system of claim 16 , wherein the one or more processors are further configured by the executable instructions to:

generate first hidden layer data using at least a portion of the text data and at least a portion of the first parameter data, wherein the NLU model comprises an artificial neural network;

generate second hidden layer data using at least a portion of the context content data and third parameter data representing values for a third portion of parameters of the NLU model;

generate long short term memory layer data using the first hidden layer data, the second hidden layer data, the second vector data, and fourth parameter data representing values for a fourth portion of parameters of the NLU model; and

generate entity score data using the long short term memory data and at least a portion of the second parameter data.

18. The system of claim 12 , wherein the one or more processors are further configured by the executable instructions to receive second contextual data representing at least one of: a background process of the computing device; an internal state of the computing device, or a capability of the computing device.

19. The system of claim 12 , wherein data representing the second element comprises a value indicating that the utterance did not occur during a display associated with the second domain.

20. The system of claim 12 , wherein the one or more processors are further configured by the executable instructions to:

obtain training data representing a first plurality of utterances made during display of content and second plurality of utterances made without display of content;

generate first randomized training data comprising data representing a first utterance of the first plurality of utterances and first randomized contextual data, wherein the first randomized contextual data is based at least partly on randomizing an order of individual items of content displayed during the first utterance;

generate second randomized training data comprising data representing a second utterance of the second plurality utterances and second randomized contextual data, wherein the second randomized contextual data is based at least partly on content displayed during a randomly-selected utterance of the first plurality of utterances and not displayed during the second utterance; and

train the NLU model using the first randomized training data and the second randomized training data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2019
From: METALLINOU, ANGELIKI; GOEL, RAHUL; ISHWAR, VISHAL
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 050466/0211 →
Continuity (1)
Provisional Application 62553066 · Aug 31, 2017
Cited By (49)
US 12,198,413 US 12,198,697 US 12,205,594 US 12,211,486 US 12,211,490 US 12,211,493 US 12,212,945 US 12,217,748 US 12,217,765 US 12,230,261 US 12,230,291 US 12,236,932 US 12,254,886 US 12,260,856 US 12,277,934 US 12,283,269 US 12,288,558 US 12,314,633 US 12,315,508 US 12,322,390 US 12,322,410 US 12,327,549 US 12,327,556 US 12,360,734 US 12,367,066 US 12,374,097 US 12,374,325 US 12,375,052 US 12,387,716 US 12,406,316 US 12,424,220 US 12,438,977 US 12,462,802 US 12,475,698 US 12,498,899 US 12,505,832 US 12,513,479 US 12,518,755 US 12,518,756 US 12,561,336 US 12,579,973 US 12,579,978 US 12,626,717 US 12,699,543 US 12,711,962 US 12,732,547 US 12,744,035 US 12,744,039 US 12,748,566