IP Library › Granted Patent US 11,842,727
Granted Patent B2
US 11,842,727 · App. 17/659,612 · Granted Dec 12, 2023

Natural language processing with contextual data representing displayed content

Inventors: Angeliki Metallinou (Mountain View, CA); Rahul Goel (Mountain View, CA); Vishal Ishwar (Tempe, AZ)
Assignee: Amazon Technologies, Inc.
G10L15/16G06F3/167G10L15/02G10L15/144G10L15/197G10L15/26G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,842,727
App. No.
17/659,612
Granted
Dec 12, 2023
Kind
B2
Abstract

Multi-modal natural language processing systems are provided. Some systems are context-aware systems that use multi-modal data to improve the accuracy of natural language understanding as it is applied to spoken language input. Machine learning architectures are provided that jointly model spoken language input (“utterances”) and information displayed on a visual display (“on-screen information”). Such machine learning architectures can improve upon, and solve problems inherent in, existing spoken language understanding systems that operate in multi-modal contexts.

Claims (51)

1. A computer-implemented method comprising:

as performed by a computing system comprising one or more computer processors configured to execute specific instructions,

receiving, from a computing device, audio data representing an utterance;

receiving contextual data, wherein the contextual data represents a plurality of content items displayed by the computing device when the utterance occurred, and wherein a first content item of the plurality of content items is associated with domain data representing a first domain of a plurality of domains;

generating automatic speech recognition (“ASR”) data using the audio data and a language model;

generating natural language understanding (“NLU”) input data for an NLU subsystem using the ASR data and the contextual data, wherein the NLU input data comprises:

first input data representing at least a portion of the utterance;

second input data representing a vector that comprises a first vector element representing the first domain and a second vector element representing a second domain of the plurality of domains, wherein the second input data indicates content associated with the first domain was displayed when the utterance occurred; and

third input data comprising a plurality of elements, wherein a first element of the plurality of elements represents the first content item, and wherein a second element of the plurality of elements represents a second content item of the plurality of content items;

generating NLU output data using the NLU subsystem, the first input data, the second input data, and the third input data, wherein the NLU output data represents intent data associated with the utterance; and

sending the intent data to the first domain.

2. The computer-implemented method of claim 1 , wherein generating the NLU output data comprises processing the first input data, the second input data, and the third input data using a neural network component.

3. The computer-implemented method of claim 1 , wherein generating the ASR data comprises generating the ASR data using an ASR subsystem, and wherein receiving the contextual data comprises receiving the contextual data from an application subsystem after at least a portion of the ASR data is generated by the ASR subsystem.

4. The computer-implemented method of claim 1 , further comprising determining that the utterance includes a wakeword, wherein generating the NLU input data using the ASR data and the contextual data is performed in response to determining that the utterance includes the wakeword.

5. The computer-implemented method of claim 1 , further comprising:

generating second NLU output data using the NLU subsystem, the first input data, the second input data, the third input data, wherein the second NLU output data represents second intent data associated with a second domain distinct from the first domain; and

selecting the intent data based at least partly on an analysis of the NLU output data with respect to the second NLU output data.

6. The computer-implemented method of claim 1 , further comprising:

prior to receiving the audio data, generating display data using the first domain, wherein the display data represents the first content item, wherein the first content item is to be displayed by the computing device, and wherein the first content item is associated with the first domain; and

sending the display data to the computing device.

7. The computer-implemented method of claim 1 , wherein receiving the contextual data comprises receiving data representing at least one of: a background process of the computing device; an internal state of the computing device, or a capability of the computing device.

8. The computer-implemented method of claim 1 , wherein generating the second input data comprises:

generating first value data for the first vector element, the first value data representing an association of the utterance with the first domain; and

generating second value data for the second vector element, the second value data representing lack of an association of the utterance with the second domain.

9. The computer-implemented method of claim 1 , wherein generating the second input data comprises:

generating first value data for the first vector element, the first value data representing an association of the utterance with the first domain; and

generating second value data for the second vector element, the second value data representing an association of the utterance with the second domain.

10. The computer-implemented method of claim 1 , further comprising generating second NLU output data using the NLU subsystem, the first input data, the second input data, and the third input data, wherein the second NLU output data represents a label of the portion of the utterance as a named entity.

11. The computer-implemented method of claim 1 , further comprising generating second NLU output data using the NLU subsystem, the first input data, the second input data, and the third input data, wherein the second NLU output data represents a content slot associated with the intent data.

12. A system comprising:

computer-readable memory storing executable instructions; and

one or more processors in communication with the computer-readable memory and configured by the executable instructions to at least:

receive, from a computing device, audio data representing an utterance;

receive contextual data, wherein the contextual data represents a plurality of content items displayed by the computing device when the utterance occurred, and wherein a first content item of the plurality of content items is associated with domain data representing a first domain of a plurality of domains;

generate automatic speech recognition (“ASR”) data using the audio data and a language model;

generate natural language understanding (“NLU”) input data for an NLU subsystem using the ASR data and the contextual data, wherein the NLU input data comprises:

first input data representing at least a portion of the utterance;

second input data representing a vector that comprises a first vector element representing the first domain and a second vector element representing a second domain of the plurality of domains, wherein the second input data indicates content associated with the first domain was displayed when the utterance occurred; and

third input data comprising a plurality of elements, wherein a first element of the plurality of elements represents the first content item, and wherein a second element of the plurality of elements represents a second content item of the plurality of content items;

generate NLU output data using the NLU subsystem, the first input data, the second input data, and the third input data, wherein the NLU output data represents intent data associated with the utterance; and

send the intent data to the first domain.

13. The system of claim 12 , wherein to generate the NLU output data, the one or more processors are further configured by the executable instructions to process the first input data, the second input data, and the third input data using a neural network component.

14. The system of claim 12 , further comprising an ASR subsystem configured to generate the ASR data, wherein the contextual data is received from an application subsystem after at least a portion of the ASR data is generated by the ASR subsystem.

15. The system of claim 12 , wherein the one or more processors are further configured by the executable instructions to:

prior to receiving the audio data, generate display data using the first domain, wherein the display data represents the first content item, wherein the first content item is to be displayed by the computing device, and wherein the first content item is associated with the first domain; and

send the display data to the computing device.

16. The system of claim 12 , wherein to receive the contextual data, the one or more processors are further configured by the executable instructions to receive data representing at least one of: a background process of the computing device; an internal state of the computing device, or a capability of the computing device.

17. The system of claim 12 , wherein to generate the second input data, the one or more processors are further configured by the executable instructions to:

generate first value data for the first vector element, the first value data representing an association of the utterance with the first domain; and

generate second value data for the second vector element, the second value data representing an association of the utterance with the second domain.

18. The system of claim 12 , wherein the one or more processors are further configured by the executable instructions to generate second NLU output data using the NLU subsystem, the first input data, the second input data, and the third input data, wherein the second NLU output data represents a label of the portion of the utterance as a named entity.

Continuity (4)
Continuation 16723762 · Dec 20, 2019
Continuation 15828174 · Nov 30, 2017
Provisional Application 62553066 · Aug 31, 2017
Related Publication 20220246139A1 · Aug 4, 2022
Cited By (3)
US 12,368,931 US 12,579,973 US 12,699,727