IP Library › Patent Application 16723762
Patent Application
App. No. 16/723,762

MULTI-MODAL NATURAL LANGUAGE PROCESSING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/723,762
Filed
Dec 20, 2019
Art Unit
2657
USPC
704/232
Abstract

Multi-modal natural language processing systems are provided. Some systems are context-aware systems that use multi-modal data to improve the accuracy of natural language understanding as it is applied to spoken language input. Machine learning architectures are provided that jointly model spoken language input (“utterances”) and information displayed on a visual display (“on-screen information”). Such machine learning architectures can improve upon, and solve problems inherent in, existing spoken language understanding systems that operate in multi-modal contexts.

Claims (54)

1 . (canceled)

2 . A computer-implemented method comprising:

as performed by a computing system comprising one or more computer processors configured to execute specific instructions,

receiving, from a computing device, audio data representing an utterance;

receiving contextual data, wherein the contextual data represents content displayed by the computing device when the utterance occurred, and wherein the content is associated with domain data representing a first domain of a plurality of domains;

generating natural language understanding (“NLU”) input data for an NLU subsystem, wherein the input data comprises:

first input data representing at least a portion of the utterance; and

second input data comprising a first element and a second element, wherein the first element corresponds to the first domain and the second element corresponds to a second domain of the plurality of domains, and wherein the second input data indicates that content associated with the first domain was displayed when the utterance was made;

generating NLU output data using the NLU subsystem, the first input data, and the second input data, wherein the NLU output data represents a correspondence of the utterance to intent data associated with the first domain; and

sending the intent data to the first domain.

3 . The computer-implemented method of claim 2 , further comprising:

generating second NLU output data using the NLU subsystem, the first input data, and the second input data, wherein the second NLU output data represents a correspondence of the utterance to second intent data associated with a second domain distinct from the first domain; and

selecting the intent data based at least partly on an analysis of the first NLU output data with respect to the second NLU output data.

4 . The computer-implemented method of claim 2 , further comprising:

prior to receiving the audio data, generating display data using the first domain, wherein the display data represents an item to be displayed by the computing device, and wherein the item is associated with the first domain; and

sending the display data to the computing device.

5 . The computer-implemented method of claim 2 , wherein receiving the contextual data comprises receiving data representing at least one of: a background process of the computing device; an internal state of the computing device, or a capability of the computing device.

6 . The computer-implemented method of claim 2 , wherein generating the NLU input data comprises generating the second input data as vector data representing a vector comprising the first element and the second element.

7 . The computer-implemented method of claim 6 , wherein generating the second input data comprises:

generating first value data for the first element, the first value data representing an association of the utterance with the first domain; and

generating second value data for the second element, the second value data representing lack of an association of the utterance with the second domain.

8 . The computer-implemented method of claim 6 , wherein generating the second input data comprises:

generating first value data for the first element, the first value data representing an association of the utterance with the first domain; and

generating second value data for the second element, the second value data representing an association of the utterance with the second domain.

9 . The computer-implemented method of claim 2 , further comprising generating data representing an internal state of an NLU model using the first input data, the second input data, and parameter data, wherein the parameter data represents parameters of the NLU model, and wherein generating the NLU output data is based at least partly on the data representing the internal state of the NLU model.

10 . The computer-implemented method of claim 2 , further comprising generating second NLU output data using the NLU subsystem, the first input data, and the second input data, wherein the second NLU output data represents label of the portion of the utterance as a named entity.

11 . The computer-implemented method of claim 2 , further comprising generating second NLU output data using the NLU subsystem, the first input data, and the second input data, wherein the second NLU output data represents a correspondence of the portion of the utterance to a content slot associated with the intent data.

12 . A system comprising:

computer-readable memory storing executable instructions; and

one or more processors in communication with the computer-readable memory and configured by the executable instructions to at least:

receive, from a computing device, audio data representing an utterance;

receive contextual data, wherein the contextual data represents content displayed by the computing device when the utterance occurred, and wherein the content is associated with domain data representing a first domain of a plurality of domains;

generate natural language understanding (“NLU”) input data for an NLU subsystem, wherein the input data comprises:

first input data representing at least a portion of the utterance; and

second input data comprising a first element and a second element, wherein the first element corresponds to the first domain and the second element corresponds to a second domain of the plurality of domains, and wherein the second input data indicates that content associated with the first domain was displayed when the utterance was made;

generate NLU output data using the NLU subsystem, the first input data, and the second input data, wherein the NLU output data represents a correspondence of the utterance to intent data associated with the first domain; and

send the intent data to the first domain.

13 . The system of claim 12 , wherein the one or more processors are further configured by the executable instructions to:

generate second NLU output data using the NLU subsystem, the first input data, and the second input data, wherein the second NLU output data represents a correspondence of the utterance to second intent data associated with a second domain distinct from the first domain; and

select the intent data based at least partly on an analysis of the first NLU output data with respect to the second NLU output data.

14 . The system of claim 12 , wherein the one or more processors are further configured by the executable instructions to:

generate display data using the first domain prior to receiving the audio data, wherein the display data represents an item to be displayed by the computing device, and wherein the item is associated with the first domain; and

send the display data to the computing device.

15 . The system of claim 12 , wherein the contextual data further represents at least one of: a background process of the computing device; an internal state of the computing device, or a capability of the computing device.

16 . The system of claim 12 , wherein to generate the NLU input data, the one or more processors are further configured by the executable instructions to generate the second input data as vector data representing a vector comprising the first element and the second element.

17 . The system of claim 16 , wherein to generate the second input data, the one or more processors are further configured by the executable instructions to:

generate first value data for the first element, the first value data representing an association of the utterance with the first domain; and

generate second value data for the second element, the second value data representing lack of an association of the utterance with the second domain.

18 . The system of claim 16 , wherein to generate the second input data, the one or more processors are further configured by the executable instructions to:

generate first value data for the first element, the first value data representing an association of the utterance with the first domain; and

generate second value data for the second element, the second value data representing an association of the utterance with the second domain.

19 . The system of claim 12 , wherein the one or more processors are further configured by the executable instructions to generate data representing an internal state of an NLU model using the first input data, the second input data, and parameter data, wherein the parameter data represents parameters of the NLU model, and wherein generating the NLU output data is based at least partly on the data representing the internal state of the NLU model.

20 . The system of claim 12 , wherein the one or more processors are further configured by the executable instructions to generate second NLU output data using the NLU subsystem, the first input data, and the second input data, wherein the second NLU output data represents label of the portion of the utterance as a named entity.

21 . The system of claim 12 , wherein the one or more processors are further configured by the executable instructions to generate second NLU output data using the NLU subsystem, the first input data, and the second input data, wherein the second NLU output data represents a correspondence of the portion of the utterance to a content slot associated with the intent data.