IP Library Granted Patent US 11,687,770
Granted Patent B2
US 11,687,770 · App. 16/417,554 · Granted Jun 27, 2023

Recurrent multimodal attention system based on expert gated networks

Inventors: Francesco Nesta (Aliso Viejo, CA); Lijiang Guo (Bloomington, IN); Minje Kim (Bloomington, IN)
Assignees: SYNAPTICS INCORPORATED; THE TRUSTEES OF INDIANA UNIVERSITY
G06N3/08G06F18/2113G06F18/254G06N3/045G06V10/40G06V10/764G06V10/811G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,687,770
App. No.
16/417,554
Granted
Jun 27, 2023
Kind
B2
Abstract

Systems and methods for multimodal classification include a plurality of expert modules, each expert module configured to receive data corresponding to one of a plurality of input modalities and extract associated features, a plurality of class prediction modules, each class prediction module configured to receive extracted features from a corresponding one of the expert modules and predict an associated class, a gate expert configured to receive the extracted features from the plurality of expert modules and output a set of weights for the input modalities, and a fusion module configured to generate a weighted prediction based on the class predictions and the set of weights. Various embodiments include one or more of an image expert, a video expert, an audio expert, class prediction modules, a gate expert, and a co-learning framework.

Claims (37)

1. A system comprising:

a plurality of feature extraction modules each configured to receive data associated with a respective input modality of a plurality of input modalities and extract features associated with the respective input modality;

a plurality of recurrent neural network (RNN) modules each configured to receive the features extracted by a respective one of the feature extraction modules and predict a class associated with the features extracted by the respective feature extraction module;

a gate expert configured to receive the features extracted by each of the plurality of feature extraction modules and output a set of weights for the plurality of input modalities based on a first neural network, the first neural network being trained to produce the set of weights based on the features extracted by each of the plurality of feature extraction modules; and

a fusion module configured to generate a weighted prediction based on the predicted classes and the set of weights.

2. The system of claim 1 , wherein the plurality of input modalities includes images, video and audio.

3. The system of claim 1 , wherein the feature extraction modules comprise an image expert, a video expert and an audio expert.

4. The system of claim 1 , wherein each feature extraction module comprises a respective neural network different than the first neural network.

5. The system of claim 1 , wherein at least one of the RNN modules comprises a long short-term memory network.

6. The system of claim 1 , wherein the gate expert comprises a long short-term memory network.

7. The system of claim 1 , further comprising a co-learning framework.

8. A method comprising:

receiving a plurality of data streams associated with a plurality of input modalities, respectively;

extracting, from each data stream of the plurality of data streams, features associated with a respective input modality of the plurality of input modalities;

predicting a plurality of classes based on a plurality of recurrent neural networks (RNNs), respectively, each RNN of the plurality of RNNs being configured to predict the respective class based on the features associated a respective input modality of the plurality of input modalities;

generating a set of weights for the plurality of input modalities based on a first neural network, the first neural network being trained to produce the set of weights based on the features associated with each input modality of the plurality of input modalities; and

generating a weighted prediction based on the predicted classes and the set of weights.

9. The method of claim 8 , wherein receiving the plurality of data streams comprises sensing, using a plurality of sensor types, one or more conditions in an environment, and wherein each of the plurality of sensor types has a corresponding input modality.

10. The method of claim 9 , wherein each of the plurality of data streams contributes to the weighted prediction with a different degree of confidence depending on sensed conditions.

11. The method of claim 9 , wherein the set of weights is generated dynamically to combine information from the plurality of sensor types in accordance with sensed conditions.

12. The method of claim 8 , wherein the first neural network comprises a gating modular neural network trained to dynamically generate a set of weights for outputs of a plurality of sensor networks, at least in part, by balancing a utility of each data stream.

13. The method of claim 8 , wherein the generating of the weighted prediction comprises:

detecting voice activity in the plurality of data streams; and

providing corresponding input frames to one or more applications for further processing.

14. The method of claim 8 , wherein the plurality of input modalities includes images, video and audio.

15. A system comprising:

a memory storing instructions;

a processor coupled to the memory and configured to execute the instructions to cause the system to perform operations comprising:

receiving a plurality of data streams associated with a plurality of input modalities, respectively;

extracting, from each data stream of the plurality of data streams, features associated with a respective input modality of the plurality of input modalities;

predicting a plurality of classes based on a plurality of recurrent neural networks (RNNs), respectively, each RNN of the plurality of RNNs being configured to predict the respective class based on the features associated with a respective input modality of the plurality of input modalities;

generating a set of weights for the plurality of input modalities based on a first neural network, the first neural network being trained to produce the set of weights based on the features associated with each input modality of the plurality of input modalities; and

generating a weighted prediction based on the plurality of classes and the set of weights.

16. The system of claim 15 , wherein the plurality of input modalities includes images, video and audio.

17. The system of claim 15 , wherein extracting features associated with the respective input modality comprises inputting the data stream to a trained neural network.

18. The system of claim 15 , wherein each of the plurality of RNNs comprises a long short-term memory network.

19. The system of claim 15 , wherein the first neural network comprises a long short-term memory network.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2023
From: GUO, LIJIANG
To: THE TRUSTEES OF INDIANA UNIVERSITY
Reel/Frame 064063/0601 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 16, 2022
From: NESTA, FRANCESCO
To: SYNAPTICS INCORPORATED
Reel/Frame 059287/0575 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 2, 2020
From: KIM, MINJE
To: THE TRUSTEES OF INDIANA UNIVERSITY
Reel/Frame 053668/0192 →
SECURITY INTEREST Recorded Feb 14, 2020
From: SYNAPTICS INCORPORATED
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 051936/0103 →