IP Library Granted Patent US 10,936,346
Granted Patent B2
US 10,936,346 · App. 16/053,600 · Granted Mar 2, 2021

Processing multimodal user input for assistant systems

Inventors: Vivek Natarajan (Palo Alto, CA); Shawn C. P. Mei (San Francisco, CA); Zhengping Zuo (Medina, WA)
Assignee: Facebook, Inc.
G06F9/453G06F3/011G06F3/017G06F3/167G06F7/14G06F16/176G06F16/2255G06F16/2365G06F16/24575G06F16/338G06F16/3323G06F16/3344G06F16/904G06F16/9038G06F16/90332G06F16/9535G06F40/30G06F40/40G06K9/00355G06K9/00664G06K9/6269G06N3/08G06N20/00G06Q50/01G10L15/063G10L15/16G10L15/183G10L15/1815G10L15/1822G10L15/22G10L15/26H04L43/0882H04L43/0894H04L51/02H04L67/22H04L67/2828H04L67/306G10L13/00G10L13/04G10L2015/223H04L51/046H04L67/10H04L67/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,936,346
App. No.
16/053,600
Granted
Mar 2, 2021
Kind
B2
Abstract

In one embodiment, a method includes receiving from a client system associated with a first user a user input based on one or more modalities, at least one of which is a visual modality, identifying one or more subjects associated with the user input based on the visual modality based on one or more machine-learning models, determining one or more attributes associated with the one or more subjects respectively based on the one or more machine-learning models, resolving one or more entities corresponding to the one or more subjects based on the determined one or more attributes, executing one or more tasks associated with the one or more resolved entities, and sending instructions for presenting a communication content including information associated with the executed one or more tasks responsive to user input to the client system associated with the first user.

Claims (51)

1. A method comprising, by one or more computing systems:

receiving, from a client system associated with a first user, a user input based on a plurality of modalities, wherein at least one of the modalities of the user input is a visual modality;

identifying, based on one or more machine-learning models, one or more subjects associated with the user input based on the visual modality;

determining, based on the one or more machine-learning models, one or more attributes associated with the one or more subjects, respectively;

resolving, based on the determined one or more attributes, one or more entities corresponding to the one or more subjects;

executing one or more tasks associated with the one or more resolved entities; and

sending, to the client system associated with the first user, instructions for presenting a communication content responsive to the user input, wherein the communication content comprises information associated with the executed one or more tasks.

2. The method of claim 1 , wherein the user input comprises two or more of:

a character string;

an audio clip;

an image; or

a video clip.

3. The method of claim 1 , wherein the one or more subjects associated with the user input comprise one or more of a person, a location, a business, or an object.

4. The method of claim 3 , wherein identifying the one or more people is based on facial recognition.

5. The method of claim 3 , wherein identifying the one or more objects is based on object detection.

6. The method of claim 1 , further comprising generating a feature representation for the user input based on the visual modality.

7. The method of claim 1 , wherein the one or more machine-learning models comprise one or more of:

a support vector machine;

a regression model; or

a convolutional neural network.

8. The method of claim 1 , further comprising identifying one or more intents and one or more slots based on the user input.

9. The method of claim 8 , wherein executing the one or more tasks associated with the one or more resolved entities is based on the identified intents and slots.

10. The method of claim 1 , wherein the communication content comprises one or more of:

a character string;

an audio clip;

an image; or

a video clip.

11. The method of claim 1 , further comprising determining one or more modalities for the communication content.

12. The method of claim 11 , wherein determining the one or more modalities for the communication content comprises:

identifying contextual information associated with the first user;

identifying contextual information associated with the client system; and

determining the one or more modalities based on the contextual information associated with the first user and the contextual information associated with the client system.

13. The method of claim 1 , further comprising:

generating a plurality of tasks based on the visual modality of the user input; and

receiving, from the client system associated with the first user, a user selection of the one or more tasks from the plurality of tasks by the first user.

14. The method of claim 1 , further comprising storing the identified one or more subjects in a dialog state.

15. The method of claim 1 , wherein the user input comprises a user interaction with a media content object.

16. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

receive, from a client system associated with a first user, a user input based on a plurality of modalities, wherein at least one of the modalities of the user input is a visual modality;

identify, based on one or more machine-learning models, one or more subjects associated with the user input based on the visual modality;

determine, based on the one or more machine-learning models, one or more attributes associated with the one or more subjects, respectively;

resolve, based on the determined one or more attributes, one or more entities corresponding to the one or more subjects;

execute one or more tasks associated with the one or more resolved entities; and

send, to the client system associated with the first user, instructions for presenting a communication content responsive to user input, wherein the communication content comprises information associated with the executed one or more tasks.

17. A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:

receive, from a client system associated with a first user, a user input based on a plurality of modalities, wherein at least one of the modalities of the user input is a visual modality;

identify, based on one or more machine-learning models, one or more subjects associated with the user input based on the visual modality;

determine, based on the one or more machine-learning models, one or more attributes associated with the one or more subjects, respectively;

resolve, based on the determined one or more attributes, one or more entities corresponding to the one or more subjects;

execute one or more tasks associated with the one or more resolved entities; and

send, to the client system associated with the first user, instructions for presenting a communication content responsive to user input, wherein the communication content comprises information associated with the executed one or more tasks.

Assignments (2)
CHANGE OF NAME Recorded Dec 20, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058553/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 2, 2020
From: NATARAJAN, VIVEK; MEI, SHAWN C.P.; ZUO, ZHENGPING
To: FACEBOOK, INC.
Reel/Frame 051400/0120 →
Continuity (2)
Provisional Application 62660876 · Apr 20, 2018
Related Publication 20190325080A1 · Oct 24, 2019
Cited By (2)
US 12,651,599 US 12,705,273