IP Library Granted Patent US 10,802,848
Granted Patent B2
US 10,802,848 · App. 16/388,130 · Granted Oct 13, 2020

Personalized gesture recognition for user interaction with assistant systems

Inventors: Xiaohu Liu (Bellevue, WA); Paul Anthony Crook (Newcastle, WA); Francislav P. Penov (Kirkland, WA); Rajen Subba (San Carlos, CA)
Assignee: Facebook Technologies, LLC
G06F9/453G06F3/011G06F3/017G06F3/167G06F16/176G06F16/24575G06F16/338G06F16/3323G06F16/3344G06F16/904G06F16/9038G06F16/90332G06F16/9535G06F40/30G06F40/40G06K9/00355G06K9/00664G06K9/6269G06N3/08G06N20/00G06Q50/01G10L15/063G10L15/16G10L15/183G10L15/1815G10L15/1822G10L15/22G10L15/26H04L51/02H04L67/22H04L67/306G10L13/043G10L2015/223H04L51/046H04L67/10H04L67/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,802,848
App. No.
16/388,130
Granted
Oct 13, 2020
Kind
B2
Abstract

In one embodiment, a method includes accessing a plurality of input tuples associated with a first user from a data store, wherein each input tuple comprises a gesture-input and a corresponding speech-input, determining a plurality of intents corresponding to the plurality of speech-inputs, respectively, by a natural-language understanding (NLU) module, generating a plurality of feature representations for the plurality of gesture-inputs based on one or more machine-learning models, determining a plurality of gesture identifiers for the plurality of gesture-inputs, respectively, based on their respective feature representations, associating the plurality of intents with the plurality of gesture identifiers, respectively, and training a personalized gesture-classification model for the first user based on the plurality of feature representations of their respective gesture-inputs and the associations between the plurality of intents and their respective gesture identifiers.

Claims (49)

1. A method comprising, by one or more computing systems:

accessing, from a data store, a plurality of input tuples associated with a first user, wherein each input tuple comprises a gesture-input and a corresponding speech-input;

determining, by a natural-language understanding (NLU) module, a plurality of intents corresponding to the plurality of speech-inputs, respectively;

generating, for the plurality of gesture-inputs, a plurality of feature representations based on one or more machine-learning models;

determining a plurality of gesture identifiers for the plurality of gesture-inputs, respectively, based on their respective feature representations;

associating the plurality of intents with the plurality of gesture identifiers, respectively; and

training, for the first user, a personalized gesture-classification model based on the plurality of feature representations of their respective gesture-inputs and the associations between the plurality of intents and their respective gesture identifiers.

2. The method of claim 1 , further comprising:

accessing, from the data store, a general gesture-classification model corresponding to a general user population, wherein training the personalized gesture-classification model is further based on the general gesture-classification model.

3. The method of claim 2 , wherein the general gesture-classification model is trained based on a plurality of gesture-inputs from the general user population.

4. The method of claim 1 , further comprising:

generating, by one or more automatic speech recognition (ASR) modules, a plurality of text-inputs for the plurality of speech-inputs, respectively.

5. The method of claim 4 , wherein determining the plurality of intents corresponding to the plurality of speech-inputs, respectively, is based on the plurality of text-inputs of the respective speech-inputs.

6. The method of claim 1 , wherein the one or more machine-learning models are based on one or more of a neural network model or a long-short term memory (LSTM) model.

7. The method of claim 1 , wherein the personalized gesture-classification model is based on convolutional neural networks.

8. The method of claim 1 , wherein generating each feature representation for each gesture-input comprises:

dividing the gesture-input into one or more components; and

modeling the one or more components into the feature representation for the gesture-input.

9. The method of claim 1 , wherein generating each feature representation for each gesture-input comprises:

determining temporal information associated with the gesture-input; and

modeling the temporal information into the feature representation for the gesture-input.

10. The method of claim 1 , further comprising:

receiving, from a client system associated with the first user, a new gesture-input from the first user; and

determining, for the new gesture-input, an intent corresponding to the new gesture-input based on the personalized gesture-classification model.

11. The method of claim 10 , further comprising:

executing one or more tasks based on the determined intent.

12. The method of claim 1 , wherein training the personalized gesture-classification model is further based on user feedback data from the first user.

13. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

access, from a data store, a plurality of input tuples associated with a first user, wherein each input tuple comprises a gesture-input and a corresponding speech-input;

determine, by a natural-language understanding (NLU) module, a plurality of intents corresponding to the plurality of speech-inputs, respectively;

generate, for the plurality of gesture-inputs, a plurality of feature representations based on one or more machine-learning models;

determine a plurality of gesture identifiers for the plurality of gesture-inputs, respectively, based on their respective feature representations;

associate the plurality of intents with the plurality of gesture identifiers, respectively; and

train, for the first user, a personalized gesture-classification model based on the plurality of feature representations of their respective gesture-inputs and the associations between the plurality of intents and their respective gesture identifiers.

14. The media of claim 13 , wherein the software is further operable when executed to:

access, from the data store, a general gesture-classification model corresponding to a general user population, wherein training the personalized gesture-classification model is further based on the general gesture-classification model.

15. The media of claim 14 , wherein the general gesture-classification model is trained based on a plurality of gesture-inputs from the general user population.

16. The media of claim 13 , wherein the software is further operable when executed to:

generate, by one or more automatic speech recognition (ASR) modules, a plurality of text-inputs for the plurality of speech-inputs, respectively.

17. The media of claim 16 , wherein determining the plurality of intents corresponding to the plurality of speech-inputs, respectively, is based on the plurality of text-inputs of the respective speech-inputs.

18. The media of claim 13 , wherein the one or more machine-learning models are based on one or more of a neural network model or a long-short term memory (LSTM) model.

19. The media of claim 13 , wherein the personalized gesture-classification model is based on convolutional neural networks.

20. A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:

access, from a data store, a plurality of input tuples associated with a first user, wherein each input tuple comprises a gesture-input and a corresponding speech-input;

determine, by a natural-language understanding (NLU) module, a plurality of intents corresponding to the plurality of speech-inputs, respectively;

generate, for the plurality of gesture-inputs, a plurality of feature representations based on one or more machine-learning models;

determine a plurality of gesture identifiers for the plurality of gesture-inputs, respectively, based on their respective feature representations;

associate the plurality of intents with the plurality of gesture identifiers, respectively; and

train, for the first user, a personalized gesture-classification model based on the plurality of feature representations of their respective gesture-inputs and the associations between the plurality of intents and their respective gesture identifiers.

Assignments (3)
CHANGE OF NAME Recorded Jul 6, 2022
From: FACEBOOK TECHNOLOGIES, LLC
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 060591/0848 →
CHANGE OF NAME Recorded May 19, 2022
From: FACEBOOK TECHNOLOGIES, LLC
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 060121/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2019
From: LIU, XIAOHU; CROOK, PAUL ANTHONY; PENOV, FRANCISLAV P.; SUBBA, RAJEN
To: FACEBOOK TECHNOLOGIES, LLC
Reel/Frame 048959/0935 →
Continuity (2)
Provisional Application 62660876 · Apr 20, 2018
Related Publication 20190324553A1 · Oct 24, 2019
Cited By (2)
US 12,542,136 US 12,572,327