IP Library Granted Patent US 11,727,677
Granted Patent B2
US 11,727,677 · App. 17/566,308 · Granted Aug 15, 2023

Personalized gesture recognition for user interaction with assistant systems

Inventors: Xiaohu Liu (Bellevue, WA); Paul Anthony Crook (Newcastle, WA); Francislav P Penov (Kirkland, WA); Rajen Subba (San Carlos, CA)
Assignee: Meta Platforms Technologies, LLC
G06V10/82G06F3/011G06F3/013G06F3/017G06F3/167G06F7/14G06F9/453G06F16/176G06F16/2255G06F16/2365G06F16/243G06F16/248G06F16/24552G06F16/24575G06F16/24578G06F16/338G06F16/3323G06F16/3329G06F16/3344G06F16/904G06F16/9038G06F16/90332G06F16/90335G06F16/951G06F16/9535G06F18/2411G06F40/205G06F40/295G06F40/30G06F40/40G06N3/006G06N3/08G06N7/01G06N20/00G06Q50/01G06V10/764G06V20/10G06V40/28G10L15/02G10L15/063G10L15/07G10L15/16G10L15/183G10L15/187G10L15/1815G10L15/1822G10L15/22G10L15/26G10L17/06G10L17/22H04L12/2816H04L41/20H04L41/22H04L43/0882H04L43/0894H04L51/02H04L51/18H04L51/216H04L51/52H04L67/306H04L67/535H04L67/5651H04L67/75H04W12/08G06F2216/13G10L13/00G10L13/04G10L2015/223G10L2015/225H04L51/046H04L67/10H04L67/53
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,727,677
App. No.
17/566,308
Granted
Aug 15, 2023
Kind
B2
Abstract

In one embodiment, a method includes receiving a user request from a first user from a client system associated with a first user, wherein the user request comprise a gesture-input from the first user and a speech-input from the first user, determining an intent corresponding to the user request based on the gesture-input by a personalized gesture-classification model associated with the first user, executing one or more tasks based on the determined intent and the speech-input, and sending instructions for presenting execution results of the one or more tasks to the client system responsive the user request.

Claims (36)

1. A system comprising:

a microphone configured to receive a speech-input;

a camera configured to receive a gesture-input; and

circuitry configured to:

communicate a user input comprising at least one of the speech-input and the gesture-input to an external assistant system to cause the external assistant system to determine an output user intent from the user input;

in response to communicating the user input, receive information from the external assistant system for the output user intent; and

execute a task based at least in part on the received information for the output user intent determined from at least one of the speech-input and the gesture-input.

2. The system of claim 1 , wherein the circuitry is further configured to execute the task by displaying the received information.

3. The system of claim 1 , wherein the circuitry is further configured to execute the task by outputting audio for the received information.

4. The system of claim 1 , wherein the information is determined at least in part by performing automatic speech recognition on the speech-input.

5. The system of claim 1 , wherein the information is determined at least in part based on the user input and personal information of a user providing the user input.

6. The system of claim 1 , wherein the circuitry is further configured to communicate the user input to the external assistant system via a network.

7. The system of claim 6 , wherein the circuitry is further configured to receive the information from the external assistant system via the network.

8. The system of claim 1 , wherein the circuitry is further configured to determine a modality of the received information from the external assistant system at least in part based on a user profile associated with a user providing the user input.

9. The system of claim 1 , wherein the circuitry is further configured to determine a structure of the received information from the external assistant system at least in part based on a user profile associated with a user providing the user input.

10. The system of claim 1 , wherein the circuitry is further configured to execute the task at least in part based on a user profile associated with a user of the system.

11. The system of claim 10 , wherein the circuitry is further configured to determine the task at least in part based on a machine-learning model that is trained using the user profile.

12. The system of claim 11 , wherein the task comprises recommending an action to the user.

13. The system of claim 1 , wherein the circuitry is further configured to determine an intent of the gesture-input.

14. The system of claim 1 , wherein the user input comprises both the speech-input and the gesture-input.

15. The system of claim 14 , wherein the circuitry is further configured to determine an intent of the gesture-input at least in part based on the speech-input.

16. The system of claim 15 , wherein the circuitry is further configured to determine the intent of the gesture-input using a personalized gesture-classification model associated with a user providing the user input.

17. The system of claim 1 , wherein the microphone and the camera are disposed in augmented reality glasses or a virtual reality headset.

18. The system of claim 1 , wherein the microphone and the camera are enclosed in a computing device corresponding to the circuitry.

19. A method comprising:

receiving a speech-input from a microphone;

receiving a gesture-input from a camera;

communicating a user input comprising at least one of the speech-input and the gesture-input to an external assistant system to cause the external assistant system to determine an output user intent from the user input;

in response to communicating the user input, receiving information from the external assistant system for the output user intent; and

execute a task based at least in part on the received information for the output user intent determined from at least one of the speech-input and the gesture-input.

20. A non-transitory computer-readable medium comprising software that, when executed by a processor, is operable to:

receive a speech-input from a microphone;

receive a gesture-input from a camera;

communicate a user input comprising at least one of the speech-input and the gesture-input to an external assistant system to cause the external assistant system to determine an output user intent from the user input;

in response to communicating the user input, receive information from the external assistant system for the output user intent; and

execute a task based at least in part on the received information for the output user intent determined from at least one of the speech-input and the gesture-input.

Assignments (2)
CHANGE OF NAME Recorded May 19, 2022
From: FACEBOOK TECHNOLOGIES, LLC
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 060121/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2022
From: LIU, XIAOHU; CROOK, PAUL ANTHONY; PENOV, FRANCISLAV P; SUBBA, RAJEN
To: FACEBOOK TECHNOLOGIES, LLC
Reel/Frame 058611/0880 →