IP Library Patent Application 17368066
Patent Application
App. No. 17/368,066

Auto-completion for Multi-modal User Input in Assistant Systems

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/368,066
Abstract

In one embodiment, a method includes receiving an initial input in a first modality from a first user at a client system, determining intents and slots corresponding to the initial input, wherein the slots are conditioned on the intents, generating one or more candidate continuation-inputs based on the intents and slots, where the one or more candidate continuation-inputs are in one or more candidate modalities, respectively, wherein the candidate modalities are different from the first modality, and wherein each of the candidate continuation-inputs references entities represented by the slots, and presenting one or more suggested inputs corresponding to one or more of the candidate continuation-inputs at the client system.

Claims (42)

1 . A method comprising, by a client system:

receiving, at the client system, an initial input from a first user, wherein the initial input is in a first modality;

determining one or more intents and one or more slots corresponding to the initial input, wherein the one or more slots are conditioned on the one or more intents;

generating, based on the one or more intents and the one or more slots, one or more candidate continuation-inputs, where the one or more candidate continuation-inputs are in one or more candidate modalities, respectively, wherein the candidate modalities are different from the first modality, and wherein each of the candidate continuation-inputs references one or more entities represented by the one or more slots; and

presenting, at the client system, one or more suggested inputs corresponding to one or more of the candidate continuation-inputs.

2 . The method of claim 1 , wherein the first modality comprises one of audio, text, image, video, motion, or orientation.

3 . The method of claim 2 , wherein the first modality comprises motion, and wherein the initial input comprises a gesture.

4 . The method of claim 2 , wherein the first modality comprises orientation, and wherein the initial input comprises a gaze on an object.

5 . The method of claim 1 , further comprising:

determining that the first user needs one or more suggested inputs.

6 . The method of claim 5 , wherein determining that the first user needs one or more suggested inputs is based on a wake-up input from the first user.

7 . The method of claim 6 , wherein the wake-up input comprises one or more of a voice utterance, a character string, an image, a video clip, a gesture, or a gaze.

8 . The method of claim 5 , wherein the initial input comprises a gaze on an object, and wherein determining that the first user needs one or more suggested inputs is further based on the gaze on the object.

9 . The method of claim 8 , wherein generating the one or more candidate continuation-inputs is further based on the object.

10 . The method of claim 5 , wherein determining that the first user needs one or more suggested inputs is further based on contextual information associated with the initial input.

11 . The method of claim 5 , wherein determining that the first user needs one or more suggested inputs is further based on the one or more intents.

12 . The method of claim 1 , further comprising:

identifying one or more entities associated with the one or more intents.

13 . The method of claim 12 , wherein generating the one or more candidate continuation-inputs is further based the one or more entities.

14 . The method of claim 1 , further comprising:

receiving, at the client system, a user-selected input from the first user, wherein the user-selected input comprises one of the suggested inputs; and

executing one or more tasks based on the user-selected input.

15 . The method of claim 1 , further comprising:

receiving, at the client system, a first user-selected input from the first user, wherein the first user-selected input comprises one of the suggested inputs, and wherein the first user-selected input is associated with a first intent;

generating, based on the first user-selected input, one or more additional candidate continuation-inputs, wherein each of the one or more additional candidate continuation-inputs is associated with the first intent;

presenting, at the client system, one or more additional suggested inputs corresponding to one or more of the additional candidate continuation-inputs;

receiving, at the client system, a second user-selected input from the first user, wherein the second user-selected input comprises one of the additional suggested inputs; and

executing one or more tasks based on the second user-selected input.

16 . The method of claim 1 , further comprising:

determining the initial input comprises an incomplete input for triggering an execution of one or more tasks corresponding to the one or more intents.

17 . The method of claim 16 , wherein each of the candidate continuation-inputs comprises a complete instruction to trigger the execution of a respective task of the one or more tasks.

18 . The method of claim 17 , wherein each of the presented suggested inputs comprises a guidance to the first user for executing the complete instruction corresponding to the respective candidate continuation-input.

19 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

receive, at a client system, an initial input from a first user, wherein the initial input is in a first modality;

determine one or more intents and one or more slots corresponding to the initial input, wherein the one or more slots are conditioned on the one or more intents;

generate, based on the one or more intents and the one or more slots, one or more candidate continuation-inputs, where the one or more candidate continuation-inputs are in one or more candidate modalities, respectively, wherein the candidate modalities are different from the first modality, and wherein each of the candidate continuation-inputs references one or more entities represented by the one or more slots; and

present, at the client system, one or more suggested inputs corresponding to one or more of the candidate continuation-inputs.

20 . A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:

receive, at a client system, an initial input from a first user, wherein the initial input is in a first modality;

determine one or more intents and one or more slots corresponding to the initial input, wherein the one or more slots are conditioned on the one or more intents;

generate, based on the one or more intents and the one or more slots, one or more candidate continuation-inputs, where the one or more candidate continuation-inputs are in one or more candidate modalities, respectively, wherein the candidate modalities are different from the first modality, and wherein each of the candidate continuation-inputs references one or more entities represented by the one or more slots; and

present, at the client system, one or more suggested inputs corresponding to one or more of the candidate continuation-inputs.

Assignments (1)
CHANGE OF NAME Recorded Jul 6, 2022
From: FACEBOOK TECHNOLOGIES, LLC
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 060591/0848 →