Auto-completion for Gesture-input in Assistant Systems
In one embodiment, a method includes receiving an initial input in a first modality from a first user at a virtual-reality (VR) headset, determining intents corresponding to the initial input by an intent-understanding module, generating candidate continuation-inputs in respective candidate modalities based on the intents, wherein the candidate modalities are different from the first modality, and presenting suggested inputs corresponding to one or more of the candidate continuation-inputs at the VR headset.
1 . A method comprising, by a virtual-reality (VR) headset:
receiving, at the VR headset, an initial input from a first user, wherein the initial input is in a first modality;
determining, by an intent-understanding module, one or more intents corresponding to the initial input;
generating, based on the one or more intents, one or more candidate continuation-inputs, where the one or more candidate continuation-inputs are in one or more candidate modalities, respectively, and wherein the candidate modalities are different from the first modality; and
presenting, at the VR headset, one or more suggested inputs corresponding to one or more of the candidate continuation-inputs.
2 . The method of claim 1 , wherein the first modality comprises one of audio, text, image, video, motion, or orientation.
3 . The method of claim 2 , wherein the first modality comprises motion, and wherein the initial input comprises a gesture.
4 . The method of claim 2 , wherein the first modality comprises orientation, and wherein the initial input comprises a gaze on an object.
5 . The method of claim 1 , further comprising:
determining that the first user needs one or more suggested inputs.
6 . The method of claim 5 , wherein determining that the first user needs one or more suggested inputs is based on a wake-up input from the first user.
7 . The method of claim 6 , wherein the wake-up input comprises one or more of a voice utterance, a character string, an image, a video clip, a gesture, or a gaze.
8 . The method of claim 5 , wherein the initial input comprises a gaze on an object, and wherein determining that the first user needs one or more suggested inputs is further based on the gaze on the object.
9 . The method of claim 8 , wherein generating the one or more candidate continuation-inputs is further based on the object.
10 . The method of claim 5 , wherein determining that the first user needs one or more suggested inputs is further based on contextual information associated with the initial input.
11 . The method of claim 5 , wherein determining that the first user needs one or more suggested inputs is further based on the one or more intents.
12 . The method of claim 1 , further comprising:
identifying one or more entities associated with the one or more intents.
13 . The method of claim 12 , wherein generating the one or more candidate continuation-inputs is further based the one or more entities.
14 . The method of claim 1 , further comprising:
receiving, at the VR headset, a user-selected input from the first user, wherein the user-selected input comprises one of the suggested inputs, and wherein the user-selected input corresponds to one of the candidate continuation-inputs and is in one of the one or more candidate modalities that is different from the first modality; and
executing one or more tasks based on the user-selected input.
15 . The method of claim 1 , further comprising:
receiving, at the VR headset, a first user-selected input from the first user, wherein the first user-selected input comprises one of the suggested inputs, wherein the first user-selected input corresponds to one of the candidate continuation-inputs and is in one of the one or more candidate modalities that is different from the first modality, and wherein the first user-selected input is associated with a first intent;
generating, based on the first user-selected input, one or more additional candidate continuation-inputs, wherein each of the one or more additional candidate continuation-inputs is associated with the first intent;
presenting, at the VR headset, one or more additional suggested inputs corresponding to one or more of the additional candidate continuation-inputs;
receiving, at the VR headset, a second user-selected input from the first user, wherein the second user-selected input comprises one of the additional suggested inputs; and
executing one or more tasks based on the second user-selected input.
16 . The method of claim 1 , wherein the one or more suggested inputs comprise suggestions for inputting one or more of the candidate continuation-inputs in the one or more candidate modalities that are different from the first modality.
17 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:
receive, at a virtual-reality (VR) headset, an initial input from a first user, wherein the initial input is in a first modality;
determine, by an intent-understanding module, one or more intents corresponding to the initial input;
generate, based on the one or more intents, one or more candidate continuation-inputs, where the one or more candidate continuation-inputs are in one or more candidate modalities, respectively, and wherein the candidate modalities are different from the first modality; and
present, at the VR headset, one or more suggested inputs corresponding to one or more of the candidate continuation-inputs.
18 . A system comprising: one or more processors; and a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:
receive, at a virtual-reality (VR) headset, an initial input from a first user, wherein the initial input is in a first modality;
determine, by an intent-understanding module, one or more intents corresponding to the initial input;
generate, based on the one or more intents, one or more candidate continuation-inputs, where the one or more candidate continuation-inputs are in one or more candidate modalities, respectively, and wherein the candidate modalities are different from the first modality; and
present, at the VR headset, one or more suggested inputs corresponding to one or more of the candidate continuation-inputs.