Voice Command Integration into Augmented Reality Systems and Virtual Reality Systems
In one embodiment, a method includes receiving, by a XR display device, a gesture-based input from a first user of the XR display device, processing, using a gesture-detection model, the gesture-based input to identify a first gesture, receiving, by the XR display device, an audio input from the first user, where the audio input includes a first voice command, processing, using a natural-language model, the audio input to identify one or more intents or one or more slots associated with the first voice command, determining whether the identified first gesture matches the first voice command, and executing, responsive to the identified first gesture matching the first voice command and by the XR display device, a first task corresponding to the first voice command based on the identified first gesture and the identified one or more intents or one or more slots.
1 . A method comprising, by an extended reality (XR) display device:
receiving, by the XR display device, a gesture-based input from a first user of the XR display device;
processing, using a gesture-detection model, the gesture-based input to identify a first gesture;
receiving, by the XR display device, an audio input from the first user, wherein the audio input comprises a first voice command;
processing, using a natural-language model, the audio input to identify one or more intents or one or more slots associated with the first voice command;
determining whether the identified first gesture matches the first voice command; and
executing, responsive to the identified first gesture matching the first voice command and by the XR display device, a first task corresponding to the first voice command based on the identified first gesture and the identified one or more intents or one or more slots.
2 . The method of claim 1 , further comprising:
receiving a tertiary input, wherein the tertiary input comprises one or more of a touch input, gaze input, or pose input; and
determining whether the tertiary input matches the first voice command, wherein executing the first task is further based on the tertiary input.
3 . The method of claim 2 , wherein executing the first task is further based on an order of receiving the gesture-based input, the audio input, and the tertiary input.
4 . The method of claim 1 , further comprising:
placing one or more microphones of the XR display device into a listening mode responsive to identifying the first gesture.
5 . The method of claim 1 , wherein the processing of the audio input is responsive to both identifying the first gesture and receiving the audio input from the first user.
6 . The method of claim 1 , further comprising:
rendering, for one or more displays of the XR display device, a visual feedback responsive to executing the first task.
7 . The method of claim 1 , further comprising:
rendering, for one or more displays of the XR display device, a user interface comprising a menu of one or more activatable user interface elements responsive to executing the first task.
8 . The method of claim 1 , further comprising:
rendering, for one or more displays of the XR display device, information corresponding to frequently asked questions responsive to executing the first task.
9 . The method of claim 1 , further comprising:
capturing, by one or more cameras of the XR display device, one or more images corresponding to a real-world environment of the first user; and
processing, using a machine-learning model, the one or more images to identify one or more real-world objects within the one or more images.
10 . The method of claim 9 , further comprising:
analyzing one or more of the identified real-world objects to perform the first task.
11 . The method of claim 10 , wherein the first task comprises:
responsive to analyzing one or more of the identified real-world objects, identifying text corresponding to the one or more real-world objects, wherein the text is in a first language;
translating the text from the first language to a second language; and
presenting, by the XR display device, the translated text to the first user.
12 . The method of claim 10 , wherein the first task comprises:
responsive to analyzing one or more of the identified real-world objects, identifying one or more online media corresponding to the one or more real-world objects; and
presenting, by the XR display device, content corresponding to the identified online media to the first user.
13 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:
receive, by a XR display device, a gesture-based input from a first user of the XR display device;
process, using a gesture-detection model, the gesture-based input to identify a first gesture;
receive, by the XR display device, an audio input from the first user, wherein the audio input comprises a first voice command;
process, using a natural-language model, the audio input to identify one or more intents or one or more slots associated with the first voice command;
determine whether the identified first gesture matches the first voice command; and
execute, responsive to the identified first gesture matching the first voice command and by the XR display device, a first task corresponding to the first voice command based on the identified first gesture and the identified one or more intents or one or more slots.
14 . The media of claim 13 , wherein the software is further operable when executed to:
receive a tertiary input, wherein the tertiary input comprises one or more of a touch input, gaze input, or pose input; and
determine whether the tertiary input matches the first voice command, wherein executing the first task is further based on the tertiary input.
15 . The media of claim 14 , wherein executing the first task is further based on an order of receiving the gesture-based input, the audio input, and the tertiary input.
16 . The media of claim 13 , wherein the software is further operable when executed to:
place one or more microphones of the XR display device into a listening mode responsive to identifying the first gesture.
17 . A system comprising:
one or more processors; and
a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:
receive, by a XR display device, a gesture-based input from a first user of the XR display device;
process, using a gesture-detection model, the gesture-based input to identify a first gesture;
receive, by the XR display device, an audio input from the first user, wherein the audio input comprises a first voice command;
process, using a natural-language model, the audio input to identify one or more intents or one or more slots associated with the first voice command;
determine whether the identified first gesture matches the first voice command; and
execute, responsive to the identified first gesture matching the first voice command and by the XR display device, a first task corresponding to the first voice command based on the identified first gesture and the identified one or more intents or one or more slots.
18 . The system of claim 17 , wherein the processors are further operable when executing the instructions to:
receive a tertiary input, wherein the tertiary input comprises one or more of a touch input, gaze input, or pose input; and
determine whether the tertiary input matches the first voice command, wherein executing the first task is further based on the tertiary input.
19 . The system of claim 18 , wherein executing the first task is further based on an order of receiving the gesture-based input, the audio input, and the tertiary input.
20 . The system of claim 17 , wherein the processors are further operable when executing the instructions to:
place one or more microphones of the XR display device into a listening mode responsive to identifying the first gesture.