Attention system design for augmented reality systems and virtual reality systems
In one embodiment, a method includes rendering, for one or more displays of a XR display device, a first output image of an XR object within an XR environment in a field of view (FOV) of a first user, where the XR object has a first form, detecting a change in a context of the user with respect to the XR object, determining, based on the detected change in the context of the user, whether to invoke an attention system with respect to the XR object, and rendering, for the one or more displays of the XR display device, a second output image of the XR object responsive to invoking the attention system, where the XR object is morphed to have a second form indicating a first attention state, where the first attention state indicates a status of the XR object to interact with one or more first voice commands.
1 . A method comprising, by an extended reality (XR) display device:
rendering, for one or more displays of the XR display device, a first output image of a first XR object and a second XR object within an XR environment in a field of view (FOV) of a first user, wherein the first XR object and the second XR object are interactable via voice commands by the first user, and wherein the first XR object and the second XR object each have a first form from a plurality of forms each corresponding to a corresponding attention state from a plurality of attention states;
detecting an action of an avatar of the first user with respect to the first XR object;
determining, based on the detected action of the avatar of the first user with respect to the first XR object and a position of the first XR object within the field of view of the first user, to invoke an attention system with respect to the first XR object; and
rendering, for the one or more displays of the XR display device, a second output image of the first XR object responsive to invoking the attention system, wherein the second output image comprises a second form from the plurality of forms of the first XR object indicating a first attention state from the plurality of attention states and the second output image further comprises an indication identifying the first attention state and the first form of the second XR object indicating that the attention system with respect to the second XR object is not invoked,
wherein the first XR object is morphed from the first form to the second form responsive to invoking the attention system,
wherein the first attention state indicates a status of the first XR object to interact with one or more first voice commands for one or more first functions enabled by the XR display device, and the attention system is configured to provide audio-visual cues to inform the first user of a current attention state from the plurality of attention states of the first XR object.
2 . The method of claim 1 , further comprising:
receiving, by one or more microphones of the XR display device, an audio input comprising the one or more first voice commands;
processing, using a natural-language understanding (NLU) model, the audio input to identify one or more intents or one or more slots associated with the one or more first voice commands; and
executing, by the XR display device, a first task corresponding to the one or more first voice commands based on the identified the one or more intents or the one or more slots.
3 . The method of claim 1 , further comprising:
rendering, for one or more speakers of the XR display device, an audio feedback indicative of the first XR object morphing to the second form.
4 . The method of claim 1 , wherein the first attention state is a microphone on state in which the first XR object is active and is ready to interact with the first user, and a change in context further comprises an interaction of the avatar of the first user with the first XR object.
5 . The method of claim 1 , wherein the first attention state is a microphone off state in which the first XR object is not active and is not ready to interact with the first user, and a change in context further comprises movement of the avatar of the first user moving beyond a threshold distance of the first XR object.
6 . The method of claim 1 , wherein the first attention state is a listening state in which the first XR object is in a continuous state of receiving audio inputs from the first user, and a change in context further comprises movement of the avatar of the first user moving to within a threshold distance of the first XR object.
7 . The method of claim 1 , wherein the first attention state is a response state in which the first XR object is in a state of outputting a response to the first user.
8 . The method of claim 1 , wherein the first attention state is an error state in which the first XR object is unable to understand a received input from the first user.
9 . The method of claim 1 , further comprising:
rendering, for the one or more displays of the XR display device, a user interface overlay in the FOV of the first user.
10 . The method of claim 9 , wherein the user interface overlay comprises a search capability to search for the one or more first voice commands for the one or more first functions enabled by the XR display device.
11 . The method of claim 10 , wherein the one or more first voice commands are categorized in a plurality of different categories, and wherein the different categories comprise one or more of avatar movement, avatar actions, object movement, object actions, application settings, or system settings.
12 . The method of claim 10 , wherein the one or more first voice commands are labeled as one type of a plurality of different types, and wherein the different types comprise commands or queries.
13 . The method of claim 10 , wherein the one or more first voice commands available for search are based on a characteristic of the first user.
14 . The method of claim 10 , wherein the one or more first voice commands available for search are categorized based on one or more XR objects in the XR environment.
15 . The method of claim 9 , wherein the user interface overlay comprises a help hub to receive one or more queries from the first user.
16 . The method of claim 9 , wherein the user interface overlay comprises an index of the one or more first voice commands.
17 . The method of claim 1 , further comprising:
determining, based on an absence of detected action of the avatar of the first user with respect to the second XR object and a position of the second XR object within the field of view of the first user, to not invoke the attention system with respect to the second XR object.
18 . One or more computer-readable non-transitory non-volatile storage media embodying software that is operable when executed to:
render, for one or more displays of an XR display device, a first output image of first XR object and a second XR object within an XR environment in a field of view (FOV) of a first user, wherein the first XR object and the second XR object are interactable via one or more first voice commands by the first user, and wherein the first XR object and the second XR object each have a first form from a plurality of forms each corresponding to a corresponding attention state from a plurality of attention states;
detect an action of an avatar of the first user with respect to the first XR object;
determine, based on the detected action of the avatar of the first user with respect to the first XR object and a position of the first XR object within the field of view of the first user, to invoke an attention system with respect to the first XR object; and
render, for the one or more displays of the XR display device, a second output image of the first XR object responsive to invoking the attention system, wherein the second output image comprises a second form from the plurality of forms of the first XR object indicating a first attention state from the plurality of attention states and the second output image further comprises an indication identifying the first attention state and the first form of the second XR object indicating that the attention system with respect to the second XR object is not invoked,
wherein the first XR object is morphed from the first form to the second form responsive to invoking the attention system,
wherein the first attention state indicates a status of the first XR object to interact with the one or more first voice commands for one or more first functions enabled by the XR display device, and the attention system is configured to provide audio-visual cues to inform the first user of a current attention state from the plurality of attention states of the first XR object.
19 . The computer-readable non-transitory non-volatile storage media embodying software of claim 18 that is further operable when executed to:
determine, based on an absence of detected action of the avatar of the first user with respect to the second XR object and a position of the second XR object within the field of view of the first user, to not invoke the attention system with respect to the second XR object.
20 . A system comprising:
one or more processors; and
a non-transitory non-volatile memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:
render, for one or more displays of an XR display device, a first output image of a first XR object and a second XR object within an XR environment in a field of view (FOV) of a first user, wherein the first XR object and the second XR object are interactable via one or more first voice commands by the first user, and wherein the first XR object has a first form from a plurality of forms each corresponding to a corresponding attention state from a plurality of attention states;
detect an action of an avatar of the first user with respect to the first XR object;
determine, based on the detected action of the avatar of the first user with respect to the first XR object and a position of the first XR object within the field of view of the first user, to invoke an attention system with respect to the first XR object; and
render, for the one or more displays of the XR display device, a second output image of the first XR object responsive to invoking the attention system, wherein the second output image comprises a second form from the plurality of forms of the first XR object indicating a first attention state from the plurality of attention states and the second output image further comprises an indication identifying the first attention state and the first form of the second XR object indicating that the attention system with respect to the second XR object is not invoked,
wherein the first XR object is morphed from the first form to the second form responsive to invoking the attention system,
wherein the first attention state indicates a status of the first XR object to interact with the one or more first voice commands for one or more first functions enabled by the XR display device, and the attention system is configured to provide audio-visual cues to inform the first user of a current attention state from the plurality of attention states of the first XR object.