IP Library › Granted Patent US 12,436,736
Granted Patent B2
US 12,436,736 · App. 18/397,786 · Granted Oct 7, 2025

Voice controlled UIs for AR wearable devices

Inventors: Sharon Moll (Lachen, CH); Piotr Gurgul (Hergiswil, CH)
Assignee: Snap Inc.
G06F3/167G06F3/04817G06F3/0484G10L15/16G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,436,736
App. No.
18/397,786
Granted
Oct 7, 2025
Kind
B2
Abstract

Systems, methods, and computer readable media for voice-controlled user interfaces (UIs) for augmented reality (AR) wearable devices are disclosed. Embodiments are disclosed that enable a user to interact with the AR wearable device without using physical user interface devices. An application has a non-voice-controlled UI mode and a voice-controlled UI mode. The user selects the mode of the UI. The application running on the AR wearable device displays UI elements on a display of the AR wearable device. The UI elements have types. Predetermined actions are associated with each of the UI element types. The predetermined actions are displayed with other information and used by the user to invoke the corresponding UI element.

Claims (77)

1. An augmented reality (AR), mixed reality (MR), or virtual reality (VR) wearable device comprising:

one or more processors; and

a memory storing instructions that, when executed by the one or more processors, configure the one or more processors to perform operations comprising:

sending a notification to an application that voice user interface (UI) mode is enabled;

receiving, from the application, an indication of a UI element, the UI element indicating a tag and a UI element type;

retrieving an action based on the UI element type from a data structure associating actions with UI element types;

causing to be displayed on a display of the AR, MR, or VR wearable device, the UI element and a voice UI label for the UI element, the voice UI label comprising the action and the tag, the action indicating how to activate the UI element with the tag;

accessing audio data captured from a microphone of the AR, MR, or VR wearable device, the audio data generated by utterances of a user;

processing the audio data to determine whether the audio data comprises the displayed action and the displayed tag; and

in response to the audio data comprising the displayed action and the displayed tag, sending, to the application, an indication that the UI element is activated.

2. The AR, MR, or VR wearable device of claim 1 , further comprising:

disabling the voice UI mode; and

removing the voice UI label for the UI element.

3. The AR, MR, or VR wearable device of claim 1 , wherein the processing the audio data further comprises:

sending the audio data to a host computer with an instruction to process the audio data; and

receiving a transcription of the audio data from the host computer.

4. The AR, MR, or VR wearable device of claim 1 , wherein the processing further comprising:

processing, using a first neural network trained to recognize the actions, a first portion of the audio data to determine the audio data comprises the action; and

processing, using a machine learning model, a second portion of the audio data to determine the audio data comprises the tag.

5. The AR, MR, or VR wearable device of claim 1 , wherein the UI element type is one of a button, a slider, a text field, a check box, or an option group, wherein the tag is text.

6. The AR, MR, or VR wearable device of claim 1 , wherein the tag is a first tag, the UI element type is a first UI element type, the UI element is a first UI element, and wherein the operations further comprise:

determining a second tag and a second UI element type for a second UI element displayed on the display by a legacy application;

associating the second tag with a second action corresponding to the second UI element type;

capturing, from a microphone of the AR, MR, or VR wearable device, audio data, the audio data generated by utterances of a user;

processing the audio data to determine that the audio data comprises the second action and the second tag; and

sending, to the legacy application, a second indication that the second UI element is activated.

7. The AR, MR, or VR wearable device of claim 1 , wherein the operations further comprise:

capturing, from a microphone, audio data, the audio data generated by utterances of a user;

processing the audio data to determine the audio data indicates a command to enable voice UI mode; and

enabling the voice UI mode.

8. The AR, MR, or VR wearable device of claim 7 , wherein the audio data is processed with a first machine learning model and the first audio data is processed by a second machine learning model.

9. The AR, MR, or VR wearable device of claim 1 , wherein the voice UI label further comprises:

tags of selection buttons, the UI element comprising the selection buttons.

10. The AR, MR, or VR wearable device of claim 1 , wherein the operations further comprise:

displaying, on the display of the AR, MR, or VR wearable device, an icon to indicate that the user is to speak the voice UI label to select the UI element.

11. The AR, MR, or VR wearable device of claim 1 , wherein the operations further comprise:

receiving, from an application, indications of UI elements, the UI elements indicating tags and UI element types;

associating the tags with actions corresponding to the UI element types;

capturing, from a microphone of the AR, MR, or VR wearable device, audio data, the audio data generated by utterances of the user;

processing the audio data to generate a transcription of the audio data;

matching the tags and actions with the transcription; and

in response to multiple tags and actions matching the transcription, displaying on the display for each of the multiple tags and actions an indication of what the user should say to select a UI element corresponding to a tag and an action of the multiple tags and actions.

12. The AR, MR, or VR wearable device of claim 1 , wherein the operations further comprise:

performing an event associated with the UI element.

13. The AR, MR, or VR wearable device of claim 1 , wherein the operations further comprise:

receiving, from the application, indications of UI elements; and

sending, to the application, predetermined actions corresponding to the UI elements.

14. The AR, MR, or VR wearable device of claim 1 , wherein the operations further comprise:

switching from a normal display mode to a voice UI mode, wherein in the normal display mode the UI element is displayed without the voice UI label, and in the voice UI mode the UI element is displayed with the voice UI label.

15. A method performed on an augmented reality (AR), mixed reality (MR), or virtual reality (VR) wearable device, the method comprising:

sending a notification to an application that voice user interface (UI) mode is enabled;

receiving, from the application, an indication of a UI element, the UI element indicating a tag and a UI element type;

retrieving an action based on the UI element type from a data structure associating actions with UI element types;

causing to be displayed on a display of the AR, MR, or VR wearable device, the UI element and a voice UI label for the UI element, the voice UI label comprising the action and the tag, the action indicating how to activate the UI element with the tag;

accessing audio data captured from a microphone of the AR, MR, or VR wearable device, the audio data generated by utterances of a user;

processing the audio data to determine whether the audio data comprises the displayed action and the displayed tag; and

in response to the audio data comprising the displayed action and the displayed tag, sending, to the application, an indication that the UI element is activated.

16. The method of claim 15 , further comprising:

accessing audio data captured from a microphone of the AR, MR, or VR wearable device, the audio data generated by utterances of a user;

processing the audio data to determine that the audio data comprises the action and the tag; and

sending, to the application, an indication that the UI element is activated.

17. The method of claim 15 , wherein the method further comprises:

switching from a normal display mode to a voice UI mode, wherein in the normal display mode the UI element is displayed without the voice UI label, and in the voice UI mode the UI element is displayed with the voice UI label.

18. A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by an apparatus of an augmented reality (AR), mixed reality (MR), or virtual reality (VR) wearable device, cause the apparatus of the AR, MR, or VR wearable device to perform operations comprising:

sending a notification to an application that voice user interface (UI) mode is enabled;

receiving, from the application, an indication of a UI element, the UI element indicating a tag and a UI element type;

retrieving an action based on the UI element type from a data structure associating actions with UI element types;

causing to be displayed on a display of the AR, MR, or VR wearable device, the UI element and a voice UI label for the UI element, the voice UI label comprising the action and the tag, the action indicating how to activate the UI element with the tag;

accessing audio data captured from a microphone of the AR, MR, or VR wearable device, the audio data generated by utterances of a user;

processing the audio data to determine whether the audio data comprises the displayed action and the displayed tag; and

in response to the audio data comprising the displayed action and the displayed tag, sending, to the application, an indication that the UI element is activated.

19. The non-transitory computer-readable storage medium of claim 18 , wherein the operations further comprise:

accessing audio data captured from a microphone of the AR, MR, or VR wearable device, the audio data generated by utterances of a user;

processing the audio data to determine that the audio data comprises the action and the tag; and

sending, to the application, an indication that the UI element is activated.

20. The non-transitory computer-readable storage medium of claim 18 , wherein the operations further comprise:

switching from a normal display mode to a voice UI mode, wherein in the normal display mode the UI element is displayed without the voice UI label, and in the voice UI mode the UI element is displayed with the voice UI label.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2023
From: MOLL, SHARON; GURGUL, PIOTR
To: SNAP INC.
Reel/Frame 065964/0421 →
Continuity (2)
Continuation 17823169 · Aug 30, 2022
Related Publication 20240126502A1 · Apr 18, 2024
References Cited (25)
US 9135914B1 · Bringert et al. · 2015 [cited by applicant]
US 10157042B1 · Jayakumar · 2018 [cited by examiner]
US 11922096B1 · Moll et al. · 2024 [cited by applicant]
US 20030156130A1 · James · 2003 [cited by examiner]
US 20130173270A1 · Han et al. · 2013 [cited by applicant]
US 20130288753A1 · Jacobsen · 2013 [cited by examiner]
US 20140196087A1 · Park et al. · 2014 [cited by applicant]
US 20150199017A1 · Murillo · 2015 [cited by examiner]
US 20160225371A1 · Agrawal · 2016 [cited by examiner]
US 20180286402A1 · Lebeau et al. · 2018 [cited by applicant]
US 20190103109A1 · Du · 2019 [cited by examiner]
US 20220093098A1 · Samal · 2022 [cited by examiner]
US 20240069856A1 · Moll et al. · 2024 [cited by applicant]
CN 119895379A · 2025 [cited by applicant]
WO WO2024049696A1 · 2024 [cited by applicant]
“U.S. Appl. No. 17/823,169, Amendment Under 37 C.F.R. § 1.312 filed Dec. 27, 2023”, 7 pgs. [cited by applicant]
“U.S. Appl. No. 17/823,169, Non Final Office Action mailed Mar. 30, 2023”, 35 pgs. [cited by applicant]
“U.S. Appl. No. 17/823,169, Notice of Allowance mailed Sep. 27, 2023”, 8 pgs. [cited by applicant]
“U.S. Appl. No. 17/823,169, Response filed Jun. 27, 2023 to Non Final Office Action mailed Mar. 30, 2023”, 9 pgs. [cited by applicant]
“International Application Serial No. PCT/US2023/031018, International Search Report mailed Dec. 13, 2023”, 3 pgs. [cited by applicant]
“International Application Serial No. PCT/US2023/031018, Written Opinion mailed Dec. 13, 2023”, 4 pgs. [cited by applicant]
“U.S. Appl. No. 17/823,169, Notice of Allowability mailed Feb. 6, 2024”, 3 pgs. [cited by applicant]
“U.S. Appl. No. 17/823,169, PTO Response to Rule 312 Communication mailed Jan. 17, 2024”, 2 pgs. [cited by applicant]
“International Application Serial No. PCT/US2023/031018, International Preliminary Report on Patentability mailed Mar. 13, 2025”, 6 pgs. [cited by applicant]
U.S. Appl. No. 17/823,169, filed Aug. 30, 2022, Voice Controlled UIS for AR Wearable Devices. [cited by applicant]