IP Library Granted Patent US 12,293,762
Granted Patent B2
US 12,293,762 · App. 18/083,315 · Granted May 6, 2025

Multi-mode guard for voice commands

Inventors: Michael J. LeBeau (New York, NY); Mat Balez (San Francisco, CA)
Assignee: GOOGLE LLC
G10L15/22G10L2015/088G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,293,762
App. No.
18/083,315
Granted
May 6, 2025
Kind
B2
Abstract

Embodiments may be implemented by a computing device, such as a head-mountable display, in order to use a single guard phrase to enable different voice commands in different interface modes. An example device includes an audio sensor and a computing system configured to analyze audio data captured by the audio sensor to detect speech that includes a predefined guard phrase, and to operate in a plurality of different interface modes comprising at least a first and a second interface mode. During operation in the first interface mode, the computing system may initially disable one or more first-mode speech commands, and respond to detection of the guard phrase by enabling the one or more first-mode speech commands. During operation in the second interface mode, the computing system may initially disable a second-mode speech command, and to respond to the guard phrase by enabling the second-mode speech command.

Claims (88)

1. A method implemented by one or more processors associated with a head mounted display, the method comprising:

transitioning the head mounted display into a second interface mode from a first interface mode in response to receiving a spoken utterance from a user captured by at least one sensor of the head mounted display,

wherein the first interface mode has a first set of supported spoken commands and the second interface mode has a second set of supported spoken commands, the first set of supported spoken commands different than the second set of supported spoken commands;

wherein the second interface mode is operable to process the second set of supported spoken commands specific to the second interface mode;

enabling, in response to transitioning the head mounted display into the second interface mode, the second set of supported spoken commands that are specific to the second interface mode,

wherein the second interface mode corresponds to a current state of a user interface of the head mounted display; and

wherein the second interface mode and the first interface mode are each one of multiple interface modes each corresponding to alternate states of the user interface of the head mounted display, each of the multiple interface modes having a different set of supported spoken commands;

while the head mounted display is in the second interface mode, and responsive to enabling the second set of supported spoken commands;

receiving audio data captured by at least one audio sensor of the head mounted display;

analyzing the audio data to determine whether any of the second set of supported spoken commands that are specific to the second interface mode, are included in the audio data;

determining whether the audio data contains a given spoken command, of the second set of supported spoken commands and that the audio data is received within a threshold period of time; and

causing, in response to determining the audio data being received within the threshold period of time contains the given spoken command, performance of one or more actions, via the head mounted display, that correspond to the given spoken command.

2. The method of claim 1 , further comprising:

while the head mounted display is in the second interface mode and while the second set of supported spoken commands are enabled:

determining the threshold period of time has lapsed; and

in response to determining the threshold period of time has lapsed, disabling the second set of supported spoken commands that are specific to the second interface mode.

3. The method of claim 2 , further comprising:

while the head mounted display is in the second interface mode and while the second set of supported spoken commands are disabled:

in response to determining the given spoken command, of the second set of supported spoken commands, is included in the audio data:

refraining from performing one or more of the actions that correspond to the given spoken command.

4. The method of claim 3 , further comprising:

subsequent to refraining from performing one or more of the actions that correspond to the given spoken command:

receiving additional audio data captured by the at least one audio sensor of the head mounted display; and

re-enabling the second set of supported spoken commands that are specific to the second interface mode.

5. The method of claim 1 , further comprising:

while the head mounted display is in the second interface mode and while the second set of supported spoken commands are enabled:

displaying a visual cue for one or more of the second set of supported spoken commands via a display of the head mounted display.

6. The method of claim 5 , further comprising:

while the head mounted display is in the second interface mode and while the second set of supported spoken commands are enabled:

determining the threshold period of time has lapsed; and

in response to determining the threshold period of time has lapsed, causing the visual cue for the second set of supported spoken commands to be removed from the display of the head mounted display.

7. The method of claim 1 ,

wherein enabling the second set of supported spoken commands that are specific to the second interface mode comprises loading, at the head mounted display, at least one hotword process for the second set of supported spoken commands; and

wherein analyzing the audio data to determine whether any of the second set of supported spoken commands are included in the audio data comprises analyzing the audio data using the at least one hotword process.

8. The method of claim 1 , wherein in addition to the spoken utterance, the at least one sensor of the head mounted display captures preceding audio data, the preceding audio data including at least a predefined guard phrase.

9. The method of claim 1 , wherein the threshold period of time is specific to the second interface mode.

10. The method of claim 1 , wherein enabling the second set of supported spoken commands is based on a first instance of audio data.

11. The method of claim 1 , wherein the head mounted display includes a microphone.

12. The method of claim 1 wherein the first interface mode is a home screen and the second interface mode is a phone.

13. The method of claim 1 wherein the first interface mode is a phone and the second interface mode is a camera.

14. The method of claim 1 wherein head mounted display communicates with at least one computing device remote from the head mounted display.

15. A head mounted display comprising memory and one or more processors configured to execute instructions stored in memory, the instructions operable to:

transition the head mounted display into a second interface mode from a first interface mode in response to receiving a spoken utterance from a user captured by at least one sensor of the head mounted display,

wherein the first interface mode has a first set of supported spoken commands and the second interface mode has a second set of supported spoken commands, the first set of supported spoken commands different than the second set of supported spoken commands

wherein the second interface mode is operable to process the second set of supported spoken commands specific to the second interface mode;

enable, in response to transitioning the head mounted display into the second interface mode, the second set of supported spoken commands that are specific to the second interface mode,

wherein the second interface mode corresponds to a current state of a user interface of the head mounted display; and

wherein the second interface mode and the first interface mode are each one of multiple interface modes each corresponding to alternate states of the user interface of the head mounted display, each of the multiple interface modes having a different set of supported spoken commands;

while the head mounted display is in the second interface mode, and responsive to the second set of supported spoken commands being enabled:

receive audio data captured by at least one audio sensor of the head mounted display;

analyze the audio data to determine whether any of the second set of supported spoken commands that are specific to the second interface mode are included in the audio data;

determine whether the audio data contains a given spoken command, of the second set of supported spoken commands and that the audio data is received within a threshold period of time; and

cause, in response to the determination that the audio data, received within the threshold period of time contains the given spoken command, performance of one or more actions, via the head mounted display, that correspond to the given spoken command.

16. The head mounted display of claim 15 , wherein the instructions further comprise instructions to:

while the head mounted display is in the second interface mode and while the second set of supported spoken commands are enabled:

determine the threshold period of time has lapsed; and

in response to determining the threshold period of time has lapsed, disable the second set of supported spoken commands that are specific to the second interface mode.

17. The head mounted display of claim 16 , wherein the instructions further comprise instructions to:

while the head mounted display is in the second interface mode and while the second set of supported spoken commands are disabled:

in response to the determination that the given spoken command, of the second set of supported spoken commands, is included in the audio data:

refrain from performing one or more of the actions that correspond to the given spoken command.

18. The head mounted display of claim 17 , wherein the instructions further comprise instructions to:

subsequent to refraining from performing one or more of the actions that correspond to the given spoken command:

receive additional audio data captured by the at least one audio sensor of the head mounted display; and

re-enable the second set of supported spoken commands that are specific to the second interface mode.

19. The head mounted display of claim 15 , wherein the instructions further comprise instructions to:

while the head mounted display is in the second interface mode and while the second set of supported spoken commands are enabled:

display a visual cue for one or more of the second set of supported spoken commands via a display of the head mounted display.

20. The head mounted display of claim 19 , wherein the instructions further comprise instructions to:

while the head mounted display is in the second interface mode and while the second set of supported spoken commands are enabled:

determine the threshold period of time has lapsed; and

in response to determining the threshold period of time has lapsed, cause the visual cue for the second set of supported spoken commands to be removed from the display of the head mounted display.

21. The head mounted display of claim 15 ,

wherein the instructions to enable the second set of supported commands that are specific to the second interface mode comprise instructions to load, at the head mounted display, at least one hotword process for the second set of supported spoken commands; and

wherein the instructions to analyze the audio data to determine whether any of the second set of supported spoken commands are included in the audio data comprise instructions to analyze the audio data using the at least one hotword process.

22. The head mounted display of claim 15 , wherein in addition to the spoken utterance, the at least one sensor of the head mounted display captures preceding audio data, the preceding audio data including at least a predefined guard phrase.

23. A non-transitory computer-readable storage medium comprising instructions executable by at least one processor, the instructions operable to:

transition a head mounted display into a second interface mode from a first interface mode in response to receiving a spoken utterance from a user captured by at least one sensor of the head mounted display,

wherein the first interface mode has a first set of supported spoken commands and the second interface mode has a second set of supported spoken commands, the first set of supported spoken commands different than the second set of supported spoken commands;

wherein the second interface mode is operable to process the second set of supported spoken commands specific to the second interface mode;

enable, in response to transitioning the head mounted display into the second interface mode, the second set of supported spoken commands that are specific to the second interface mode,

wherein the second interface mode corresponds to a current state of a user interface of the head mounted display; and

wherein the second interface mode and the first interface mode are each one of multiple interface modes each corresponding to alternate states of the user interface of the head mounted display, each of the multiple interface modes having a different set of supported spoken commands;

while the head mounted display is in the second interface mode, and responsive to the second set of supported spoken commands being enabled:

receive audio data captured by at least one audio sensor of the head mounted display;

analyze the audio data to determine whether any of the second set of supported spoken commands that are specific to the second interface mode are included in the audio data;

determine whether the audio data contains a given spoken command, of the second set of supported spoken commands and that the audio data is received within a threshold period of time; and

cause, in response to the determination that the audio data, received within the threshold period of time, contains the given spoken command, performance of one or more actions, via the head mounted display, that correspond to the given spoken command.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2023
From: LEBEAU, MICHAEL J.; BALEZ, MAT
To: GOOGLE INC.
Reel/Frame 062494/0160 →
CHANGE OF NAME Recorded Jan 26, 2023
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 062513/0550 →
Continuity (5)
Continuation 17109265 · Dec 2, 2020
Continuation 16216516 · Dec 11, 2018
Continuation 15371027 · Dec 6, 2016
Continuation 13859588 · Apr 9, 2013
Related Publication 20230120601A1 · Apr 20, 2023
References Cited (72)
US 6081782A · Rabin · 2000 [cited by applicant]
US 6127990A · Zwern · 2000 [cited by applicant]
US 6233560B1 · Tannenbaum · 2001 [cited by applicant]
US 6456977B1 · Wang · 2002 [cited by applicant]
US 6928614B1 · Everhart · 2005 [cited by applicant]
US 7023498B2 · Ishihara · 2006 [cited by applicant]
US 7324947B2 · Jordan et al. · 2008 [cited by applicant]
US 7720683B1 · Vermeulen et al. · 2010 [cited by applicant]
US 8019606B2 · Badger et al. · 2011 [cited by applicant]
US 8161172B2 · Reisman · 2012 [cited by applicant]
US 8165886B1 · Gagnon et al. · 2012 [cited by applicant]
US 8195467B2 · Mozer et al. · 2012 [cited by applicant]
US 8234119B2 · Dhawan et al. · 2012 [cited by applicant]
US 8452602B1 · Bringert et al. · 2013 [cited by applicant]
US 8515766B1 · Bringert et al. · 2013 [cited by applicant]
US 8620667B2 · Andrew · 2013 [cited by applicant]
US 8650036B2 · Han · 2014 [cited by examiner]
US 8924219B1 · Bringert et al. · 2014 [cited by applicant]
US 8977255B2 · Freeman · 2015 [cited by examiner]
US 9275637B1 · Salvador et al. · 2016 [cited by applicant]
US 9530410B1 · LeBeau · 2016 [cited by examiner]
US 9686595B2 · Fabian-Issacs et al. · 2017 [cited by applicant]
US 10181324B2 · LeBeau · 2019 [cited by examiner]
US 10891953B2 · LeBeau · 2021 [cited by examiner]
US 20020193989A1 · Geilhufe et al. · 2002 [cited by applicant]
US 20030005076A1 · Koch et al. · 2003 [cited by applicant]
US 20030028382A1 · Chambers et al. · 2003 [cited by applicant]
US 20040225499A1 · Wang et al. · 2004 [cited by applicant]
US 20060074658A1 · Chadha · 2006 [cited by applicant]
US 20070069976A1 · Willins et al. · 2007 [cited by applicant]
US 20070213984A1 · Ativanichayaphong et al. · 2007 [cited by applicant]
US 20080065486A1 · Vincent et al. · 2008 [cited by applicant]
US 20090177477A1 · Nenov et al. · 2009 [cited by applicant]
US 20090262205A1 · Smith · 2009 [cited by applicant]
US 20090328101A1 · Suomela et al. · 2009 [cited by applicant]
US 20100031150A1 · Andrew · 2010 [cited by applicant]
US 20100076850A1 · Parekh et al. · 2010 [cited by applicant]
US 20100079356A1 · Hoellwarth · 2010 [cited by applicant]
US 20100312547A1 · Van Os et al. · 2010 [cited by applicant]
US 20100332236A1 · Tan · 2010 [cited by applicant]
US 20110075818A1 · Vance et al. · 2011 [cited by applicant]
US 20110187640A1 · Jacobsen et al. · 2011 [cited by applicant]
US 20120182205A1 · Gamst · 2012 [cited by applicant]
US 20120235896A1 · Jacobsen et al. · 2012 [cited by applicant]
US 20120253824A1 · Talavera · 2012 [cited by applicant]
US 20130018659A1 · Chi · 2013 [cited by applicant]
US 20130035941A1 · Kim et al. · 2013 [cited by applicant]
US 20130090930A1 · Monson et al. · 2013 [cited by applicant]
US 20130226591A1 · Ahn et al. · 2013 [cited by applicant]
US 20130253937A1 · Cho et al. · 2013 [cited by applicant]
US 20130275875A1 · Gruber et al. · 2013 [cited by applicant]
US 20130289986A1 · Graylin · 2013 [cited by applicant]
US 20130325484A1 · Chakladar et al. · 2013 [cited by applicant]
US 20130339028A1 · Rosner et al. · 2013 [cited by applicant]
US 20140006022A1 · Yoon et al. · 2014 [cited by applicant]
US 20140074481A1 · Newman · 2014 [cited by applicant]
US 20140122087A1 · Macho · 2014 [cited by applicant]
US 20140201639A1 · Savolainen et al. · 2014 [cited by applicant]
US 20140222436A1 · Binder et al. · 2014 [cited by applicant]
US 20140244253A1 · Bringert et al. · 2014 [cited by applicant]
US 20140270258A1 · Wang · 2014 [cited by applicant]
US 20140278419A1 · Bishop et al. · 2014 [cited by applicant]
US 20140337028A1 · Wang et al. · 2014 [cited by applicant]
US 20150279839A1 · Lebeau et al. · 2015 [cited by applicant]
US 20210082435A1 · LeBeau · 2021 [cited by examiner]
US 20210314282A1 · Sharma · 2021 [cited by examiner]
WO 2009013518 · 2009 [cited by applicant]
WO 2012040086 · 2012 [cited by applicant]
Bringert et al., U.S. Appl. No. 13/620,987, filed Sep. 15, 2012. 29 pages. [cited by applicant]
Lebeau et al., U.S. Appl. No. 13/622,180, filed Sep. 18, 2012. 31 pages. [cited by applicant]
Office Action for U.S. Appl. No. 13/754,488 mailed Mar. 23, 2015. 14 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 13/754,488 mailed Jul. 13, 2015. 9 pages. [cited by applicant]