IP Library Granted Patent US 10,891,953
Granted Patent B2
US 10,891,953 · App. 16/216,516 · Granted Jan 12, 2021

Multi-mode guard for voice commands

Inventors: Michael J. LeBeau (New York, NY); Mat Balez (San Francisco, CA)
Assignee: GOOGLE LLC
G10L15/22G10L2015/088G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,891,953
App. No.
16/216,516
Granted
Jan 12, 2021
Kind
B2
Abstract

Embodiments may be implemented by a computing device, such as a head-mountable display, in order to use a single guard phrase to enable different voice commands in different interface modes. An example device includes an audio sensor and a computing system configured to analyze audio data captured by the audio sensor to detect speech that includes a predefined guard phrase, and to operate in a plurality of different interface modes comprising at least a first and a second interface mode. During operation in the first interface mode, the computing system may initially disable one or more first-mode speech commands, and respond to detection of the guard phrase by enabling the one or more first-mode speech commands. During operation in the second interface mode, the computing system may initially disable a second-mode speech command, and to respond to the guard phrase by enabling the second-mode speech command.

Claims (57)

1. A method implemented by one or more processors of a computing device, the method comprising:

receiving, at the computing device, incoming data, wherein the incoming data is one of: a notification, an incoming message, or an incoming call;

in response to receiving the incoming data:

causing the computing device to transition to a given interface mode associated with the incoming data, wherein one or more speech commands associated with the given interface mode are disabled when the computing device transitions to the given interface mode;

receiving, at the computing device, and while the computing device is in the given interface mode, a predefined guard phrase;

in response to receiving the predefined guard phrase while the computing device is in the given interface mode:

activating one or more speech commands that are specific to the given interface mode,

wherein the given interface mode is one of multiple interface modes each corresponding to corresponding alternate states of the user interface; and

displaying a visual cue for one or more of the speech commands; and

subsequent to displaying the visual cue for one or more of the speech commands:

receiving audio data captured by at least one audio sensor of the computing device;

analyzing, based on the one or more speech commands being activated, the audio data to determine whether any of the one or more speech commands are included in the audio data; and

in response to determining a given speech command, of the one or more speech commands, is included in the audio data:

performing one or more actions, via the computing device, that correspond to the given speech command.

2. The method of claim 1 , wherein activating one or more of the speech commands is based on the incoming data.

3. The method of claim 1 ,

wherein activating the one or more speech commands that are specific to the given interface mode comprises loading, at the computing device, at least one hotword process for the one or more speech commands; and

wherein analyzing, based on the one or more speech commands being activated, the audio data to determine whether any of the one or more speech commands are included in the audio data comprises:

analyzing the audio data using the at least one hotword process.

4. The method of claim 1 , wherein the computing device is a head-mountable device.

5. A computing device including memory and one or more processors configured to execute instructions stored in the memory, comprising instructions to:

receive, at the computing device, incoming data, wherein the incoming data is one of: a notification, an incoming message, or an incoming call;

in response to receiving the incoming data:

cause the computing device to transition to a given interface mode associated with the incoming data, wherein one or more speech commands associated with the given interface mode are disabled when the computing device transitions to the given interface mode;

receive, at the computing device, and while the computing device is in the given interface mode, a predefined guard phrase;

in response to receiving the predefined guard phrase while the computing device is in the given interface mode:

activate one or more speech commands that are specific to the given interface mode,

wherein the given interface mode is one of multiple interface modes each corresponding to corresponding alternate states of the user interface; and

display a visual cue for one or more of the speech commands; and

subsequent to displaying the visual cue for one or more of the speech commands:

receive audio data captured by at least one audio sensor of the computing device;

analyze, based on the one or more speech commands being activated, the audio data to determine whether any of the one or more speech commands are included in the audio data; and

in response to determining a given speech command, of the one or more speech commands, is included in the audio data:

perform one or more actions, via the computing device, that correspond to the given speech command.

6. The computing device of claim 5 , wherein activating one or more of the speech commands is based on the incoming data.

7. The computing device of claim 5 ,

wherein the instructions to activate the one or more speech commands that are specific to the given interface mode comprise instructions to load, at the computing device, at least one hotword process for the one or more speech commands; and

wherein the instructions to analyze, based on the one or more speech commands being activated, the audio data to determine whether any of the one or more speech commands are included in the audio data comprise instructions to:

analyze the audio data using the at least one hotword process.

8. The computing device of claim 5 , wherein the computing device is a head-mountable device.

9. A non-transitory computer readable storage medium storing instructions executable by a processor, the instructions including instructions to:

receive, at the computing device, incoming data, wherein the incoming data is one of: a notification, an incoming message, or an incoming call;

in response to receiving the incoming data:

cause the computing device to transition to a given interface mode associated with the incoming data, wherein one or more speech commands associated with the given interface mode are disabled when the computing device transitions to the given interface mode;

receive, at the computing device, and while the computing device is in the given interface mode, a predefined guard phrase;

in response to receiving the predefined guard phrase while the computing device is in the given interface mode:

activate one or more of the speech commands that are specific to the given interface mode,

wherein the given interface mode is one of multiple interface modes each corresponding to corresponding alternate states of the user interface; and

display a visual cue for one or more of the speech commands; and

subsequent to displaying the visual cue for one or more of the speech commands:

receive audio data captured by at least one audio sensor of the computing device;

analyze, based on the one or more speech commands being activated, the audio data to determine whether any of the one or more speech commands are included in the audio data; and

in response to determining a given speech command, of the one or more speech commands, is included in the audio data:

perform one or more actions, via the computing device, that correspond to the given speech command.

10. The method of claim 1 , further comprising:

capturing, via the at least one audio sensor of the computing device, additional audio while the computing device is in a given one of the corresponding alternate states of the user interface; and

performing one or more further actions based on an additional speech command that is determined based on processing the additional audio data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2018
From: LEBEAU, MICHAEL J.; BALEZ, MAT
To: GOOGLE INC.
Reel/Frame 047760/0336 →
CHANGE OF NAME Recorded Dec 12, 2018
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 047815/0324 →
Continuity (3)
Continuation 15371027 · Dec 6, 2016
Continuation 13859588 · Apr 9, 2013
Related Publication 20190237073A1 · Aug 1, 2019
Cited By (1)
US 12,293,762