IP Library Granted Patent US 11,915,706
Granted Patent B2
US 11,915,706 · App. 18/150,561 · Granted Feb 27, 2024

Hotword detection on multiple devices

Inventor: Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/285G10L15/01G10L15/08G10L15/22G10L15/32G10L17/22G06F3/167G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,915,706
App. No.
18/150,561
Granted
Feb 27, 2024
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for hotword detection on multiple devices are disclosed. In one aspect, a method includes the actions of receiving, by a first computing device, audio data that corresponds to an utterance. The actions further include determining a first value corresponding to a likelihood that the utterance includes a hotword. The actions further include receiving a second value corresponding to a likelihood that the utterance includes the hotword, the second value being determined by a second computing device. The actions further include comparing the first value and the second value. The actions further include based on comparing the first value to the second value, initiating speech recognition processing on the audio data.

Claims (40)

1. A computer-implemented method executed on data processing hardware of a first computing device that causes the data processing hardware to perform operations comprising:

receiving audio data that corresponds to an utterance of a voice command and a hotword preceding the voice command, the utterance of the voice command and the hotword preceding the voice command captured by the first computing device and a second computing device, the first computing device and the second computing device each configured to respond to voice commands that are preceded by the hotword;

receiving, from the second computing device after the second computing device captured the utterance of the hotword preceding the voice command, a message; and

based on the message received from the second computing device, causing the first computing device to not respond to the voice command despite receiving the audio data that corresponds to the voice command and the hotword preceding the voice command.

2. The computer-implemented method of claim 1 , wherein the message received from the second computing device indicates an audio metric related to the utterance of the hotword captured by the second computing device.

3. The computer-implemented method of claim 2 , wherein the audio metric comprises a loudness of the utterance captured by the second computing device.

4. The computer-implemented method of claim 2 , wherein the operations further comprise, after receiving the audio data that corresponds to the utterance of the voice command and the hotword preceding the voice command:

processing the audio data to generate an additional audio metric; and

transmitting an additional message to the second computing device, the additional message comprising the additional audio metric.

5. The computer-implemented method of claim 4 , wherein the additional audio metric comprises a loudness of the utterance captured by the first computing device.

6. The computer-implemented method of claim 1 , wherein the operations further comprise, in response to receiving the audio data that corresponds to the utterance of the voice command and the hotword preceding the voice command, determining that the audio data likely includes the utterance of the hotword preceding the voice command without performing automated speech recognition on the audio data.

7. The computer-implemented method of claim 6 , wherein determining that the audio data likely includes the utterance of the hotword preceding the voice command comprises:

determining a hotword score that reflects a likelihood that the audio data includes the utterance of the hotword preceding the voice command; and

determining that the hotword score satisfies a threshold.

8. The computer-implemented method of claim 1 , wherein:

receiving the audio data comprises receiving the audio data while the first computing device is in a low power mode; and

determining to not respond to the voice command comprises maintaining the first computing device in the low power mode.

9. The computer-implemented method of claim 1 , wherein the first computing device and the second computing device are in communication via a local network.

10. The computer-implemented method of claim 1 , wherein the operations further comprise determining not to perform speech recognition on the audio data based on the message received from the second computing device.

11. A first computing device comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving audio data that corresponds to an utterance of a voice command and a hotword preceding the voice command, the utterance of the voice command and the hotword preceding the voice command captured by the first computing device and a second computing device, the first computing device and the second computing device each configured to respond to voice commands that are preceded by the hotword;

receiving, from the second computing device after the second computing device captured the utterance of the hotword preceding the voice command, a message; and

based on the message received from the second computing device, causing the first computing device to not respond to the voice command despite receiving the audio data that corresponds to the voice command and the hotword preceding the voice command.

12. The first computing device of claim 11 , wherein the message received from the second computing device indicates an audio metric related to the utterance of the hotword captured by the second computing device.

13. The first computing device of claim 12 , wherein the audio metric comprises a loudness of the utterance captured by the second computing device.

14. The first computing device of claim 12 , wherein the operations further comprise, after receiving the audio data that corresponds to the utterance of the voice command and the hotword preceding the voice command:

processing the audio data to generate an additional audio metric; and

transmitting an additional message to the second computing device, the additional message comprising the additional audio metric.

15. The first computing device of claim 14 , wherein the additional audio metric comprises a loudness of the utterance captured by the first computing device.

16. The first computing device of claim 11 , wherein the operations further comprise, in response to receiving the audio data that corresponds to the utterance of the voice command and the hotword preceding the voice command, determining that the audio data likely includes the utterance of the hotword preceding the voice command without performing automated speech recognition on the audio data.

17. The first computing device of claim 16 , wherein determining that the audio data likely includes the utterance of the hotword preceding the voice command comprises:

determining a hotword score that reflects a likelihood that the audio data includes the utterance of the hotword preceding the voice command; and

determining that the hotword score satisfies a threshold.

18. The first computing device of claim 11 , wherein:

receiving the audio data comprises receiving the audio data while the first computing device is in a low power mode; and

determining to not respond to the voice command comprises maintaining the first computing device in the low power mode.

19. The first computing device of claim 11 , wherein the first computing device and the second computing device are in communication via a local network.

20. The first computing device of claim 11 , wherein the operations further comprise determining not to perform speech recognition on the audio data based on the message received from the second computing device.

Continuity (8)
Continuation 17137157 · Dec 29, 2020
Continuation 16553883 · Aug 28, 2019
Continuation 16171495 · Oct 26, 2018
Continuation 15346914 · Nov 9, 2016
Continuation 15088477 · Apr 1, 2016
Continuation 14675932 · Apr 1, 2015
Provisional Application 62061830 · Oct 9, 2014
Related Publication 20230147222A1 · May 11, 2023