IP Library › Granted Patent US 12,400,654
Granted Patent B2
US 12,400,654 · App. 18/446,420 · Granted Aug 26, 2025

Adapting hotword recognition based on personalized negatives

Inventors: Aleksandar Kracun (New York, NY); Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/22G10L15/197G10L15/30G10L17/06G10L17/24G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,654
App. No.
18/446,420
Granted
Aug 26, 2025
Kind
B2
Abstract

A method for adapting hotword recognition includes receiving audio data characterizing a hotword event detected by a first stage hotword detector in streaming audio captured by a user device. The method also includes processing, using a second stage hotword detector, the audio data to determine whether a hotword is detected by the second stage hotword detector in a first segment of the audio data. When the hotword is not detected by the second stage hotword detector, the method includes, classifying the first segment of the audio data as containing a negative hotword that caused a false detection of the hotword event in the streaming audio by the first stage hotword detector. Based on the first segment of the audio data classified as containing the negative hotword, the method includes updating the first stage hotword detector to prevent triggering the hotword event in subsequent audio data that contains the negative hotword.

Claims (74)

1. A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving audio data characterizing a hotword detected by a first stage hotword detector in a first segment of the audio data;

detecting, using a second stage hotword detector, the hotword in the first segment of the audio data, the second stage hotword detector different from the first stage hotword detector;

based on detecting the hotword using the second stage hotword detector, processing a second segment of the audio data that follows the first segment of the audio data to determine the second segment of the audio data is not indicative of a spoken query-type utterance, the second segment of the audio data characterizing one or more other terms spoken after the hotword; and

based on determining the second segment of the audio data is not indicative of the spoken query-type utterance:

determining a negative hotword confidence score classifying the first segment of the audio data as a negative hotword; and

updating the first stage hotword detector to prevent triggering a hotword event in subsequent audio data that contains the negative hotword based on the negative hotword confidence score.

2. The computer-implemented method of claim 1 , wherein the operations further comprise:

after receiving the audio data characterizing the hotword detected by the first stage hotword detector, determining that no follow-up query was provided by a user of a user device,

wherein determining the negative hotword confidence score is further based on determining that no follow-up query was provided by the user of the user device.

3. The computer-implemented method of claim 1 , wherein the operations further comprise:

after receiving the audio data characterizing the hotword detected by the first stage hotword detector, receiving a user interaction indicating user suppression of a wake-up process on a user device,

wherein determining the negative hotword confidence score is further based on the user interaction indicating user suppression of the wake-up process.

4. The computer-implemented method of claim 1 , wherein updating the first stage hotword detector to prevent triggering the hotword event in subsequent audio data comprises providing the first segment of the audio data classified as containing the negative hotword to a user device, the user device configured to retrain the first stage hotword detector using the first segment of audio data classified as containing the negative hotword.

5. The computer-implemented method of claim 4 , wherein the user device is configured to retrain the first stage hotword detector by:

storing, in memory hardware of the user device, each instance of the first segment of the audio data classified as containing the negative hotword in memory hardware of the user device; and

retraining the first stage hotword detector based on an aggregation of a number of instances of the first segment of the audio data classified as containing the negative hotword stored in the memory hardware.

6. The computer-implemented method of claim 5 , wherein the user device is further configured to, prior to retraining the first stage hotword detector:

determine that a corresponding confidence score associated each instance of the first segment of the audio data classified as containing the negative hotword fails to satisfy a negative hotword threshold score; and

determine that the number of instances exceeds a threshold number of instances.

7. The computer-implemented method of claim 1 , wherein:

updating the first stage hotword detector to prevent triggering the hotword event in subsequent audio data comprises providing the first segment of the audio data classified as containing the negative hotword to a user device, the user device configured to:

obtain an embedding representation of the first segment of the audio data; and

store, in memory hardware of the user device, the embedding representation of the first segment of the audio data; and

the user device is configured to determine when subsequent audio data characterizing the hotword event detected by the first stage hotword detector includes the negative hotword by:

computing an evaluation embedding representation for the subsequent audio data;

determining a similarity score between the embedding representation of the first segment of the audio data classified as the negative hotword and the evaluation embedding representation for the subsequent audio data; and

when the similarity score satisfies a similarity score threshold, determining that the subsequent audio data includes the negative hotword.

8. The computer-implemented method of claim 1 , wherein:

the data processing hardware resides on a server in communication with a user device.

9. The computer-implemented method of claim 1 , wherein:

the data processing hardware resides on a user device associated with the user; and

the first stage hotword detector executes on the data processing hardware.

10. The computer-implemented method of claim 1 , wherein:

the first stage hotword detector that executes on a digital signal processor (DSP) of the data processing hardware; and

the second stage hotword detector that executes on an application processor of the data processing hardware.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving audio data characterizing a hotword detected by a first stage hotword detector in a first segment of the audio data;

detecting, using a second stage hotword detector, the hotword in the first segment of the audio data, the second stage hotword detector different from the first stage hotword detector;

based on detecting the hotword using the second stage hotword detector, processing a second segment of the audio data that follows the first segment of the audio data to determine the second segment of the audio data is not indicative of a spoken query-type utterance, the second segment of the audio data characterizing one or more other terms spoken after the hotword; and

based on determining the second audio segment of the audio data is not indicative of the spoken query-type utterance:

determining a negative hotword confidence score classifying the first segment of the audio data as a negative hotword; and

updating the first stage hotword detector to prevent triggering a hotword event in subsequent audio data that contains the negative hotword based on the negative hotword confidence score.

12. The system of claim 11 , wherein the operations further comprise:

after receiving the audio data characterizing the hotword detected by the first stage hotword detector, determining that no follow-up query was provided by a user of a user device,

wherein determining the negative hotword confidence score is further based on determining that no follow-up query was provided by the user of the user device.

13. The system of claim 11 , wherein the operations further comprise:

after receiving the audio data characterizing the hotword detected by the first stage hotword detector, receiving a user interaction indicating user suppression of a wake-up process on a user device,

wherein determining the negative hotword confidence score is further based on the user interaction indicating user suppression of the wake-up process.

14. The system of claim 11 , wherein updating the first stage hotword detector to prevent triggering the hotword event in subsequent audio data comprises providing the first segment of the audio data classified as containing the negative hotword to a user device, the user device configured to retrain the first stage hotword detector using the first segment of audio data classified as containing the negative hotword.

15. The system of claim 14 , wherein the user device is configured to retrain the first stage hotword detector by:

storing, in memory hardware of the user device, each instance of the first segment of the audio data classified as containing the negative hotword in memory hardware of the user device; and

retraining the first stage hotword detector based on an aggregation of a number of instances of the first segment of the audio data classified as containing the negative hotword stored in the memory hardware.

16. The system of claim 15 , wherein the user device is further configured to, prior to retraining the first stage hotword detector:

determine that a corresponding confidence score associated each instance of the first segment of the audio data classified as containing the negative hotword fails to satisfy a negative hotword threshold score; and

determine that the number of instances exceeds a threshold number of instances.

17. The system of claim 11 , wherein:

updating the first stage hotword detector to prevent triggering the hotword event in subsequent audio data comprises providing the first segment of the audio data classified as containing the negative hotword to a user device, the user device configured to:

obtain an embedding representation of the first segment of the audio data; and

store, in memory hardware of the user device, the embedding representation of the first segment of the audio data; and

the user device is configured to determine when subsequent audio data characterizing the hotword event detected by the first stage hotword detector includes the negative hotword by:

computing an evaluation embedding representation for the subsequent audio data;

determining a similarity score between the embedding representation of the first segment of the audio data classified as the negative hotword and the evaluation embedding representation for the subsequent audio data; and

when the similarity score satisfies a similarity score threshold, determining that the subsequent audio data includes the negative hotword.

18. The system of claim 11 , wherein:

the data processing hardware resides on a server in communication with a user device.

19. The system of claim 11 , wherein:

the data processing hardware resides on a user device associated with the user; and

the first stage hotword detector executes on the data processing hardware.

20. The system of claim 11 , wherein:

the first stage hotword detector that executes on a digital signal processor (DSP) of the data processing hardware; and

the second stage hotword detector that executes on an application processor of the data processing hardware.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2023
From: KRACUN, ALEKSANDAR; SHARIFI, MATTHEW
To: GOOGLE LLC
Reel/Frame 064530/0590 →
Continuity (2)
Continuation 16953510 · Nov 20, 2020
Related Publication 20230386468A1 · Nov 30, 2023
References Cited (19)
US 9443198B1 · Vitaladevuni · 2016 [cited by applicant]
US 9443517B1 · Foerster · 2016 [cited by examiner]
US 9600231B1 · Sun et al. · 2017 [cited by applicant]
US 10204624B1 · Knudson · 2019 [cited by examiner]
US 20130080167A1 · Mozer · 2013 [cited by applicant]
US 20160125877A1 · Foerster et al. · 2016 [cited by applicant]
US 20190051297A1 · Knudson · 2019 [cited by examiner]
US 20190311715A1 · Pfeffinger · 2019 [cited by examiner]
US 20200184966A1 · Yavagal · 2020 [cited by applicant]
US 20200395013A1 · Smith · 2020 [cited by examiner]
CN 111640426A · 2020 [cited by applicant]
EP 3923272A1 · 2021 [cited by applicant]
FI 128000B · 2019 [cited by applicant]
JP 201492750A · 2014 [cited by applicant]
JP 2015520409A · 2015 [cited by applicant]
JP 2016505888A · 2016 [cited by applicant]
JP 202016784A · 2020 [cited by applicant]
WO 2016157782A1 · 2016 [cited by applicant]
JPO. Office Action relating to applicaiton No. 2023-530712 dated Aug. 26, 2024, mailed Sep. 2, 2024. [cited by applicant]