IP Library Granted Patent US 12,394,416
Granted Patent B2
US 12,394,416 · App. 18/384,764 · Granted Aug 19, 2025

Detecting near matches to a hotword or phrase

Inventors: Matthew Sharifi (Kilchberg, CH); Victor Carbune (Zurich, CH)
Assignee: GOOGLE LLC
G10L15/22G10L15/08G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,416
App. No.
18/384,764
Granted
Aug 19, 2025
Kind
B2
Abstract

Techniques are described herein for identifying a failed hotword attempt. A method includes: receiving first audio data; processing the first audio data to generate a first predicted output; determining that the first predicted output satisfies a secondary threshold but does not satisfy a primary threshold; receiving second audio data; processing the second audio data to generate a second predicted output; determining that the second predicted output satisfies the secondary threshold but does not satisfy the primary threshold; in response to the first predicted output and the second predicted output satisfying the secondary threshold but not satisfying the primary threshold, and in response to the first spoken utterance and the second spoken utterance satisfying one or more temporal criteria relative to one another, identifying a failed hotword attempt; and in response to identifying the failed hotword attempt, providing a hint that is responsive to the failed hotword attempt.

Claims (59)

1. A method implemented by one or more processors, the method comprising:

receiving, via one or more microphones of a client device, first audio data that captures a first spoken utterance of a user;

processing the first audio data using one or more machine learning models to generate a first predicted output that indicates a probability of one or more hotwords being present in the first audio data;

determining that the first predicted output satisfies a secondary threshold that is less indicative of the one or more hotwords being present in audio data than is a primary threshold but does not satisfy the primary threshold;

receiving, via the one or more microphones of the client device, second audio data that captures a second spoken utterance of a user;

processing the second audio data using the one or more machine learning models to generate a second predicted output that indicates a probability of the one or more hotwords being present in the second audio data;

determining that the second predicted output satisfies the secondary threshold but does not satisfy the primary threshold;

in response to the first predicted output and the second predicted output satisfying the secondary threshold but not satisfying the primary threshold, and in response to the first spoken utterance and the second spoken utterance satisfying one or more temporal criteria relative to one another, identifying a failed hotword attempt; and

in response to identifying the failed hotword attempt:

determining an intended hotword corresponding to the failed hotword attempt, wherein neither the intended hotword nor another supported hotword is included in the first spoken utterance and the second spoken utterance;

providing a hint, comprising displaying the intended hotword on a display of the client device or providing, by the client device, an audio response that includes the intended hotword; and

performing an action corresponding to the intended hotword.

2. The method according to claim 1 , wherein identifying the failed hotword attempt is further in response to determining that a similarity between the first spoken utterance and the second spoken utterance exceeds a similarity threshold.

3. The method according to claim 1 , wherein identifying the failed hotword attempt is further in response to determining that the probability indicated by the first predicted output and the probability indicated by the second predicted output correspond to a same hotword of the one or more hotwords.

4. The method according to claim 1 , further comprising determining, using a model conditioned on acoustic features, that the first audio data and the second audio data comprise a command,

wherein identifying the failed hotword attempt is further in response to the first audio data and the second audio data comprising the command.

5. The method according to claim 1 , wherein the intended hotword is determined based on acoustic similarity between at least a portion of the first audio data, at least a portion of the second audio data, and the intended hotword.

6. The method according to claim 1 , wherein the intended hotword is determined based on semantic similarity between the intended hotword and text generated based on at least a portion of the first audio data or based on at least a portion of the second audio data.

7. The method according to claim 1 , further comprising, in response to identifying the failed hotword attempt, updating the one or more machine learning models based on the first audio data and the second audio data.

8. A computer program product comprising one or more non-transitory computer-readable storage media having program instructions collectively stored on the one or more non-transitory computer-readable storage media, the program instructions executable to:

receive, via one or more microphones of a client device, first audio data that captures a first spoken utterance of a user;

process the first audio data using one or more machine learning models to generate a first predicted output that indicates a probability of one or more hotwords being present in the first audio data;

determine that the first predicted output satisfies a secondary threshold that is less indicative of the one or more hotwords being present in audio data than is a primary threshold but does not satisfy the primary threshold;

receive, via the one or more microphones of the client device, second audio data that captures a second spoken utterance of a user;

process the second audio data using the one or more machine learning models to generate a second predicted output that indicates a probability of the one or more hotwords being present in the second audio data;

determine that the second predicted output satisfies the secondary threshold but does not satisfy the primary threshold;

in response to the first predicted output and the second predicted output satisfying the secondary threshold but not satisfying the primary threshold, and in response to the first spoken utterance and the second spoken utterance satisfying one or more temporal criteria relative to one another, identify a failed hotword attempt; and

in response to identifying the failed hotword attempt:

determine an intended hotword corresponding to the failed hotword attempt, wherein neither the intended hotword nor another supported hotword is included in the first spoken utterance and the second spoken utterance;

provide a hint, comprising displaying the intended hotword on a display of the client device or providing, by the client device, an audio response that includes the intended hotword; and

perform an action corresponding to the intended hotword.

9. The computer program product according to claim 8 , wherein identifying the failed hotword attempt is further in response to determining that a similarity between the first spoken utterance and the second spoken utterance exceeds a similarity threshold.

10. The computer program product according to claim 8 , wherein identifying the failed hotword attempt is further in response to determining that the probability indicated by the first predicted output and the probability indicated by the second predicted output correspond to a same hotword of the one or more hotwords.

11. The computer program product according to claim 8 , wherein:

the program instructions are further executable to determine, using a model conditioned on acoustic features, that the first audio data and the second audio data comprise a command; and

identifying the failed hotword attempt is further in response to the first audio data and the second audio data comprising the command.

12. The computer program product according to claim 8 , wherein the intended hotword is determined based on acoustic similarity between at least a portion of the first audio data, at least a portion of the second audio data, and the intended hotword.

13. The computer program product according to claim 8 , wherein the intended hotword is determined based on semantic similarity between the intended hotword and text generated based on at least a portion of the first audio data or based on at least a portion of the second audio data.

14. The computer program product according to claim 8 , wherein the program instructions are further executable to, in response to identifying the failed hotword attempt, update the one or more machine learning models based on the first audio data and the second audio data.

15. A system comprising:

a processor, a computer-readable memory, one or more computer-readable storage media, and program instructions collectively stored on the one or more computer-readable storage media, the program instructions executable to:

receive, via one or more microphones of a client device, first audio data that captures a first spoken utterance of a user;

process the first audio data using one or more machine learning models to generate a first predicted output that indicates a probability of one or more hotwords being present in the first audio data;

determine that the first predicted output satisfies a secondary threshold that is less indicative of the one or more hotwords being present in audio data than is a primary threshold but does not satisfy the primary threshold;

receive, via the one or more microphones of the client device, second audio data that captures a second spoken utterance of a user;

process the second audio data using the one or more machine learning models to generate a second predicted output that indicates a probability of the one or more hotwords being present in the second audio data;

determine that the second predicted output satisfies the secondary threshold but does not satisfy the primary threshold;

in response to the first predicted output and the second predicted output satisfying the secondary threshold but not satisfying the primary threshold, and in response to the first spoken utterance and the second spoken utterance satisfying one or more temporal criteria relative to one another, identify a failed hotword attempt; and

in response to identifying the failed hotword attempt:

determine an intended hotword corresponding to the failed hotword attempt, wherein neither the intended hotword nor another supported hotword is included in the first spoken utterance and the second spoken utterance;

provide a hint, comprising displaying the intended hotword on a display of the client device or providing, by the client device, an audio response that includes the intended hotword; and

perform an action corresponding to the intended hotword.

16. The system according to claim 15 , wherein identifying the failed hotword attempt is further in response to determining that a similarity between the first spoken utterance and the second spoken utterance exceeds a similarity threshold.

17. The system according to claim 15 , wherein identifying the failed hotword attempt is further in response to determining that the probability indicated by the first predicted output and the probability indicated by the second predicted output correspond to a same hotword of the one or more hotwords.

18. The system according to claim 15 , wherein:

the program instructions are further executable to determine, using a model conditioned on acoustic features, that the first audio data and the second audio data comprise a command; and

identifying the failed hotword attempt is further in response to the first audio data and the second audio data comprising the command.

19. The system according to claim 15 , wherein the intended hotword is determined based on acoustic similarity between at least a portion of the first audio data, at least a portion of the second audio data, and the intended hotword.

20. The system according to claim 15 , wherein the intended hotword is determined based on semantic similarity between the intended hotword and text generated based on at least a portion of the first audio data or based on at least a portion of the second audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2023
From: SHARIFI, MATTHEW; CARBUNE, VICTOR
To: GOOGLE LLC
Reel/Frame 065650/0014 →
Continuity (3)
Continuation 17081645 · Oct 27, 2020
Provisional Application 63091237 · Oct 13, 2020
Related Publication 20240055002A1 · Feb 15, 2024
References Cited (13)
US 5255386A · Prager · 1993 [cited by examiner]
US 5737724A · Atal · 1998 [cited by examiner]
US 6697782B1 · Iso-Sipila · 2004 [cited by examiner]
US 9729690B2 · Byrne · 2017 [cited by examiner]
US 10650802B2 · Kunitake · 2020 [cited by examiner]
US 20150161990A1 · Sharifi · 2015 [cited by applicant]
US 20210012770A1 · Choudhary · 2021 [cited by examiner]
US 20220115011A1 · Sharifi et al. · 2022 [cited by applicant]
Intellectual Property India; Examination Report issued in Application No. 202227064647; 7 pages; dated Jul. 27, 2023. [cited by applicant]
European Patent Office; International Search Report and Written Opinion of PCT Application No. PCT/US2021/054606; 10 pages; dated Feb. 2, 2022. [cited by applicant]
European Patent Office, Intention to Grant issued in Application No. 21801796.0; 54 pages; dated Sep. 27, 2024. [cited by applicant]
Intellectual Property India; Hearing Notice issued in Application No. 202227064647; 2 pages; dated Oct. 7, 2024. [cited by applicant]
European Patent Office; Extended European Search Report issued in Application No. 25153790.8-1207; 6 pages; dated Apr. 17, 2025. [cited by applicant]