IP Library › Granted Patent US 12,555,577
Granted Patent B2
US 12,555,577 · App. 18/323,725 · Granted Feb 17, 2026

Hotphrase triggering based on a sequence of detections

Inventors: Victor Carbune (Zürich, CH); Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L15/22G06F1/3231G06F16/24522G10L15/16G10L15/285G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,577
App. No.
18/323,725
Granted
Feb 17, 2026
Kind
B2
Abstract

A method includes receiving audio data corresponding to an utterance spoken by the user and captured by the user device. The utterance includes a command for a digital assistant to perform an operation. The method also includes determining, using a hotphrase detector configured to detect each trigger word in a set of trigger words associated with a hotphrase, whether any of the trigger words in the set of trigger words are detected in the audio data during the corresponding fixed-duration time window. The method also includes determining identifying, in the audio corresponding to the utterance, the hotphrase when each other trigger word in the set of trigger words was also detected in the audio data. The method also includes triggering an automated speech recognizer to perform speech recognition on the audio data when the hotphrase is identified in the audio data corresponding to the utterance.

Claims (64)

1 . A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

receiving audio data corresponding to a single utterance spoken by a user in streaming audio captured by a microphone, the single utterance comprising:

a command for a digital assistant to perform an operation;

a set of trigger words associated with a hotphrase; and

one or more other words that are not associated with the hotphrase, the one or more other words spoken between a first trigger word in the set of trigger words and a last trigger word in the set of trigger words, each of the one or more other words is different than each trigger word of the set of trigger words;

before triggering an automated speech recognizer (ASR) to perform speech recognition on the audio data:

determining that the first trigger word in the set of trigger words is detected in the audio data characterizing the single utterance;

after determining that the first trigger word in the set of trigger words is detected in the audio data, determining that each other trigger word in the set of trigger words is also detected in the audio data characterizing the single utterance in a sequence that matches a predefined sequential order associated with the hotphrase; and

based on determining that the first trigger word and each other trigger word in the set of trigger words is detected in the audio data characterizing the single utterance in the sequence that matches the predefined sequential order associated with the hotphrase, identifying the hotphrase in the audio data; and

based on identifying the hotphrase in the audio data, triggering the ASR to perform speech recognition on the audio data.

2 . The method of claim 1 , wherein triggering the ASR to perform speech recognition on the audio data comprises:

generating a transcription of the single utterance by processing the audio data; and

performing query interpretation on the transcription to identify that the transcription includes the command for the digital assistant to perform the operation.

3 . The method of claim 2 , wherein generating the transcription comprises:

rewinding the audio data buffered in memory hardware in communication with the data processing to a time at or before the first trigger word in the set of trigger words was detected in the audio data; and

processing the audio data commencing at the time at or before the first trigger word in the set of trigger words to generate the transcription of the single utterance.

4 . The method of claim 2 , wherein the transcription comprises, between the first trigger word in the set of trigger words and the last trigger word in the set of trigger words, the one or more other words.

5 . The method of claim 1 , wherein the operations further comprise:

determining that each other trigger word in the set of trigger words is detected in the audio data during a fixed-duration time window commencing when the first trigger word in the set of trigger words was detected in the audio data,

wherein triggering the ASR to perform speech recognition processing is based on determining that each other trigger word in the set of trigger words is detected in the audio data during the fixed-duration time window.

6 . The method of claim 1 , wherein determining that the first trigger word in the set of trigger words is detected in the audio data characterizing the single utterance comprises:

generating, using a hotphrase detector, a trigger word confidence score indicating a likelihood that the first trigger word is present in the audio data;

detecting the first trigger word in the audio data when the trigger word confidence score satisfies a trigger word confidence threshold; and

buffering, in memory hardware in communication with the data processing hardware, the audio data and a trigger event for the first trigger word detected in the audio data, the trigger event indicating the trigger word confidence score and a timestamp indicating when the first trigger word was detected in the audio data.

7 . The method of claim 6 , wherein the operations further comprise, based on determining that the first trigger word in the set of trigger words is detected in the audio data characterizing the single utterance, executing a trigger word aggregation routine configured to:

determine that a respective trigger event for each other corresponding trigger word in the set of trigger words is also buffered in the memory hardware; and

when the respective trigger event for each other corresponding trigger word in the set of trigger words is also buffered in the memory hardware, determine a hotphrase confidence score indicating a likelihood that the single utterance spoken by the user includes the set of trigger words,

wherein triggering the ASR to perform speech recognition on the audio data comprises triggering the ASR to perform speech recognition on the audio data when the hotphrase confidence score satisfies a hotphrase confidence threshold.

8 . The method of claim 7 , wherein executing the trigger word aggregation routine comprises executing a neural network-based model.

9 . The method of claim 7 , wherein executing the trigger word aggregation routine comprises executing a heuristic-based model.

10 . The method of claim 1 , wherein the data processing hardware resides on a user device.

11 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving audio data corresponding to a single utterance spoken by a user in streaming audio captured by a microphone, the single utterance comprising:

a command for a digital assistant to perform an operation;

a set of trigger words associated with a hotphrase; and

one or more other words that are not associated with the hotphrase, the one or more other words spoken between a first trigger word in the set of trigger words and a last trigger word in the set of trigger words, each of the one or more other words is different than each trigger word of the set of trigger words;

before triggering an automated speech recognizer (ASR) to perform speech recognition on the audio data:

determining that the first trigger word in the set of trigger words is detected in the audio data characterizing the single utterance;

after determining that the first trigger word in the set of trigger words is detected in the audio data, determining that each other trigger word in the set of trigger words is also detected in the audio data characterizing the single utterance in a sequence that matches a predefined sequential order associated with the hotphrase; and

based on determining that the first trigger word and each other trigger word in the set of trigger words is detected in the audio data characterizing the single utterance in the sequence that matches the predefined sequential order associated with the hotphrase, identifying the hotphrase in the audio data; and

based on identifying the hotphrase in the audio data, triggering the ASR to perform speech recognition on the audio data.

12 . The system of claim 11 , wherein triggering the ASR to perform speech recognition on the audio data comprises:

generating a transcription of the single utterance by processing the audio data; and

performing query interpretation on the transcription to identify that the transcription includes the command for the digital assistant to perform the operation.

13 . The system of claim 12 , wherein generating the transcription comprises:

rewinding the audio data buffered in memory hardware in communication with the data processing to a time at or before the first trigger word in the set of trigger words was detected in the audio data; and

processing the audio data commencing at the time at or before the first trigger word in the set of trigger words to generate the transcription of the single utterance.

14 . The system of claim 12 , wherein the transcription comprises, between the first trigger word in the set of trigger words and the last trigger word in the set of trigger words, the one or more other words.

15 . The system of claim 11 , wherein the operations further comprise:

determining that each other trigger word in the set of trigger words is detected in the audio data during a fixed-duration time window commencing when the first trigger word in the set of trigger words was detected in the audio data,

wherein triggering the ASR to perform speech recognition processing is based on determining that each other trigger word in the set of trigger words is detected in the audio data during the fixed-duration time window.

16 . The system of claim 11 , wherein determining that the first trigger word in the set of trigger words is detected in the audio data characterizing the single utterance comprises:

generating, using a hotphrase detector, a trigger word confidence score indicating a likelihood that the first trigger word is present in the audio data;

detecting the first trigger word in the audio data when the trigger word confidence score satisfies a trigger word confidence threshold; and

buffering, in memory hardware in communication with the data processing hardware, the audio data and a trigger event for the first trigger word detected in the audio data, the trigger event indicating the trigger word confidence score and a timestamp indicating when the first trigger word was detected in the audio data.

17 . The system of claim 16 , wherein the operations further comprise, based on determining that the first trigger word in the set of trigger words is detected in the audio data characterizing the single utterance, executing a trigger word aggregation routine configured to:

determine that a respective trigger event for each other corresponding trigger word in the set of trigger words is also buffered in the memory hardware; and

when the respective trigger event for each other corresponding trigger word in the set of trigger words is also buffered in the memory hardware, determine a hotphrase confidence score indicating a likelihood that the single utterance spoken by the user includes the set of trigger words,

wherein triggering the ASR to perform speech recognition on the audio data comprises triggering the ASR to perform speech recognition on the audio data when the hotphrase confidence score satisfies a hotphrase confidence threshold.

18 . The system of claim 17 , wherein executing the trigger word aggregation routine comprises executing a neural network-based model.

19 . The system of claim 17 , wherein executing the trigger word aggregation routine comprises executing a heuristic-based model.

20 . The system of claim 11 , wherein the data processing hardware resides on a user device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2023
From: CARBUNE, VICTOR; SHARIFI, MATTHEW
To: GOOGLE LLC
Reel/Frame 063763/0845 →
Continuity (2)
Continuation 17118251 · Dec 10, 2020
Related Publication 20230298588A1 · Sep 21, 2023
References Cited (25)
US 7487091B2 · Miyazaki · 2009 [cited by applicant]
US 9916831B2 · Panin · 2018 [cited by examiner]
US 10276175B1 · Garcia · 2019 [cited by examiner]
US 11488600B2 · Neumann · 2022 [cited by examiner]
US 20040199388A1 · Armbruster · 2004 [cited by examiner]
US 20150095027A1 · Parada San Martin et al. · 2015 [cited by applicant]
US 20190043488A1 · Bocklet et al. · 2019 [cited by applicant]
US 20190266240A1 · Georges et al. · 2019 [cited by applicant]
US 20200175962A1 · Thomson · 2020 [cited by examiner]
US 20200273447A1 · Zhou · 2020 [cited by applicant]
US 20200302913A1 · Marcinkiewicz · 2020 [cited by applicant]
US 20200395006A1 · Smith et al. · 2020 [cited by applicant]
US 20210225363A1 · Iwase et al. · 2021 [cited by applicant]
JP 2006215499A · 2006 [cited by applicant]
JP 2019133156A · 2019 [cited by applicant]
JP 2019185011A · 2019 [cited by applicant]
JP 2020012954A · 2020 [cited by applicant]
JP 2020160387A · 2020 [cited by applicant]
WO 0101389A2 · 2001 [cited by applicant]
WO 2019079957A1 · 2019 [cited by applicant]
WO 2019239656A1 · 2019 [cited by applicant]
International Search Report and Written Opinion for the related Application No. PCT/US2021/060233, Dated: Mar. 4, 2022, 119 pages. [cited by applicant]
USPTO. Office Action relating to U.S. Appl. No. 17/118,251, dated Dec. 13, 2022. [cited by applicant]
Japanese Patent Office, Office Action for Application No. 2023-535570 dated Aug. 5, 2024. [cited by applicant]
Japanese Patent Office, Office Action for Application No. 2024-215360 dated Jul. 29, 2025. [cited by applicant]