IP Library › Granted Patent US 11,688,392
Granted Patent B2
US 11,688,392 · App. 17/115,742 · Granted Jun 27, 2023

Freeze words

Inventors: Matthew Sharifi (Kilchberg, CH); Aleksandar Kracun (New York, NY)
Assignee: Google LLC
G10L15/16G10L15/05G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,688,392
App. No.
17/115,742
Granted
Jun 27, 2023
Kind
B2
Abstract

A method for detecting freeze words includes receiving audio data that corresponds to an utterance spoken by a user and captured by a user device associated with the user. The method also includes processing, using a speech recognizer, the audio data to determine that the utterance includes a query for a digital assistant to perform an operation. The speech recognizer is configured to trigger endpointing of the utterance after a predetermined duration of non-speech in the audio data. Before the predetermined duration of non-speech, the method includes detecting a freeze word in the audio data. In response to detecting the freeze word in the audio data, the method also includes triggering a hard microphone closing event at the user device. The hard microphone closing event prevents the user device from capturing any audio subsequent to the freeze word.

Claims (106)

1. A method for detecting freeze words, the method comprising:

receiving, at data processing hardware, audio data corresponding to an utterance spoken by a user and captured by a user device associated with the user;

processing, by the data processing hardware, using a speech recognizer, the audio data to determine that the utterance includes a query for a digital assistant to perform an operation, wherein the speech recognizer is configured to trigger endpointing of the utterance after a predetermined duration of non-speech in the audio data; and

before the predetermined duration of non-speech in the audio data:

detecting, by the data processing hardware, a freeze word in the audio data, the freeze word following the query in the utterance spoken by the user and captured by the user device; and

in response to detecting the freeze word in the audio data:

triggering, by the data processing hardware, a hard microphone closing event at the user device to prevent the user device from capturing any further audio data for the utterance subsequent to the freeze word;

modifying, by the data processing hardware, a speech recognition result for the audio data by stripping the freeze word from the speech recognition result; and

providing, by the data processing hardware, for output from the user device, the speech recognition result.

2. The method of claim 1 , wherein the freeze word comprises one of:

a predefined freeze word comprising one or more fixed terms across all users in a given language;

a user-selected freeze word comprising one or more terms specified by the user of the user device; or

an action-specific freeze word associated with the operation to be performed by the digital assistant.

3. The method of claim 1 , wherein detecting the freeze word in the audio data comprises:

extracting audio features from the audio data;

generating, using a freeze word detection model, a freeze word confidence score by processing the extracted audio features, the freeze word detection model executing on the data processing hardware; and

determining that the audio data corresponding to the utterance includes the freeze word when the freeze word confidence score satisfies a freeze word confidence threshold.

4. The method of claim 1 , wherein detecting the freeze word in the audio data comprises recognizing, using the speech recognizer executing on the data processing hardware, the freeze word in the audio data.

5. The method of claim 1 , further comprising, in response to detecting the freeze word in the audio data:

instructing, by the data processing hardware, the speech recognizer to cease any active processing on the audio data; and

instructing, by the data processing hardware, the digital assistant to fulfill performance of the operation.

6. The method of claim 1 , wherein processing the audio data to determine that the utterance includes the query for the digital assistant to perform the operation comprises:

processing, using the speech recognizer, the audio data to generate a speech recognition result for the audio data; and

performing semantic interpretation on the speech recognition result for the audio data to determine that the audio data includes the query to perform the operation.

7. The method of claim 6 , further comprising, in response to detecting the freeze word in the audio data

instructing, by the data processing hardware, using the speech recognition result, the digital assistant to perform the operation requested by the query.

8. The method of claim 1 , further comprising, prior to processing the audio data using the speech recognizer:

detecting, by the data processing hardware, using a hotword detection model, a hotword in the audio data that precedes the query; and

in response to detecting the hotword, triggering, by the data processing hardware, the speech recognizer to process the audio data by performing speech recognition on the hotword and/or one or more terms following the hotword in the audio data.

9. The method of claim 8 , further comprising verifying, by the data processing hardware, a presence of the hotword detected by the hotword detection model based on detecting the freeze word in the audio data.

10. The method of claim 8 , wherein:

detecting the freeze word in the audio data comprises executing a freeze word detection model on the data processing hardware that is configured to detect the freeze word in the audio data without performing speech recognition on the audio data; and

the freeze word detection model and the hotword detection model each comprise the same or different neural network-based models.

11. A method for detecting a freeze word, the method comprising:

receiving, at data processing hardware, a first instance of audio data corresponding to a dictation-based query for a digital assistant to dictate audible contents spoken by a user, the dictation-based query spoken by the user and captured by an assistant-enabled device associated with the user;

receiving, at the data processing hardware, a second instance of the audio data corresponding to an utterance of the audible contents spoken by the user and captured by the assistant-enabled device;

processing, by the data processing hardware, using a speech recognizer, the second instance of the audio data to generate a transcription of the audible contents; and

during the processing of the second instance of the audio data:

detecting, by the data processing hardware, a freeze word in the second instance of the audio data, the freeze word following the audible contents in the utterance spoken by the user and captured by the assistant-enabled device; and

in response to detecting the freeze word in the second instance of the audio data:

stripping, by the data processing hardware, the freeze word from an end of the transcription of the audible contents spoken by the user; and

providing, by the data processing hardware, for output from the assistant-enabled device, the transcription of the audible contents spoken by the user.

12. The method of claim 11 , further comprising, in response to detecting the freeze word in the second instance of the audio data:

initiating, by the data processing hardware, a hard microphone closing event at the assistant-enabled device to prevent the assistant-enabled device from capturing any further audio for the utterance subsequent to the freeze word; and

ceasing, by the data processing hardware, any active processing on the second instance of the audio data.

13. The method of claim 11 , further comprising:

processing, by the data processing hardware, using the speech recognizer, the first instance of the audio data to generate a speech recognition result; and

performing, by the data processing hardware, semantic interpretation on the speech recognition result for the first instance of the audio data to determine that the first instance of the audio data comprises the dictation-based query to dictate the audible contents spoken by the user.

14. The method of claim 13 , further comprising, prior to initiating processing on the second instance of the audio data to generate the transcription:

determining, by the data processing hardware, that the dictation-based query specifies the freeze word based on the semantic interpretation performed on the speech recognition result for the first instance of the audio data; and

instructing, by the data processing hardware, an endpointer to increase an endpointing timeout duration for endpointing the utterance of the audible contents.

15. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving audio data corresponding to an utterance spoken by a user and captured by a user device associated with the user;

processing, using a speech recognizer, the audio data to determine that the utterance includes a query for a digital assistant to perform an operation, wherein the speech recognizer is configured to trigger endpointing of the utterance after a predetermined duration of non-speech in the audio data; and

before the predetermined duration of non-speech in the audio data:

detecting a freeze word in the audio data, the freeze word following the query in the utterance spoken by the user and captured by the user device; and

in response to detecting the freeze word in the audio data;

triggering a hard microphone closing event at the user device to prevent the user device from capturing any further audio data for the utterance subsequent to the freeze word;

modifying, by the data processing hardware, a speech recognition result for the audio data by stripping the freeze word from the speech recognition result; and

providing, for output by the user device, the speech recognition result.

16. The system of claim 15 , wherein the freeze word comprises one of:

a predefined freeze word comprising one or more fixed terms across all users in a given language;

a user-selected freeze word comprising one or more terms specified by the user of the user device; or

an action-specific freeze word associated with the operation to be performed by the digital assistant.

17. The system of claim 15 , wherein detecting the freeze word in the audio data comprises:

extracting audio features from the audio data;

generating, using a freeze word detection model, a freeze word confidence score by processing the extracted audio features, the freeze word detection model executing on the data processing hardware; and

determining that the audio data corresponding to the utterance includes the freeze word when the freeze word confidence score satisfies a freeze word confidence threshold.

18. The system of claim 15 , wherein detecting the freeze word in the audio data comprises recognizing, using the speech recognizer executing on the data processing hardware, the freeze word in the audio data.

19. The system of claim 15 , wherein the operations further comprise, in response to detecting the freeze word in the audio data:

instructing the speech recognizer to cease any active processing on the audio data; and

instructing the digital assistant to fulfill performance of the operation.

20. The system of claim 15 , wherein processing the audio data to determine that the utterance includes the query for the digital assistant to perform the operation comprises:

processing, using the speech recognizer, the audio data to generate a speech recognition result for the audio data; and

performing semantic interpretation on the speech recognition result for the audio data to determine that the audio data includes the query to perform the operation.

21. The system of claim 20 , wherein the operations further comprise, in response to detecting the freeze word in the audio data

instructing, using the speech recognition result, the digital assistant to perform the operation requested by the query.

22. The system of claim 15 , wherein the operations further comprise, prior to processing the audio data using the speech recognizer:

detecting, using a hotword detection model, a hotword in the audio data that precedes the query; and

in response to detecting the hotword, triggering the speech recognizer to process the audio data by performing speech recognition on the hotword and/or one or more terms following the hotword in the audio data.

23. The system of claim 22 , wherein the operations further comprise verifying a presence of the hotword detected by the hotword detection model based on detecting the freeze word in the audio data.

24. The system of claim 22 , wherein:

detecting the freeze word in the audio data comprises executing a freeze word detection model on the data processing hardware that is configured to detect the freeze word in the audio data without performing speech recognition on the audio data; and

the freeze word detection model and the hotword detection model each comprise the same or different neural network-based models.

25. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a first instance of audio data corresponding to a dictation-based query for a digital assistant to dictate audible contents spoken by a user, the dictation-based query spoken by the user and captured by an assistant-enabled device associated with the user;

receiving a second instance of the audio data corresponding to an utterance of the audible contents spoken by the user and captured by the assistant-enabled device;

processing, using a speech recognizer, the second instance of the audio data to generate a transcription of the audible contents; and

during the processing of the second instance of the audio data:

detecting a freeze word in the second instance of the audio data, the freeze word following the audible contents in the utterance spoken by the user and captured by the assistant-enabled device; and

in response to detecting the freeze word in the second instance of the audio data:

stripping, by the data processing hardware, the freeze word from an end of the transcription of the audible contents spoken by the user; and

providing, for output from the assistant-enabled device, the transcription of the audible contents spoken by the user.

26. The system of claim 25 , wherein the operations further comprise, in response detecting the freeze word in the second instance of the audio data:

initiating a hard microphone closing event at the assistant-enabled device to prevent the assistant-enabled device from capturing any further audio for the utterance subsequent to the freeze word; and

ceasing any active processing on the second instance of the audio data.

27. The system of claim 25 , wherein the operations further comprise:

processing, using the speech recognizer, the first instance of the audio data to generate a speech recognition result; and

performing semantic interpretation on the speech recognition result for the first instance of the audio data to determine that the first instance of the audio data comprises the dictation-based query to dictate the audible contents spoken by the user.

28. The system of claim 27 , wherein the operations further comprise, prior to initiating processing on the second instance of the audio data to generate the transcription:

determining that the dictation-based query specifies the freeze word based on the semantic interpretation performed on the speech recognition result for the first instance of the audio data; and

instructing an endpointer to increase an endpointing timeout duration for endpointing the utterance of the audible contents.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2020
From: SHARIFI, MATTHEW; KRACUN, ALEKSANDAR
To: GOOGLE LLC
Reel/Frame 054587/0211 →
Continuity (1)
Related Publication 20220180862A1 · Jun 9, 2022
Cited By (1)
US 12,315,497