IP Library Granted Patent US 11,676,594
Granted Patent B2
US 11,676,594 · App. 17/111,467 · Granted Jun 13, 2023

Decaying automated speech recognition processing results

Inventors: Matthew Sharifi (Kilchberg, CH); Victor Carbune (Zürich, CH)
Assignee: Google LLC
G10L15/22G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,676,594
App. No.
17/111,467
Granted
Jun 13, 2023
Kind
B2
Abstract

A method for decaying speech processing includes receiving, at a voice-enabled device, an indication of a microphone trigger event indicating a possible interaction with the device through speech where the device has a microphone that, when open, is configured to capture speech for speech recognition. In response to receiving the indication of the microphone trigger event, the method also includes instructing the microphone to open or remain open for a duration window to capture an audio stream in an environment of the device and providing the audio stream captured by the open microphone to a speech recognition system. During the duration window, the method further includes decaying a level of the speech recognition processing based on a function of the duration window and instructing the speech recognition system to use the decayed level of speech recognition processing over the audio stream captured by the open microphone.

Claims (78)

1. A method comprising:

receiving, at data processing hardware of a voice-enabled device, an indication of a microphone trigger event indicating a possible user interaction with the voice-enabled device through speech, the voice-enabled device having a microphone that, when open, is configured to capture speech for recognition by an automated speech recognition (ASR) system;

in response to receiving the indication of the microphone trigger event:

instructing, by the data processing hardware, the microphone to open or remain open for an open microphone duration window to capture an audio stream in an environment of the voice-enabled device; and

providing, by the data processing hardware, the audio stream captured by the open microphone to the ASR system to perform ASR processing over the audio stream; and

while the ASR system is performing the ASR processing over the audio stream captured by the open microphone:

decaying, by the data processing hardware, a level of ASR processing that the ASR system performs over the audio stream based on a function of the open microphone duration window, the level of ASR processing being decayed from a first level of ASR processing to a second level of ASR processing different from the first level of ASR processing, the second level of ASR processing performing speech recognition over the audio stream; and

instructing, by the data processing hardware, the ASR system to use the second level of ASR processing for performing speech recognition over the audio stream captured by the open microphone.

2. The method of claim 1 , further comprising, while the ASR system is performing the ASR processing over the audio stream captured by the open microphone:

determining, by the data processing hardware, whether voice activity is detected in the audio stream captured by the open microphone,

wherein decaying the level of ASR processing the ASR system performs over the audio stream is further based on the determination of whether any voice activity is detected in the audio stream.

3. The method of claim 1 , wherein:

the ASR system initially uses the first level of ASR processing to perform the ASR processing over the audio stream upon commencement of the open microphone duration window, the first level of ASR processing associated with full processing capabilities of the ASR system, and

decaying the level of ASR processing the ASR system performs over the audio stream based on the function of the open microphone duration window comprises:

determining whether a first interval of time has elapsed since commencing the open microphone duration window; and

when the first interval of time has elapsed, decaying the level of ASR processing the ASR system performs over the audio stream by reducing the level of ASR processing from the first level of ASR processing to the second level of ASR processing, the second level of ASR processing less than the first level of ASR processing.

4. The method of claim 1 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to switch from performing the ASR processing on a remote server in communication with the voice-enabled device to performing the ASR processing on the data processing hardware of the voice-enabled device.

5. The method of claim 1 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to switch from using a first ASR model to a second ASR model for performing the ASR processing over the audio stream, the second ASR model comprising fewer parameters than the first ASR model.

6. The method of claim 1 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to reduce a number of ASR processing steps performed over the audio stream.

7. The method of claim 1 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to adjust beam search parameters to reduce a decoding search space of the ASR system.

8. The method of claim 1 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to perform quantization and/or sparsification on one or more parameters of the ASR system.

9. The method of claim 1 , further comprising:

obtaining, by the data processing hardware, a current context when the indication of the microphone trigger event is received,

wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to bias speech recognition results based on the current context.

10. The method of claim 1 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to switch from system on a chip-based (SOC-based) processing to perform the ASR processing on the audio stream to digital signal processor-based (DSP-based) processing to perform the ASR processing on the audio stream.

11. The method of claim 1 , wherein, while the ASR system is using the second level of ASR processing over the audio stream captured by the open microphone, the ASR system is configured to:

generate a speech recognition result for audio data corresponding to a query spoken by a user; and

provide the speech recognition result to an application to perform an action specified by the query.

12. The method of claim 1 , further comprising, after instructing the ASR system to use the second level of the ASR processing over the audio stream:

receiving, at the data processing hardware, an indication that a confidence for a speech recognition result for a voice query that is output by the ASR system fails to satisfy a confidence threshold; and

instructing, by the data processing hardware, the ASR system to:

increase the level of ASR processing from the second level of ASR processing to a third level of ASR processing; and

reprocess the voice query using the third level of ASR processing.

13. The method of claim 1 , further comprising, while the ASR system is performing the ASR processing over the audio stream captured by the open microphone:

decaying, by the data processing hardware, the level of ASR processing from the second level of ASR processing to a third level of ASR processing;

determining, by the data processing hardware, when the third level of ASR processing the ASR system performs over the audio stream based on the function of the open microphone duration window is equal to zero; and

when the third level of ASR processing the ASR system performs is equal to zero, instructing, by the data processing hardware, the microphone to close.

14. The method of claim 1 , further comprising displaying, by the data processing hardware, in a graphical user interface of the voice-enabled device, a graphical indicator indicating the second level of ASR processing is being performed by the ASR system on the audio stream.

15. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware, cause the data processing hardware to perform operations comprising:

receiving, at a voice-enabled device, an indication of a microphone trigger event indicating a possible user interaction with the voice-enabled device through speech, the voice-enabled device having a microphone that, when open, is configured to capture speech for recognition by an automated speech recognition (ASR) system;

in response to receiving the indication of the microphone trigger event:

instructing the microphone to open or remain open for an open microphone duration window to capture an audio stream in an environment of the voice-enabled device; and

providing the audio stream captured by the open microphone to the ASR system to perform ASR processing over the audio stream; and

while the ASR system is performing the ASR processing over the audio stream captured by the open microphone:

decaying, by the data processing hardware, a level of ASR processing that the ASR system performs over the audio stream based on a function of the open microphone duration window, the level of ASR processing being decayed from a first level of ASR processing to a second level of ASR processing different from the first level of ASR processing, the second level of ASR processing performing speech recognition over the audio stream; and

instructing, by the data processing hardware, the ASR system to use the second level of ASR processing for performing speech recognition over the audio stream captured by the open microphone.

16. The system of claim 15 , wherein the operations further comprise, while the ASR system is performing the ASR processing over the audio stream captured by the open microphone:

determining whether voice activity is detected in the audio stream captured by the open microphone,

wherein decaying the level of ASR processing the ASR system performs over the audio stream is further based on the determination of whether any voice activity is detected in the audio stream.

17. The system of claim 15 , wherein:

the ASR system initially uses the first level of ASR processing to perform the ASR processing over the audio stream upon commencement of the open microphone duration window, the first level of ASR processing associated with full processing capabilities of the ASR system, and

decaying the level of ASR processing the ASR system performs over the audio stream based on the function of the open microphone duration window comprises:

determining whether a first interval of time has elapsed since commencing the open microphone duration window; and

when the first interval of time has elapsed, decaying the level of ASR processing the ASR system performs over the audio stream by reducing the level of ASR processing from the first level of ASR processing to the second level of ASR processing, the second level of ASR processing less than the first level of ASR processing.

18. The system of claim 15 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to switch from performing the ASR processing on a remote server in communication with the voice-enabled device to performing the ASR processing on the data processing hardware of the voice-enabled device.

19. The system of claim 15 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to switch from using a first ASR model to a second ASR model for performing the ASR processing over the audio stream, the second ASR model comprising fewer parameters than the first ASR model.

20. The system of claim 15 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to reduce a number of ASR processing steps performed over the audio stream.

21. The system of claim 15 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to adjust beam search parameters to reduce a decoding search space of the ASR system.

22. The system of claim 15 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to perform quantization and/or sparsification on one or more parameters of the ASR system.

23. The system of claim 15 , wherein the operations further comprise:

obtaining a current context when the indication of the microphone trigger event is received,

wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to bias speech recognition results based on the current context.

24. The system of claim 15 , wherein instructing the ASR system to use the second level of ASR processing comprises instructing the ASR system to switch from system on a chip-based (SOC-based) processing to perform the ASR processing on the audio stream to digital signal processor-based (DSP-based) processing to perform the ASR processing on the audio stream.

25. The system of claim 15 , wherein, while the ASR system is using the second level of ASR processing over the audio stream captured by the open microphone, the ASR system is configured to:

generate a speech recognition result for audio data corresponding to a query spoken by a user; and

provide the speech recognition result to an application to perform an action specified by the query.

26. The system of claim 15 , wherein the operations further comprise, after instructing the ASR system to use the second level of ASR processing over the audio stream:

receiving, at the data processing hardware, an indication that a confidence for a speech recognition result for a voice query that is output by the ASR system fails to satisfy a confidence threshold; and

instructing, by the data processing hardware, the ASR system to:

increase the level of ASR processing from the second level of ASR processing to a third level of ASR processing; and

reprocess the voice query using the third level of ASR processing.

27. The system of claim 15 , wherein the operations further comprise, while the ASR system is performing the ASR processing over the audio stream captured by the open microphone:

decaying, by the data processing hardware, the level of ASR processing from the second level of ASR processing to a third level of ASR processing;

determining, by the data processing hardware, when the third level of ASR processing the ASR system performs over the audio stream based on the function of the open microphone duration window is equal to zero; and

when the third level of ASR processing the ASR system performs is equal to zero, instructing, by the data processing hardware, the microphone to close.

28. The system of claim 15 , wherein the operations further comprise displaying, by the data processing hardware, in a graphical user interface of the voice-enabled device, a graphical indicator indicating the second level of ASR processing is being performed by the ASR system on the audio stream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2020
From: SHARIFI, MATTHEW; CARBUNE, VICTOR
To: GOOGLE LLC
Reel/Frame 054560/0153 →
Continuity (1)
Related Publication 20220180866A1 · Jun 9, 2022