IP Library Granted Patent US 12,334,071
Granted Patent B2
US 12,334,071 · App. 18/307,736 · Granted Jun 17, 2025

Decaying automated speech recognition processing results

Inventors: Matthew Sharifi (Kilchberg, CH); Victor Carbune (Zürich, CH)
Assignee: Google LLC
G10L15/22G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,334,071
App. No.
18/307,736
Granted
Jun 17, 2025
Kind
B2
Abstract

A method for decaying speech processing includes receiving, at a voice-enabled device, an indication of a microphone trigger event indicating a possible interaction with the device through speech where the device has a microphone that, when open, is configured to capture speech for speech recognition. In response to receiving the indication of the microphone trigger event, the method also includes instructing the microphone to open or remain open for a duration window to capture an audio stream in an environment of the device and providing the audio stream captured by the open microphone to a speech recognition system. During the duration window, the method further includes decaying a level of the speech recognition processing based on a function of the duration window and instructing the speech recognition system to use the decayed level of speech recognition processing over the audio stream captured by the open microphone.

Claims (36)

1. A computer-implemented method when executed by data processing hardware causes the data processing hardware to perform operations comprising:

receiving audio data corresponding to an utterance and captured by an open microphone;

generating, by processing the audio data using a first level of automated speech recognition (ASR) processing, a first-pass speech recognition result for the audio data;

determining a confidence level for the first-pass speech recognition result;

based on the confidence level for the first-pass speech recognition result, generating, by processing the audio data using a second level of ASR processing, a second-pass speech recognition result for the audio data, the second level of ASR processing greater than the first level of ASR processing, the second level of ASR processing decaying over time while processing the audio data using the second level of ASR processing; and

when the second level of ASR processing is equal to zero, instructing the microphone to close.

2. The computer-implemented method of claim 1 , wherein the operations further comprise:

determining that the confidence level for the first-pass speech recognition result fails to satisfy a confidence threshold; and

increasing the first level of ASR processing to the second level of ASR processing based on determining that the confidence level for the first-pass speech recognition result fails to satisfy the confidence threshold.

3. The computer-implemented method of claim 1 , wherein the first level of ASR processing comprises a partial processing capability of an ASR system.

4. The computer-implemented method of claim 1 , wherein generating, by processing the audio data using the first level of ASR processing, the first-pass speech recognition result for the audio data comprises performing speech recognition at a voice-enabled device.

5. The computer-implemented method of claim 1 , wherein the second level of ASR processing comprises a full processing capability of an ASR system.

6. The computer-implemented method of claim 1 , wherein generating, by processing the audio data using the second level of ASR processing, the second-pass speech recognition result for the audio data comprises performing speech recognition at a remote server in communication with a voice-enabled device.

7. The computer-implemented method of claim 1 , wherein the operations further comprise receiving an indication of a microphone trigger event indicating a possible user interaction with a voice-enabled device through speech, the voice-enabled device comprising the microphone, the microphone configured to capture speech for recognition by an ASR system in an open state.

8. The computer-implemented method of claim 7 , wherein the operations further comprise instructing the microphone to open or remain open for an open microphone duration window to capture the utterance.

9. The computer-implemented method of claim 8 , wherein the operations further comprise, while generating the second-pass speech recognition result, decaying the second level of ASR processing to a third level of ASR processing less than the second level of ASR processing.

10. The computer-implemented method of claim 1 , wherein the first and second levels of ASR processing each correspond to a respective amount of computing resources used to process the audio data.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving audio data corresponding to an utterance and captured by an open microphone;

generating, by processing the audio data using a first level of automated speech recognition (ASR) processing, a first-pass speech recognition result for the audio data;

determining a confidence level for the first-pass speech recognition result;

based on the confidence level for the first-pass speech recognition result, generating, by processing the audio data using a second level of ASR processing, a second-pass speech recognition result for the audio data, the second level of ASR processing greater than the first level of ASR processing, the second level of ASR processing decaying over time while processing the audio data using the second level of ASR processing; and

when the second level of ASR processing is equal to zero, instructing the microphone to close.

12. The system of claim 11 , wherein the operations further comprise:

determining that the confidence level for the first-pass speech recognition result fails to satisfy a confidence threshold; and

increasing the first level of ASR processing to the second level of ASR processing based on determining that the confidence level for the first-pass speech recognition result fails to satisfy the confidence threshold.

13. The system of claim 11 , wherein the first level of ASR processing comprises a partial processing capability of an ASR system.

14. The system of claim 11 , wherein generating, by processing the audio data using the first level of ASR processing, the first-pass speech recognition result for the audio data comprises performing speech recognition at a voice-enabled device.

15. The system of claim 11 , wherein the second level of ASR processing comprises a full processing capability of an ASR system.

16. The system of claim 11 , wherein generating, by processing the audio data using the second level of ASR processing, the second-pass speech recognition result for the audio data comprises performing speech recognition at a remote server in communication with a voice-enabled device.

17. The system of claim 11 , wherein the operations further comprise receiving an indication of a microphone trigger event indicating a possible user interaction with a voice-enabled device through speech, the voice-enabled device comprising the microphone, the microphone configured to capture speech for recognition by an ASR system in an open state.

18. The system of claim 17 , wherein the operations further comprise instructing the microphone to open or remain open for an open microphone duration window to capture the utterance.

19. The system of claim 18 , wherein the operations further comprise, while generating the second-pass speech recognition result, decaying the second level of ASR processing to a third level of ASR processing less than the second level of ASR processing.

20. The system of claim 11 , wherein the first and second levels of ASR processing each correspond to a respective amount of computing resources used to process the audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2023
From: SHARIFI, MATTHEW; CARBUNE, VICTOR
To: GOOGLE LLC
Reel/Frame 063454/0312 →
Continuity (2)
Continuation 17111467 · Dec 3, 2020
Related Publication 20240096320A1 · Mar 21, 2024
References Cited (25)
US 8463610B1 · Bourke et al. · 2013 [cited by applicant]
US 11087750B2 · Ganong, III et al. · 2021 [cited by applicant]
US 20090210230A1 · Schwarz · 2009 [cited by examiner]
US 20140244272A1 · Shao · 2014 [cited by examiner]
US 20150025890A1 · Jagatheesan et al. · 2015 [cited by applicant]
US 20150379987A1 · Panainte · 2015 [cited by examiner]
US 20180025731A1 · Lovitt · 2018 [cited by applicant]
US 20180025732A1 · Lepauloux et al. · 2018 [cited by applicant]
US 20180358019A1 · Mont-Reynaud · 2018 [cited by examiner]
US 20190027130A1 · Tsunoo · 2019 [cited by examiner]
US 20190043503A1 · Bauer et al. · 2019 [cited by applicant]
US 20190251960A1 · Maker · 2019 [cited by examiner]
US 20190318724A1 · Chao · 2019 [cited by examiner]
US 20190325862A1 · Shankar et al. · 2019 [cited by applicant]
US 20200098359A1 · Nakamae · 2020 [cited by examiner]
US 20200134151A1 · Magi · 2020 [cited by examiner]
US 20200175961A1 · Thomson et al. · 2020 [cited by applicant]
US 20200243094A1 · Thomson et al. · 2020 [cited by applicant]
US 20200365148A1 · Ji et al. · 2020 [cited by applicant]
US 20200410991A1 · Jost et al. · 2020 [cited by applicant]
US 20210183379A1 · Gharpure · 2021 [cited by examiner]
US 20210264899A1 · Kohara · 2021 [cited by examiner]
US 20220157318A1 · Sharifi et al. · 2022 [cited by applicant]
US 20220269762A1 · Zhao et al. · 2022 [cited by applicant]
USPTO. Office Action relating U.S. Appl. No. 17/111,467, dated Sep. 15, 2022. [cited by applicant]